A multi-module parallel acceleration method, system and platform for processing images suitable for GPUs
By introducing a multi-module parallel acceleration method for image processing, GPU performance is detected and optimized in real time. This solves the problems of single acceleration methods and lack of performance analysis in existing technologies, and achieves a significant improvement in GPU performance as well as system stability and efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-22
- Publication Date
- 2026-04-10
AI Technical Summary
Existing GPU acceleration systems rely on a single acceleration method and lack performance analysis and optimization modules, which fails to significantly improve GPU performance and hinders its application potential in high-performance computing, big data processing, and artificial intelligence.
The system incorporates modules for automatic data acquisition, GPU self-detection, code optimization, parallel data processing, and performance analysis and optimization. By using these multiple modules in parallel to accelerate image processing and monitor and optimize GPU performance in real time, the system ensures the stability and effectiveness of the acceleration effect.
By working collaboratively across multiple modules, the acceleration effect of the GPU and the actual application effect of the system are significantly improved, thereby enhancing the performance of the GPU and the overall efficiency and reliability of the system.
Smart Images

Figure CN119478168B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of GPU image processing, and particularly relates to a multi-module parallel acceleration processing image method, system and platform suitable for a GPU. BACKGROUND
[0002] A GPU (Graphics Processing Unit) is a microprocessor specially used for image and graphics processing on personal computers, workstations, game consoles and mobile devices. It reduces the dependence on the central processing unit (CPU) and is specially used for processing image and graphics tasks originally processed by the CPU. The GPU uses core technologies such as hardware T&L (transformation and lighting), cubic environment texture mapping, vertex blending, texture compression and bump mapping to significantly improve the efficiency and quality of image rendering in 3D graphics processing.
[0003] With the increasing demand for computing, the performance improvement of the GPU has become the focus of research. In order to improve the processing capacity of the GPU, the existing GPU acceleration system has carried out various optimizations, such as the design of parallel computing architecture, the application of hardware acceleration technology and the optimization of software algorithm. However, the current GPU acceleration system still has some deficiencies in realizing performance improvement. For example, the acceleration method in the prior art is often single, lacking multi-dimensional performance optimization means, resulting in insufficient acceleration effect.
[0004] In addition, many existing systems lack effective performance analysis and optimization modules, making it difficult to detect and adjust the acceleration strategy in real time, limiting the further improvement of the overall performance of the system.
[0005] At present, although the performance of the GPU has been improved to some extent, the acceleration method is still relatively single, which fails to fully utilize the parallel processing advantage of the GPU, and the efficient performance analysis and optimization module is not integrated, which cannot monitor and dynamically adjust the acceleration effect in real time. These problems limit the application potential of the GPU acceleration system in the fields of high-performance computing, big data processing and artificial intelligence.
[0006] That is, there are the following technical problems and defects: the existing acceleration method is single: the existing GPU acceleration system has limited acceleration methods and cannot significantly improve the performance of the GPU. Lack of performance analysis and optimization module: unable to detect and optimize the performance of the accelerated GPU, affecting the actual application effect of the overall system.
[0007] Therefore, in view of the above technical problems and defects, it is urgent to design and develop a multi-module parallel acceleration processing image method, system and platform suitable for a GPU. SUMMARY
[0008] In order to overcome the deficiencies and difficulties existing in the prior art, the purpose of the present application is to solve the problems of single acceleration mode and lack of performance analysis optimization module in the prior art, and provide a GPU suitable multi-module parallel acceleration image processing method, system, platform and storage medium, so as to realize multi-module parallel acceleration: introduce data automatic acquisition module, GPU self-detection module, code optimization processing module, data parallel processing module and performance analysis optimization module, and multiple mode synchronous acceleration GPU. Performance analysis and optimization: increase the performance analysis optimization module, analyze the performance of the accelerated GPU, further optimize the acceleration process according to the results, and ensure the stability and effectiveness of the GPU acceleration effect.
[0009] The first purpose of the present application is to provide a GPU suitable multi-module parallel acceleration image processing method; the second purpose of the present application is to provide a GPU suitable multi-module parallel acceleration image processing system; the third purpose of the present application is to provide a GPU suitable multi-module parallel acceleration image processing platform; and the fourth purpose of the present application is to provide a computer readable storage medium.
[0010] The first purpose of the present application is achieved by the method comprising the following steps:
[0011] Obtain first information data corresponding to the GPU, detect the running performance of the GPU in real time based on the first information data, and generate corresponding second information data; wherein the first information data is the basic data of the GPU, including: GPU parameter data, running environment information data, power consumption state data and running state data; the second information data is the original running performance index data corresponding to the GPU;
[0012] According to the second information data, the running performance corresponding to the GPU is optimized in real time, and corresponding third information data is generated; wherein the third information data is the index data corresponding to the GPU and optimized by the running performance;
[0013] Based on the GPU and in combination with the third information data, the image data to be processed is processed and rendered in multi-module parallel acceleration.
[0014] Further, the first information data corresponding to the GPU is obtained, the running performance of the GPU is detected in real time based on the first information data, and the corresponding second information data is generated, which further comprises:
[0015] Generate and obtain model data, architecture data corresponding to the GPU, and GPU current task data requiring acceleration;
[0016] Real-time detection of working environment data corresponding to the GPU, and monitoring of power consumption state data corresponding to the GPU;
[0017] Generate and obtain running state information data corresponding to the GPU; wherein the running state information data includes GPU usage rate data and memory occupation data.
[0018] Further, the first information data corresponding to the GPU is acquired, the running performance of the GPU is detected in real time based on the first information data, and corresponding second information data is generated.
[0019] Detect the to-be-detected performance index data corresponding to the GPU, and generate first detection data corresponding to the to-be-detected performance index data; wherein the to-be-detected performance index data includes the calculation running performance, the memory capacity, the GPU running frequency, the GPU power consumption, the GPU running frequency, and the GPU running frequency corresponding to the GPU.
[0020] Based on the first detection data, it is determined whether the performance index data corresponding to the GPU reaches the preset standard data, if yes, the next step is executed; otherwise, the self-detection is continued.
[0021] Further, the running performance corresponding to the GPU is optimized in real time according to the second information data, and corresponding third information data is generated.
[0022] Based on the second information data, first code data corresponding to the first information data is generated and obtained, the first code data is optimized, and second code data corresponding to the first information data is generated, wherein the first code data is code data that needs to be optimized; the second code data is optimized code data.
[0023] Based on the second information data, first data set corresponding to the first information data is generated and obtained, and the first data set is divided and processed, and second data set corresponding to the first information data is generated, wherein the first data set is the current data set data of the GPU; the second data set is the sub data set corresponding to the first data set.
[0024] Further, the running performance corresponding to the GPU is optimized in real time according to the second information data, and corresponding third information data is generated.
[0025] Generate and obtain state data corresponding to the GPU, and allocate and process the second data set corresponding to the first information data;
[0026] Based on the state data corresponding to the GPU, the second data set corresponding to the GPU is calculated in parallel.
[0027] Further, the image acceleration processing unit, based on the GPU and in combination with the third information data, multi-module and parallel accelerated processing and rendering of the image data to be processed, further comprises:
[0028] In combination with the performance analysis tool, real-time estimation and generation of performance data corresponding to the GPU, and generation of a corresponding evaluation report according to the performance data;
[0029] Based on the performance data and the evaluation report, the code data and model data corresponding to the GPU are respectively optimized and adjusted.
[0030] The second object of the application is achieved in that the system is used to implement the image method suitable for multi-module and parallel accelerated processing of the GPU, and the system comprises:
[0031] The data acquisition and generation unit is used to acquire first information data corresponding to the GPU, to detect the running performance of the GPU in real time based on the first information data, and to generate corresponding second information data; wherein the first information data is the basic data of the GPU, including: GPU parameter data, running environment information data, power consumption state data and running state data; and the second information data is the original running performance index data corresponding to the GPU;
[0032] The data parallel processing unit is used to optimize the running performance corresponding to the GPU in real time according to the second information data, and to generate corresponding third information data; wherein the third information data is the index data corresponding to the GPU and processed by the running performance optimization;
[0033] The image acceleration processing unit is used to, based on the GPU and in combination with the third information data, multi-module and parallel accelerated processing and rendering of the image data to be processed.
[0034] Further, the data acquisition and generation unit further comprises:
[0035] The first data generation module is used to generate and acquire model data, architecture data corresponding to the GPU, and GPU current acceleration task data;
[0036] The data monitoring and processing module is used to detect the working environment data corresponding to the GPU in real time, and to monitor the power consumption state data corresponding to the GPU;
[0037] The second data generation module is used to generate and acquire running state information data corresponding to the GPU; wherein the running state information data includes GPU usage rate data and memory occupation data;
[0038] And / or, the data acquisition and generation unit further comprises:
[0039] The third data generation module is configured to detect to-be-detected performance index data corresponding to the GPU, and generate first detection data corresponding to the to-be-detected performance index data; wherein the to-be-detected performance index data includes calculation running performance, memory capacity, GPU memory bit width, GPU power consumption, GPU running frequency, and GPU memory type corresponding to the GPU.
[0040] The first data determination module is configured to determine, based on the first detection data, whether the performance index data corresponding to the GPU reaches preset standard data.
[0041] And / or, the data parallelization processing unit further includes:
[0042] The fourth data generation module is configured to generate and acquire first code data corresponding to the first information data based on the second information data, optimize the first code data, and generate second code data corresponding to the first information data; wherein the first code data is code data that needs to be optimized, and the second code data is optimized code data.
[0043] The fifth data generation module is configured to generate and acquire a first data set corresponding to the first information data based on the second information data, divide and process the first data set, and generate a second data set corresponding to the first information data; wherein the first data set is a current data set of the GPU, and the second data set is a sub-data set corresponding to the first data set.
[0044] And / or, the data parallelization processing unit further includes:
[0045] The sixth data generation module is configured to generate and acquire state data corresponding to the GPU, and allocate and process the second data set corresponding to the first information data.
[0046] The data parallelization calculation processing module is configured to parallelize and calculate the second data set corresponding to the GPU based on the state data corresponding to the GPU.
[0047] And / or, the image acceleration processing unit further includes:
[0048] The seventh data generation module is configured to combine a performance analysis tool, estimate and generate performance data corresponding to the GPU in real time, and generate a corresponding evaluation report based on the performance data.
[0049] The data optimization adjustment processing module is configured to optimize and adjust code data and model data corresponding to the GPU based on the performance data and the evaluation report, respectively.
[0050] A third object of the present application is achieved by comprising a processor, a memory and a multi-module parallel acceleration image processing platform control program suitable for GPU; wherein the processor executes the multi-module parallel acceleration image processing platform control program suitable for GPU, the multi-module parallel acceleration image processing platform control program suitable for GPU is stored in the memory, and the multi-module parallel acceleration image processing platform control program suitable for GPU implements the multi-module parallel acceleration image processing method suitable for GPU.
[0051] A fourth object of the present application is achieved by storing a multi-module parallel acceleration image processing platform control program suitable for GPU in the computer readable storage medium, and the multi-module parallel acceleration image processing platform control program suitable for GPU implements the multi-module parallel acceleration image processing method suitable for GPU.
[0052] The present application obtains first information data corresponding to the GPU by a method, detects the running performance of the GPU in real time based on the first information data, and generates corresponding second information data; wherein the first information data is the basic data of the GPU, including GPU parameter data, running environment information data, power consumption state data and running state data; the second information data is the original running performance index data corresponding to the GPU; according to the second information data, the running performance corresponding to the GPU is optimized in real time, and corresponding third information data is generated; wherein the third information data is the index data corresponding to the GPU and optimized by the running performance; based on the GPU and in combination with the third information data, the image data to be processed is processed and rendered in parallel by multiple modules, and the system, platform and storage medium corresponding to the method are used to realize multi-module parallel acceleration: a data automatic acquisition module, a GPU self-detection module, a code optimization processing module, a data parallel processing module and a performance analysis optimization module are introduced, and multiple modes are synchronized to accelerate the GPU. Performance analysis and optimization: the performance analysis optimization module is added to analyze the performance of the accelerated GPU, the acceleration process is further optimized according to the result, and the stability and effectiveness of the GPU acceleration effect are ensured.
[0053] That is to say, in the scheme of the present application, by setting the data automatic acquisition module and the GPU self-detection module in the system, the various parameter data about the GPU can be acquired and detected, the GPU itself is ensured to be in a normal state, the success rate of the GPU acceleration is ensured, and through the code optimization processing module, the data parallelization processing module and the model parallelization processing module, the GPU is simultaneously accelerated through the three modes, the acceleration effect and effectiveness of the GPU are effectively ensured, and the performance analysis optimization module is further set in the system, the performance of the GPU after acceleration is analyzed, the acceleration process is further optimized according to the analysis result, the optimization of the GPU acceleration effect is helped, and the actual application effect of the system is improved. BRIEF DESCRIPTION OF DRAWINGS
[0054] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed to be used in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative effort.
[0055] Figure 1 A module structure schematic diagram of an embodiment of the present application of a multi-module parallel acceleration processing image method suitable for a GPU;
[0056] Figure 2 A sub-module structure schematic diagram of a data automatic acquisition module of an embodiment of the present application of a multi-module parallel acceleration processing image method suitable for a GPU;
[0057] Figure 3 A sub-module structure schematic diagram of a GPU self-detection module of an embodiment of the present application of a multi-module parallel acceleration processing image method suitable for a GPU;
[0058] Figure 4 A sub-module structure schematic diagram of a code optimization processing module of an embodiment of the present application of a multi-module parallel acceleration processing image method suitable for a GPU;
[0059] Figure 5 A sub-module structure schematic diagram of a data parallelization processing module of an embodiment of the present application of a multi-module parallel acceleration processing image method suitable for a GPU;
[0060] Figure 6 A sub-module structure schematic diagram of a performance analysis optimization module of an embodiment of the present application of a multi-module parallel acceleration processing image method suitable for a GPU;
[0061] Figure 7A flow step schematic diagram of a multi-module parallel acceleration image processing method suitable for a GPU according to the present application;
[0062] Figure 8 A system architecture schematic diagram of a multi-module parallel acceleration image processing system suitable for a GPU according to the present application;
[0063] Figure 9 A platform architecture schematic diagram of a multi-module parallel acceleration image processing platform suitable for a GPU according to the present application;
[0064] Figure 10 A computer readable storage medium architecture schematic diagram in an embodiment of the present application;
[0065] In the figure: 1 - data automatic acquisition module; 2 - GPU self-detection module; 3 - code optimization processing module; 4 - data parallel processing module; 5 - performance analysis optimization module.
[0066] The purposes, functional features and advantages of the present application will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION
[0067] In order to better understand the purposes, technical solutions and advantages of the present application, the present application will be further described below with reference to the accompanying drawings and specific embodiments, and those skilled in the art can easily understand other advantages and effects of the present application from the disclosed content.
[0068] The present application can also be implemented or applied through other different specific examples, and various modifications and changes can be made to the details in the specification based on different views and applications without departing from the spirit of the present application.
[0069] It should be noted that if the embodiments of the present application involve directional indications (such as up, down, left, right, front, back, etc.), the directional indications are only used to explain the relative positional relationship, movement condition, etc. between the components in a certain specific posture (as shown in the drawings), and if the specific posture changes, the directional indications will also change accordingly.
[0070] In addition, if the embodiments of the present application involve descriptions of "first", "second", etc., the descriptions of "first", "second", etc. are only for description purposes, and cannot be understood as indicating or implying the relative importance of the indicated technical features or implicitly indicating the number of the indicated technical features. Therefore, the features limited by "first", "second" can explicitly or implicitly include at least one of the features. Secondly, the technical solutions of the various embodiments can be combined with each other, but must be based on the fact that a person skilled in the art can implement it, and when the combination of technical solutions appears to be contradictory or unimplementable, it should be considered that the combination of technical solutions does not exist and is not within the protection scope required by the present application.
[0071] Preferably, the GPU-adapted multi-module parallel image processing method is applied in one or more terminals or servers. The terminal is a device capable of automatically performing numerical calculation and / or information processing according to pre-set or stored instructions, and the hardware thereof includes but is not limited to microprocessors, application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.
[0072] The terminal can be a desktop computer, a notebook computer, a palm computer, a cloud server, or the like. The terminal can interact with a user through a keyboard, a mouse, a remote controller, a touchpad, a voice control device, or the like.
[0073] The present application provides a GPU-adapted multi-module parallel image processing method, system, platform, and storage medium.
[0074] As shown in FIG. 1, which is a flowchart of the GPU-adapted multi-module parallel image processing method provided by the present application. Figure 7
[0075] In the present embodiment, the GPU-adapted multi-module parallel image processing method can be applied in a terminal with display function or a fixed terminal, and the terminal is not limited to a personal computer, a smart phone, a tablet computer, a desktop computer with a camera, an all-in-one computer, or the like.
[0076] The GPU-adapted multi-module parallel image processing method can also be applied in a hardware environment composed of a terminal and a server connected to the terminal through a network. The network includes but is not limited to a wide area network, a metropolitan area network, or a local area network. The GPU-adapted multi-module parallel image processing method of the present embodiment can be executed by the server, or by the terminal, or by both the server and the terminal.
[0077] For example, for a terminal that needs to perform GPU-adapted multi-module parallel acceleration processing image, the GPU-adapted multi-module parallel acceleration processing image function provided by the method of the present application can be integrated directly on the terminal, or a client for implementing the method of the present application can be installed. For another example, the method provided by the present application can also run on a server or the like in the form of a software development kit (SDK), and the GPU-adapted multi-module parallel acceleration processing image function is provided in the form of an SDK, so that a terminal or other device can implement the GPU-adapted multi-module parallel acceleration processing image function through the provided interface. The present application is further described below with reference to the accompanying drawings.
[0078] As shown in Figures 1-7 , the present application provides a GPU-adapted multi-module parallel acceleration processing image method, which comprises the following steps:
[0079] S1, acquiring first information data corresponding to a GPU, detecting the running performance of the GPU in real time based on the first information data, and generating corresponding second information data; wherein the first information data is basic data of the GPU, including GPU parameter data, running environment information data, power consumption state data and running state data; the second information data is original running performance index data corresponding to the GPU;
[0080] S2, according to the second information data, real-time optimization processing of the running performance corresponding to the GPU, and generating corresponding third information data; wherein the third information data is index data corresponding to the GPU and processed by running performance optimization;
[0081] S3, based on the GPU, and in combination with the third information data, multi-module parallel acceleration processing and rendering of the image data to be processed.
[0082] The first information data corresponding to the GPU is acquired, the running performance of the GPU is detected in real time based on the first information data, and the corresponding second information data is generated, which further comprises:
[0083] S11, generating and acquiring model data, architecture data corresponding to the GPU, and GPU current acceleration task data;
[0084] S12, real-time detection of working environment data corresponding to the GPU, and monitoring of power consumption state data corresponding to the GPU;
[0085] S13, generating and acquiring running state information data corresponding to the GPU; wherein the running state information data includes GPU usage rate data and memory occupation data.
[0086] The first information data corresponding to the GPU is acquired, the running performance of the GPU is detected in real time based on the first information data, and corresponding second information data is generated, and the method further comprises:
[0087] S14, detecting the to-be-detected performance index data corresponding to the GPU, and generating first detection data corresponding to the to-be-detected performance index data; wherein the to-be-detected performance index data comprises the calculation running performance, the memory capacity, the GPU running frequency, the GPU power consumption, the GPU running frequency, and the GPU running frequency corresponding to the GPU.
[0088] S15, based on the first detection data, determining whether the performance index data corresponding to the GPU reaches the preset standard data, if yes, executing the next step; otherwise, continuing to execute the self-detection.
[0089] The running performance corresponding to the GPU is optimized in real time according to the second information data, and corresponding third information data is generated, and the method further comprises:
[0090] S21, based on the second information data, generating and acquiring first code data corresponding to the first information data, and optimizing the first code data, and generating second code data corresponding to the first information data, wherein the first code data is code data that needs to be optimized; and the second code data is the optimized code data.
[0091] S22, based on the second information data, generating and acquiring a first data set corresponding to the first information data, and dividing and processing the first data set, and generating a second data set corresponding to the first information data, wherein the first data set is the current data set data of the GPU; and the second data set is a sub-data set corresponding to the first data set.
[0092] The running performance corresponding to the GPU is optimized in real time according to the second information data, and corresponding third information data is generated, and the method further comprises:
[0093] S23, generating and acquiring state data corresponding to the GPU, and distributing and processing the second data set corresponding to the first information data;
[0094] S24, based on the state data corresponding to the GPU, parallel computing and processing the second data set corresponding to the GPU.
[0095] The running performance corresponding to the GPU is optimized in real time according to the second information data, and corresponding third information data is generated, and the method further comprises:
[0096] S31, in combination with a performance analysis tool, real-time estimation and generation of performance data corresponding to the GPU, and generation of a corresponding evaluation report according to the performance data;
[0097] S32, based on the performance data and the evaluation report, respectively optimizing and adjusting the code data and the model data corresponding to the GPU.
[0098] Specifically, in the embodiment of the present application, a new distributed GPU acceleration system is provided, which comprises a data automatic acquisition module, the output end of the data automatic acquisition module is connected with the input end of the GPU self-detection module, the output end of the GPU self-detection module is connected with the input end of the code optimization processing module, and the output end of the code optimization processing module is connected with the input end of the data parallelization processing module.
[0099] The output end of the data parallelization processing module is connected with the input end of the model parallelization processing module, the output end of the model parallelization processing module is connected with the input end of the data feedback module, the output end of the data feedback module is connected with the input end of the data receiving terminal, the output end of the data receiving terminal is connected with the input end of the performance analysis optimization module, and the data receiving terminal is one or more of a mobile phone and a notebook computer.
[0100] The data automatic acquisition module is used to comprehensively and systematically acquire various parameters and state information of the GPU, to ensure the integrity and real-time performance of the data. The GPU self-detection module is used to ensure that the GPU is in the best state in all aspects through a comprehensive self-detection module, to improve the reliability and stability of the overall system. The code optimization processing module is used to significantly improve the efficiency and accuracy of code optimization through automatic code retrieval, extraction, optimization, verification and merging. The data parallelization processing module is used to improve the efficiency and accuracy of data parallelization processing through automatic detection and distribution. The performance analysis optimization module is used to realize the integration of performance evaluation, result analysis and optimization adjustment, to improve the real-time performance and accuracy of optimization. Compared with the prior art, the present application adopts automatic and systematic processing methods in data acquisition, self-detection, code optimization, data parallelization processing and performance analysis optimization, to improve the efficiency and performance of the overall system.
[0101] The data automatic acquisition module comprises a GPU parameter acquisition module, the output end of the GPU parameter acquisition module is connected with the input end of the acceleration project data acquisition module, and the output end of the acceleration project data acquisition module is connected with the input end of the GPU running environment information acquisition module.
[0102] The output end of the GPU running environment information acquisition module is connected with the input end of the GPU running power consumption data acquisition module, and the output end of the GPU running power consumption data acquisition module is connected with the input end of the GPU running state information acquisition module.
[0103] The GPU self-detection module comprises a GPU computing power self-checking module and a GPU memory size self-checking module, the output end of the GPU computing power self-checking module is connected with the input end of the GPU memory size self-checking module, and the output end of the GPU memory size self-checking module is connected with the input end of the GPU video memory bit width self-checking module.
[0104] The output end of the GPU video memory bit width self-checking module is connected with the input end of the GPU power consumption self-checking module, the output end of the GPU power consumption self-checking module is connected with the input end of the GPU frequency self-checking module, and the output end of the GPU frequency self-checking module is connected with the input end of the GPU video memory type and capacity self-checking module.
[0105] The code optimization processing module comprises an execution code retrieval module and an execution code extraction module, the output end of the execution code retrieval module is connected with the input end of the execution code extraction module, and the output end of the execution code extraction module is connected with the input end of the programming model automatic import module.
[0106] The output end of the programming model automatic import module is connected with the input end of the code rewriting module, the output end of the code rewriting module is connected with the input end of the code automatic verification module, the output end of the code automatic verification module is connected with the input end of the code re-merging module, and the programming model imported by the programming model automatic import module is one or more of CUDA and OpenCL.
[0107] The data parallelization processing module comprises a data set automatic detection module, the output end of the data set automatic detection module is connected with the input end of the data subset automatic division module, the output end of the data subset automatic division module is connected with the input end of the GPU automatic detection module, the output end of the GPU automatic detection module is connected with the input end of the data subset automatic allocation module, and the output end of the data subset automatic allocation module is connected with the input end of the GPU parallel computing module.
[0108] The performance analysis optimization module comprises a performance analysis tool calling module, an output end of the performance analysis tool calling module being connected with an input end of a GPU accelerated performance evaluation module, an output end of the GPU accelerated performance evaluation module being connected with an input end of an evaluation result automatic analysis module, an output end of the evaluation result automatic analysis module being connected with an input end of a code optimization adjustment module, an output end of the code optimization adjustment module being connected with an input end of a data parallelization adjustment module, an output end of the data parallelization adjustment module being connected with an input end of a model parallelization adjustment module.
[0109] Specifically, the automatic acquisition module acquires GPU parameters: automatically collects basic information of the GPU, such as model, architecture, etc. Acquire task data: collect data of the current task that needs to be accelerated, such as dataset size and type. Collect running environment information: detect the current working environment of the GPU, including temperature, fan speed, etc. Monitor power consumption: monitor the power consumption of the GPU to ensure safe operation within a safe range. Get running status: collect running status information such as GPU usage and memory usage.
[0110] In the prior art, only part of the information is usually collected. This scheme collects all relevant data comprehensively and systematically through multi-module cooperation, ensuring data completeness and real-time updating.
[0111] GPU self-detection module: computing power detection: run standard computing tasks to detect whether the computing power of the GPU meets the requirements. Memory size detection: detect the size of the GPU memory to ensure that it meets the task requirements. Video memory bit width detection: detect the video memory bit width to ensure that the data transmission speed meets the requirements. Power consumption detection: monitor the power consumption of the GPU to ensure safe operation. Frequency detection: detect the running frequency of the GPU to ensure stable frequency. Video memory type and capacity detection: detect the type and capacity of the video memory to ensure compatibility and sufficient capacity.
[0112] The prior art usually only detects part of the performance indicators. This scheme comprehensively detects various performance indicators of the GPU to ensure that it operates in the best state.
[0113] Code optimization processing module: code retrieval: automatically retrieve the code part that needs to be optimized. Code extraction: extract the code segment that needs to be optimized. Import programming model: import a suitable programming model (such as CUDA, OpenCL) to rewrite the code. Code rewriting: optimize and rewrite the code. Code verification: automatically verify whether the function of the optimized code is correct. Code merging: merge the optimized code back into the project.
[0114] The prior art usually relies on manual code optimization. This scheme improves the efficiency and accuracy of code optimization through an automated process.
[0115] Data parallelization processing module: data set detection: automatically detect the current data set. Data subset division: divide large data sets into multiple subsets for parallel processing. GPU detection: detect the status and performance of each GPU. Data distribution: distribute data subsets to different GPUs for parallel computing according to the status of the GPUs. Parallel computing: each GPU processes its own data subset in parallel to improve computing efficiency.
[0116] While the prior art usually manually allocates data, the present scheme improves the efficiency and accuracy of data processing through automatic detection and distribution.
[0117] Performance analysis and optimization module: call performance analysis tool: obtain performance data after acceleration. Performance evaluation: evaluate the performance of the GPU after acceleration and generate an evaluation report. Result analysis: automatically analyze the evaluation results to find performance bottlenecks and optimization points. Optimization adjustment: optimize and adjust the code and data processing method according to the analysis results. Adjust parallelization: adjust the parallelization processing method of data and models according to the analysis results.
[0118] In the prior art, performance analysis and optimization are usually separate. The present scheme integrates performance evaluation, result analysis, and optimization adjustment through the integrated performance analysis and optimization module, improving the real-time and accuracy of optimization.
[0119] That is, through the coordinated work of data automatic acquisition, GPU self-detection, code optimization processing, data parallelization processing, and performance analysis and optimization module, the present scheme realizes comprehensive optimization and acceleration of the GPU, significantly improving the performance of the GPU and the actual application effect of the system. Compared with the prior art, the present scheme has significant advantages in automation, systematization, and real-time performance.
[0120] Example 1
[0121] As shown in the accompanying Figures 1-6 , the present scheme provides a novel distributed GPU acceleration system, which includes a data automatic acquisition module 1, the output end of which is connected to the input end of a GPU self-detection module 2, the output end of which is connected to the input end of a code optimization processing module 3, the output end of which is connected to the input end of a data parallelization processing module 4, the output end of which is connected to the input end of a model parallelization processing module, the output end of which is connected to the input end of a data feedback module, the output end of which is connected to the input end of a data receiving terminal, the output end of which is connected to the input end of a performance analysis and optimization module 5, and the data receiving terminal is a notebook computer.
[0122] The data automatic acquisition module 1 comprises a GPU parameter acquisition module, an output end of the GPU parameter acquisition module is connected with an input end of an acceleration item data acquisition module, an output end of the acceleration item data acquisition module is connected with an input end of a GPU running environment information acquisition module, an output end of the GPU running environment information acquisition module is connected with an input end of a GPU running power consumption data acquisition module, and an output end of the GPU running power consumption data acquisition module is connected with an input end of a GPU running state information acquisition module.
[0123] The GPU running environment information acquisition module is used for collecting the environment information of the current running of the GPU, such as temperature, fan speed, etc. The GPU running power consumption data acquisition module is used for monitoring the power consumption of the GPU. The GPU running state information acquisition module is used for acquiring the current use state of the GPU, including use rate, memory occupation, etc.
[0124] In the prior art, only part of the parameters are usually acquired, and there is no systematic data acquisition process. The new scheme comprehensively and systematically acquires various parameters and state information of the GPU through cooperation of multiple modules, and ensures the completeness and real-time performance of the data.
[0125] The GPU self-detection module 2 comprises a GPU computing power self-detection module and a GPU memory size self-detection module, an output end of the GPU computing power self-detection module is connected with an input end of the GPU memory size self-detection module, an output end of the GPU memory size self-detection module is connected with an input end of a GPU video memory bit width self-detection module, an output end of the GPU video memory bit width self-detection module is connected with an input end of a GPU power consumption self-detection module, an output end of the GPU power consumption self-detection module is connected with an input end of a GPU frequency self-detection module, and an output end of the GPU frequency self-detection module is connected with an input end of a GPU video memory type and capacity self-detection module.
[0126] The GPU self-detection is realized through the following modules: the GPU computing power self-detection module is used for detecting the computing power of the GPU, ensuring that the GPU reaches the expected performance. The GPU memory size self-detection module is used for detecting the memory size of the GPU, ensuring that the GPU meets the task requirements. The GPU video memory bit width self-detection module is used for detecting the video memory bit width, ensuring the data transmission speed. The GPU power consumption self-detection module is used for monitoring the power consumption, ensuring that the GPU runs within a safe range. The GPU frequency self-detection module is used for detecting the running frequency of the GPU, ensuring the stability of the frequency. The GPU video memory type and capacity self-detection module is used for detecting the type and capacity of the video memory, ensuring the compatibility and sufficient capacity.
[0127] In the prior art, only part of the performance indicators are detected, while the present scheme ensures that the GPU reaches the best state in all aspects through a comprehensive self-checking module, thereby improving the reliability and stability of the overall system.
[0128] The GPU memory bit width self-checking module is implemented by the following steps: running a memory bandwidth test program, such as memory copy and data transmission test; monitoring the transmission rate and bandwidth utilization; comparing with the standard value to detect whether the memory bit width meets the standard; and generating a self-checking report to record the memory bit width self-checking result as the basis for subsequent optimization. In the prior art, memory bit width self-checking usually relies on external tools, and the present scheme realizes autonomous detection and analysis of memory bit width through an integrated self-checking module.
[0129] The code optimization processing module 3 comprises an execution code retrieval module and an execution code extraction module, the output end of the execution code retrieval module is connected with the input end of the execution code extraction module, the output end of the execution code extraction module is connected with the input end of the programming model automatic import module, the output end of the programming model automatic import module is connected with the input end of the code rewriting module, the output end of the code rewriting module is connected with the input end of the code automatic verification module, the output end of the code automatic verification module is connected with the input end of the code re-merging module, and the programming model imported by the programming model automatic import module is OpenCL.
[0130] The code optimization processing module is implemented by the following steps: the execution code retrieval module is used to retrieve the code part that needs to be optimized; the execution code extraction module is used to extract the code segment that needs to be optimized; the programming model automatic import module is used to import a suitable programming model (such as CUDA or OpenCL) for code rewriting; the code rewriting module is used to optimize and rewrite the code; the code automatic verification module is used to verify whether the function of the rewritten code is correct; and the code re-merging module is used to merge the optimized code into the overall project. In the prior art, code optimization is usually performed manually, while the present scheme significantly improves the efficiency and accuracy of code optimization through automatic code retrieval, extraction, optimization, verification, and merging.
[0131] The data parallelization processing module 4 comprises a data set automatic detection module, the output end of the data set automatic detection module is connected with the input end of the data subset automatic division module, the output end of the data subset automatic division module is connected with the input end of the GPU automatic detection module, the output end of the GPU automatic detection module is connected with the input end of the data subset automatic allocation module, and the output end of the data subset automatic allocation module is connected with the input end of the GPU parallel computing module.
[0132] The data parallelization processing module is realized by the following steps: the data set automatic detection module is used to detect the current data set condition. The data subset automatic division module is used to divide the large data set into multiple subsets for parallel processing. The GPU automatic detection module is used to detect the state and performance of each GPU. The data subset automatic allocation module is used to allocate the data subsets to different GPUs for parallel calculation according to the state of the GPUs. The GPU parallel calculation module is used to parallel process the respective data subsets by each GPU, thereby improving the calculation efficiency. The prior art usually manually allocates data, while the scheme improves the efficiency and accuracy of data parallelization processing by automatic detection and allocation.
[0133] The performance analysis optimization module 5 comprises a performance analysis tool calling module, an output end of the performance analysis tool calling module being connected with an input end of a GPU acceleration performance evaluation module, an output end of the GPU acceleration performance evaluation module being connected with an input end of an evaluation result automatic analysis module, an output end of the evaluation result automatic analysis module being connected with an input end of a code optimization adjustment module, an output end of the code optimization adjustment module being connected with an input end of a data parallelization adjustment module, and an output end of the data parallelization adjustment module being connected with an input end of a model parallelization adjustment module.
[0134] The performance analysis optimization module is realized by the following steps: the performance analysis tool calling module is used to call the performance analysis tool to obtain the performance data after acceleration. The GPU acceleration performance evaluation module is used to evaluate the performance after GPU acceleration to generate an evaluation report. The evaluation result automatic analysis module is used to automatically analyze the evaluation result to find out the performance bottleneck and optimization point. The code optimization adjustment module is used to further optimize and adjust the code according to the analysis result. The data parallelization adjustment module is used to adjust the data parallelization processing mode according to the analysis result. The model parallelization adjustment module is used to adjust the model parallelization processing mode according to the analysis result. In the prior art, performance analysis and optimization are usually separated, and the scheme realizes the integrated processing of performance evaluation, result analysis and optimization adjustment by the integrated performance analysis optimization module, thereby improving the real-time performance and accuracy of optimization.
[0135] In the embodiment, the data automatic acquisition module and the GPU self-detection module are arranged in the system to acquire and detect various parameter data of the GPU, to ensure that the GPU is in a normal state, to ensure the success rate of GPU acceleration, and to simultaneously arrange the code optimization processing module, the data parallelization processing module and the model parallelization processing module to simultaneously accelerate the GPU by the three modes, thereby effectively ensuring the acceleration effect and effectiveness of the GPU.
[0136] The GPU computing power self-checking module is realized by the following steps: running a standard computing task, such as matrix multiplication, large-scale parallel computing, etc. The computing time and result are monitored for comparison with a standard result to detect whether the computing performance meets the standard. A self-checking report is generated to record the computing power self-checking result as a basis for subsequent optimization. In the prior art, computing power self-checking usually relies on external tools, and the present scheme realizes self-detection and analysis of computing power through an integrated self-checking module.
[0137] Embodiment 2
[0138] As shown in the accompanying Figures 1-6 The present application provides a technical solution: a new distributed GPU acceleration system, comprising a data automatic acquisition module 1, the output end of the data automatic acquisition module 1 is connected with the input end of the GPU self-detection module 2, the output end of the GPU self-detection module 2 is connected with the input end of the code optimization processing module 3, the output end of the code optimization processing module 3 is connected with the input end of the data parallelization processing module 4, the output end of the data parallelization processing module 4 is connected with the input end of the model parallelization processing module, the output end of the model parallelization processing module is connected with the input end of the data feedback module, the output end of the data feedback module is connected with the input end of the data receiving terminal, the output end of the data receiving terminal is connected with the input end of the performance analysis optimization module 5, and the data receiving terminal is a mobile phone.
[0139] The data automatic acquisition module 1 comprises a GPU parameter acquisition module, the output end of the GPU parameter acquisition module is connected with the input end of the acceleration project data acquisition module, the output end of the acceleration project data acquisition module is connected with the input end of the GPU running environment information acquisition module, the output end of the GPU running environment information acquisition module is connected with the input end of the GPU running power consumption data acquisition module, and the output end of the GPU running power consumption data acquisition module is connected with the input end of the GPU running state information acquisition module.
[0140] The GPU self-detection module 2 comprises a GPU computing power self-checking module and a GPU memory size self-checking module, the output end of the GPU computing power self-checking module is connected with the input end of the GPU memory size self-checking module, the output end of the GPU memory size self-checking module is connected with the input end of the GPU video memory bit width self-checking module, the output end of the GPU video memory bit width self-checking module is connected with the input end of the GPU power consumption self-checking module, the output end of the GPU power consumption self-checking module is connected with the input end of the GPU frequency self-checking module, and the output end of the GPU frequency self-checking module is connected with the input end of the GPU video memory type and capacity self-checking module.
[0141] The code optimization processing module 3 comprises an execution code retrieval module and an execution code extraction module, the output end of the execution code retrieval module is connected with the input end of the execution code extraction module, the output end of the execution code extraction module is connected with the input end of the programming model automatic import module, the output end of the programming model automatic import module is connected with the input end of the code rewriting module, the output end of the code rewriting module is connected with the input end of the code automatic verification module, the output end of the code automatic verification module is connected with the input end of the code re-merging module, and the programming model imported by the programming model automatic import module is CUDA.
[0142] The data parallelization processing module 4 comprises a data set automatic detection module, the output end of the data set automatic detection module is connected with the input end of the data subset automatic division module, the output end of the data subset automatic division module is connected with the input end of the GPU automatic detection module, the output end of the GPU automatic detection module is connected with the input end of the data subset automatic allocation module, and the output end of the data subset automatic allocation module is connected with the input end of the GPU parallel computing module.
[0143] The performance analysis optimization module 5 comprises a performance analysis tool calling module, the output end of the performance analysis tool calling module is connected with the input end of the GPU acceleration performance evaluation module, the output end of the GPU acceleration performance evaluation module is connected with the input end of the evaluation result automatic analysis module, the output end of the evaluation result automatic analysis module is connected with the input end of the code optimization adjustment module, the output end of the code optimization adjustment module is connected with the input end of the data parallelization adjustment module, and the output end of the data parallelization adjustment module is connected with the input end of the model parallelization adjustment module.
[0144] In the embodiment, the performance analysis optimization module is further arranged in the system, which can analyze the performance after GPU acceleration, and further optimize the acceleration process according to the analysis result, thereby helping to stabilize the optimization of the GPU acceleration effect and improving the actual application effect of the system.
[0145] Preferably, in the application scheme, the actual effects of multi-module synchronous acceleration and performance optimization are achieved by: increasing specific experimental data and case analysis, showing the performance improvement effect of the modules in different application scenarios, and providing comparison data to prove the effectiveness.
[0146] Implementation method: experimental data and case analysis
[0147] Step 1: design experiment
[0148] Application scenarios: image processing, big data analysis, and machine learning model training.
[0149] Test dataset: Choose public standard datasets such as ImageNet (image processing), Kaggle dataset (big data analysis), MNIST (machine learning).
[0150] 2. Perform benchmarking: Run without optimization, record processing time, power consumption, and GPU usage.
[0151] 3. Enable multi-module synchronous acceleration: Enable data automatic acquisition module, record basic data. Enable GPU self-detection module, detect and optimize GPU state. Enable code optimization processing module, optimize the code to be processed. Enable data parallel processing module, allocate data tasks. Enable performance analysis optimization module, analyze and adjust optimization strategies.
[0152] 4. Analyze results: Compare performance data at different stages, generate charts and analysis reports, and show the contribution of each module and overall improvement effect.
[0153] Virtual results: In the image processing scenario, processing time is reduced from 200 seconds to 120 seconds, power consumption is reduced from 250W to 180W, and GPU usage is increased from 60% to 90%.
[0154] In the big data analysis scenario, processing time is reduced from 500 seconds to 300 seconds, power consumption is reduced from 300W to 220W, and GPU usage is increased from 55% to 85%.
[0155] In the machine learning model training scenario, training time is reduced from 1000 seconds to 600 seconds, power consumption is reduced from 350W to 250W, and GPU usage is increased from 50% to 80%.
[0156] Comparison between automatic code optimization and manual optimization: The invention details the specific steps and logic of automatic code optimization, including code retrieval, extraction, rewriting, verification, and merging implementation methods.
[0157] Implementation method: The specific steps and logic of automatic code optimization are described in detail as follows:
[0158] Step 1: Code retrieval: Use Clang Static Analyzer to automatically retrieve the code part that needs to be optimized.
[0159] 2. Code extraction: Develop Python scripts to automatically extract the code segment that needs to be optimized.
[0160] 3. Programming model import: Use CUDA automatic conversion tool to import the appropriate programming model.
[0161] 4. Code rewriting: Use LLVM Pass for code optimization and rewriting.
[0162] 5. Code verification: Automatically verify the functionality of the optimized code using the Google Test framework.
[0163] 6. Code merging: Automatically merge the optimized code into the main project using Git and perform continuous integration testing (Jenkins).
[0164] Virtual results: Through automated optimization, the code running time is reduced from 50ms to 30ms, and the code size is reduced by 20%. During the automated optimization process, 5 potential performance bottlenecks were discovered and fixed, which were not found during manual optimization. The automated optimization took 1 hour, while the manual optimization took 8 hours.
[0165] Real-time performance detection and optimization of GPU: The present application provides a real-time performance detection and optimization process for GPU, which includes the following steps:
[0166] Implementation method: The specific steps and logic of automated code optimization are described in detail as follows:
[0167] Step 1: Code retrieval: Automatically retrieve the code sections that need to be optimized using the Clang Static Analyzer.
[0168] 2. Code extraction: Develop a Python script to automatically extract the code segments that need to be optimized.
[0169] 3. Import programming model: Import the appropriate programming model using the CUDA automatic conversion tool.
[0170] 4. Code rewriting: Use LLVM Pass to optimize and rewrite the code.
[0171] 5. Code verification: Automatically verify the functionality of the optimized code using the Google Test framework.
[0172] 6. Code merging: Automatically merge the optimized code into the main project using Git and perform continuous integration testing (Jenkins).
[0173] Virtual results: Through automated optimization, the code running time is reduced from 50ms to 30ms, and the code size is reduced by 20%. During the automated optimization process, 5 potential performance bottlenecks were discovered and fixed, which were not found during manual optimization. The automated optimization took 1 hour, while the manual optimization took 8 hours.
[0174] Efficiency of data parallel processing module: The invention scheme supplements the specific technical implementation details of the data parallel processing module and provides performance comparison analysis in actual application, which illustrates the advantages and specific improvements of the module compared with the traditional manual allocation method.
[0175] Implementation method: The technical implementation details and performance comparison analysis are as follows:
[0176] Step 1: Automatic detection: Use Pandas to detect the size and structure of the current data set.
[0177] 2, Subset division: Develop a division algorithm based on data blocks to divide large data sets into small subsets.
[0178] 3, GPU detection: Real-time detection of the status and performance of each GPU (usage, temperature, power consumption).
[0179] 4, Data allocation: Use load balancing algorithms to allocate data subsets to different GPUs.
[0180] 5, Parallel computing: Each GPU parallel processes the allocated data subset, records the processing time and resource utilization.
[0181] 6, Performance comparison: Compare the performance data of manual allocation and automatic allocation, generate charts and reports, and show the optimization effect.
[0182] Virtual results: In the case of automatic allocation, the data processing time is reduced from 600 seconds to 400 seconds. The GPU usage is balanced at more than 80%, and the usage fluctuates greatly in the case of manual allocation. The automatic allocation algorithm reduces data transmission overhead and improves overall computing efficiency by 25%.
[0183] Integration of performance analysis tools: The invention scheme details the integration method and working mechanism of the performance analysis tools, including the data interaction process between modules, to ensure the coherence and efficiency of the optimization process.
[0184] Implementation method: Tool integration method and collaboration mechanism:
[0185] Step 1: Tool selection: Select NVIDIA Nsight, VTune, Perfetto for integration.
[0186] 2, Unified interface: Develop a unified interface module to ensure that each tool can obtain and provide performance data through the interface. 3, Data interaction: Design data format and transmission protocol to ensure smooth data exchange between tools.
[0187] 4, Collaboration mechanism: Design a tool collaboration workflow, first perform performance evaluation, generate a report, then analyze the results, and finally optimize the code and data processing.
[0188] 5. Testing and Verification: Test the integration effect in actual application to ensure that the tools can work together, and make adjustments and optimizations based on feedback.
[0189] Virtual Results: After integration, performance analysis time was reduced from 30 minutes to 10 minutes, and the accuracy of analysis results improved by 15%. With the collaboration of various tools, the success rate of optimization strategy execution increased from 70% to 95%. Real-time performance monitoring and optimization improved overall system performance by 20% and concurrent processing capacity by 30%.
[0190] To achieve the above objectives, the present invention also provides a multi-module parallel acceleration image processing system suitable for GPUs, such as... Figure 8 As shown, the system specifically includes:
[0191] The data acquisition and generation unit is used to acquire first information data corresponding to the GPU, detect the GPU's operating performance in real time based on the first information data, and generate corresponding second information data; wherein, the first information data is the GPU's basic data, including: GPU parameter data, operating environment information data, power consumption status data, and operating status data; the second information data is the raw operating performance index data corresponding to the GPU;
[0192] A data parallelization processing unit is used to optimize the running performance corresponding to the GPU in real time based on the second information data, and generate corresponding third information data; wherein, the third information data is the indicator data corresponding to the GPU and after running performance optimization processing;
[0193] The image acceleration processing unit is used to accelerate the processing and rendering of the image data to be processed in parallel using multiple modules based on the GPU and in combination with the third information data.
[0194] The data acquisition and generation unit further includes:
[0195] The first data generation module is used to generate and obtain model data, architecture data, and data on the tasks that the GPU currently needs to accelerate, which are corresponding to the GPU.
[0196] The data monitoring and processing module is used to detect the working environment data corresponding to the GPU in real time, and to monitor the power consumption status data corresponding to the GPU.
[0197] The second data generation module is used to generate and acquire runtime status information data corresponding to the GPU; wherein, the runtime status information data includes GPUD utilization data and memory usage data;
[0198] And / or, the data acquisition and generation unit further includes:
[0199] The third data generation module is configured to detect to-be-detected performance index data corresponding to the GPU and generate first detection data corresponding to the to-be-detected performance index data; wherein the to-be-detected performance index data includes calculation running performance, memory capacity, display memory bit width, GPU power consumption, GPU running frequency and display memory type corresponding to the GPU.
[0200] The first data determination module is configured to determine, based on the first detection data, whether the performance index data corresponding to the GPU reaches preset standard data.
[0201] And / or, the data parallelization processing unit further includes:
[0202] The fourth data generation module is configured to generate and acquire first code data corresponding to the first information data based on the second information data, optimize the first code data and generate second code data corresponding to the first information data; wherein the first code data is code data that needs to be optimized; and the second code data is optimized code data.
[0203] The fifth data generation module is configured to generate and acquire a first data set corresponding to the first information data based on the second information data, divide and process the first data set, and generate a second data set corresponding to the first information data; wherein the first data set is a current data set of the GPU; and the second data set is a sub-data set corresponding to the first data set.
[0204] And / or, the data parallelization processing unit further includes:
[0205] The sixth data generation module is configured to generate and acquire state data corresponding to the GPU and allocate and process the second data set corresponding to the first information data.
[0206] The data parallelization calculation processing module is configured to parallelize and calculate the second data set corresponding to the GPU based on the state data corresponding to the GPU.
[0207] And / or, the image acceleration processing unit further includes:
[0208] The seventh data generation module is configured to combine a performance analysis tool, estimate and generate performance data corresponding to the GPU in real time, and generate a corresponding evaluation report based on the performance data.
[0209] The data optimization adjustment processing module is configured to optimize and adjust code data and model data corresponding to the GPU based on the performance data and the evaluation report, respectively.
[0210] In the system embodiment of the present application, the steps of the method for processing images in parallel by multiple modules suitable for GPUs are described above in detail, that is, the functional modules in the system are used to implement the steps or sub-steps in the above method embodiment, which will not be described here.
[0211] To achieve the above object, the present application further provides a platform for processing images in parallel by multiple modules suitable for GPUs, as shown in the accompanying drawings, comprising a processor, a memory and a platform control program for processing images in parallel by multiple modules suitable for GPUs; wherein the processor executes the platform control program for processing images in parallel by multiple modules suitable for GPUs, the platform control program for processing images in parallel by multiple modules suitable for GPUs is stored in the memory, and the platform control program for processing images in parallel by multiple modules suitable for GPUs implements the steps of the method for processing images in parallel by multiple modules suitable for GPUs. Figure 9
[0212] S1, obtaining first information data corresponding to a GPU, detecting the running performance of the GPU in real time based on the first information data, and generating corresponding second information data; wherein the first information data is the basic data of the GPU, including GPU parameter data, running environment information data, power consumption state data and running state data; the second information data is the original running performance index data corresponding to the GPU;
[0213] S2, optimizing the running performance corresponding to the GPU in real time according to the second information data, and generating corresponding third information data; wherein the third information data is the index data corresponding to the GPU after running performance optimization processing;
[0214] S3, processing and rendering the image data to be processed in parallel by multiple modules based on the GPU and in combination with the third information data.
[0215] The specific steps have been described above, which will not be described here.
[0216] In the embodiment of the present application, the processor built-in in the image processing platform suitable for GPU multi-module parallel acceleration can be composed of integrated circuits, for example, can be composed of a single packaged integrated circuit, or can be composed of multiple packaged integrated circuits with same or different functions, including one or more combinations of central processing units (CPU), microprocessors, digital processing chips, graphic processors and various control chips. The processor connects various components through various interfaces and lines, executes programs or units stored in the memory, and calls data stored in the memory, to perform various functions and process data of the image processing platform suitable for GPU multi-module parallel acceleration.
[0217] The memory is used for storing program codes and various data, is installed in the image processing platform suitable for GPU multi-module parallel acceleration, and realizes high-speed and automatic access to programs or data during running.
[0218] The memory includes read-only memory (ROM), random access memory (RAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), one-time programmable read-only memory (OTPROM), electrically-erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disc memory, magnetic disc memory, magnetic tape memory, or any other computer-readable medium capable of carrying or storing data.
[0219] To achieve the above object, the present application further provides a computer readable storage medium, such as Figure 10 As shown in the figure, the computer readable storage medium stores a platform control program suitable for GPU multi-module parallel acceleration image processing, which realizes the steps of the method of the image processing platform suitable for GPU multi-module parallel acceleration, for example:
[0220] S1, acquire first information data corresponding to the GPU, detect the running performance of the GPU in real time based on the first information data, and generate corresponding second information data; wherein the first information data is the basic data of the GPU, including: GPU parameter data, running environment information data, power consumption state data and running state data; the second information data is the original running performance index data corresponding to the GPU;
[0221] S2, according to the second information data, real-time optimization processing of the running performance corresponding to the GPU, and generating corresponding third information data; wherein the third information data is the index data corresponding to the GPU and after the running performance optimization processing;
[0222] S3, based on the GPU, and combined with the third information data, multi-module parallel acceleration processing and rendering of the image data to be processed.
[0223] The specific details of the steps have been described above, and will not be repeated here.
[0224] In the description of the embodiments of the application, it should be noted that any process or method described in the flowchart or otherwise described herein can be understood as representing a module, a segment or a part of code including one or more executable instructions for implementing a specific logical function or process, and the scope of the preferred embodiments of the application includes additional implementations in which the functions can be performed in an order other than that shown or discussed, including in a substantially simultaneous manner or in a reverse order according to the functions involved, which should be understood by those skilled in the art to which the embodiments of the application belong.
[0225] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a list of ordered executable instructions for implementing the logical function, which can be specifically implemented in any computer readable medium for use by or in conjunction with an instruction execution system, device or apparatus, such as a computer-based system, a system including a processing module or other system that can fetch and execute instructions from the instruction execution system, device or apparatus. For the purpose of this specification, "computer readable medium" can be any device that can contain, store, communicate, propagate or transport programs for use by or in conjunction with an instruction execution system, device or apparatus. More specific examples (non-exhaustive list) of computer readable medium include the following: electrical connections having one or more wires (electronic devices), portable computer disk boxes (magnetic devices), random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memories), optical fiber devices, and portable compact disc read-only memories (CDROMs).
[0226] In addition, a computer readable medium can even be paper or other suitable medium upon which the program is printed, since the program can be electronically obtained, for example via an optical scanner, then edited, interpreted or otherwise processed in an electronic manner, and then stored in a computer memory.
[0227] In the embodiments of the present application, in order to achieve the above-mentioned purpose, the present application also provides a chip system, which comprises at least one processor, and when program instructions are executed in the at least one processor, the chip system executes the steps of the image processing method suitable for GPU multi-module parallel acceleration, for example:
[0228] S1, acquiring first information data corresponding to the GPU, detecting the running performance of the GPU in real time based on the first information data, and generating corresponding second information data; wherein the first information data is the basic data of the GPU, including: GPU parameter data, running environment information data, power consumption state data and running state data; the second information data is the original running performance index data corresponding to the GPU;
[0229] S2, according to the second information data, real-time optimization processing the running performance corresponding to the GPU, and generating corresponding third information data; wherein the third information data is the index data corresponding to the GPU and processed by the running performance optimization;
[0230] S3, based on the GPU and combined with the third information data, multi-module parallel acceleration processing and rendering the image data to be processed.
[0231] The specific details of the steps have been described above, and will not be repeated here.
[0232] Those skilled in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be realized in electronic hardware or a combination of computer software and electronic hardware. Whether the functions are realized in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to realize the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application. Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the above-described system, device and unit can refer to the corresponding process in the foregoing method embodiments, which will not be repeated here.
[0233] The application obtains first information data corresponding to the GPU by a method, detects the running performance of the GPU in real time based on the first information data, and generates corresponding second information data; wherein the first information data is basic data of the GPU, including GPU parameter data, running environment information data, power consumption state data and running state data; the second information data is original running performance index data corresponding to the GPU; the running performance corresponding to the GPU is optimized in real time according to the second information data, and corresponding third information data is generated; wherein the third information data is index data corresponding to the GPU and processed after running performance optimization; based on the GPU and in combination with the third information data, the image data to be processed is processed and rendered in parallel by multiple modules, and a system, a platform and a storage medium corresponding to the method are used to realize multi-module parallel acceleration: a data automatic acquisition module, a GPU self-detection module, a code optimization processing module, a data parallel processing module and a performance analysis optimization module are introduced, and the GPU is accelerated in multiple modes simultaneously. Performance analysis and optimization: the performance analysis optimization module is added to analyze the performance of the accelerated GPU, the acceleration process is further optimized according to the result, and the stability and effectiveness of the GPU acceleration effect are ensured.
[0234] That is, in the application scheme, by setting the data automatic acquisition module and the GPU self-detection module in the system, various parameter data about the GPU can be acquired and detected, the GPU itself is ensured to be in a normal state, the success rate of GPU acceleration is ensured, and by setting the code optimization processing module, the data parallel processing module and the model parallel processing module, the GPU is accelerated in three modes simultaneously, the acceleration effect and effectiveness of the GPU are effectively ensured, and the performance analysis optimization module is further set in the system, which can analyze the performance of the accelerated GPU and further optimize the acceleration process according to the analysis result, which helps to stabilize the optimization of the GPU acceleration effect and improve the actual application effect of the system.
[0235] In other words, by the application scheme, the acceleration effect can be improved: by the synchronous work of multiple acceleration modules, the GPU performance is significantly improved. The detection and optimization capability is improved: by the performance analysis optimization module, the acceleration effect is detected and further optimized, and the effectiveness and stability of the GPU acceleration are ensured. The actual application effect of the system is enhanced: multiple optimization measures are combined to improve the effect and performance of the system in actual application.
[0236] That is, multiple acceleration modules work synchronously: through the cooperative work of data automatic acquisition, GPU self-detection, code optimization processing, data parallel processing and performance analysis optimization modules, the performance of the GPU is significantly improved. The performance analysis optimization module detects acceleration: through performance analysis tool calling, GPU acceleration performance evaluation, evaluation result automatic analysis, code optimization adjustment and other steps, the acceleration effect is detected and further optimized. Multiple optimization measures are integrated: through the comprehensive optimization of data acquisition, self-detection, code optimization, parallel processing and performance analysis, the effect and performance of the system in actual application are improved.
[0237] The above-described embodiments only express several embodiments of the present application, and the description is relatively specific and detailed, but it should not be understood as a limitation on the scope of the patent of the present application. It should be noted that, for ordinary skilled persons in the art, without departing from the concept of the present application, several modifications and improvements can be made, which are all within the protection scope of the present application. Therefore, the protection scope of the patent of the present application should be subject to the appended claims.
Claims
1. A multi-module parallel accelerated processing method for images suitable for GPUs, characterized in that, The method comprises the steps of: acquiring first information data corresponding to the GPU, detecting the running performance of the GPU in real time based on the first information data, and generating corresponding second information data; wherein the first information data is basic data of the GPU, including GPU parameter data, running environment information data, power consumption state data and running state data; the second information data is original running performance index data corresponding to the GPU; optimizing the running performance corresponding to the GPU in real time according to the second information data, and generating corresponding third information data; wherein the third information data is index data corresponding to the GPU and processed by running performance optimization; further comprising the steps of: based on the second information data, generating and acquiring first code data corresponding to the first information data, optimizing the first code data, and generating second code data corresponding to the first information data, wherein the first code data is code data that needs to be optimized; the second code data is code data processed by optimization; based on the second information data, generating and acquiring a first data set corresponding to the first information data, dividing and processing the first data set, and generating a second data set corresponding to the first information data, wherein the first data set is the current data set data of the GPU; the second data set is a sub-data set corresponding to the first data set; generating and acquiring state data corresponding to the GPU, and distributing and processing the second data set corresponding to the first information data; based on the state data corresponding to the GPU, parallel computing and processing the second data set corresponding to the GPU; based on the GPU and in combination with the third information data, multi-module parallel acceleration processing and rendering the image data to be processed; further comprising the steps of: in combination with the performance analysis tool, real-time estimation and generation of performance data corresponding to the GPU, and generation of a corresponding evaluation report according to the performance data; based on the performance data and the evaluation report, respectively optimizing and adjusting the code data and model data corresponding to the GPU.
2. The method of claim 1, wherein the method is adapted for GPU-based multi-module parallel acceleration processing of images, and wherein the method further comprises: The acquisition of the first information data corresponding to the GPU, the real-time detection of the running performance of the GPU based on the first information data, and the generation of the corresponding second information data further comprises: generating and acquiring model data, architecture data corresponding to the GPU, and current acceleration task data required by the GPU; real-time detection of working environment data corresponding to the GPU, and monitoring of power consumption state data corresponding to the GPU; generating and acquiring running state information data corresponding to the GPU; wherein the running state information data includes GPU usage rate data and memory occupation data.
3. The method according to claim 1 or 2, wherein, The acquisition of the first information data corresponding to the GPU, the real-time detection of the running performance of the GPU based on the first information data, and the generation of the corresponding second information data further comprises: detecting to-be-detected performance index data corresponding to the GPU, and generating first detection data corresponding to the to-be-detected performance index data; wherein the to-be-detected performance index data includes calculation running performance, memory capacity, video memory bit width, GPU power consumption, GPU running frequency, and video memory type corresponding to the GPU; based on the first detection data, determining whether performance index data corresponding to the GPU reaches preset standard data, if yes, performing the next step; otherwise, continuing to perform self-detection.
4. A multi-module parallel accelerated processing image system suitable for GPU, characterized in that, The system is used to implement the GPU-adapted multi-module parallel acceleration processing image method according to any one of claims 1-3, and the system comprises: a data acquisition and generation unit configured to acquire first information data corresponding to the GPU, detect running performance of the GPU in real time based on the first information data, and generate corresponding second information data; wherein the first information data is basic data of the GPU, including GPU parameter data, running environment information data, power consumption state data, and running state data; and the second information data is original running performance index data corresponding to the GPU; a data parallelization processing unit configured to optimize running performance corresponding to the GPU in real time according to the second information data, and generate corresponding third information data; wherein the third information data is index data corresponding to the GPU and processed after running performance optimization; an image acceleration processing unit configured to perform multi-module parallel acceleration processing and rendering on to-be-processed image data based on the GPU and in combination with the third information data.
5. The multi-module parallel accelerated image processing system for GPU of claim 4, wherein, The data acquisition and generation unit further comprises: a first data generation module configured to generate and acquire model data, architecture data corresponding to the GPU, and GPU current acceleration task data; a data monitoring and processing module configured to detect working environment data corresponding to the GPU in real time, and monitor power consumption state data corresponding to the GPU; a second data generation module configured to generate and acquire running state information data corresponding to the GPU; wherein the running state information data includes GPU usage rate data and memory occupation data; And / or, the data acquisition and generation unit further comprises: a third data generation module configured to detect to-be-detected performance index data corresponding to the GPU, and generate first detection data corresponding to the to-be-detected performance index data; wherein the to-be-detected performance index data includes calculation running performance, memory capacity, video memory bit width, GPU power consumption, GPU running frequency, and video memory type corresponding to the GPU; a first data determination module configured to determine whether performance index data corresponding to the GPU reaches preset standard data based on the first detection data; And / or, the data parallelization processing unit further comprises: The fourth data generation module is configured to generate and acquire first code data corresponding to the first information data based on the second information data, and to optimize the first code data and generate second code data corresponding to the first information data, wherein the first code data is code data that needs to be optimized, and the second code data is optimized code data. The fifth data generation module is configured to generate and acquire a first data set corresponding to the first information data based on the second information data, and to divide and process the first data set and generate a second data set corresponding to the first information data, wherein the first data set is a current data set of the GPU, and the second data set is a sub-data set corresponding to the first data set. The data parallelization processing unit further includes: The sixth data generation module is configured to generate and acquire state data corresponding to the GPU, and to allocate and process the second data set corresponding to the first information data. The data parallelization calculation processing module is configured to parallelize and calculate the second data set corresponding to the GPU based on the state data corresponding to the GPU. The image acceleration processing unit further includes: The seventh data generation module is configured to estimate and generate performance data corresponding to the GPU in real time in combination with a performance analysis tool, and to generate a corresponding evaluation report based on the performance data. The data optimization adjustment processing module is configured to optimize and adjust code data and model data corresponding to the GPU based on the performance data and the evaluation report.
6. A multi-module parallel accelerated processing image platform suitable for GPU, characterized in that, The computer readable storage medium stores a GPU multi-module parallel acceleration image processing platform control program, and the GPU multi-module parallel acceleration image processing platform control program implements the GPU multi-module parallel acceleration image processing method in any one of claims 1 to 3.
7. A computer readable storage medium, characterized in that, The computer readable storage medium stores a GPU multi-module parallel acceleration image processing platform control program, and the GPU multi-module parallel acceleration image processing platform control program implements the GPU multi-module parallel acceleration image processing method in any one of claims 1 to 3.
Citation Information
Patent Citations
High-performance server GPU performance bottleneck tuning method and device and storage medium
CN112000472A
Intelligent interior design image rendering system
CN117830489A