Data processing method and device, equipment and storage medium

By determining the identification of the server device in the preset address mapping relationship and using the GPU resources of the device to process data, the problem of waste of GPU resources when there are fewer applications on the physical machine is solved, and efficient utilization of GPU resources is achieved and waste of computing resources is reduced.

CN120104293APending Publication Date: 2025-06-06CHINA UNITED NETWORK COMM GRP CO LTD +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311648924.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-04
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

In the prior art, when there are fewer applications running on the physical machine, the GPU cannot fully utilize, resulting in waste of GPU resources; when there are many applications running, it may cause insufficient CPU resources or insufficient GPU resources, which will lead to waste of computing resources.

Method used

By determining the identification of the server device in the preset address mapping relationship, the server device has GPU resources running on the server device, the client device sends a data processing request to the server device, uses the server device's GPU resources to process the data, and returns the processing result to the client device.

Benefits of technology

Remotely call GPU resources on server devices, reducing the use of memory resources on client devices and reducing waste of computing resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120104293A_ABST
    Figure CN120104293A_ABST
Patent Text Reader

Abstract

According to the data processing method and device, the equipment and the storage medium provided by the invention, the identifier of the server-side equipment is determined in the preset address mapping relation, the GPU resources run on the server-side equipment, the data processing request is sent to the server-side equipment, the data processing request carries the to-be-processed data, and the to-be-processed data is sent to the server-side equipment; and after the GPU resource completes related processing of the to-be-processed data, receiving a processing result corresponding to the to-be-processed data sent by the server-side equipment, and in the technical scheme, processing of the data is realized by remotely calling the GPU resource on the server-side equipment, so that occupation of memory resources of the client-side equipment is reduced, and waste of computing resources is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of Internet technology, and in particular to a data processing method, device, equipment and storage medium. Background Art

[0002] The Central Processing Unit (CPU) is mainly used to process horizontal calculations, while the Graphics Processing Unit (GPU) is mainly used for parallel computing. With the explosive growth of the artificial intelligence industry, computing has become more complicated and computing power is insufficient. In addition, the parallel computing capability of the CPU is inferior to that of the GPU, making the general computing advantage of the GPU increasingly obvious.

[0003] In the prior art, a CPU and a GPU are often provided on a physical machine so that the GPU can be called in real time to realize data processing during data processing.

[0004] However, when fewer applications are running on a physical machine, the GPU cannot be fully utilized, resulting in a waste of GPU resources; when more applications are running on a physical machine, it may cause insufficient CPU resources and surplus GPU resources, or it may cause surplus CPU resources and insufficient GPU resources. In either scenario, it is a waste of computing resources. Summary of the invention

[0005] The present application provides a data processing method, apparatus, device and storage medium to solve the technical problem of wasting computing resources.

[0006] In a first aspect, the present application provides a data processing method, applied to a client device, the method comprising:

[0007] Determine an identifier of a server device in a preset address mapping relationship, wherein a graphics processor GPU resource runs on the server device;

[0008] Sending a data processing request to a server device, wherein the data processing request carries data to be processed;

[0009] After the GPU resource completes the relevant processing of the data to be processed, a processing result corresponding to the data to be processed sent by the server device is received.

[0010] Optionally, the method further includes:

[0011] Obtaining the identification of the server device;

[0012] The address mapping relationship is generated according to the identifier of the server device and the identifier of the global variable in the device program in the parallel computing platform and programming model.

[0013] Optionally, determining the identifier of the server device in a preset address mapping relationship includes:

[0014] In response to the GPU call request, the identifier of the server device is determined in the address mapping relationship according to the identifiers of global variables in the device program in the parallel computing platform and the programming model.

[0015] In a second aspect, the present application provides a data processing method, which is applied to a server device, and the method includes:

[0016] Receiving a data processing request sent by a client device, wherein the data processing request carries data to be processed;

[0017] Perform relevant processing on the data to be processed according to GPU resources to obtain processing results corresponding to the data to be processed;

[0018] The processing result corresponding to the data to be processed is sent to the client device.

[0019] Optionally, the data processing request also carries an identifier of a global variable in a device program in a parallel computing platform and a programming model;

[0020] The performing relevant processing on the data to be processed according to the GPU resources to obtain a processing result corresponding to the data to be processed includes:

[0021] According to the identifiers of global variables in the device program in the parallel computing platform and the programming model, the identifier of the server device is determined in a preset address mapping relationship;

[0022] Determine the address information of the GPU resource according to the identifier of the server device;

[0023] Based on the address information of the GPU resource, the data to be processed is processed in accordance with the GPU resource to obtain a processing result corresponding to the data to be processed.

[0024] Optionally, the method further includes:

[0025] Obtaining the identification of the server device and the identification of global variables in the device program in the parallel computing platform and programming model;

[0026] The address mapping relationship is generated according to the identifier of the server device and the identifier of the global variable in the device program in the parallel computing platform and programming model.

[0027] In a third aspect, the present application provides a data processing device, applied to a client device, the device comprising:

[0028] A determination module, used to determine an identifier of a server device in a preset address mapping relationship, wherein a graphics processing unit (GPU) resource is running on the server device;

[0029] A sending module, used to send a data processing request to a server device, wherein the data processing request carries data to be processed;

[0030] The acquisition module is used to receive the processing result corresponding to the data to be processed sent by the server device after the GPU resource completes the relevant processing of the data to be processed.

[0031] Optionally, the acquisition module is further used to:

[0032] Obtaining the identification of the server device;

[0033] The address mapping relationship is generated according to the identifier of the server device and the identifier of the global variable in the device program in the parallel computing platform and programming model.

[0034] Optionally, the determining module is specifically used to:

[0035] In response to the GPU call request, the identifier of the server device is determined in the address mapping relationship according to the identifiers of global variables in the device program in the parallel computing platform and the programming model.

[0036] In a fourth aspect, the present application provides a data processing device, applied to a server device, the device comprising:

[0037] An acquisition module, configured to receive a data processing request sent by a client device, wherein the data processing request carries data to be processed;

[0038] A processing module, used for performing relevant processing on the data to be processed according to GPU resources to obtain processing results corresponding to the data to be processed;

[0039] The sending module is used to send the processing result corresponding to the data to be processed to the client device.

[0040] Optionally, the data processing request also carries an identifier of a global variable in a device program in a parallel computing platform and a programming model;

[0041] The processing module is specifically used for:

[0042] According to the identifiers of global variables in the device program in the parallel computing platform and the programming model, the identifier of the server device is determined in a preset address mapping relationship;

[0043] Determine the address information of the GPU resource according to the identifier of the server device;

[0044] Based on the address information of the GPU resource, the data to be processed is processed in accordance with the GPU resource to obtain a processing result corresponding to the data to be processed.

[0045] Optionally, the acquisition module is further used to:

[0046] Obtaining the identification of the server device and the identification of global variables in the device program in the parallel computing platform and programming model;

[0047] The address mapping relationship is generated according to the identifier of the server device and the identifier of the global variable in the device program in the parallel computing platform and programming model.

[0048] In a fifth aspect, the present application provides an electronic device, comprising: a processor, and a memory communicatively connected to the processor;

[0049] The memory stores computer-executable instructions;

[0050] The processor executes the computer-executable instructions stored in the memory to implement the methods described in the first and second aspects and various possible designs.

[0051] In a sixth aspect, the present application provides a computer-readable storage medium, in which computer-executable instructions are stored. When the computer-executable instructions are executed by a processor, they are used to implement the methods described in the first and second aspects and various possible designs mentioned above.

[0052] The data processing method, apparatus, device and storage medium provided by the present application determine the identifier of the server device in a preset address mapping relationship, GPU resources are running on the server device, and a data processing request is sent to the server device, the data processing request carries the data to be processed, and then after the GPU resources complete the relevant processing of the data to be processed, the processing result corresponding to the data to be processed sent by the server device is received. In this technical solution, data processing is realized by remotely calling the GPU resources on the server device to reduce the occupancy of the memory resources of the client device, thereby reducing the waste of computing resources. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0054] Figure 1 Schematic diagram of the data processing method provided in the embodiment of the present application Figure 1 ;

[0055] Figure 2 Schematic diagram of the data processing method provided in the embodiment of the present application Figure 2 ;

[0056] Figure 3 A schematic diagram of the structure of the data processing device provided in the embodiment of the present application Figure 1 ;

[0057] Figure 4 A schematic diagram of the structure of the data processing device provided in the embodiment of the present application Figure 2 ;

[0058] Figure 5 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application.

[0059] The above drawings have shown clear embodiments of the present application, which will be described in more detail later. These drawings and text descriptions are not intended to limit the scope of the present application in any way, but to illustrate the concept of the present application to those skilled in the art by referring to specific embodiments. DETAILED DESCRIPTION

[0060] Exemplary embodiments will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The implementations described in the following exemplary embodiments do not represent all implementations consistent with the present application. Instead, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.

[0061] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant laws, regulations and standards, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0062] First, the professional terms involved in the embodiments of the present application are explained:

[0063] Graphics Processing Unit (GPU): Also known as display core, visual processor, display chip, it is a microprocessor that specializes in image and graphics related computing on personal computers, workstations, game consoles and some mobile devices (such as tablets, smartphones, etc.);

[0064] Kernel function: a function executed on the GPU. The kernel function source code will be compiled into a binary executable by the GPU, and the GPU driver will put the binary on the GPU for execution;

[0065] The parallel computing platform and programming model (Compute Unified Device Architecture, CUDA) can significantly improve computing performance by leveraging the processing power of graphics processing units (GPUs);

[0066] The central processing unit (CPU) is the computing and control core of the computer system and the final execution unit for information processing and program running.

[0067] Remote Procedure Call (RPC): A third-party client program calls a standard or custom function within SAP through an interface, obtains the data returned by the function, processes it, and then displays or prints it.

[0068] Secondly, the technical background involved in the embodiments of the present application is described:

[0069] GPU chips, commonly known as "graphics cards", are good at parallel computing, while CPUs are good at horizontal computing; the two form a golden pair for heterogeneous computing. With the explosive growth of the artificial intelligence industry, computing has become more complicated and computing power is insufficient. In addition, the parallel computing capabilities of CPUs are inferior to those of GPUs, making the general computing advantages of GPUs more obvious.

[0070] CUDA is a general-purpose parallel computing platform and programming model proposed by NVIDIA that uses GPU acceleration. A typical CUDA program execution steps are as follows:

[0071] 1. Allocate GPU memory (i.e. video memory) through cudaMalloc;

[0072] 2. Copy data from CPU memory to video memory through application programming interfaces (APIs) such as cudaMemcpy / cudaMemcpyToSymbol;

[0073] 3. Call cudaLaunchKernel to call the kernel function to calculate the data stored in the video memory;

[0074] 4. Call cudaMemcpy / cudaMemcpyFromSymbol and other APIs to transfer data from video memory back to CPU memory. ;

[0075] 5. Call cudaFree to release video memory.

[0076] CUDA is a closed-source development tool developed by NVIDIA that needs to be bound to GPU card hardware. That is, it can only be used on servers equipped with GPUs. In other words, the GPU is limited to local use on physical machines, which has the following disadvantages:

[0077] Resource configuration is not flexible enough. When there are fewer applications running on the physical machine, the GPU cannot be fully utilized, resulting in a waste of GPU resources. When there are more applications running on the physical machine, it may cause insufficient CPU resources and surplus GPU resources, or it may cause surplus CPU resources and insufficient GPU resources. In either scenario, it is a waste of computing resources.

[0078] Some APIs do not have a feasible solution to implement remote calls, such as:

[0079] cudaMemcpyToSymbol / cudaMemcpyFromSymbol, their function prototypes are:

[0080] cudaError_t cudaMemcpyToSymbol(const void*symbol,const void*src,size_t count,size_t offset,enum cudaMemcpyKind kind)

[0081] cudaError_t cudaMemcpyFromSymbol(void*dst,const char*symbol,size_tcount,size_t offset,enum cudaMemcpyKind kind);

[0082] Among them, cudaMemcpyToSymbol is used to copy the data in the host memory pointed to by src to the video memory associated with symbol, and cudaMemcpyFromSymbol is used to copy the data in the video memory associated with symbol to the host memory pointed to by dst.

[0083] Symbol is the address value of the video memory variable registered in the memory virtual address space. When the two APIs cudaMemcpyToSymbol / cudaMemcpyFromSymbol are called locally, the internal implementation of the API will automatically translate the symbol address value into the video memory address (the video memory address is invisible) and complete the data copy.

[0084] When the two APIs cudaMemcpyToSymbol / cudaMemcpyFromSymbol are implemented remotely, the address value saved by the symbol is located on the client side (i.e., the client side), and the server side (the server side) cannot recognize the symbol address in the client side address space; the video memory corresponding to the symbol is located on the server side, and the address value is unknown, so it cannot be directly copied.

[0085] The data processing method provided in this application is intended to solve the above technical problems in the prior art.

[0086] The technical solution of the present application and how the technical solution of the present application solves the above-mentioned technical problems are described in detail below with specific embodiments. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below in conjunction with the accompanying drawings.

[0087] Figure 1 Schematic diagram of the data processing method provided in the embodiment of the present application Figure 1 ,like Figure 1 As shown, the data processing method includes the following steps:

[0088] Step 11: The client device determines the identifier of the server device in a preset address mapping relationship;

[0089] Among them, GPU resources are running on the server device;

[0090] In this step, an application is running on the client device. The application needs to use GPU resources to process some data, but there are no corresponding GPU resources on the client device. It is necessary to call the server device to use the GPU resources in the server device to process the data accordingly so that the application can be applied.

[0091] Prior to this step, the identification of the server device is obtained, and an address mapping relationship is generated according to the identification of the server device and the identification of the global variables in the device program in the parallel computing platform and the programming model.

[0092] In one implementation, Symbol refers to a global variable in the device program in CUDA programming, which is stored in the global memory of the client device. Symbols can be accessed on the device side and accessed under the host virtual address space through a specific address.

[0093] That is, under this implementation, the address mapping relationship between the client-side (client device) symbol and the server-side (server device) video memory (i.e. GPU memory) is established:

[0094] Specifically, if you want to copy the data in src to the video memory associated with symbol by calling the api cudaMemcpyFromSymboll(symbol, src, size_t count, size_t offset) on the client side, you must first establish a mapping relationship between the symbol address value and the server-side video memory.

[0095] The address mapping relationship is established by hijacking the __cudaRegisterVar api, and its function prototype is: __cudaRegisterVar(void**fatCubinHandle,char*symbol,char*deviceAddress,constchar*deviceName,int ext,size_t size,int constant,int global);

[0096] Among them, fatCubinHandle points to fatbindata (the executable binary file under the GPU system compiled by the nvcc compiler) in the client virtual address space, deviceAddress and deviceName (which can be the identifier of the server device) are the address value and variable name of the video memory variable respectively, and symbol is the address of the video memory variable in the host virtual address space.

[0097] Optionally, this step may be implemented by: in response to the GPU call request, determining the identifier of the server device in the address mapping relationship according to the identifier of the global variable in the device program in the parallel computing platform and the programming model.

[0098] Under this implementation, the identifier of the server device corresponding to the use of GPU resources is determined.

[0099] Step 12: The client device sends a data processing request to the server device;

[0100] The data processing request carries data to be processed.

[0101] In this implementation, it is implemented by the CUDA Libraries module (a dynamic link library used to hijack cuda runtime API calls); in the cuda environment, applications usually implement video memory application / release / data copy (copy between video memory and host memory) by calling runtime APIs such as cudaMalloc / cudaFree / cudaMemcpyFromSymbol, etc.; in order to be compatible with the existing cuda ecosystem and maintain the unchanged usage of the App, the client-side CUDALibraries re-implemented all runtime APIs in the original API format.

[0102] In the specific implementation of the API, the RPC-CLIENT function is called to forward the received API call request (that is, it can be a data processing request in implementation) to the server. Taking the cudaMemcpyFromSymbol API as an example, as shown below:

[0103] / / API prototype, its function is to copy count size data from the offset offset of the starting address of the video memory associated with symbol to the address pointed to by dst.

[0104] cudaError_t cudaMemcpyFromSymbol(void*dst,const char*symbol,size_tcount,size_t offset)

[0105] {

[0106] 1. Prepare rpc buffer

[0107] buf_len=sizeof(symbol)+sizeof(count)+sizeof(kind)+sizeof(offset);

[0108] buffer=start=malloc(buf_len); ...

[0110] 2. Fill the client's symbol address into the rpc buffer and send it to the server

[0111] memcpy(buffer,&symbol,sizeof(symbol));

[0112] buffer=(uint8_t*)buffer+sizeof(symbol); ...

[0114] 3. Execute rpc call

[0115] / / rpc call

[0116] auto response=rpc_launch_call(__func__,start,buf_len); ...

[0118] 4. Get the data returned by the server from the rpc response buffer and copy it to the host memory pointed to by dst const char*resp = response.rcuda_resp_buffer().c_str();

[0119] Get the execution result from the rpc responsebuffer

[0120] memcpy(&ret,resp,sizeof(cudaError_t));

[0121] Take the data from the rpc responsebuffer and copy it to the memory pointed to by dst.

[0122] resp+=sizeof(cudaError_t);

[0123] memcpy(dst,resp,count);

[0124] return ret;

[0125] }

[0126] In a possible implementation, the RPC-CLIENT on the client side is used to receive the call request of CUDA Libraries and send the request to the server side.

[0127] Step 13: The server device performs relevant processing on the data to be processed according to the GPU resources to obtain the processing result corresponding to the data to be processed;

[0128] In this step, after obtaining the data to be processed, the GPU resources calculate the data to be processed to obtain a processing result corresponding to the data to be processed.

[0129] Step 14: The server device sends the processing result corresponding to the data to be processed to the client device.

[0130] In a possible implementation, the RPC-CLIENT on the client side receives the return information from the server side (ie, the processing result corresponding to the data to be processed), and provides the return information to the CUDA Libraries.

[0131] The data processing method provided in the embodiment of the present application determines the identifier of the server device in a preset address mapping relationship, a GPU resource runs on the server device, and sends a data processing request to the server device, the data processing request carries the data to be processed, and then after the GPU resource completes the relevant processing of the data to be processed, receives the processing result corresponding to the data to be processed sent by the server device. In this technical solution, the data is processed by remotely calling the GPU resources on the server device to reduce the occupancy of the memory resources of the client device, thereby reducing the waste of computing resources.

[0132] Based on the above embodiment, the data processing request also carries the identifiers of global variables in the device program in the parallel computing platform and the programming model; then Figure 2 Schematic diagram of the data processing method provided in the embodiment of the present application Figure 2 ,like Figure 2 As shown, step 13 in the data processing method may include the following steps:

[0133] Step 21: According to the identifiers of global variables in the device program in the parallel computing platform and the programming model, the identifier of the server device is determined in the preset address mapping relationship;

[0134] In this step, it is first necessary to determine whether the server device can provide GPU resource related services to the client device.

[0135] Before this step, you can also: obtain the identification of the server device and the identification of the global variables in the device program in the parallel computing platform and programming model; generate an address mapping relationship based on the identification of the server device and the identification of the global variables in the device program in the parallel computing platform and programming model.

[0136] In one possible implementation, by hijacking __cudaRegisterVar, a set of information (fatbindata, symbol value, deviceName) is sent to the server through RPC, and two maps are established on the server:

[0137] client_symbol_var_name_map[symbol value]=deviceName

[0138] client_symbol_fatbindata_map[symbol value]=fatbindata

[0139] They are used to save the address mapping relationship of (symbol value, fatbindata) and (symbol value, deviceName) respectively.

[0140] Among them, fatbindata (also known as fat binary data) is used to describe binary files containing binary codes for multiple platforms. In CUDA programming, fatbindata is often used in applications that support multiple GPU architectures at the same time.

[0141] Step 22: Determine the address information of the GPU resource according to the identifier of the server device;

[0142] In a possible implementation, the server can obtain fatbindata and deviceName through the symbol value and two maps respectively. After obtaining these two pieces of information, the video memory address devPtr, that is, the address information of the GPU resource, can be obtained through the following two steps.

[0143] cuModuleLoadFatBinary(&module,fatbindata);

[0144] cuModuleGetGlobal(&devPtr,...,module,deviceName).

[0145] Step 23: Based on the address information of the GPU resources, the data to be processed is processed in accordance with the GPU resources to obtain a processing result corresponding to the data to be processed.

[0146] In a possible implementation, the RPC-SERVER on the server side includes two main functions:

[0147] 1) Execute the service processing function to receive and process the call request from the client;

[0148] 2) Call the original CUDA API to manage video memory.

[0149] / / Server-side cudaMemcpyFromSymbol service processing function

[0150] cudaError_tRpcSvcCudaMemcpyFromSymbol(constrcuda::RemoteCudaRequest*request,rcuda::RemoteCudaResponse*response){

[0151] 1. Parse the symbol information filled by the client from the RPC request buffer

[0152] tmpptr=(void*)request->rcuda_request_buffer().c_str();

[0153] memcpy(&clientSymbol,tmpptr,sizeof(clientSymbol)); ...

[0155] 2. Get varName and fatbindata through symbol information and the map established in step 1

[0156] varName=client_symbol_server_var_map[clientSymbol];

[0157] fatbindata=client_symbol_fatbindata_map[clientSymbol];

[0158] 3. Get the video memory address detPtr associated with clientSymbol

[0159] cuModuleLoadFatBinary(&module,fatbindata);

[0160] cuModuleGetGlobal(&devPtr,...,module,varName);

[0161] 4. Allocate the rpc response buffer and fill the buffer with the video memory data at devPtr+offset and the copy result

[0162] result_buf=malloc(result_buf_len);

[0163] dst=(char*)result_buf+sizeof(cudaError_t);

[0164] curesult=cuMemcpyDtoH(dst,devPtr+offset,count);

[0165] memcpy(result_buf,&curesult,sizeof(cudaError_t));

[0166] 5. Return video memory data and API call results to the client

[0167] response->set_rcuda_resp_buffer(result_buf,result_buf_len);

[0168] response->set_rcuda_resp_buffer_len(result_buf_len);

[0169] }.

[0170] The data processing method provided in the embodiment of the present application determines the identification of the server device in the preset address mapping relationship according to the identification of the global variables in the device program in the parallel computing platform and the programming model; determines the address information of the GPU resource according to the identification of the server device; based on the address information of the GPU resource, performs relevant processing on the data to be processed according to the GPU resource to obtain the processing result corresponding to the data to be processed. This technical solution realizes the processing of the data to be processed by the GPU resource in the server device.

[0171] The following are device embodiments of the present application, which can be used to execute the method embodiments of the present application. For details not disclosed in the device embodiments of the present application, please refer to the method embodiments of the present application.

[0172] Figure 3 A schematic diagram of the structure of the data processing device provided in the embodiment of the present application Figure 1 .like Figure 3 As shown, the data processing device is applied to a client device, including:

[0173] A determination module 31 is used to determine the identifier of the server device in a preset address mapping relationship, where a graphics processor GPU resource runs on the server device;

[0174] A sending module 32, used to send a data processing request to a server device, wherein the data processing request carries data to be processed;

[0175] The acquisition module 33 is used to receive the processing result corresponding to the data to be processed sent by the server device after the GPU resource completes the relevant processing of the data to be processed.

[0176] Optionally, the acquisition module 33 is further used to:

[0177] Obtaining the identification of the server device;

[0178] The address mapping relationship is generated according to the identifier of the server device and the identifier of the global variable in the device program in the parallel computing platform and programming model.

[0179] Optionally, the determining module 31 is specifically configured to:

[0180] In response to the GPU call request, the identifier of the server device is determined in the address mapping relationship according to the identifiers of global variables in the device program in the parallel computing platform and the programming model.

[0181] The data processing device provided in the embodiment of the present application can be used to execute the data processing method applied to the client device in any of the above embodiments. Its implementation principle and technical effects are similar and will not be repeated here.

[0182] Figure 4 A schematic diagram of the structure of the data processing device provided in the embodiment of the present application Figure 2 .like Figure 4 As shown, the data processing device is applied to a server device, including:

[0183] The acquisition module 41 is used to receive a data processing request sent by a client device, wherein the data processing request carries data to be processed;

[0184] The processing module 42 is used to perform relevant processing on the data to be processed according to the GPU resources to obtain processing results corresponding to the data to be processed;

[0185] The sending module 43 is used to send the processing result corresponding to the data to be processed to the client device.

[0186] Optionally, the data processing request also carries an identifier of a global variable in a device program in a parallel computing platform and a programming model;

[0187] The processing module 42 is specifically used for:

[0188] According to the identifiers of global variables in the device program in the parallel computing platform and the programming model, the identifier of the server device is determined in a preset address mapping relationship;

[0189] Determine the address information of the GPU resource according to the identifier of the server device;

[0190] Based on the address information of the GPU resource, the data to be processed is processed in accordance with the GPU resource to obtain a processing result corresponding to the data to be processed.

[0191] Optionally, the acquisition module 41 is further used for:

[0192] Obtaining the identification of the server device and the identification of global variables in the device program in the parallel computing platform and programming model;

[0193] The address mapping relationship is generated according to the identifier of the server device and the identifier of the global variable in the device program in the parallel computing platform and programming model.

[0194] The data processing device provided in the embodiment of the present application can be used to execute the data processing method applied to the server device in any of the above embodiments. Its implementation principle and technical effects are similar and will not be repeated here.

[0195] It should be noted that it should be understood that the division of the various modules of the above device is only a division of logical functions. In actual implementation, they can be fully or partially integrated into one physical entity, or they can be physically separated. And these modules can all be implemented in the form of software calling through processing elements; they can also be all implemented in the form of hardware; some modules can also be implemented in the form of software calling through processing elements, and some modules can be implemented in the form of hardware. In addition, all or part of these modules can be integrated together, or they can be implemented independently. The processing element described here can be an integrated circuit with signal processing capabilities. In the implementation process, each step of the above method or each module above can be completed by an integrated logic circuit of hardware in the processor element or instructions in the form of software.

[0196] Figure 5 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application, such as Figure 5 As shown, the electronic device is a client device or a server device.

[0197] It may include: a processor 51, a memory 52, and computer program instructions stored in the memory 52 and executable on the processor 51, and the processor 51 implements the method provided in any of the above-mentioned embodiments when executing the computer program instructions.

[0198] Optionally, the above-mentioned components of the electronic device may be connected via a system bus.

[0199] The memory 52 may be a separate storage unit or a storage unit integrated in the processor 51. The number of the processor 51 may be one or more.

[0200] It should be understood that the processor 51 can be a central processing unit (CPU), or other general-purpose processors 51, digital signal processors 51 (DSP), application-specific integrated circuits (ASIC), etc. The general-purpose processor 51 can be a microprocessor 51 or the processor 51 can also be any conventional processor 51, etc. The steps of the method disclosed in the present application can be directly embodied as being executed by the hardware processor 51, or can be executed by a combination of hardware and software modules in the processor 51.

[0201] The system bus may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The system bus may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, only one thick line is used in the figure, but it does not mean that there is only one bus or one type of bus. The memory 52 may include a random access memory 52 (RAM), and may also include a non-volatile memory 52 (NVM), such as at least one disk storage 52.

[0202] All or part of the steps of the above-mentioned method embodiments can be completed by hardware related to program instructions. The above-mentioned program can be stored in a readable memory 52. ​​When the program is executed, the steps of the above-mentioned method embodiments are executed; and the above-mentioned memory 52 (storage medium) includes: read-only memory 52 (ROM), RAM, flash memory 52, hard disk, solid state hard disk, magnetic tape, floppy disk, optical disc and any combination thereof.

[0203] The electronic device provided in the embodiment of the present application can be used to execute the method provided in any of the above method embodiments. Its implementation principles and technical effects are similar and will not be repeated here.

[0204] An embodiment of the present application provides a computer-readable storage medium, in which computer instructions are stored. When the computer instructions are executed on a computer, the computer executes the above method.

[0205] The computer-readable storage medium mentioned above can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory, electrically erasable programmable read-only memory, erasable programmable read-only memory, programmable read-only memory, read-only memory, magnetic memory, flash memory, magnetic disk or optical disk. The readable storage medium can be any available medium that can be accessed by a general or special-purpose computer.

[0206] Optionally, a readable storage medium is coupled to a processor so that the processor can read information from the readable storage medium and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can be located in an application specific integrated circuit (ASIC). Of course, the processor and the readable storage medium can also exist in the device as discrete components.

[0207] An embodiment of the present application also provides a computer program product, which includes a computer program. The computer program is stored in a computer-readable storage medium. At least one processor can read the computer program from the computer-readable storage medium. When the at least one processor executes the computer program, the above method can be implemented.

[0208] It should be understood that the present application is not limited to the precise structures that have been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present application is limited only by the appended claims.

Claims

1. A data processing method, It is characterized in that Applied to a client device, the method comprises: Determine an identifier of a server device in a preset address mapping relationship, wherein a graphics processor GPU resource runs on the server device; Sending a data processing request to the server device, wherein the data processing request carries data to be processed; After the GPU resource completes the relevant processing of the data to be processed, a processing result corresponding to the data to be processed sent by the server device is received.

2. The method according to claim 1, It is characterized in that The method further comprises: Obtaining the identification of the server device; The address mapping relationship is generated according to the identifier of the server device and the identifier of the global variable in the device program in the parallel computing platform and programming model.

3. The method according to claim 2, It is characterized in that Determining the identifier of the server device in the preset address mapping relationship includes: In response to the GPU call request, the identifier of the server device is determined in the address mapping relationship according to the identifiers of global variables in the device program in the parallel computing platform and the programming model.

4. A data processing method, It is characterized in that Applied to a server device, the method comprises: Receiving a data processing request sent by a client device, wherein the data processing request carries data to be processed; Perform relevant processing on the data to be processed according to GPU resources to obtain processing results corresponding to the data to be processed; The processing result corresponding to the data to be processed is sent to the client device.

5. The method according to claim 4, It is characterized in that The data processing request also carries the identification of global variables in the device program in the parallel computing platform and programming model; The performing relevant processing on the data to be processed according to the GPU resources to obtain a processing result corresponding to the data to be processed includes: According to the identifiers of global variables in the device program in the parallel computing platform and the programming model, the identifier of the server device is determined in a preset address mapping relationship; Determine the address information of the GPU resource according to the identifier of the server device; Based on the address information of the GPU resource, the data to be processed is processed in accordance with the GPU resource to obtain a processing result corresponding to the data to be processed.

6. The method according to claim 5, It is characterized in that The method further comprises: Obtaining the identification of the server device and the identification of global variables in the device program in the parallel computing platform and programming model; The address mapping relationship is generated according to the identifier of the server device and the identifier of the global variable in the device program in the parallel computing platform and programming model.

7. A data processing device, It is characterized in that Applied to a client device, the device comprises: A determination module, used to determine an identifier of a server device in a preset address mapping relationship, wherein a graphics processing unit (GPU) resource is running on the server device; A sending module, used to send a data processing request to a server device, wherein the data processing request carries data to be processed; The acquisition module is used to receive the processing result corresponding to the data to be processed sent by the server device after the GPU resource completes the relevant processing of the data to be processed.

8. A data processing device, It is characterized in that Applied to a server device, the device comprises: An acquisition module, configured to receive a data processing request sent by a client device, wherein the data processing request carries data to be processed; A processing module, used for performing relevant processing on the data to be processed according to GPU resources to obtain processing results corresponding to the data to be processed; The sending module is used to send the processing result corresponding to the data to be processed to the client device.

9. An electronic device, It is characterized in that include: A processor, and a memory communicatively connected to the processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory to implement the method according to any one of claims 1 to 6.

10. A computer-readable storage medium, It is characterized in that The computer-readable storage medium stores computer-executable instructions, which are used to implement the method according to any one of claims 1 to 6 when executed by a processor.