Data processing method, device, electronic device and storage medium

By determining the machine learning model and hardware configuration information in the calculation graph processing and allocating processing threads according to the optimization strategy, the inefficiency of data processing caused by improper distribution of processing nodes in the prior art is solved, and more efficient calculation graph processing is achieved.

CN114764372BActive Publication Date: 2025-06-06ALIBABA GROUP HOLDING LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110058060.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-01-15
Publication Date
2025-06-06
Estimated Expiration
2041-01-15

AI Technical Summary

Technical Problem

In the prior art, when processing calculation diagrams, due to hardware resource limitations, it is impossible to allocate one processing thread to each processing node, resulting in processing nodes with a longer execution time being randomly allocated to the same thread, resulting in reduced data processing efficiency.

Method used

By determining the machine learning model and hardware configuration information, the processing threads corresponding to each processing unit of the machine learning model are determined based on the optimization strategy and hardware configuration information, and the processing unit is scheduled to the scheduling queue of the processing thread to complete the calculation of the processing unit through the processing thread.

Benefits of technology

Improve data processing efficiency, optimize the allocation and use of processing threads, avoid serial execution between processing nodes, and improve the processing performance of computing graphs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114764372B_ABST
    Figure CN114764372B_ABST
Patent Text Reader

Abstract

The embodiments of the present application provide a data processing method, device, electronic device and storage medium, wherein the method comprises: determining a machine learning model and hardware configuration information; determining a processing thread corresponding to each processing unit of the machine learning model based on an optimization strategy and the hardware configuration information; scheduling the processing unit to a scheduling queue of the processing thread so as to complete the calculation corresponding to the processing unit through the processing thread to obtain a data processing result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a data processing method, a data processing device, an electronic device and a storage medium. Background Art

[0002] Machine learning is a multi-disciplinary subject that involves probability theory, statistics, approximation theory, convex analysis, algorithm complexity theory, etc. Machine learning models are usually used to solve some reasoning problems. Common machine learning models include algorithm models, decision models, neural network models, etc. At present, multiple machine learning models are usually connected to form a computational graph of multiple compute nodes to provide services.

[0003] Currently, computational graphs are usually processed by processing threads. However, due to hardware resource limitations, it is usually impossible to assign a processing thread to each processing node. Therefore, the existing method is usually to randomly assign two or more nodes to the same processing thread for data processing.

[0004] However, in this way, two processing nodes with longer execution times may be assigned to the same thread for serial execution (execution of one node after the other is completed), which may reduce data processing efficiency. Summary of the invention

[0005] The embodiment of the present application provides a data processing method to improve data processing efficiency.

[0006] Correspondingly, an embodiment of the present application also provides a data processing device, an electronic device and a storage medium to ensure the implementation and application of the above system.

[0007] In order to solve the above problems, an embodiment of the present application discloses a data processing method, which method includes: determining a machine learning model and hardware configuration information; determining a processing thread corresponding to each processing unit of the machine learning model based on an optimization strategy and hardware configuration information; scheduling the processing unit to a scheduling queue of the processing thread to complete the calculation corresponding to the processing unit through the processing thread to obtain a data processing result.

[0008] In order to solve the above problems, an embodiment of the present application discloses a data processing method, including: analyzing the processing unit of the machine learning model to determine the corresponding optimization strategy; providing hardware configuration reference information based on the optimization strategy; receiving selection information of the hardware configuration reference information to determine the corresponding hardware configuration information.

[0009] In order to solve the above problems, an embodiment of the present application discloses a data processing method, including: providing an interactive page, the interactive page including multiple model upload interfaces corresponding to different programming methods; receiving a trigger on the model upload interface to obtain corresponding uploaded data, and uploading it to determine the machine learning model; receiving hardware configuration reference information corresponding to the machine learning model, and displaying it in the interactive page; receiving a selection operation on the hardware configuration reference information to determine the selection information and upload it to determine the hardware configuration information.

[0010] In order to solve the above problems, an embodiment of the present application discloses a data processing device, including: a processing model acquisition module, used to determine the machine learning model and hardware configuration information; a processing thread acquisition module, used to determine the processing threads corresponding to each processing unit of the machine learning model based on the optimization strategy and hardware configuration information; a processing unit scheduling module, used to schedule the processing unit to the scheduling queue of the processing thread, so as to complete the calculation corresponding to the processing unit through the processing thread to obtain the data processing result.

[0011] In order to solve the above problems, an embodiment of the present application discloses a data processing device, including: an optimization strategy determination module, used to analyze the processing unit of the machine learning model and determine the corresponding optimization strategy; a reference information determination module, used to provide hardware configuration reference information based on the optimization strategy; a selection information determination module, used to receive selection information of the hardware configuration reference information and determine the corresponding hardware configuration information.

[0012] In order to solve the above problems, an embodiment of the present application discloses a data processing device, including: an interactive page display module, used to provide an interactive page, the interactive page includes multiple model upload interfaces corresponding to different programming methods; a model data upload module, used to receive a trigger on the model upload interface to obtain the corresponding uploaded data, and upload it to determine the machine learning model; a reference information display module, used to receive the hardware configuration reference information corresponding to the machine learning model, and display it in the interactive page; a selection information upload module, used to receive a selection operation on the hardware configuration reference information to determine the selection information and upload it to determine the hardware configuration information.

[0013] In order to solve the above problems, an embodiment of the present application discloses an electronic device, including: a processor; and a memory, on which executable code is stored. When the executable code is executed, the processor executes one or more methods described in the above embodiments.

[0014] In order to solve the above problems, the embodiments of the present application disclose one or more machine-readable media on which executable codes are stored. When the executable codes are executed, the processor executes one or more methods described in the above embodiments.

[0015] Compared with the prior art, the embodiments of the present application have the following advantages:

[0016] In an embodiment of the present application, the machine learning model and hardware configuration information can be obtained, and according to the optimization strategy of the machine learning model, the processing unit of the machine learning model can be scheduled to the scheduling queue of the processing thread, so as to complete the calculation corresponding to the processing unit through the processing thread. In an embodiment of the present application, each processing unit in the machine learning model can be analyzed in advance to determine the optimization strategy corresponding to the machine learning model. In the process of data processing by the machine learning model, the processing thread can be allocated to the processing unit of the machine learning model according to the optimization strategy, which can improve the data processing efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1A It is a flowchart of a data processing method according to an embodiment of the present application;

[0018] Figure 1B is a flowchart of a data processing method according to another embodiment of the present application;

[0019] Figure 2 is a flowchart of a data processing method according to another embodiment of the present application;

[0020] Figure 3 is a flowchart of a data processing method according to another embodiment of the present application;

[0021] Figure 4 is a flowchart of a data processing method according to another embodiment of the present application;

[0022] Figure 5 is a flowchart of a data processing method according to another embodiment of the present application;

[0023] Figure 6 is a flowchart of a data processing method according to another embodiment of the present application;

[0024] Figure 7 is a structural schematic diagram of a data processing device according to an embodiment of the present application;

[0025] Figure 8 is a structural schematic diagram of a data processing device according to another embodiment of the present application;

[0026] Fig. 9 is a structural schematic diagram of a data processing device according to another embodiment of the present application;

[0027] Fig.10 It is a schematic diagram of the structure of a device provided in one embodiment of the present application. DETAILED DESCRIPTION

[0028] In order to make the above-mentioned objects, features and advantages of the present application more obvious and easy to understand, the present application is further described in detail below in conjunction with the accompanying drawings and specific implementation methods.

[0029] The embodiments of the present application can be applied to the field of machine learning models, and can optimize and analyze models based on machine learning. Machine learning is a multi-disciplinary interdisciplinary subject involving probability theory, statistics, approximation theory, convex analysis, algorithm complexity theory and other disciplines. Common machine learning models include algorithm models, decision models, neural network (NN) models, etc.

[0030] The following describes the method of the embodiment of the present application by taking the application of the method of the embodiment of the present application in the optimization scenario of the neural network model as an example, wherein the neural network model is a complex network system formed by a large number of simple neurons (or operators, processing nodes, computing nodes, etc.) widely connected to each other. In addition, the rapid development of services / applications / training based on deep learning has led to the fact that a single neural network model can no longer meet the needs of the scene. More scenarios will use multiple neural network models in series and parallel to form a calculation graph of multiple computing nodes (Compute Node) to provide services. Therefore, in the embodiment of the present application, the calculation graph of multiple computing nodes can be regarded as a neural network model.

[0031] like Figure 1A As shown, the embodiment of the present application can be executed by the server, and the server can provide a plurality of model upload interfaces corresponding to different programming methods to the terminal, and the model provider can upload the relevant data of the neural network model through the terminal model upload interface. After the server determines the neural network model based on the relevant data, it can analyze the neural network model to obtain an optimization strategy, wherein the optimization strategy includes the correspondence between the processing unit and the processing thread of the processor. Then, according to the optimization strategy, the corresponding hardware configuration reference information is determined, and then the hardware configuration reference information is sent to the terminal. The model provider can select the hardware configuration reference information at the terminal, and then upload the selection information to the server. After receiving the selection information, the server can determine the hardware configuration information selected by the model provider, so as to perform data processing of the neural network model according to the optimization strategy through the device corresponding to the hardware configuration information. In one example, the hardware configuration information may include device information of an accelerator device for applying the neural network model.

[0032] Specifically, if Figure 1BAs shown, in an embodiment of the present application, the server can provide the terminal with multiple model upload interfaces corresponding to different programming methods. The model provider can upload the upload data corresponding to the neural network model through the terminal model upload interface. After receiving the uploaded data, the server can compile the uploaded data into a neural network model by compiling the optimization model. For example, the uploaded data can be code data, and the compilation optimization model can compile the code data to determine the corresponding operator, and further obtain the neural network model; for another example, the uploaded data can be a model architecture diagram, and the model architecture diagram can include nodes and connecting edges. The compilation optimization model can analyze the node elements of the nodes and the edge elements of the connecting edges in the model architecture diagram, and compile and process them to obtain the neural network model. Then, at least one neural network model can be converted into a computational graph (Compute Directed acyclic graph) through the compilation framework. The computational graph is composed of computational nodes (Compute Node), and the computational nodes can be understood as processing nodes and operators.

[0033] Afterwards, the server can analyze each processing unit in the calculation graph to determine the execution time required for each processing unit to process data, wherein the processing unit includes at least one computing node, such as configuring a processing thread of a processor for the processing unit, and setting the input data volume threshold of the processing thread to determine the execution time required for the processing thread to execute the processing unit alone, wherein the processor may include a central processing unit (CPU), a graphics processing unit (GPU) and other processors. The processing thread can also be called a thread, which is the smallest unit that can perform calculation scheduling. It is included in the process and is the actual operation unit in the process. A thread refers to a single sequential control flow in a process. Multiple threads can be concurrent in a process, and each thread executes different tasks in parallel. The input data volume threshold can also be called a batch size. The role of the batch size is to limit the number of samples selected by the processing unit for a data processing (such as training). The size of the batch size affects the optimization degree and speed of the model. For example, when the total number of training data remains unchanged, if the amount of data input at a time (batch size) increases, the number of iterations decreases; if the amount of data input at a time (batch size) decreases, the number of iterations increases. By adjusting the batch size, the optimization speed of the model can be controlled. For example, when using 1,000 sets of training data to train a neural network model, if the batch size is 100 sets, the corresponding neural network model needs to be iterated 10 times; if the batch size is 10 sets, the corresponding neural network model needs to be iterated 100 times. By adjusting the batch size, the optimization speed of the model can be controlled.

[0034] After determining the execution time required by the processing unit, the server can adjust the processing thread and input data volume threshold of the processing unit according to the execution time corresponding to each processing unit, and then obtain an optimization strategy, wherein the optimization strategy includes the correspondence between the processing unit and the processing thread of the processor. For example, the server can bind the two processing units with the shortest execution time, and configure the corresponding processing threads and the input data volume threshold of the processing threads to form an optimization strategy; it can also adjust the input data volume threshold of the processing unit with the shortest execution time, and then determine the optimization strategy.

[0035] After determining the optimization strategy, the corresponding hardware configuration reference information can be determined based on the optimization strategy, and the hardware configuration reference information can be sent to the terminal of the model provider. The model provider can select the hardware configuration reference information at the terminal, and then upload the selection information to the server. After the server receives the selection information, it can determine the accelerator device selected by the model provider and obtain the hardware configuration information so as to perform data processing of the neural network model through the accelerator device. In the embodiment of the present application, the neural network model can be trained by the accelerator device, and the data processing of the trained neural network model can also be performed by the accelerator device, which can be specifically set according to the needs.

[0036] Specifically, taking data processing through a trained neural network model as an example, the embodiment of the present application can obtain the optimization strategy corresponding to the neural network model, and the optimization strategy includes the corresponding relationship between the processing unit and the processing thread of the processor in the accelerator device. Then, the server can allocate the processing unit in the calculation graph to the reasoning request queue by running the system, and read the processing unit from the reasoning request queue through the corresponding reasoning model, and allocate it to the processing unit queue, wherein the processing unit queue can be divided according to the occupancy of the processing thread by the processing unit, and specifically may include: a computing-intensive queue, a memory-intensive queue, and a data transmission queue. After the reasoning model schedules the processing unit to the processing unit queue, the execution sequence (or scheduling sequence) between the processing units can be determined according to the dependency relationship between the processing units, and the processing units in the processing unit queue can be mapped to the processing thread queue according to the order, so as to complete the calculation corresponding to the processing unit through the corresponding processing thread. In the embodiment of the present application, the execution time of each processing unit in the neural network model can be analyzed in advance to determine the optimization strategy corresponding to the neural network model. In the process of data processing by the neural network model, the processing thread can be allocated to the processing unit of the neural network model according to the optimization strategy, which can improve the processing efficiency of the data.

[0037] The method of the embodiment of the present application is to optimize the basic level (hardware configuration) of the machine learning model. Therefore, the method of the embodiment of the present application can be applied to various data processing scenarios of machine learning models, for example, it can be applied to the training scenario of the machine learning model, and it can also be applied to the application scenario of the trained machine learning model; for another example, it can also be applied to the machine learning model for processing audio, image, video and other data, such as the machine learning model can be used for voice recognition of audio, depth recognition of image or video, face recognition of image, body movement recognition of characters in image, etc. For another example, it can also be applied to the training and application scenarios of deep neural network (Deep Neural Networks, DNN) model. For another example, the embodiment of the present application can be applied to the machine learning model in the scenes of live broadcast, e-commerce, finance, logistics, social networking, automatic driving, etc., and can optimize and analyze the machine learning model in the training stage and / or reasoning stage of the machine learning model in the above scene to improve the data processing efficiency of the machine learning model in the above scene. For example, it can be applied to optimize the machine learning model used to complete object recognition, face recognition or trajectory tracking in the live broadcast scene. For another example, it can be applied to optimize the machine learning model used to complete product recognition or product recommendation in e-commerce scenarios. For another example, it can be applied to optimize the machine learning model used to complete object recognition, object trajectory tracking or object distance recognition in autonomous driving scenarios.

[0038] The embodiment of the present application provides a data processing method that can be applied to a server. The server can be understood as a device that performs machine learning model training and / or applies the trained machine learning model to perform data processing. The server can interact with the terminal of the model provider. The server can analyze the machine learning model provided by the model provider, determine the corresponding optimization strategy, and provide the model provider with hardware configuration reference information according to the optimization strategy, so that the model provider can select hardware suitable for the machine learning model. Specifically, Figure 2 As shown, the method includes:

[0039] Step 202, analyze the processing unit of the machine learning model and determine the corresponding optimization strategy, wherein the optimization strategy includes the processing thread and the input data volume threshold corresponding to the processing unit. The server can provide multiple interfaces corresponding to different programming methods to the terminal, and the model provider can upload the relevant data of the machine learning model through the corresponding interface. Specifically, as an optional embodiment, the method also includes: providing an interactive page, wherein the interactive page includes multiple model upload interfaces corresponding to different programming methods; obtaining the uploaded data according to the model upload interface, and compiling and processing the uploaded data to obtain the machine learning model. The server can send an interactive page to the terminal, wherein the interactive page includes multiple types of model upload interfaces, and the model provider can select the corresponding model upload interface through the interactive page of the terminal to upload the uploaded data of the machine learning model, wherein the uploaded data can be a code segment corresponding to different programming languages, or a model architecture diagram, wherein the model architecture diagram can include nodes and connecting edges, wherein the nodes of the model architecture diagram correspond to the computing nodes (or operators), and the connecting edges of the model architecture diagram correspond to the dependencies (sequence) between the computing nodes, and the computing nodes corresponding to the nodes and the dependencies between the computing nodes can be determined by configuring the node elements of the nodes in the model architecture diagram and the edge elements of the connecting edges. After receiving the model architecture diagram, the server can analyze the node elements of the nodes and the edge elements of the connecting edges in the model architecture diagram, and compile and process them to obtain a machine learning model.

[0040] After determining the machine learning model, the server may analyze the execution time of each processing unit in the machine learning model to determine an optimization strategy for the corresponding machine learning model. Specifically, as an optional embodiment, the processing unit of the machine learning model is analyzed to determine the corresponding optimization strategy, including: allocating a processing thread to each processing unit of the machine learning model, and configuring an input data volume threshold of the processing thread; determining the execution time required for the processing unit based on the input data volume threshold and processing thread corresponding to the processing unit; determining whether the machine learning model meets the preset operating conditions based on the execution time required for the processing unit, and obtaining a model operation analysis result; when the model operation analysis result is the first result, determining the optimization strategy based on the processing thread and the input data volume threshold of the processing unit.

[0041] In an optional example, the input data volume threshold can be set according to the input data volume upper limit of the processing thread. Specifically, the configuration of the input data volume threshold of the processing thread includes: obtaining the input data volume upper limit corresponding to the processing thread assigned to the processing unit; configuring the input data volume threshold of the processing thread according to the input data volume upper limit. The input data volume threshold of the processing thread can be set according to the corresponding proportion based on the input data volume upper limit of each processing thread. For example, the input data flow threshold of the processing thread can be set according to the proportion of 100%, 80%, 60%, etc. After the processing thread and the input data volume threshold are configured for the processing unit, the execution time required for the processing unit to process data can be obtained. Then, based on the execution time required for the processing unit, it can be determined whether the machine learning model meets the preset operating conditions. The preset operating conditions may include the number of processing threads, the model execution time of the machine learning model, etc. The preset operating conditions can be determined based on the complexity of the machine learning model and the computing power of the accelerator device. For example, the ratio information can be determined based on the complexity of the machine learning model and the computing power of the accelerator device, and the preset operating conditions can be determined based on the ratio information and the execution time required by the processing unit with the longest execution time in the machine learning model. The server can determine whether the machine learning model meets the preset operating conditions based on the execution time required by the processing unit of the machine learning model, and then obtain the model operation analysis results.

[0042] The model operation analysis result may include a first result and a second result, wherein the first result may be understood as the machine learning model meets the preset operation conditions, and the second result may be understood as the machine learning model does not meet the preset operation conditions. When the model operation analysis result is the first result, the server may record the processing thread and input data volume threshold of the processing unit to obtain an optimization strategy. When the model operation result is the second result, the server may filter out the target processing unit from the processing unit and optimize the target processing unit for the next analysis. Specifically, as an optional embodiment, the method further includes: when the model operation analysis result is the second result, filter out at least two target processing units with the shortest execution time; obtain the execution time difference between the target processing units, and determine whether the execution time difference exceeds the preset threshold; when the execution time difference exceeds the preset threshold, adjust the input data volume threshold of the target processing unit with the short execution time; when the execution time difference does not exceed the preset threshold, bind at least two target processing units, and reconfigure the processing thread and input data volume threshold to determine the execution time corresponding to the processing unit.

[0043] When the model operation analysis result is the second result, at least two target processing units with the shortest execution time can be screened out, and compared two by two to determine the execution time difference, and then the execution time difference is compared with the preset threshold, and when the execution time difference exceeds the preset threshold, the input data volume threshold of the target processing unit with the short execution time is adjusted. For example, the input data volume threshold of the target processing unit with the short execution time can be shortened so that the target processing unit with the short execution time can be executed earlier. When the execution time difference does not exceed the preset threshold, the two target processing units can be bound so that the two target processing units are processed by the same processing thread. After the two target processing units are bound, the processing thread and the input data volume threshold can be reconfigured to perform the next iterative analysis until the optimization strategy corresponding to the machine learning model is determined when the model operation analysis result is the first result.

[0044] In the actual application process, multiple machine learning models are usually connected in series and parallel to form a large machine learning model (including multiple computing nodes). Therefore, in an optional embodiment, the embodiment of the present application can determine the optimization strategy of a machine learning model composed of multiple sub-machine learning models. In addition, in another optional embodiment, the embodiment of the present application can also divide the machine learning model including multiple computing nodes, and analyze the divided parts as a separate machine learning model to determine the corresponding optimization strategy. Specifically, the method also includes: obtaining a machine learning model, dividing the machine model, obtaining a sub-machine learning model, and analyzing the processing unit of the sub-machine learning model to determine the corresponding optimization strategy. In this embodiment, the machine learning model can be divided into multiple sub-machine learning models, and the sub-machine learning model is analyzed to obtain the optimization strategy corresponding to each sub-machine learning model, so that according to different optimization strategies, the corresponding hardware is configured for each divided part to perform data processing corresponding to each divided part. The segmentation method of the machine learning model can be configured according to demand. For example, according to the process of data processing, it can be divided into multiple different data processing stages, so as to segment the machine learning model according to the data processing stage.

[0045] In addition, in an optional embodiment, in the embodiment of the present application, after determining the optimization strategy corresponding to the machine learning model, the type of the machine learning model and the corresponding optimization strategy can be recorded to form a corresponding relationship. Specifically, the method also includes: obtaining the model type of the machine learning model, and determining the optimization strategy corresponding to the machine learning model according to the preset corresponding relationship. In this embodiment, the model type of the machine learning model can be determined based on the number of processing units of the machine learning model and the way the processing unit processes data. This embodiment can continuously form new corresponding relationships by continuously recording various types of machine learning models and optimization strategies, so that after receiving the machine learning model in the future, the corresponding optimization strategy can be determined according to the type of the machine learning model, and the process of complex analysis of the machine learning model is omitted, which can improve the analysis efficiency of the optimization strategy.

[0046] After determining the optimization strategy, the server can provide hardware configuration reference information according to the optimization strategy in step 204. The server can send the configuration reference information to the terminal, and the terminal can display multiple hardware configuration reference information in the interactive page. Specifically, as an optional embodiment, the provision of hardware configuration reference information according to the optimization strategy includes: determining multiple hardware configuration reference information according to the optimization strategy; sending multiple hardware configuration reference information to be displayed in the interactive page. When setting the preset operating conditions, multiple preset operating conditions can be set to determine multiple optimization methods as optimization strategies, and provide multiple hardware configuration reference information to the model provider for selection according to the optimization strategy. After the user selects the hardware configuration reference information in the interactive page of the terminal, the selection information can be uploaded to the server. In step 206, the server receives the selection information of the hardware configuration reference information and determines the corresponding hardware configuration information. Among them, the hardware configuration information may include device information of an accelerator device for applying a machine learning model. In the embodiment of the present application, the machine learning model can be trained by an accelerator device, and the data processing of the trained machine learning model can also be performed by the accelerator device, which can be set according to the needs.

[0047] In an embodiment of the present application, the processing unit of the machine learning model can be analyzed to determine the execution time required for each processing unit to process data, and the corresponding optimization strategy can be determined. Then, based on the optimization strategy, hardware configuration reference information is provided to the model provider, and the model provider can make a choice based on the hardware configuration reference information to determine the corresponding hardware configuration information. In an embodiment of the present application, the processing unit of the machine learning model can be analyzed to determine multiple hardware configurations that match the applied machine learning model, and the model provider's choice can be obtained to obtain a hardware configuration suitable for the machine learning model, thereby completing the data processing process corresponding to the machine learning model through the corresponding hardware device according to the optimization strategy, thereby improving data processing efficiency.

[0048] The present application also provides a data processing method, which can be applied on the server side, such as Figure 3 As shown, the method includes:

[0049] Step 302: Provide an interactive page, which includes multiple model upload interfaces corresponding to different programming methods, obtain upload data according to the model upload interface, and compile and process the upload data to obtain a machine learning model.

[0050] Step 304: allocate a processing thread to each processing unit of the machine learning model; and obtain the input data volume upper limit value corresponding to the processing thread allocated to the processing unit, and configure the input data volume threshold value of the processing thread according to the input data volume upper limit value.

[0051] Step 306: Determine the execution time required for the processing unit based on the input data volume threshold and processing thread corresponding to the processing unit; and determine whether the machine learning model meets the preset operating conditions based on the execution time required for the processing unit to obtain the model operation analysis results.

[0052] Step 308: When the model operation analysis result is the first result, an optimization strategy is determined according to the processing thread of the processing unit and the input data volume threshold, wherein the optimization strategy includes the processing thread and the input data volume threshold corresponding to the processing unit.

[0053] Step 310: When the model operation analysis result is the second result, select at least two target processing units with the shortest execution time; and obtain the execution time difference between the target processing units to determine whether the execution time difference exceeds a preset threshold.

[0054] Step 312: When the execution time difference exceeds a preset threshold, adjust the input data volume threshold of the target processing unit with a shorter execution time.

[0055] Step 314: When the execution time difference does not exceed a preset threshold, at least two target processing units are bound, and the processing thread and input data volume threshold are reconfigured to determine the execution time corresponding to the processing unit.

[0056] Step 316: Determine multiple hardware configuration reference information based on the optimization strategy, and send multiple hardware configuration reference information to be displayed in the interactive page.

[0057] Step 318: Receive selection information for hardware configuration reference information and determine corresponding hardware configuration information.

[0058] In an embodiment of the present application, an interactive page can be provided to the terminal of the model provider, and the model provider uploads the upload data of the machine learning model through the model upload interface in the interactive page, and the server determines the machine learning model based on the uploaded data. Then, the processing thread and the input data volume threshold can be configured for the processing unit of the machine learning model to obtain the execution time required for each processing unit to process the data, and determine the corresponding optimization strategy, and then provide the model provider with hardware configuration reference information based on the optimization strategy. The model provider can make a choice based on the hardware configuration reference information to determine the hardware configuration information.

[0059] On the basis of the above embodiments, the embodiments of the present application provide a data processing method that can be applied to a terminal. The terminal can be understood as a device that interacts with a server to upload a machine learning model to the server. The model provider can upload the relevant data of the machine learning model to the server through the terminal. The server can compile the received relevant data to obtain the machine learning model and determine the corresponding optimization strategy. Then, according to the optimization strategy, the hardware configuration reference information is provided to the terminal so that the model provider can select hardware suitable for the machine learning model. Specifically, Figure 4 As shown, the method includes:

[0060] Step 402: provide an interactive page, wherein the interactive page includes a plurality of model upload interfaces corresponding to different programming methods.

[0061] Step 404: Receive a trigger on the model upload interface to obtain corresponding upload data, and upload it to determine the machine learning model.

[0062] Step 406: Receive hardware configuration reference information corresponding to the machine learning model and display it on the interactive page.

[0063] Step 408: Receive a selection operation on the hardware configuration reference information to determine the selected information and upload it to determine the hardware configuration information.

[0064] The implementation of this embodiment is similar to that of the above embodiment. The specific implementation may refer to the specific implementation of the above embodiment, which will not be described again here.

[0065] In an embodiment of the present application, the terminal can receive an interactive page provided by the server, and the interactive page includes multiple model upload interfaces. The model provider can upload the upload data of the machine learning model through the model upload interface to determine the machine learning model based on the uploaded data. Then, the server can analyze the processing unit of the machine learning model, determine the execution time required for each processing unit to process the data, and determine the corresponding optimization strategy, and then provide hardware configuration reference information to the terminal based on the optimization strategy. The model provider can make a choice based on the hardware configuration reference information to determine the corresponding hardware configuration information.

[0066] On the basis of the above embodiments, the embodiments of the present application also provide a data processing method, which can be applied to a server. The server can be understood as a device that performs machine learning model training and / or applies the trained machine learning model to perform data processing. The server can interact with the terminal of the model provider. The server can allocate the processing units of the machine learning model to the corresponding processing threads according to the optimization strategy to perform data processing more efficiently. Specifically, Figure 5 As shown, the method includes:

[0067] Step 502, determine the machine learning model and hardware configuration information, the machine learning model is composed of processing units, the hardware configuration information may include device information of the accelerator device, and the accelerator device includes at least one processor. The server can obtain at least one machine learning model and combine it into a calculation graph to provide services. Before the machine learning model performs data processing (such as the training process of model training, the application process of the trained model), the server can analyze (the execution time of) each processing unit in the machine learning model to obtain an optimization strategy, and provide hardware configuration reference information to the model provider based on the optimization strategy. The model provider can select the corresponding hardware configuration reference information to determine the accelerator device corresponding to the calculation graph, and determine the corresponding optimization strategy, so as to determine the processing thread corresponding to each processing unit of the machine learning model in step 504 according to the optimization strategy. Among them, the optimization strategy includes the corresponding relationship between the processing unit and the processing thread of the processor, and the optimization strategy is determined based on the execution time required for the processing unit to perform data processing alone in the processing thread. After determining the processing threads corresponding to each processing unit of the machine learning model, the server can schedule the processing unit to the scheduling queue of the processing thread in step 506, so as to complete the calculation corresponding to the processing unit through the processing thread and obtain the data processing result.

[0068] The server can obtain each processing unit in the machine learning model and add it to the inference request queue, and then allocate the processing unit in the inference request queue to the scheduling queue of the processing thread of the processor through the corresponding inference model. Specifically, as an optional embodiment, the scheduling of the processing unit to the scheduling queue of the processing thread includes: scheduling the processing unit to the inference request queue; reading the processing unit from the inference request queue through the inference model, and allocating the processing unit to the processing unit queue according to the processing unit type; mapping the processing unit in the processing unit queue to the scheduling queue of the processing thread. Among them, as an optional embodiment, the processing unit can be divided into three types according to the occupancy of the processing unit to the processing thread, including: computing intensive, memory access intensive and data transmission type. Accordingly, the processing unit queue includes a computing intensive queue, a memory access intensive queue and a data transmission queue. Accordingly, the scheduling queue corresponds to the processing unit queue, and the scheduling queue includes: computing intensive queue, memory access intensive queue and data transmission queue. This embodiment can allocate processing units to corresponding processing unit queues according to the types of processing units, and then map them to the scheduling queue, so that the inference model can allocate processing units of different machine learning models, further improving the utilization of hardware on the basis of model scheduling.

[0069] Among them, in the process of mapping the processing units from the processing unit queue to the scheduling queue, the processing units in the processing unit queue can be mapped to the scheduling queue according to the order in which the processing units are executed. Specifically, as an optional embodiment, the mapping of the processing units in the processing unit queue to the scheduling queue of the processing thread includes: configuring the processing unit marking information for the processing units in the processing unit queue, the processing unit marking information includes the machine learning model corresponding to the processing unit and the dependency relationship between the processing unit and other processing units; determining the scheduling order of the processing units based on the processing unit marking information, and mapping the processing units in the processing unit queue to the scheduling sequence based on the scheduling order. After the processing units are assigned to the processing unit queue according to the type, the server can mark the machine learning model to which the processing units in the processing unit queue belong and the dependency relationship between the processing units, and then determine the scheduling order between the processing units based on the dependency relationship between the processing units, and map the processing units in the processing unit queue to the scheduling queue according to the scheduling order. The processing thread of the processor can sequentially call the processing units from the scheduling queue and perform corresponding data processing.

[0070] In addition to configuring a processing thread for a processing unit, the embodiment of the present application can also configure an input data volume threshold of the processing thread to control the input data volume of the processing thread so as to perform data processing more efficiently. Specifically, as an optional embodiment, the optimization strategy also includes an input data volume threshold, and the calculation corresponding to the processing unit is completed by the processing thread, including: obtaining the input data corresponding to the processing unit according to the input data volume threshold; performing data processing corresponding to the processing unit on the input data to obtain the data processing result. In an optional embodiment, a fixed input data volume threshold can be configured for the processing thread, so that the processing thread can perform data processing corresponding to the processing unit when the input data volume reaches the input data volume threshold. In another optional embodiment, the execution of the processing unit executed later needs to depend on the data processing result of the processing unit executed earlier. Therefore, if the input data volume threshold remains unchanged, it may cause the processing unit executed later to need a longer time to wait for the input data volume to reach the input data volume threshold, which will cause the corresponding processing thread to idle for too long. Therefore, the method may also include: adjusting the input data volume threshold according to the utilization rate of the processing unit to the processing thread. The server can also detect the utilization of the processing thread corresponding to the processing unit, and dynamically adjust the input data volume threshold according to the utilization of the processing thread to further improve data processing efficiency.

[0071] In order to synchronize data to the processing thread more efficiently, the embodiment of the present application can also adopt an asynchronous parallel method to synchronize data. Specifically, as an optional embodiment, the acquisition of the input data corresponding to the processing unit includes: calling the input data into the first cache; adopting an asynchronous parallel method to synchronize the input data to the second cache of the processing thread, so that the processing thread obtains the input data from the second cache. The input data can be data stored in a database or data obtained from other devices. The first cache and the second cache can be understood as caches, which refer to a high-speed memory with an access speed faster than a general random access memory (RAM). It exchanges data with the CPU before the memory, so the rate is very fast. The first cache can also be called device memory, the first cache is the cache corresponding to the input data, the second cache can also be called host memory, and the second cache is the cache corresponding to the processing thread. The embodiment of the present application can schedule the input data to the first cache, and adopt an asynchronous parallel method to synchronize the input data to the second cache, which can improve the efficiency of data transmission.

[0072] In an embodiment of the present application, the machine learning model and hardware configuration information can be obtained, and according to the optimization strategy of the machine learning model, the processing unit of the machine learning model can be scheduled to the scheduling queue of the processing thread, so as to complete the calculation corresponding to the processing unit through the processing thread. In an embodiment of the present application, each processing unit in the machine learning model can be analyzed in advance to determine the optimization strategy corresponding to the machine learning model. In the process of data processing by the machine learning model, the processing thread can be allocated to the processing unit of the machine learning model according to the optimization strategy, which can improve the data processing efficiency.

[0073] Based on the above embodiments, the present application embodiment provides a data processing method that can be applied to the server. Specifically, Figure 6 As shown, the method includes:

[0074] Step 602: Determine the machine learning model and hardware configuration information.

[0075] Step 604: Determine the processing threads corresponding to each processing unit of the machine learning model based on the optimization strategy and hardware configuration information. The optimization strategy includes the correspondence between the processing unit and the processing thread of the processor. The optimization strategy is determined based on the execution time required for the processing unit to perform data processing in the processing thread alone.

[0076] Step 606: Schedule the processing unit to the inference request queue.

[0077] Step 608: Read the processing units from the inference request queue through the inference model, and assign the processing units to the processing unit queues according to the processing unit types. As an optional embodiment, the processing unit queues include a computation-intensive queue, a memory-intensive queue, and a data transmission queue.

[0078] Step 610: Configure processing unit marking information for the processing units in the processing unit queue, where the processing unit marking information includes a machine learning model corresponding to the processing unit and a dependency relationship between the processing unit and other processing units.

[0079] Step 612: Determine the scheduling order of the processing units according to the processing unit tag information, and map the processing units in the processing unit queue to the scheduling sequence according to the scheduling order.

[0080] Step 614, based on the input data volume threshold, obtain the input data corresponding to the processing unit. As an optional embodiment, the input data volume threshold can be dynamically adjusted based on the utilization of the processing thread. Specifically, the method further includes: adjusting the input data volume threshold based on the utilization of the processing unit to the processing thread. This embodiment can also adopt a parallel synchronization method to synchronize input data. Specifically, as an optional embodiment, the acquisition of the input data corresponding to the processing unit includes: retrieving the input data into the first cache; adopting an asynchronous parallel method to synchronize the input data to the second cache of the processing thread so that the processing thread obtains the input data from the second cache.

[0081] Step 616: Perform data processing corresponding to the processing unit on the input data to obtain a data processing result.

[0082] In an embodiment of the present application, a machine learning model and corresponding hardware configuration information can be obtained, and the processing thread corresponding to the processing unit of the machine learning model can be determined based on the optimization strategy of the machine learning model and the hardware configuration information. After determining the processing thread corresponding to the processing unit, the processing unit can be scheduled to the reasoning request queue, and then the processing unit can be read from the reasoning request queue through the reasoning model, and the processing unit can be assigned to the corresponding processing unit queue according to the type of processing unit to which the processing unit belongs, and then the processing unit in the processing unit queue can be configured with processing unit tag information to determine the scheduling order between the processing units. After that, according to the scheduling order, the processing units in the processing unit queue are mapped to the scheduling queue so that the processing thread obtains the processing unit from the scheduling queue, executes the data processing corresponding to the processing unit, and determines the corresponding data processing result.

[0083] Based on the above embodiments, the present application provides a data processing method, which can be applied on the server side and in the live broadcast scene, and can calculate and process the live broadcast data according to the machine learning model to obtain the corresponding processing results. Specifically, the method includes:

[0084] Obtain live data and determine the corresponding machine learning model and hardware configuration information.

[0085] According to the optimization strategy and hardware configuration information, the processing thread corresponding to each processing unit of the machine learning model is determined, and the processing unit is used to process the live broadcast data.

[0086] The processing unit is dispatched to the dispatch queue of the processing thread, so that the calculation corresponding to the processing unit can be completed through the processing thread, and the analysis and processing of the live broadcast data can be completed to obtain the data processing result.

[0087] In the embodiment of the present application, the machine learning model can analyze and process the live broadcast data, for example, face recognition, object recognition, target tracking and other processing can be performed on the targets in the live broadcast data. In this embodiment, the live broadcast data, the machine learning model and the hardware configuration information can be obtained, and according to the optimization strategy of the machine learning model, the processing unit of the machine learning model can be scheduled to the scheduling queue of the processing thread, so as to complete the calculation corresponding to the processing unit through the processing thread, thereby completing the analysis and processing of the live broadcast data and obtaining the data analysis result. After determining the data analysis result, the target can be processed according to the data analysis result, such as adding virtual information (such as sunglasses special effects) to the face to achieve the effect of combining virtual and reality, such as beautifying and whitening the face of the anchor. In the embodiment of the present application, each processing unit in the machine learning model can be analyzed in advance to determine the optimization strategy corresponding to the machine learning model. In the process of data processing by the machine learning model, the processing thread can be allocated to the processing unit of the machine learning model according to the optimization strategy, which can improve the data processing efficiency.

[0088] Based on the above embodiments, the embodiments of the present application provide a data processing method, which can be applied on the server side and in the autonomous driving scenario, and can calculate and process the driving data according to the machine learning model, so as to obtain the corresponding processing results to control the vehicle driving. Specifically, the method includes:

[0089] Obtain driving data and determine the corresponding machine learning model and hardware configuration information.

[0090] According to the optimization strategy and hardware configuration information, the processing thread corresponding to each processing unit of the machine learning model is determined, and the processing unit is used to process the driving data.

[0091] The processing unit is dispatched to the dispatch queue of the processing thread, so that the calculation corresponding to the processing unit is completed through the processing thread, the driving video data is analyzed and processed, and the data processing result is obtained.

[0092] Based on the data processing results, driving instructions are determined to control the vehicle.

[0093] In an embodiment of the present application, the driving data may include driving videos. The present embodiment may obtain driving data, machine learning models, and hardware configuration information, and according to the optimization strategy of the machine learning model, schedule the processing unit of the machine learning model to the scheduling queue of the processing thread, so as to complete the calculation corresponding to the processing unit through the processing thread, thereby completing the analysis and processing of the driving data and obtaining the data processing results, so as to determine the driving instructions (such as acceleration, deceleration, turning, etc.) according to the data processing results, and then control the vehicle according to the driving instructions. In an embodiment of the present application, each processing unit in the machine learning model may be analyzed in advance to determine the optimization strategy corresponding to the machine learning model. In the process of data processing by the machine learning model, the processing thread may be allocated to the processing unit of the machine learning model according to the optimization strategy, which can improve the data processing efficiency.

[0094] It should be noted that, for the method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the embodiments of the present application are not limited by the described order of actions, because according to the embodiments of the present application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily required by the embodiments of the present application.

[0095] Based on the above embodiment, this embodiment further provides a data processing device, referring to Figure 7 , specifically it can include the following modules:

[0096] The processing model acquisition module 702 is used to determine the machine learning model and hardware configuration information.

[0097] The processing thread acquisition module 704 is used to determine the processing threads corresponding to each processing unit of the machine learning model based on the optimization strategy and hardware configuration information.

[0098] The processing unit scheduling module 706 is used to schedule the processing unit to the scheduling queue of the processing thread, so as to complete the calculation corresponding to the processing unit through the processing thread and obtain the data processing result.

[0099] In summary, in an embodiment of the present application, the machine learning model and hardware configuration information can be obtained, and according to the optimization strategy of the machine learning model, the processing unit of the machine learning model can be scheduled to the scheduling queue of the processing thread, so as to complete the calculation corresponding to the processing unit through the processing thread. In an embodiment of the present application, each processing unit in the machine learning model can be analyzed in advance to determine the optimization strategy corresponding to the machine learning model. In the process of data processing by the machine learning model, the processing thread can be allocated to the processing unit of the machine learning model according to the optimization strategy, which can improve the data processing efficiency.

[0100] On the basis of the above embodiment, this embodiment further provides a data processing device, which may specifically include the following modules:

[0101] The processing model acquisition processing module is used to determine the machine learning model and hardware configuration information.

[0102] The processing thread acquisition processing module is used to determine the processing threads corresponding to each processing unit of the machine learning model based on the optimization strategy and hardware configuration information. The optimization strategy includes the correspondence between the processing unit and the processing thread of the processor. The optimization strategy is determined based on the execution time required for the processing unit to perform data processing in the processing thread alone.

[0103] The processing unit scheduling processing module is used to schedule the processing unit to the inference request queue.

[0104] The processing unit allocation processing module is used to read the processing unit from the inference request queue through the inference model, and allocate the processing unit to the processing unit queue according to the processing unit type. As an optional embodiment, the processing unit queue includes a computing intensive queue, a memory access intensive queue and a data transmission queue.

[0105] The processing unit marking processing module is used to configure processing unit marking information for the processing units in the processing unit queue, and the processing unit marking information includes the machine learning model corresponding to the processing unit and the dependency relationship between the processing unit and other processing units.

[0106] The processing unit mapping processing module is used to determine the scheduling order of the processing units according to the processing unit tag information, and map the processing units in the processing unit queue to the scheduling sequence according to the scheduling order.

[0107] An input data acquisition processing module is used to acquire the input data corresponding to the processing unit according to an input data volume threshold. As an optional embodiment, the input data volume threshold can be dynamically adjusted according to the utilization rate of the processing thread. Specifically, the device further includes: an input data volume threshold adjustment processing module, which is used to adjust the input data volume threshold according to the utilization rate of the processing unit for the processing thread. This embodiment can also adopt a parallel synchronization method to synchronize input data. Specifically, as an optional embodiment, the input data volume threshold adjustment processing module specifically includes: retrieving input data into a first cache; adopting an asynchronous parallel method to synchronize input data to a second cache of the processing thread so that the processing thread obtains input data from the second cache.

[0108] The data processing result acquisition processing module is used to perform data processing corresponding to the processing unit on the input data to obtain the data processing result.

[0109] In an embodiment of the present application, a machine learning model and corresponding hardware configuration information can be obtained, and the processing thread corresponding to the processing unit of the machine learning model can be determined based on the optimization strategy of the machine learning model and the hardware configuration information. After determining the processing thread corresponding to the processing unit, the processing unit can be scheduled to the reasoning request queue, and then the processing unit can be read from the reasoning request queue through the reasoning model, and the processing unit can be assigned to the corresponding processing unit queue according to the type of processing unit to which the processing unit belongs, and then the processing unit in the processing unit queue can be configured with processing unit tag information to determine the scheduling order between the processing units. After that, according to the scheduling order, the processing units in the processing unit queue are mapped to the scheduling queue so that the processing thread obtains the processing unit from the scheduling queue, executes the data processing corresponding to the processing unit, and determines the corresponding data processing result.

[0110] Based on the above embodiment, this embodiment further provides a data processing device, referring to Figure 8 , specifically it can include the following modules:

[0111] The optimization strategy determination module 802 is used to analyze the processing unit of the machine learning model and determine the corresponding optimization strategy.

[0112] The reference information determination module 804 is used to provide hardware configuration reference information according to the optimization strategy.

[0113] The selection information determination module 806 is used to receive selection information for hardware configuration reference information and determine corresponding hardware configuration information.

[0114] In summary, in an embodiment of the present application, the processing unit of the machine learning model can be analyzed to determine the execution time required for each processing unit to process data, and the corresponding optimization strategy can be determined. Then, based on the optimization strategy, the hardware configuration reference information is provided to the model provider, and the model provider can make a choice based on the hardware configuration reference information to determine the corresponding hardware configuration information. In an embodiment of the present application, the processing unit of the machine learning model can be analyzed to determine multiple hardware configurations that match the applied machine learning model, and the model provider's choice can be obtained to obtain a hardware configuration suitable for the machine learning model, thereby completing the data processing process corresponding to the machine learning model through the corresponding hardware device according to the optimization strategy, thereby improving data processing efficiency.

[0115] On the basis of the above embodiment, this embodiment further provides a data processing device, which may specifically include the following modules:

[0116] The interactive page sending processing module is used to provide an interactive page, which includes multiple model uploading interfaces corresponding to different programming methods. The uploaded data is obtained according to the model uploading interface, and the uploaded data is compiled and processed to obtain a machine learning model.

[0117] The processing unit configures a processing module, which is used to assign a processing thread to each processing unit of the machine learning model; and obtains the input data volume upper limit value corresponding to the processing thread assigned to the processing unit, and configures the input data volume threshold of the processing thread according to the input data volume upper limit value.

[0118] The execution time acquisition processing module is used to determine the execution time required for the processing unit based on the input data volume threshold and processing thread corresponding to the processing unit; and based on the execution time required by the processing unit, determine whether the machine learning model meets the preset operating conditions and obtain the model operation analysis results.

[0119] The optimization strategy acquisition processing module is used to determine the optimization strategy based on the processing thread and input data volume threshold of the processing unit when the model operation analysis result is the first result. The optimization strategy includes the processing thread and input data volume threshold corresponding to the processing unit.

[0120] The execution time difference judgment processing module is used to screen out at least two target processing units with the shortest execution time when the model operation analysis result is the second result; and obtain the execution time difference between the target processing units to determine whether the execution time difference exceeds a preset threshold.

[0121] The first adjustment processing module is used to adjust the input data volume threshold of the target processing unit with a shorter execution time when the execution time difference exceeds a preset threshold.

[0122] The second adjustment processing module is used to bind at least two target processing units and reconfigure the processing thread and the input data volume threshold to determine the execution time corresponding to the processing unit when the execution time difference does not exceed the preset threshold.

[0123] The reference information acquisition processing module is used to determine multiple hardware configuration reference information according to the optimization strategy, and send multiple hardware configuration reference information to be displayed in the interactive page.

[0124] The selection information receiving and processing module is used to receive selection information of hardware configuration reference information and determine corresponding hardware configuration information.

[0125] In an embodiment of the present application, an interactive page can be provided to the terminal of the model provider, and the model provider uploads the upload data of the machine learning model through the model upload interface in the interactive page, and the server determines the machine learning model based on the uploaded data. Then, the processing thread and the input data volume threshold can be configured for the processing unit of the machine learning model to obtain the execution time required for each processing unit to process the data, and determine the corresponding optimization strategy, and then provide the model provider with hardware configuration reference information based on the optimization strategy. The model provider can make a choice based on the hardware configuration reference information to determine the hardware configuration information.

[0126] Based on the above embodiment, this embodiment further provides a data processing device, referring to Fig. 9 , specifically it can include the following modules:

[0127] The interactive page display module 902 is used to provide an interactive page, and the interactive page includes multiple model upload interfaces corresponding to different programming methods.

[0128] The model data upload module 904 is used to receive a trigger on the model upload interface to obtain the corresponding upload data and upload it to determine the machine learning model.

[0129] The reference information display module 906 is used to receive the hardware configuration reference information corresponding to the machine learning model and display it on the interactive page.

[0130] The selection information uploading module 908 is used to receive a selection operation on the hardware configuration reference information to determine the selection information and upload it to determine the hardware configuration information.

[0131] In summary, in an embodiment of the present application, the terminal can receive an interactive page provided by the server, and the interactive page includes multiple model upload interfaces. The model provider can upload the upload data of the machine learning model through the model upload interface to determine the machine learning model based on the uploaded data. Then, the server can analyze the processing unit of the machine learning model, determine the execution time required for each processing unit to process the data, and determine the corresponding optimization strategy, and then provide hardware configuration reference information to the terminal based on the optimization strategy. The model provider can make a choice based on the hardware configuration reference information to determine the corresponding hardware configuration information.

[0132] The embodiment of the present application also provides a non-volatile readable storage medium, which stores one or more modules (programs). When the one or more modules are applied to a device, the device can execute instructions (instructions) of each method step in the embodiment of the present application.

[0133] The present application embodiment provides one or more machine-readable media, on which instructions are stored, and when executed by one or more processors, an electronic device executes one or more methods as described in the above embodiments. In the present application embodiment, the electronic device includes a server, a terminal device, and the like.

[0134] The embodiments of the present disclosure may be implemented as a device configured as desired using any appropriate hardware, firmware, software, or any combination thereof, and the device may include electronic devices such as a server (cluster), a terminal, etc. Fig.10 An exemplary apparatus 1000 that can be used to implement various embodiments described in this application is schematically illustrated.

[0135] For one embodiment, Fig.10 An exemplary apparatus 1000 is shown having one or more processors 1002, a control module (chip set) 1004 coupled to at least one of the (one or more) processors 1002, a memory 1006 coupled to the control module 1004, a non-volatile memory (NVM) / storage device 1008 coupled to the control module 1004, one or more input / output devices 1010 coupled to the control module 1004, and a network interface 1012 coupled to the control module 1004.

[0136] The processor 1002 may include one or more single-core or multi-core processors, and the processor 1002 may include any combination of general-purpose processors or special-purpose processors (such as graphics processors, application processors, baseband processors, etc.). In some embodiments, the device 1000 can be used as a server, terminal, or other device described in the embodiments of the present application.

[0137] In some embodiments, the apparatus 1000 may include one or more computer-readable media (e.g., memory 1006 or NVM / storage device 1008) having instructions 1014 and one or more processors 1002 combined with the one or more computer-readable media and configured to execute the instructions 1014 to implement a module to perform the actions described in the present disclosure.

[0138] For one embodiment, the control module 1004 may include any suitable interface controller to provide any suitable interface to at least one of the processor(s) 1002 and / or any suitable device or component in communication with the control module 1004 .

[0139] The control module 1004 may include a memory controller module to provide an interface to the memory 1006. The memory controller module may be a hardware module, a software module, and / or a firmware module.

[0140] The memory 1006 may be used, for example, to load and store data and / or instructions 1014 for the device 1000. For one embodiment, the memory 1006 may include any suitable volatile memory, such as a suitable DRAM. In some embodiments, the memory 1006 may include double data rate type four synchronous dynamic random access memory (DDR4 SDRAM).

[0141] For one embodiment, control module 1004 may include one or more input / output controllers to provide an interface to NVM / storage device 1008 and input / output device(s) 1010 .

[0142] For example, NVM / storage 1008 may be used to store data and / or instructions 1014. NVM / storage 1008 may include any suitable non-volatile memory (e.g., flash memory) and / or may include any suitable non-volatile storage device(s) (e.g., one or more hard disk drives (HDDs), one or more compact disk (CD) drives, and / or one or more digital versatile disk (DVD) drives).

[0143] NVM / storage device 1008 may include storage resources that are part of the device on which apparatus 1000 is installed, or it may be accessible to the device without being part of the device. For example, NVM / storage device 1008 may be accessed via (one or more) input / output devices 1010 over a network.

[0144] (One or more) input / output devices 1010 may provide an interface for the apparatus 1000 to communicate with any other appropriate device, and the input / output device 1010 may include a communication component, an audio component, a sensor component, etc. The network interface 1012 may provide an interface for the apparatus 1000 to communicate through one or more networks, and the apparatus 1000 may wirelessly communicate with one or more components of a wireless network according to any of one or more wireless network standards and / or protocols, for example, accessing a wireless network based on a communication standard, such as WiFi, 2G, 3G, 4G, 5G, etc., or a combination thereof for wireless communication.

[0145] For one embodiment, at least one of the processor(s) 1002 may be packaged together with the logic of one or more controllers (e.g., a memory controller module) of the control module 1004. For one embodiment, at least one of the processor(s) 1002 may be packaged together with the logic of one or more controllers of the control module 1004 to form a system-in-package (SiP). For one embodiment, at least one of the processor(s) 1002 may be integrated on the same die with the logic of one or more controllers of the control module 1004. For one embodiment, at least one of the processor(s) 1002 may be integrated on the same die with the logic of one or more controllers of the control module 1004 to form a system-on-chip (SoC).

[0146] In various embodiments, the apparatus 1000 may be, but is not limited to, a terminal device such as a server, a desktop computing device, or a mobile computing device (e.g., a laptop computing device, a handheld computing device, a tablet computer, a netbook, etc.). In various embodiments, the apparatus 1000 may have more or fewer components and / or a different architecture. For example, in some embodiments, the apparatus 1000 includes one or more cameras, a keyboard, a liquid crystal display (LCD) screen (including a touch screen display), a non-volatile memory port, multiple antennas, a graphics chip, an application-specific integrated circuit (ASIC), and a speaker.

[0147] Among them, the main control chip can be used as a processor or control module in the detection device, sensor data, location information, etc. are stored in a memory or NVM / storage device, the sensor group can be used as an input / output device, and the communication interface may include a network interface.

[0148] An embodiment of the present application further provides an electronic device, comprising: a processor; and a memory, on which executable code is stored, and when the executable code is executed, the processor executes one or more methods described in the embodiments of the present application.

[0149] The embodiments of the present application also provide one or more machine-readable media on which executable codes are stored. When the executable codes are executed, the processor executes one or more methods described in the embodiments of the present application.

[0150] As for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0151] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.

[0152] The present application embodiment is described with reference to the flowchart and / or block diagram of the method, terminal device (system) and computer program product according to the embodiment of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, and the combination of the process and / or box in the flowchart and / or block diagram can be realized by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing terminal device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal device produce a device for realizing the function specified in one process or multiple processes in the flowchart and / or one box or multiple boxes in the block diagram.

[0153] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing terminal device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured product including an instruction device that implements the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.

[0154] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device so that a series of operating steps are executed on the computer or other programmable terminal device to produce computer-implemented processing, so that the instructions executed on the computer or other programmable terminal device provide steps for implementing the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.

[0155] Although the preferred embodiments of the present application have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the embodiments of the present application.

[0156] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or terminal device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or terminal device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or terminal device including the elements.

[0157] The data processing method, the data processing device, the electronic device and the storage medium provided by the present application are introduced in detail above. The principles and implementation methods of the present application are explained in this article by using specific examples. The description of the above embodiments is only used to help understand the method of the present application and its core idea. At the same time, for those skilled in the art, according to the idea of ​​the present application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.

Claims

1. A data processing method, It is characterized in that The method includes: Identify the machine learning model; Allocate a processing thread to each processing unit of the machine learning model, and configure the input data volume threshold of the processing thread; Determine the execution time required for the processing unit according to the input data volume threshold and processing thread corresponding to the processing unit; Determine whether the machine learning model meets the preset operating conditions based on the execution time required by the processing unit, and obtain the model operation analysis results; When the model operation analysis result satisfies the preset operation conditions, an optimization strategy is determined according to the processing thread of the processing unit and the input data volume threshold; According to the optimization strategy, provide hardware configuration reference information; Receiving selection information of hardware configuration reference information and determining corresponding hardware configuration information; Determine the processing threads corresponding to each processing unit of the machine learning model based on the optimization strategy and hardware configuration information; The processing unit is dispatched to the dispatch queue of the processing thread so that the calculation corresponding to the processing unit is completed through the processing thread to obtain the data processing result.

2. The method according to claim 1, It is characterized in that The optimization strategy also includes an input data volume threshold, and the processing thread is used to complete the calculation corresponding to the processing unit, including: According to the input data amount threshold, obtaining the input data corresponding to the processing unit; Perform data processing corresponding to the processing unit on the input data to obtain a data processing result.

3. The method according to claim 2, It is characterized in that The obtaining of input data corresponding to the processing unit includes: Retrieving input data into the first cache; An asynchronous and parallel approach is adopted to synchronize the input data to the second cache of the processing thread, so that the processing thread obtains the input data from the second cache.

4. The method according to claim 2, It is characterized in that Also includes: The input data volume threshold is adjusted according to the utilization rate of the processing unit on the processing thread.

5. The method according to claim 1, It is characterized in that The step of scheduling the processing unit into a scheduling queue of the processing thread includes: Scheduling processing units into inference request queues; Reading processing units from the inference request queue through the inference model, and assigning the processing units to the processing unit queue according to the processing unit type; Map the processing units in the processing unit queue to the scheduling queue of the processing thread.

6. The method according to claim 5, It is characterized in that Mapping the processing units in the processing unit queue to the scheduling queue of the processing thread includes: Configuring processing unit marking information for processing units in the processing unit queue; The scheduling order of the processing units is determined according to the processing unit tag information, and the processing units in the processing unit queue are mapped into the scheduling sequence according to the scheduling order.

7. The method according to claim 5, It is characterized in that The processing unit queues include a computation-intensive queue, a memory-access-intensive queue, and a data transmission queue.

8. The method according to claim 1, It is characterized in that Also includes: When the model operation analysis result is a result that the preset operation condition is not met, at least two target processing units with the shortest execution time are selected; Obtaining the execution time difference between the target processing units, and determining whether the execution time difference exceeds a preset threshold; When the execution time difference exceeds a preset threshold, adjusting the input data volume threshold of the target processing unit with a shorter execution time; When the execution time difference does not exceed a preset threshold, at least two target processing units are bound, and the processing threads and input data volume threshold are reconfigured to determine the execution time corresponding to the processing unit.

9. The method according to claim 1, It is characterized in that The configuration of the input data volume threshold of the processing thread includes: Obtaining an upper limit value of the amount of input data corresponding to a processing thread allocated to a processing unit; According to the upper limit of the input data volume, an input data volume threshold of the processing thread is configured.

10. The method according to claim 1, It is characterized in that Also includes: Providing an interactive page, the interactive page including a plurality of model uploading interfaces corresponding to different programming methods; The uploaded data is obtained according to the model upload interface, and the uploaded data is compiled and processed to obtain a machine learning model.

11. The method according to claim 10, It is characterized in that The providing hardware configuration reference information according to the optimization strategy includes: Determining multiple hardware configuration reference information according to the optimization strategy; Send multiple hardware configuration reference information to be displayed on the interactive page.

12. A data processing method, It is characterized in that include: Providing an interactive page, the interactive page including a plurality of model uploading interfaces corresponding to different programming methods; Receive a trigger on the model upload interface to obtain corresponding upload data and upload it to determine the machine learning model; Receive hardware configuration reference information corresponding to the machine learning model and display it on an interactive page, wherein the hardware configuration reference information is determined by the following steps: assign a processing thread to each processing unit of the machine learning model, and configure the input data volume threshold of the processing thread; determine the execution time required for the processing unit according to the input data volume threshold and the processing thread corresponding to the processing unit; determine whether the machine learning model meets the preset operating conditions based on the execution time required for the processing unit, and obtain the model operation analysis result; when the model operation analysis result is a result that meets the preset operating conditions, determine the optimization strategy based on the processing thread and the input data volume threshold of the processing unit; provide hardware configuration reference information based on the optimization strategy; A selection operation on hardware configuration reference information is received to determine the selection information and upload it to determine the hardware configuration information.

13. A data processing device, It is characterized in that include: Processing model acquisition module, used to determine the machine learning model; An optimization strategy determination module is used to allocate a processing thread to each processing unit of the machine learning model and configure the input data volume threshold of the processing thread; according to the input data volume threshold and the processing thread corresponding to the processing unit, determine the execution time required for the processing unit; Determine whether the machine learning model meets the preset operating conditions based on the execution time required by the processing unit, and obtain the model operation analysis results; When the model operation analysis result satisfies the preset operation conditions, an optimization strategy is determined according to the processing thread of the processing unit and the input data volume threshold; A reference information determination module, used to provide hardware configuration reference information according to the optimization strategy; A selection information determination module, used to receive selection information of hardware configuration reference information and determine corresponding hardware configuration information; A processing thread acquisition module is used to determine the processing thread corresponding to each processing unit of the machine learning model according to the optimization strategy and hardware configuration information; The processing unit scheduling module is used to schedule the processing unit to the scheduling queue of the processing thread, so as to complete the calculation corresponding to the processing unit through the processing thread and obtain the data processing result.

14. A data processing device, It is characterized in that include: An interactive page display module, used to provide an interactive page, wherein the interactive page includes a plurality of model upload interfaces corresponding to different programming methods; The model data upload module is used to receive a trigger on the model upload interface to obtain the corresponding upload data and upload it to determine the machine learning model; A reference information display module, used to receive hardware configuration reference information corresponding to the machine learning model and display it on an interactive page, wherein the hardware configuration reference information is determined in the following manner: a processing thread is assigned to each processing unit of the machine learning model, and an input data volume threshold of the processing thread is configured; based on the input data volume threshold and processing thread corresponding to the processing unit, the execution time required for the processing unit is determined; based on the execution time required for the processing unit, whether the machine learning model meets the preset operating conditions is determined, and the model operation analysis result is obtained; when the model operation analysis result is a result that meets the preset operating conditions, an optimization strategy is determined based on the processing thread and the input data volume threshold of the processing unit; based on the optimization strategy, hardware configuration reference information is provided; The selection information uploading module is used to receive the selection operation on the hardware configuration reference information to determine the selection information and upload it to determine the hardware configuration information.

15. An electronic device, It is characterized in that include: processor; and A memory having executable codes stored thereon, which, when executed, cause the processor to perform the method according to one or more of claims 1-12.

16. One or more machine-readable media having executable codes stored thereon, which, when executed, cause a processor to perform the method of one or more of claims 1-12.

Citation Information

Patent Citations

  • Configuration method and device of deep learning network model and storage medium

    CN111914985A

  • Cloud GPU Video memory scheduling method and device, electronic equipment and storage medium

    CN112052083A