Model online reasoning method, device and system, electronic equipment and storage medium
By aggregating the received multiple request data and performing online inference, the problem of low GPU utilization due to small amount of data per request is solved, and more efficient processor utilization and reduced computing time are achieved.
Patent Information
- Application Number
- CN202311558097.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-21
- Publication Date
- 2025-05-23
AI Technical Summary
Due to the small amount of data in a single requested data, the GPU utilization rate is lower when the model performs online inference of the requested data.
The at least two request data received are aggregated to obtain the target request data, and the target request data is inferred online, and the inference data is spliced and processed and sent in accordance with the order of execution of the inference steps.
By aggregating multiple requested data, the computing power of the processor can be fully utilized, the utilization rate of the processor can be improved, and the computing time of the online inference service can be reduced.
Smart Images

Figure CN120029746A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of data processing technology, and in particular to a method and device, system, electronic device and storage medium for online model reasoning. Background Art
[0002] With the explosive growth of data in various industries, the amount of data that needs to be processed has become huge, and more and more application scenarios require real-time and efficient model online reasoning capabilities.
[0003] In the related art, the online inference method of the model infers the request data one by one in a certain order. Since the amount of a single request data is small, the utilization rate of the graphics processing unit (GPU) is low when the model performs online inference on the request data. Summary of the invention
[0004] The present disclosure provides a method and device, system, electronic device and storage medium for online model reasoning. The main purpose is to solve the problem that the utilization rate of GPU is low when the model performs online reasoning on the requested data due to the small amount of single request data.
[0005] According to a first aspect of the present disclosure, a method for online model reasoning is provided, comprising:
[0006] Aggregate at least two received request data to obtain target request data;
[0007] Performing online reasoning on the target request data to obtain reasoning data corresponding to the at least two request data respectively;
[0008] The inference data are spliced according to the execution order of the inference steps, and the spliced inference data are sent, wherein the execution order is the time sequence of receiving the at least two request data.
[0009] Optionally, aggregating the received at least two request data to obtain target request data includes:
[0010] Obtaining the data sending time and data volume of at least two request data respectively;
[0011] In the preset dynamic library, according to the data sending time, the data volume and the maximum load of the processor, the at least two request data are aggregated in the order of the data sending time until the data volume of the aggregated request data reaches the maximum load of the processor, so as to obtain the target request data, wherein the preset dynamic library runs in the processor.
[0012] Optionally, the method for creating the preset dynamic library includes:
[0013] In response to a configuration instruction for a preset aggregation rule in the configuration file, configuring the configuration file;
[0014] After completing the configuration of the configuration file, the preset dynamic library is created based on the configuration file.
[0015] Optionally, performing online reasoning on the target request data to obtain the reasoning data corresponding to the at least two request data respectively includes:
[0016] The preset data transmission interface receives the target request data through the preset data transmission interface, and performs online reasoning on the target request data to obtain reasoning data corresponding to the at least two request data respectively, wherein the preset data transmission interface is an interface for receiving the target request data and outputting reasoning data.
[0017] Optionally, the splicing the inference data according to the execution order of the inference steps and sending the spliced inference data includes:
[0018] Outputting the current inference data based on the preset data transmission interface;
[0019] According to the reasoning steps, a preset number of previous reasoning data are obtained, where the previous reasoning data is the reasoning data output before the current reasoning data;
[0020] splicing the previous reasoning data and the current reasoning data to obtain spliced reasoning data;
[0021] The spliced inference data is sent.
[0022] Optionally, before outputting the current inference data based on the preset data transmission interface, the method further includes:
[0023] Determining whether the output order of the current inference data exceeds a preset order threshold;
[0024] If it is determined that the output order of the current reasoning data exceeds the preset order threshold, pausing the output of the current reasoning data and determining the request data for which the reasoning has not been completed;
[0025] The unfinished inference request data is re-aggregated with the newly received request data to obtain new target request data.
[0026] Optionally, outputting the current inference data based on the preset data transmission interface includes:
[0027] If it is determined that the output order of the current inference data does not exceed the preset order threshold, the current inference data is output based on the preset data transmission interface.
[0028] Optionally, the method further includes:
[0029] Receive parameter configuration instructions;
[0030] Adjust parameters according to the configuration parameters carried in the parameter configuration instruction.
[0031] According to a second aspect of the present disclosure, a device for online model reasoning is provided, comprising:
[0032] an aggregation unit, used for aggregating at least two received request data to obtain target request data;
[0033] An inference unit, configured to perform online inference on the target request data to obtain inference data corresponding to the at least two request data respectively;
[0034] A sending unit is used to splice the inference data according to the execution order of the inference steps and send the spliced inference data, wherein the execution order is the time sequence of receiving the at least two request data.
[0035] Optionally, the polymerization unit includes:
[0036] An acquisition module, used to respectively acquire the data sending time and data volume of at least two request data;
[0037] An aggregation module is used to aggregate the at least two request data in a preset dynamic library according to the data sending time, data size and the maximum load of the processor in the order of the data sending time, until the data size of the aggregated request data reaches the maximum load of the processor, so as to obtain the target request data, wherein the preset dynamic library runs in the processor.
[0038] Optionally, the aggregation module is further used for:
[0039] In response to a configuration instruction for a preset aggregation rule in the configuration file, configuring the configuration file;
[0040] After completing the configuration of the configuration file, the preset dynamic library is created based on the configuration file.
[0041] Optionally, the inference unit is also used to preset a data transmission interface to receive the target request data through a preset data transmission interface, and perform online inference on the target request data to obtain inference data corresponding to the at least two request data respectively, wherein the preset data transmission interface is an interface for receiving the target request data and outputting inference data.
[0042] Optionally, the sending unit includes:
[0043] An output module, used for outputting the current inference data based on the preset data transmission interface;
[0044] An acquisition module, used for acquiring a preset amount of previous reasoning data according to the reasoning steps, wherein the previous reasoning data is reasoning data output before the current reasoning data;
[0045] A splicing module, used for splicing the previous reasoning data and the current reasoning data to obtain spliced reasoning data;
[0046] A sending module is used to send the spliced inference data.
[0047] Optionally, the device further comprises:
[0048] A judging unit, used to judge whether the output sequence of the current inference data exceeds a preset sequence threshold;
[0049] A determination unit, configured to, when determining that the output order of the current reasoning data exceeds the preset order threshold, suspend the output of the current reasoning data and determine the request data for which the reasoning is not completed;
[0050] The aggregation unit is further used to re-aggregate the unfinished inference request data with the newly received request data to obtain new target request data.
[0051] Optionally, the output module is further used to output the current reasoning data based on the preset data transmission interface when it is determined that the output order of the current reasoning data does not exceed the preset order threshold.
[0052] Optionally, the device further comprises:
[0053] A receiving unit, used for receiving a parameter configuration instruction;
[0054] An input unit is used to adjust parameters according to the configuration parameters carried in the parameter configuration instruction.
[0055] According to a third aspect of the present disclosure, a system for online model reasoning is provided, comprising: a preset reasoning framework and an online reasoning engine; wherein:
[0056] The preset reasoning framework is used to aggregate at least two received request data to obtain target request data, and transmit the target request data to the online reasoning engine;
[0057] The online reasoning engine is used to receive the target request data transmitted by the preset reasoning framework, perform online reasoning on the target request data, obtain reasoning data corresponding to the at least two request data respectively, splice the reasoning data according to the execution order of the reasoning steps, and send the spliced reasoning data to the preset reasoning framework, wherein the execution order is the time sequence of receiving the at least two request data;
[0058] The preset reasoning framework is used to receive the spliced reasoning data sent by the online reasoning engine; and output the spliced reasoning data.
[0059] According to a fourth aspect of the present disclosure, there is provided an electronic device, including:
[0060] at least one processor; and
[0061] a memory communicatively connected to the at least one processor; wherein,
[0062] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method described in the first aspect.
[0063] According to a fifth aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute the method described in the first aspect.
[0064] According to a sixth aspect of the present disclosure, a computer program product is provided, comprising a computer program, wherein when the computer program is executed by a processor, the method described in the first aspect is implemented.
[0065] The method, device, system, electronic device and storage medium for online inference of a model provided by the present disclosure aggregate at least two received request data to obtain target request data; perform online inference on the target request data to obtain inference data corresponding to the at least two request data respectively; splice the inference data according to the execution order of the inference steps, and send the spliced inference data, wherein the execution order is the time sequence of receiving the at least two request data. Compared with the related art, during the operation of the processor, the embodiment of the present application aggregates at least two received request data to obtain target request data, performs online inference on the target request data, obtains inference results corresponding to each request data respectively, splices the inference results according to the order of inference steps, and aggregates multiple request data to make full use of the computing power of the processor and improve the utilization rate of the processor.
[0066] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present application, nor is it intended to limit the scope of the present application. Other features of the present application will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0067] The accompanying drawings are used to better understand the present solution and do not constitute a limitation of the present disclosure.
[0068] Figure 1 A flowchart of a method for online model reasoning provided by an embodiment of the present disclosure;
[0069] Figure 2 A schematic diagram of the architecture of an inference system provided by an embodiment of the present disclosure;
[0070] Figure 3 A flowchart of a method for aggregating request data provided by an embodiment of the present disclosure;
[0071] Figure 4 A schematic diagram of the architecture of another reasoning system provided by an embodiment of the present disclosure;
[0072] Figure 5 A schematic diagram of the architecture of another reasoning system provided by an embodiment of the present disclosure;
[0073] Figure 6 A schematic diagram of the structure of a device for online model inference provided by an embodiment of the present disclosure;
[0074] Figure 7 A schematic diagram of the structure of another device for online model inference provided by an embodiment of the present disclosure;
[0075] Figure 8A schematic diagram of the structure of a system for online model inference provided by an embodiment of the present disclosure;
[0076] Fig. 9 A schematic block diagram of an exemplary electronic device provided for an embodiment of the present disclosure. DETAILED DESCRIPTION
[0077] The following is a description of exemplary embodiments of the present disclosure in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be recognized by those of ordinary skill in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0078] The following describes the method and device, system, electronic device and storage medium for online model reasoning in embodiments of the present disclosure with reference to the accompanying drawings.
[0079] Figure 1 A flowchart of a method for online model inference provided in an embodiment of the present disclosure.
[0080] like Figure 1 As shown, the method is applied to a processor of a server, and the method comprises the following steps:
[0081] Step 101: Aggregate at least two received request data to obtain target request data.
[0082] In order to better understand the aggregation of request data, such as Figure 2 As shown, Figure 2A schematic diagram of the architecture of an inference system provided for an embodiment of the present application; a preset backend code based on a preset inference framework is used to generate a preset dynamic library, wherein the core code of the preset backend code integrates an online inference engine; by modifying the configuration file of the preset inference framework, the data volume parameter value of the request data in the configuration file is modified to the parameter value of the maximum load of the processor, so as to support the preset dynamic library to aggregate at least two requests received, and convert the aggregated data into data that can be input to the online inference engine through the preset dynamic library, that is, the target request data; however, it should be clear that this statement is not intended to limit the method for aggregating request data, but may also be other methods that can aggregate request data, and in addition, it does not limit the preset inference architecture, preset backend code, online inference engine and preset dynamic library used; the order in which the request data are aggregated is positively correlated with the time order in which they are transmitted to the server, and the earlier the request data is transmitted to the server, the priority is given to aggregation, wherein the number of aggregated request data is related to the data volume of the request data and the maximum load of the processor. For example, if the data volume of two request data reaches the maximum load of the processor, the two request data are aggregated; if the data volume of ten request data can reach the maximum load of the processor, the ten request data are aggregated.
[0083] Step 102: Perform online reasoning on the target request data to obtain reasoning data corresponding to the at least two request data respectively.
[0084] Please continue reading Figure 2 , the target request data is input into the custom interface of the online reasoning engine through a preset dynamic library, and the custom interface transmits the target request data to the reasoning model of the online reasoning engine. However, it should be clear that this statement is not intended to limit the target request data to be transmitted to the reasoning model only through the custom interface, and it can also be transmitted to the reasoning model through other methods; the target request data is inferred through the reasoning model to obtain the reasoning data corresponding to the at least two request data respectively. Since the data volume of the target request data reaches the maximum load of the processor, the reasoning model can make full use of the computing power of the processor when reasoning the target request data, thereby improving the utilization rate of the processor.
[0085] Step 103, splicing the inference data according to the execution order of the inference steps, and sending the spliced inference data, wherein the execution order is the time sequence of receiving the at least two request data.
[0086] Please continue reading Figure 2In order to reduce the inference delay and improve the user experience, in the process of the inference model inferring the target request data, the current inference data is spliced with the previous inference data through a custom interface, and the spliced inference data is transmitted to the preset dynamic library. The preset dynamic library splits the spliced inference data into inference data corresponding to the request data, and sends the inference data corresponding to the request data to the corresponding client, so as to achieve the purpose of inference and output at the same time; however, it should be clear that this statement is not intended to limit the transmission method of the inference data, and the inference data can also be transmitted through other methods.
[0087] It should be noted that the preset reasoning framework and the online reasoning engine may be located in the same carrier or in different carriers, which is not limited in this embodiment.
[0088] The method for online inference of a model provided by the present disclosure aggregates at least two received request data to obtain target request data; performs online inference on the target request data to obtain inference data corresponding to the at least two request data respectively; splices the inference data according to the execution order of the inference steps, and sends the spliced inference data, wherein the execution order is the time sequence of receiving the at least two request data. Compared with the related art, during the operation of the processor, the embodiment of the present application aggregates at least two received request data to obtain target request data, performs online inference on the target request data, obtains inference results corresponding to each request data respectively, splices the inference results according to the order of inference steps, and aggregates multiple request data, so as to make full use of the computing power of the processor and improve the utilization rate of the processor.
[0089] As a refinement of step 101, when aggregating the at least two received request data to obtain the target request data, the following methods may be used but are not limited to: Figure 3 As shown, Figure 3 A flowchart of a method for aggregating request data provided in an embodiment of the present application includes:
[0090] Step 201, respectively obtain the data sending time and data volume of at least two request data.
[0091] In order to determine the number of request data aggregated through the preset dynamic library, it is necessary to obtain the data sending time and data size of at least two request data. When the request data is transmitted to the server, the current time is recorded by the server's time function as the sending time of the request data. The server reads the Content-Length (content length) of the request data to determine the data size of the request data, and transmits the sending time of the request data and the data size of the request data to the preset dynamic library. However, it should be clear that this statement is not intended to limit the method of obtaining the data sending time and data size of the request data.
[0092] Step 202, in a preset dynamic library, according to the data sending time, the data volume and the maximum load of the processor, the at least two request data are aggregated in the order of the data sending time until the data volume of the aggregated request data reaches the maximum load of the processor, to obtain the target request data, wherein the preset dynamic library runs in the processor.
[0093] In order to improve user experience, when request data is aggregated through a preset dynamic library, the maximum load of the processor is used as the upper limit, and the request data is aggregated according to the time sequence of the received request data to obtain target request data, specifically including: comparing the data reception time of the request data, and the request data with the earlier data reception time is aggregated first, and the data size of the request data is compared with the maximum load of the processor. When the data size of the aggregated request data reaches the maximum load of the processor, the aggregation of the request data is stopped. At this time, the aggregated data is the target request data. The processor can be any processor that can perform reasoning, such as a GPU, but it should be clear that this statement is not intended to limit the processor to only a GPU.
[0094] As a refinement of step 202, when executing the method for creating the preset dynamic library, the configuration file of the preset reasoning architecture is configured according to the configuration instructions of the preset aggregation rules. After completing the configuration of the configuration file, based on the preset back-end code of the preset reasoning framework, the configured configuration file is compiled by the encoder, and the configured configuration file is compiled into a preset dynamic library. The function of the preset dynamic library is to aggregate the received request data, convert the aggregated data into data that can be input to the online reasoning engine, receive the reasoning data sent by the online reasoning engine, and send the reasoning data to the client. However, it should be clear that this statement is not intended to limit the method for creating the preset dynamic library.
[0095] As a refinement of step 102, please continue to refer to Figure 2When performing the online reasoning on the target request data to obtain the reasoning data corresponding to the at least two request data respectively, the target request data is received in the reasoning model of the online reasoning engine through the preset data transmission interface, the reasoning model extracts features from the target request data to obtain the target request data after feature extraction, and the reasoning model reasons on the target request data after feature extraction to obtain the reasoning data corresponding to the at least two request data respectively, wherein the preset data transmission interface is an interface for receiving the target request data and outputting reasoning data.
[0096] As a refinement of step 103, please continue to refer to Figure 2 , execute the splicing processing of the reasoning data according to the execution order of the reasoning steps, and when the spliced reasoning data is sent, since the transmission of the reasoning data adopts the method of transmitting while reasoning, the reasoning data is divided into multiple segments. When it is subsequently transmitted to the client, it may cause the loss or error of the reasoning data. Therefore, when the current reasoning data obtained by the online reasoning engine is output based on the preset data transmission interface, it is necessary to obtain a preset number of previous reasoning data according to the reasoning steps, and the previous reasoning data is the reasoning data output before the current reasoning data. The previous reasoning data and the current reasoning data are spliced to obtain the spliced reasoning data, and the spliced reasoning data is sent. However, it should be clear that this statement is not intended to limit the transmission method of the reasoning data; for example, the current reasoning data is b, the previous reasoning data is a, the spliced reasoning data is ab, the current reasoning data is c, the previous reasoning data is b, and the spliced reasoning data is bc. Both ab and bc are sent to the preset dynamic library. The preset dynamic library integrates ab and bc into abc based on the fact that both ab and bc contain b.
[0097] In practical applications, during the process of the preset inference model inferring the target request data, the amount of the target request data becomes smaller and smaller during the inference process, resulting in reduced processor utilization. In order to better understand the inference of the request data, Figure 4 As shown, Figure 4An architectural diagram of another reasoning system provided for an embodiment of the present application, before outputting the current reasoning data based on the preset data transmission interface, the preset data transmission interface is used to determine whether the output order of the reasoning data for completing the reasoning of the target request data exceeds the preset order threshold; if it is determined that the output order of the reasoning data for completing the reasoning of the target request data exceeds the preset order threshold, the output of the current reasoning data is paused, and the target request data of the unfinished reasoning is transmitted to the preset dynamic library through the scheduling module. The preset dynamic library re-aggregates the target request data of the unfinished reasoning with the newly received request data to obtain new target request data. However, it should be clear that this statement is not intended to limit the target request data of the unfinished reasoning to be transmitted only through the scheduling module.
[0098] In actual applications, during the process of the preset reasoning model reasoning the target request data, the preset data transmission interface determines that the output order of the reasoning data after the target request data completes the reasoning does not exceed the preset order threshold, then the preset data transmission interface outputs the current reasoning data.
[0099] In practical applications, it is therefore necessary to modify the transmission protocol of the preset inference architecture in order to facilitate a better understanding of the configuration parameter transmission, such as Figure 5 As shown, Figure 5 The schematic diagram of the architecture of another reasoning system provided for the embodiment of the present application supports the preset reasoning model to adjust and optimize the configuration parameters during the reasoning process, sends the parameter configuration instruction to the online reasoning engine, and the online reasoning engine receives the parameter configuration instruction for the online reasoning engine. The online reasoning engine transmits the parameter configuration through a preset transmission channel (such as context), transmits the configuration parameters carried in the parameter configuration instruction to the custom interface, and inputs the configuration parameters carried in the parameter configuration instruction into the online reasoning engine through the custom interface, so that the online reasoning engine adjusts and optimizes the parameters of the model according to the configuration parameters. In order to support dynamic parameter transmission, the Triton framework input protocol is modified to support the transmission of configuration parameters, and the configuration parameters transmitted by the protocol are transmitted to the online reasoning engine through the context to realize dynamic parameter tuning of the model, thereby realizing dynamic parameter tuning.
[0100] In summary, the embodiments of the present disclosure can achieve the following effects:
[0101] 1. The embodiment of the present disclosure obtains target request data by aggregating at least two received request data, inputs the target request data into an online reasoning engine, obtains the reasoning results corresponding to each request data respectively, and splices the reasoning results in the order of reasoning steps. By performing aggregating processing on multiple request data, the computing power of the processor can be fully utilized and the utilization rate of the processor can be improved.
[0102] 2. The disclosed embodiment uses an online reasoning engine based on a preset reasoning framework. When the output order of the current reasoning data exceeds a preset order threshold, the output of the current reasoning data is suspended, and the request data corresponding to the current reasoning data is re-aggregated with the newly received request data to obtain new target request data, thereby effectively improving the user experience, reducing the user's sense of delay caused by the long calculation time of the online reasoning service, effectively improving the GPU utilization rate, and reducing the online reasoning cost.
[0103] 3. The embodiment of the present disclosure transmits the configuration parameters carried in the parameter configuration instruction to the online inference engine through a preset transmission channel, supports dynamic parameter tuning, and can improve the efficiency of algorithm tuning.
[0104] Corresponding to the above-mentioned model online reasoning method, the present invention also proposes a model online reasoning device. Since the device embodiment of the present invention corresponds to the above-mentioned method embodiment, the details not disclosed in the device embodiment can be referred to the above-mentioned method embodiment, and will not be repeated in the present invention.
[0105] Figure 6 A schematic diagram of the structure of a device for online model inference provided by an embodiment of the present disclosure, such as Figure 6 As shown, including:
[0106] The aggregation unit 31 is used to aggregate at least two received request data to obtain target request data;
[0107] An inference unit 32, configured to perform online inference on the target request data to obtain inference data corresponding to the at least two request data respectively;
[0108] The sending unit 33 is used to splice the inference data according to the execution order of the inference steps and send the spliced inference data, wherein the execution order is the time sequence of receiving the at least two request data.
[0109] The device for online inference of the model provided by the present disclosure aggregates at least two received request data to obtain target request data; performs online inference on the target request data to obtain inference data corresponding to the at least two request data respectively; splices the inference data according to the execution order of the inference steps, and sends the spliced inference data, wherein the execution order is the time sequence of receiving the at least two request data. Compared with the related art, during the operation of the processor, the embodiment of the present application aggregates at least two received request data to obtain target request data, performs online inference on the target request data, obtains inference results corresponding to each request data respectively, splices the inference results according to the order of inference steps, and aggregates multiple request data, so as to make full use of the computing power of the processor and improve the utilization rate of the processor.
[0110] Furthermore, in a possible implementation of this embodiment, Figure 7 As shown, the aggregation unit 31 includes:
[0111] An acquisition module 311 is used to respectively acquire the data sending time and data volume of at least two request data;
[0112] The aggregation module 312 is used to aggregate the at least two request data in a preset dynamic library according to the data sending time, the data volume and the maximum load of the processor in the order of the data sending time, until the data volume of the aggregated request data reaches the maximum load of the processor, so as to obtain the target request data, wherein the preset dynamic library runs in the processor.
[0113] Furthermore, in a possible implementation of this embodiment, the aggregation module 312 is further configured to:
[0114] In response to a configuration instruction for a preset aggregation rule in the configuration file, configuring the configuration file;
[0115] After completing the configuration of the configuration file, the preset dynamic library is created based on the configuration file.
[0116] Furthermore, in a possible implementation of this embodiment, Figure 7 As shown, the inference unit 32 is also used to preset a data transmission interface to receive the target request data through the preset data transmission interface, and perform online inference on the target request data to obtain inference data corresponding to the at least two request data respectively, wherein the preset data transmission interface is an interface for receiving the target request data and outputting inference data.
[0117] Furthermore, in a possible implementation of this embodiment, Figure 7As shown, the sending unit 33 includes:
[0118] An output module 331 is used to output the current inference data based on the preset data transmission interface;
[0119] An acquisition module 332, configured to acquire a preset number of previous reasoning data according to the reasoning steps, wherein the previous reasoning data is reasoning data output before the current reasoning data;
[0120] A splicing module 333, used for splicing the previous reasoning data and the current reasoning data to obtain spliced reasoning data;
[0121] The sending module 334 is used to send the spliced inference data.
[0122] Furthermore, in a possible implementation of this embodiment, as Figure 7 As shown, the device also includes:
[0123] A judging unit 34, configured to judge whether the output sequence of the current inference data exceeds a preset sequence threshold;
[0124] The determining unit 35 is configured to, when it is determined that the output order of the current reasoning data exceeds the preset order threshold, suspend the output of the current reasoning data and determine the request data for which the reasoning is not completed;
[0125] The aggregation unit 31 is further used to re-aggregate the unfinished inference request data with the newly received request data to obtain new target request data.
[0126] Furthermore, in a possible implementation of this embodiment, the output module 331 is also used to output the current reasoning data based on the preset data transmission interface when it is determined that the output order of the current reasoning data does not exceed the preset order threshold.
[0127] Furthermore, in a possible implementation of this embodiment, as Figure 7 As shown, the device also includes:
[0128] A receiving unit 36, configured to receive a parameter configuration instruction;
[0129] The input unit 37 is used to adjust parameters according to the configuration parameters carried in the parameter configuration instruction.
[0130] It should be noted that the above explanation of the method embodiment is also applicable to the device of this embodiment, and the principle is the same, which is not limited in this embodiment.
[0131] Corresponding to the above-mentioned model online reasoning method, the present invention also proposes a model online reasoning system.
[0132] Figure 8 A schematic diagram of the structure of a system for online model inference provided by an embodiment of the present disclosure is shown in FIG. Figure 8 As shown, including:
[0133] A preset reasoning framework 41 and an online reasoning engine 42; wherein,
[0134] The preset reasoning framework is used to aggregate at least two received request data to obtain target request data, and transmit the target request data to the online reasoning engine;
[0135] The online reasoning engine is used to receive the target request data transmitted by the preset reasoning framework, perform online reasoning on the target request data, obtain reasoning data corresponding to the at least two request data respectively, splice the reasoning data according to the execution order of the reasoning steps, and send the spliced reasoning data to the preset reasoning framework, wherein the execution order is the time sequence of receiving the at least two request data;
[0136] The preset reasoning framework is used to receive the spliced reasoning data sent by the online reasoning engine; and output the spliced reasoning data.
[0137] Further, in a possible implementation of this embodiment, the preset reasoning framework 41 is further used to configure the configuration file in response to a configuration instruction for a preset aggregation rule in the configuration file; after completing the configuration of the configuration file, create the preset dynamic library based on the configuration file;
[0138] Furthermore, in a possible implementation of this embodiment, the online reasoning engine 42 is also used to receive parameter configuration instructions; and adjust parameters according to the configuration parameters carried in the parameter configuration instructions.
[0139] It should be noted that the preset reasoning framework 41 and the online reasoning engine 42 may be located in the same carrier or in different carriers, which is not limited in this embodiment.
[0140] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium and a computer program product.
[0141] Fig. 9A schematic block diagram of an example electronic device 500 that can be used to implement an embodiment of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.
[0142] like Fig. 9 As shown, the device 500 includes a computing unit 501, which can perform various appropriate actions and processes according to a computer program stored in a ROM (Read-Only Memory) 502 or a computer program loaded from a storage unit 508 to a RAM (Random Access Memory) 503. In the RAM 503, various programs and data required for the operation of the device 500 can also be stored. The computing unit 501, the ROM 502, and the RAM 503 are connected to each other via a bus 504. An I / O (Input / Output) interface 505 is also connected to the bus 504.
[0143] A number of components in the device 500 are connected to the I / O interface 505, including: an input unit 506, such as a keyboard, a mouse, etc.; an output unit 507, such as various types of displays, speakers, etc.; a storage unit 508, such as a disk, an optical disk, etc.; and a communication unit 509, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 509 allows the device 500 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0144] The computing unit 501 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 501 include, but are not limited to, a CPU (Central Processing Unit), a GPU (Graphic Processing Units), various dedicated AI (Artificial Intelligence) computing chips, various computing units running machine learning model algorithms, a DSP (Digital Signal Processor), and any appropriate processor, controller, microcontroller, etc. The computing unit 501 performs the various methods and processes described above, such as a method for online model reasoning. For example, in some embodiments, the method for online model reasoning may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 508. In some embodiments, part or all of the computer program may be loaded and / or installed on the device 500 via the ROM 502 and / or the communication unit 509. When the computer program is loaded into the RAM 503 and executed by the computing unit 501, one or more steps of the method described above may be performed. Alternatively, in other embodiments, the computing unit 501 may be configured to execute the aforementioned model online inference method in any other appropriate manner (for example, by means of firmware).
[0145] Various embodiments of the systems and techniques described above herein may be implemented in digital electronic circuit systems, integrated circuit systems, FPGAs (Field Programmable Gate Arrays), ASICs (Application-Specific Integrated Circuits), ASSPs (Application Specific Standard Products), SOCs (System On Chips), CPLDs (Complex Programmable Logic Devices), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include: being implemented in one or more computer programs that may be executed and / or interpreted on a programmable system including at least one programmable processor that may be a dedicated or general-purpose programmable processor that may receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0146] The program code for implementing the method of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that the program code, when executed by the processor or controller, enables the functions / operations specified in the flow chart and / or block diagram to be implemented. The program code may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.
[0147] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or device, or any suitable combination of the foregoing. More specific examples of machine-readable storage media may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a RAM, a ROM, an EPROM (Electrically Programmable Read-Only-Memory) or a flash memory, an optical fiber, a CD-ROM (Compact Dis sc Read-Only Memory), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0148] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (Cathode-Ray Tube) or LCD (Liquid Crystal Display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0149] The systems and techniques described herein may be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: LAN (Local Area Network), WAN (Wide Area Network), the Internet, and blockchain networks.
[0150] A computer system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The relationship between the client and the server is generated by computer programs running on the corresponding computers and having a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system to solve the defects of difficult management and weak business scalability in traditional physical hosts and VPS services ("Virtual Private Server", or "VPS" for short). The server may also be a server of a distributed system, or a server combined with a blockchain.
[0151] It should be noted that artificial intelligence is a discipline that studies how computers can simulate certain human thought processes and intelligent behaviors (such as learning, reasoning, thinking, planning, etc.), and includes both hardware-level and software-level technologies. Artificial intelligence hardware technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, and big data processing; artificial intelligence software technologies mainly include computer vision technology, speech recognition technology, natural language processing technology, as well as machine learning / deep learning, big data processing technology, knowledge graph technology, and other major directions.
[0152] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps recorded in this disclosure can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this document does not limit this.
[0153] The above specific implementations do not constitute a limitation on the protection scope of the present disclosure. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present disclosure shall be included in the protection scope of the present disclosure.
Claims
1. A method for online model inference, It is characterized in that include: Aggregate at least two received request data to obtain target request data; Performing online reasoning on the target request data to obtain reasoning data corresponding to the at least two request data respectively; The inference data are spliced according to the execution order of the inference steps, and the spliced inference data are sent, wherein the execution order is the time sequence of receiving the at least two request data.
2. The method according to claim 1, It is characterized in that The step of aggregating the at least two received request data to obtain the target request data includes: Obtaining the data sending time and data volume of at least two request data respectively; In a preset dynamic library, the at least two request data are aggregated in the order of the data sending time until the data size of the aggregated request data reaches the maximum load of the processor, thereby obtaining the target request data, wherein the preset dynamic library runs in the processor.
3. The method according to claim 2, It is characterized in that The method for creating the preset dynamic library includes: In response to a configuration instruction for a preset aggregation rule in the configuration file, configuring the configuration file; After completing the configuration of the configuration file, the preset dynamic library is created based on the configuration file.
4. The method according to claim 1, It is characterized in that The performing online reasoning on the target request data to obtain the reasoning data corresponding to the at least two request data respectively comprises: The target request data is received through a preset data transmission interface, and online reasoning is performed on the target request data to obtain reasoning data corresponding to the at least two request data respectively, wherein the preset data transmission interface is an interface for receiving the target request data and outputting the reasoning data.
5. The method according to claim 4, It is characterized in that The step of splicing the inference data according to the execution order of the inference steps and sending the spliced inference data comprises: Outputting the current inference data based on the preset data transmission interface; According to the reasoning steps, a preset number of previous reasoning data are obtained, where the previous reasoning data is the reasoning data output before the current reasoning data; splicing the previous reasoning data and the current reasoning data to obtain spliced reasoning data; The spliced inference data is sent.
6. The method according to claim 5, It is characterized in that Before outputting the current inference data based on the preset data transmission interface, the method includes: Determining whether the output order of the current inference data exceeds a preset order threshold; If it is determined that the output order of the current reasoning data exceeds the preset order threshold, pausing the output of the current reasoning data and determining the request data for which the reasoning has not been completed; The unfinished inference request data is re-aggregated with the newly received request data to obtain new target request data.
7. The method according to claim 6, It is characterized in that The outputting the current inference data based on the preset data transmission interface includes: If it is determined that the output order of the current inference data does not exceed the preset order threshold, the current inference data is output based on the preset data transmission interface.
8. The method according to any one of claims 1 to 7, It is characterized in that Before performing online reasoning on the target request data to obtain reasoning data corresponding to the at least two request data respectively, the method includes: Receive parameter configuration instructions; Adjust parameters according to the configuration parameters carried in the parameter configuration instruction.
9. A device for online model inference, It is characterized in that include: an aggregation unit, used for aggregating at least two received request data to obtain target request data; An inference unit, configured to perform online inference on the target request data to obtain inference data corresponding to the at least two request data respectively; A sending unit is used to splice the inference data according to the execution order of the inference steps and send the spliced inference data, wherein the execution order is the time sequence of receiving the at least two request data.
10. A system for online model reasoning, It is characterized in that The system includes: a preset reasoning framework and an online reasoning engine; wherein, The preset reasoning framework is used to aggregate at least two received request data to obtain target request data, and transmit the target request data to the online reasoning engine; The online reasoning engine is used to receive the target request data transmitted by the preset reasoning framework, perform online reasoning on the target request data, obtain reasoning data corresponding to the at least two request data respectively, splice the reasoning data according to the execution order of the reasoning steps, and send the spliced reasoning data to the preset reasoning framework, wherein the execution order is the time sequence of receiving the at least two request data; The preset reasoning framework is used to receive the spliced reasoning data sent by the online reasoning engine; and output the spliced reasoning data.
11. An electronic device, It is characterized in that include: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 8.
12. A non-transitory computer-readable storage medium storing computer instructions, It is characterized in that The computer instructions are used to cause the computer to execute the method according to any one of claims 1-8.
13. A computer program product, It is characterized in that The invention comprises a computer program which, when executed by a processor, implements the method according to any one of claims 1 to 8.