Inference method

By configuring the inference server and GPU in the device, and using the shared memory pool and gRPC protocol for cross-process data transmission, the problem of network timeout in remote model inference is solved, and the inference efficiency and performance is improved.

CN120196438APending Publication Date: 2025-06-24BEIJING SANKUAI NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510265414.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-06
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

The existing remote model inference method has network timeout problems, resulting in inference failure, time-consuming, low efficiency and poor performance.

Method used

By configuring an inference server and GPU in the device, using the shared memory pool and gRPC protocol, cross-process data transmission between the inference process and the business process is realized, data copying overhead is reduced, and local inference capabilities are established.

Benefits of technology

It effectively avoids network timeout problems, improves inference efficiency and performance, and reduces the time-consuming data transmission during inference.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120196438A_ABST
    Figure CN120196438A_ABST
Patent Text Reader

Abstract

The invention provides a reasoning method, and relates to the technical field of computers, a reasoning server and a GPU are configured in a device applied to reasoning operation, the reasoning method comprises the following steps: starting a service process and starting a reasoning process represented by the reasoning server, under the condition that the reasoning server has a reasoning capability, starting a reasoning process represented by the reasoning server; determining a target shared memory from a shared resource pool in the equipment; constructing a reasoning request, obtaining data corresponding to the reasoning request based on a business process, and storing the data corresponding to the reasoning request to the target shared memory; and calling a reasoning process corresponding to the GPU to obtain data corresponding to the reasoning request from the target shared memory so as to perform reasoning operation to obtain a reasoning result, and storing the reasoning result to the target shared memory. The reasoning time consumption can be reduced, and the reasoning effect is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technology, and in particular, to an inference method. Background Art

[0002] With the development of models, more and more algorithms prefer to use model technology to improve the accuracy of algorithms. In the conventional remote model inference method, there is a serious network timeout problem, which easily causes inference failure. Therefore, the inference process takes a long time, the inference efficiency is low, and the inference performance is poor. Summary of the Invention

[0003] An object of the present disclosure is to provide an inference method, so as to at least to some extent overcome the problems of inference delay and low inference efficiency caused by the limitations and defects of related technologies.

[0004] Other features and advantages of the present disclosure will become apparent through the following detailed description, or be partially learned through the practice of the present disclosure.

[0005] According to one aspect of the present disclosure, there is provided an inference method, which is applied to a device configured with an inference server and a GPU for inference operations. The inference method includes: starting a service process and starting an inference process represented by the inference server. When the inference server has inference capabilities, determining a target shared memory from a shared resource pool inside the device; constructing an inference request, obtaining data corresponding to the inference request based on the service process, and storing the data corresponding to the inference request into the target shared memory; calling an inference process corresponding to the GPU to obtain the data corresponding to the inference request from the target shared memory to perform an inference operation to obtain an inference result, and storing the inference result into the target shared memory.

[0006] In an exemplary embodiment of the present disclosure, the determining a target shared memory from a shared memory pool inside the device includes: obtaining a memory identifier from a shared memory pool inside the device, and determining the target shared memory according to the memory identifier; the memory identifier is determined according to positions corresponding to shared memories of multiple sub-blocks obtained by splitting the shared memory; wherein, the application process of the shared memory pool includes: after starting the inference process, applying for a shared memory pool according to a configuration file of a model applied by the inference request.

[0007] In an exemplary embodiment of the present disclosure, the applying for a shared memory pool according to a configuration file of a model applied by the inference request includes: determining a memory size of the shared memory in the shared memory pool according to the volume of a single request of the model, the maximum number of concurrent requests supported by the model, and a reserved buffer space, and applying for the shared memory pool according to the memory size.

[0008] In an exemplary embodiment of the present disclosure, starting the inference process represented by the inference server includes: invoking the startup script of the inference process to start the inference process. After the inference process is successfully started, writing the process identifier of the inference process into a specified file and configuring the status of the inference process to a first status; listening to the semaphore of the inference process through the process monitoring function, and configuring the status of the inference process to a second status according to the semaphore or the specified file; when the status of the inference process is configured to the second status, restarting the inference process.

[0009] In an exemplary embodiment of the present disclosure, configuring the status of the inference process to the second status according to the semaphore or the specified file includes: if the semaphore terminates or the process identifier in the specified file does not exist, configuring the status of the inference process to the second status.

[0010] In an exemplary embodiment of the present disclosure, before starting the inference process, the method further includes: when it is detected that there is an un-terminated inference process on the device, terminating the un-terminated inference process and clearing the shared memory allocated on the device.

[0011] In an exemplary embodiment of the present disclosure, invoking the inference process corresponding to the GPU to obtain the data corresponding to the inference request from the target shared memory to perform an inference operation to obtain an inference result includes: invoking the inference process to obtain the data corresponding to the inference request from the target shared memory; monitoring the status of the inference process, and when the status of the inference process is a third status, performing an inference operation based on the data corresponding to the inference request.

[0012] In an exemplary embodiment of the present disclosure, monitoring the status of the inference process and performing an inference operation when the status of the inference process is a third status includes: accessing the interface of the inference process through a timer and determining the interface access status; if the interface access status is an interface access exception, performing a restart operation of the inference process; changing the status of the inference process to the third status and recording the restart time; spin-waiting for the inference process to restart, and when the time difference between the current time and the restart time is greater than a time threshold, performing a re-restart operation on the inference process until the inference process is successfully restarted; when the inference process is successfully restarted or the interface access status is normal interface access, performing a loading operation on the model to perform an inference operation.

[0013] In an exemplary embodiment of the present disclosure, the method further includes: determining an interface for the inference request to interact with the inference process through an inference process control module, determining a model usage method, and defining a model processing method; wherein, determining an interface for the inference request to interact with the inference process through the inference process control module, determining a model usage method, and defining a model processing method includes: defining, through an interface layer, an interface for the inference request to interact with the inference process; determining, through an abstraction layer, one or more of a model loading method, an unloading method, and an inference method; and determining, based on an execution layer, a model acquisition location and a processing method for input parameters in the inference request.

[0014] In an exemplary embodiment of the present disclosure, the method further includes: monitoring the inference performance of the inference request, the system state of the GPU, the capacity of the shared memory pool, and abnormal situations; wherein, monitoring the inference performance of the inference request, the system state of the GPU, the capacity of the shared memory pool, and abnormal situations includes: monitoring at least one of the queries per second rate and the inference speed of the inference request; monitoring one or more of the usage rate, utilization rate, power consumption, and main frequency of the GPU; monitoring whether the capacity of the shared memory pool exceeds a water level line; and monitoring one or more of process exceptions, inference exceptions, shared memory capacity exceptions, machine performance exceptions, and process state exceptions.

[0015] In some embodiments of the present disclosure, through the provision of a GPU and an inference server, inference operations are implemented through an inference process and a service process inside the device, and data transmission between different processes is achieved through a target shared memory, reducing the performance overhead of data copying. By establishing an inference ability locally, the time consumed in network requests can be reduced, the data transmission efficiency in the inference request process can be improved, timeout problems can be avoided, and thus the inference efficiency, inference performance, and accuracy can be improved.

[0016] It should be understood that the above general description and the following detailed description are merely exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The accompanying drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure. Obviously, the drawings in the following description are only some embodiments of the present disclosure, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.

[0018] Figure 1 A schematic diagram showing the system architecture of an inference method in the related art is shown.

[0019] Figure 2 A schematic diagram showing an inference method according to an embodiment of the present disclosure is shown.

[0020] Figure 3 A schematic diagram showing the internal structural configuration of a device according to an embodiment of the present disclosure is shown.

[0021] Figure 4 A schematic diagram showing the internal layer structure of code according to an embodiment of the present disclosure is shown.

[0022] Figure 5 A schematic flowchart showing a subprocess management module according to an embodiment of the present disclosure is shown.

[0023] Figure 6 A schematic diagram showing the initialization of a service process according to an embodiment of the present disclosure is shown.

[0024] Figure 7 A schematic flowchart showing an inference process according to an embodiment of the present disclosure is shown.

[0025] Figure 8 A schematic diagram showing the health detection of an inference process according to an embodiment of the present disclosure is shown. Detailed implementation manners

[0026] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this disclosure will be more complete and comprehensive, and will fully convey the concept of the example embodiments to those skilled in the art. The features, structures, or characteristics described may be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of the embodiments of the present disclosure. However, those skilled in the art will recognize that the technical solutions of the present disclosure can be practiced without one or more of the specific details, or other methods, components, devices, steps, etc. may be used. In other cases, well-known technical solutions are not shown or described in detail to avoid obscuring the various aspects of the present disclosure.

[0027] In addition, the accompanying drawings are only schematic illustrations of the present disclosure and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and thus repeated descriptions thereof will be omitted. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.

[0028] Figure 1The flowchart of remote inference in some embodiments is schematically shown. Refer to Figure 1 As shown in , currently, when invoking model inference, most of the adopted solutions are to build a remote inference cluster and use RPC communication. The above deployment method is simple, but since the input parameter volume of the inference request may be relatively large, for example, reaching dozens of MB. It may cause a significant increase in the overall inference time during the transmission, serialization, and deserialization processes, resulting in low inference efficiency.

[0029] To solve the above technical problems, an inference method is provided in the embodiments of the present disclosure, which can be applied to the process of performing inference operations on any type of inference request.

[0030] In some embodiments, the inference method can be executed based on a business single machine. The business single machine refers to a device, such as a server, and in addition, it can also be a client. A GPU acceleration card can be installed inside the business single machine container and the inference server Triton Server can be deployed. Through cross-process inference and shared memory data transmission within the container, the performance overhead of data copying can be reduced, thereby maximizing the optimization of inference performance and response time.

[0031] Next, refer to Figure 2 As shown in , each step in the inference method in the embodiments of the present disclosure will be described in detail.

[0032] In step S210, start the business process and start the inference process represented by the inference server. When the inference server has the inference ability, determine the target shared memory from the shared resource pool inside the device.

[0033] In some embodiments of the present disclosure, the device can be a server or a server node. A GPU can be installed in the server node and the inference server Triton Server can be deployed. Here, the GPU refers to a GPU acceleration card, which is built based on a powerful GPU chip and is equipped with a large amount of video memory and high-speed interfaces. The inference server can be used to represent the inference process, and the inference process refers to the GPU inference process, that is, the inference process corresponding to the GPU. Specifically, Nvidia Tritonserver can be adopted. The GPU inference process can be an inference process that applies the GPU acceleration card.

[0034] In the inference process of the GPU acceleration card, it is first necessary to be correctly installed and compatible with other parts of the server. Install appropriate driver programs and deep learning frameworks to make full use of the computing power of the GPU.

[0035] The structure of the entire device can be deployed in a multi-process manner. Refer to Figure 3As shown in the figure, it may include an inference process and a business process. The inference process can be the GPU inference process Nvidia Triton server, and the business process can be a Java process. In addition, the architecture can also include shared resources. The shared resource part can include a local model repository and a shared memory pool. The shared memory pool is used to store the data corresponding to the inference request and the inference result. The local model repository is used to load the model when the business process starts the inference process. Among them, the business process part includes a client, a software development kit (SDK) for the inference process, and subprocess management. The SDK is used to implement the life cycle management and gRPC communication management of the inference process. The inference process and the business process communicate with each other using the gRPC protocol. For example, the business process can send an RPC request to the inference process, and the inference process can generate an RPC request and transmit it to the business process. The subprocess management module can start, perform a health check, and restart the inference process.

[0036] Data transmission uses shared memory to reduce the overhead of data serialization and deserialization while reducing one data copy, which can maximize the data transmission efficiency of inference requests.

[0037] For the code provided above, the internal layer structure of the code can refer to Figure 4 As shown in the figure, it mainly includes an inference process control module, a subprocess management module, an inference process status monitoring module, and a management module for the shared memory pool.

[0038] Among them, the inference process control module is used to determine the interface for the inference request to interact with the inference process, determine the model usage method, and define the model processing method. Exemplarily, the interface for the inference request to interact with the inference process can be defined through the interface layer; one or more of the model loading method, unloading method, and inference method can be determined through the abstraction layer; based on the execution layer, determine the location where the model is obtained and the processing method of the input parameters in the inference request.

[0039] For example, through the Interface layer, interfaces for interacting with the inference process can be defined, such as: model loading interface, model unloading interface, and inference interface. Through the Abstract layer, common abstract logics can be implemented, such as determining the model loading method, model unloading method, and inference method, etc. Through the Executer layer, the customized functions that the algorithm and the model need to implement can be determined, such as: the location where the loaded model is obtained and the processing method of the input parameters during inference.

[0040] Refer to Figure 4 As shown in the figure, after determining the above content, it can be transmitted to Figure 4The gRPC framework in it is used to transmit to the corresponding inference process of the inference server through the gRPC protocol. The data corresponding to the inference request can be written into the memory pool through the interface, and further the inference server reads the data from the memory pool.

[0041] When performing inference operations based on the GPU acceleration card and the inference server in the above architecture, first, the business process and the inference process represented by the inference server can be started. Exemplarily, the business process can be started first. Exemplarily, the configuration of the shared memory pool can be obtained. If the shared memory cannot be allocated on the current device, the startup of the business process can be terminated. For example, the maximum capacity or maximum occupancy of the configured shared memory can be obtained. When the calculated memory pool limit exceeds the maximum capacity, the startup of the business process can be terminated. If it does not exceed the maximum capacity, a gRPC connection pool can be created.

[0042] After the business process is started, the initialization function of the Spring framework can be used to load this function into the framework. Further, the environment can be initialized to check whether there are any un-terminated inference processes on the server. If there are, the un-terminated inference processes can be terminated, and at the same time, the shared memory allocated on the server can be cleared. Exemplarily, the inference process can be terminated according to the termination command, and the termination command can be the KILL-9PID command.

[0043] On this basis, the inference process can be started. The inference process can be started through a startup script. After the inference process is successfully started, the process identifier of the inference process can be written into a specified file, and the status of the inference process can be configured to the first status. The first status can be Startup Success, which is used to indicate that the inference process is ready to support inference.

[0044] After starting the inference process, the life cycle of the inference process can also be managed through the subprocess management module. Exemplarily, after starting the inference process, the inference process can be daemonized. A process monitoring function can be created to monitor the semaphore of the inference process through a timer. According to the semaphore or the specified file, the status of the inference process can be configured to the second status; when the status of the inference process is configured to the second status, the inference process can be restarted. Among them, it can be determined that the status of the inference process is configured to the second status according to the semaphore termination or the non-existence of the process identifier stored in the specified file. The second status Dead is used to indicate that the inference process is terminated. When the status of the inference process is the second status, the inference server does not have the inference ability, so the inference process needs to be restarted.

[0045] Figure 5 The flowchart of the subprocess management module is schematically shown in Figure 5 As shown in

[0046] Step S502, the business process starts, for example, newScore starts.

[0047] Step S504, the object Rean is initialized, that is, the environment is initialized. Specifically, it can be checked whether there is an un-terminated inference process; if so, go to step S506; if not, go to step S508.

[0048] Step S506, terminate the un-terminated inference process.

[0049] Step S508, start the inference process.

[0050] Step S510, after the inference process starts successfully, write the process identifier to the specified file.

[0051] Step S512, start the inference server through a script.

[0052] Step S514, create a process death monitoring function.

[0053] Step S516, monitor the semaphore of the inference process.

[0054] Step S518, it is monitored that the inference process terminates.

[0055] Step S520, set the state of the inference process to the second state.

[0056] Step S522, delete the process identifier of the inference process.

[0057] Step S524, restart the inference server again, and go back to step S504 to continue execution.

[0058] By using the subprocess management module to monitor the semaphore or the process identifier of the specified file in real time, the state of the inference process can be determined in real time, and then restarted immediately to avoid affecting the inference process.

[0059] The inference process status monitoring module in the embodiments of the present disclosure can monitor the inference performance of inference requests, the system status of the GPU, the capacity of the shared memory pool, and abnormal conditions. Exemplarily, at least one of the query rate per second and the inference speed of the inference request can be monitored. The query rate per second QPS refers to a measure of the amount of traffic processed by a specific query server within a specified time. The inference speed can be TP50, TP90, TP99, TP999, TP9999, MAX, AVG. The system status of the GPU includes one or more of the usage rate, utilization rate, power consumption, and main frequency of the GPU. Each model has its own shared memory pool configuration, which can show the usage of each shared memory pool. When the capacity of the shared memory pool exceeds the water level line, an alarm can be issued. The water level line can be 90%. The abnormal conditions can be various types of abnormal states, for example, it can include one or more of process exception, inference exception, shared memory capacity exception, machine performance exception, and process status exception. The abnormal condition monitoring can issue an abnormal alarm when any of the above abnormal conditions is detected.

[0060] In the embodiments of the present disclosure, after starting the inference process, the status of the inference server can be checked. If the inference server does not have the inference ability, an exception is directly thrown. If the inference server has the inference ability, the target shared memory can be determined from the shared memory pool inside the device.

[0061] In some embodiments, memory identifiers can be stored in the shared memory pool, so that the target shared memory can be determined according to the memory identifiers stored in the shared memory pool, facilitating the determination of the shared memory for storing the data corresponding to the inference request. The memory identifier can be determined according to the position corresponding to a sub-block of the shared memory applied in advance. The memory identifier stored in the shared memory pool can be any one of the multiple memory identifiers corresponding to the shared memories of multiple sub-blocks, which can be determined in advance according to actual requirements. It should be noted that the shared memory pool is applied according to the model. The application process of the shared memory pool may include: after starting the inference process, applying for shared memory according to the configuration file of the model applied by the inference request and registering the shared memory to the GPU. In the application of the shared memory pool, the mmap method is used to map a section of memory to a file. The mmap method is a method of memory mapping a file, that is, mapping a file or other object to the address space of the process, realizing the one-to-one correspondence between the file disk address and a section of virtual address in the process virtual address space. Exemplarily, the configuration file refers to the configuration file of the model, which may include the volume of a single request of the model, the maximum number of concurrent requests supported by the model, and the reserved cache space. Based on this, the volume of a single request of the model, the maximum number of concurrent requests supported by the model, and the reserved buffer space can be logically combined to determine the memory size of the shared memory. For example, the sum of the maximum number of concurrent requests supported by the model and the reserved cache space can be added to obtain a summation result, and the summation result can be multiplied by the volume of a single request to obtain the memory size of the shared memory. For example, for the KM model, the volume of a single request is 5MB, the model supports a maximum of 20 concurrent requests, and the reserved buffer space is 5. Therefore, it may be necessary to apply for (20 + 5) * 5MB of shared memory.

[0062] Furthermore, the entire shared memory can be split into multiple sub-blocks according to the volume of a single request of the model, so that multiple requests can be used simultaneously. Among them, the position of each sub-block of the shared memory will be recorded, and its position can be determined as the memory identifier of each shared memory to be used to uniquely indicate the shared memory corresponding to a sub-block. Based on this, each shared memory pool can include multiple memory identifiers, and the shared memory corresponding to each memory identifier can store a request and / or the corresponding data. The Apache Commons Pool component is used for the management of the shared memory pool. After applying for the shared memory, the shared memory can be registered to the GPU. Exemplarily, the SystemSharedMemoryRegister interface of the inference server can be called to complete the registration.

[0063] After obtaining the shared memory corresponding to the shared memory pool, the target shared memory can be determined according to the memory identifiers stored in the shared memory pool, so as to determine the shared memory for storing the data corresponding to the inference request. The usage of shared memory is restricted for each device and cannot be used arbitrarily. Therefore, it is necessary to manage it in the form of a shared memory pool. Using a shared memory pool can reduce the frequent creation of shared memory and control the usage. The shared memory pool can adopt the mmap method to map a section of memory to a file. For example, it can include multiple abstract classes FileChannel for reading operations in the file. A MappedByteBuffer object is obtained through FileChannel to map a part of the file to the memory address space, so that the operations on the file can be directly completed through memory access. Among them, each FileChannel can include multiple MappedByteBuffer objects. For each MappedByteBuffer object, its position and size can be displayed.

[0064] In the embodiments of the present disclosure, the processes of starting the service process and starting the inference process belong to the steps in the service process initialization process. Refer to Figure 6 As shown in, the service process initialization process may include the following steps:

[0065] Step S602, start the service process.

[0066] Step S604, obtain the configuration of each type of memory pool.

[0067] Step S606, obtain the maximum capacity of the configured shared memory.

[0068] Step S608, calculate whether the memory pool limit exceeds the maximum capacity.

[0069] Step S610, create a gRPC connection pool.

[0070] Step S612, terminate the un-terminated inference process.

[0071] Step S614, execute the inference process startup script to start the inference process.

[0072] Step S616, determine the status of the inference server.

[0073] Step S618, clear the content under the shared memory file directory.

[0074] Step S620, create a corresponding shared memory pool according to the model and algorithm type.

[0075] Step S614, create a shared memory file channel FileChannel.

[0076] Step S616, perform memory mapping on the file channel FileChannel.

[0077] Step S618, register the gRPC interface.

[0078] Step S620, register into the shared memory.

[0079] By applying for a shared memory pool according to the configuration file of the model after starting the inference process, the accuracy of the shared memory pool can be improved.

[0080] Next, in step S220, construct an inference request, obtain the data corresponding to the inference request based on the service process, and store the data corresponding to the inference request into the target shared memory.

[0081] In the embodiments of the present disclosure, the process of constructing an inference request includes: constructing the input parameters of the inference request, the position of the shared memory where the output needs to be written, and the model information used by the inference request. After obtaining the inference request, the inference request can be processed based on the service process to obtain the data corresponding to the inference request. The data corresponding to the inference request can be input data and features. Further, the data corresponding to the inference request can be written into the mapping file corresponding to the target shared memory. Exemplarily, an RPC connection can be obtained from the gRPC connection pool to initiate a gRPC request. The data corresponding to the inference request is written into the target shared memory specified by the shared memory pool through the gRPC request.

[0082] In step S230, call the inference process corresponding to the GPU to obtain the data corresponding to the inference request from the target shared memory, perform an inference operation to obtain an inference result, and store the inference result into the target shared memory.

[0083] In the embodiments of the present disclosure, the inference process corresponding to the GPU can be called through a gRPC request. If the request is successful, the data corresponding to the inference request can be obtained from the target shared memory to perform an inference operation to obtain an inference result. Further, the inference result can be stored into the mapping file of the target shared memory through a gRPC request. After completing the inference operation, the connection can be returned to the connection pool, and the target shared memory can be returned to the shared memory pool.

[0084] Figure 7 Schematically shows a schematic diagram of the inference process. Refer to Figure 7 as shown in, which mainly includes the following steps:

[0085] Step S702, start the inference process.

[0086] Step S704: Determine whether the inference server has the conditions for inference; if not, go to step S706. If so, go to step S708.

[0087] Step S706: Throw an exception.

[0088] Step S708: Obtain the buffer of the shared memory from the shared memory pool.

[0089] Step S710: Write the data into the target shared memory.

[0090] Step S712: Construct an inference request.

[0091] Step S714: Obtain an RPC connection from the gRPC connection pool.

[0092] Step S716: Invoke gRPC inference.

[0093] Step S718: Read the return data from the target shared memory.

[0094] Step S720: Construct a return result.

[0095] Step S722: Release the target shared memory.

[0096] Specifically, if it is checked that the inference server does not have the inference ability, an exception is directly thrown. Obtain the reference memory identifier of the shared memory object from the shared memory pool to facilitate defining the writing location. Construct an inference request, including the input parameters of the request, the location where the output needs to be written to the shared memory, and the model information, etc. Obtain an RPC connection from the gRPC connection pool. Initiate a gRPC request. If an exception occurs in the request, an exception is thrown and handed over to the upper layer for processing. After the request is successful, read the data from the specified output location of the target shared memory, return the connection to the connection pool, and return the target shared memory to the shared memory pool.

[0097] In the embodiments of the present disclosure, through cross-process inference within the container and shared memory data transmission, the performance overhead of data copying is reduced, thereby maximizing the optimization of inference performance and response time, improving inference efficiency, and enhancing inference performance.

[0098] In some embodiments, the health status of the inference process can also be detected. Exemplarily, the interface of the inference process can be accessed through a timer, and the interface access status can be determined; if the interface access status is abnormal interface access, a restart operation of the inference process is executed; the status of the inference process is changed to a third state, and the restart time is recorded; spin and wait for the inference process to restart, and when the time difference between the current time and the restart time is greater than the time threshold, perform a restart operation on the inference process again until the inference process restarts successfully; after the inference process restarts successfully, perform a loading operation on the model to perform an inference operation according to the data corresponding to the inference request and the loaded model. Among them, the timer can access the interface of the inference process according to a preset period, and the preset period can be once per second. The interface can be the health interface of the inference process. If some interfaces are abnormal, monitoring needs to be carried out through the timer. When the interface access status is abnormal, the restart operation of the inference process can be immediately executed. During the execution of the restart operation, the inference process cannot process inference requests, so all inference requests can be degraded to the standby process for execution. The status of the inference process is changed to the third state STARTING, and the restart time is recorded. The restart time can be the time when the status is changed to the third state. The system spins and waits for the inference process to restart. If the difference between the current time and the restart time is greater than the time threshold, the inference process can be restarted again, that is, when the time of the current round of restart operation is greater than the time threshold, it can be directly switched to the next round of restart operation, and the above steps are looped until the inference process restarts successfully. When the inference process restarts successfully or the interface access status is normal interface access, it can be considered that the status of the inference process is the fourth state LIVE, and in this case, the inference operation can be continued.

[0099] Figure 8 The flowchart of health detection of the inference process is schematically shown in Figure 8 As shown in, it mainly includes the following steps:

[0100] Step S802, probe the alive timer.

[0101] Step S804, determine whether the RPC of the inference process has abnormal access. If not, go to step S806; if so, go to step S808.

[0102] Step S806, determine the status of the inference process; if it is the fourth state, go to step S810; if not, go to step S808.

[0103] Step S808, execute the actions related to the restart operation of the inference process, and go to step S812.

[0104] Step S810, the loading situation of the RPC model, and go to step S822.

[0105] Step S812: Change the status of the inference process to the third status and record the restart time.

[0106] Step S814: Restart the inference process.

[0107] Step S816: Spin and wait for whether the inference process restarts successfully. If yes, go to Step S818. If no, go to Step S820.

[0108] Step S818: Set the status to the fourth status Live.

[0109] Step S820: Determine whether the difference between the current time and the restart time exceeds the difference threshold. If it exceeds, trigger the restart operation again and return to Step S808 to continue execution.

[0110] Step S822: Perform the model loading action. Specifically, obtain the model resources according to the type. For the algorithm, write the model resources required by the algorithm into model_repostor. For the model, regularly detect the mechanism of reloading to solve the problem of model resources.

[0111] Based on the above health status detection of the inference process, it can detect the abnormal status of the inference process in time and perform the restart operation on the inference process in time to quickly discover the abnormal termination of the inference process and quickly recover, avoid the impact on the model loading action, and ensure the realization of the inference operation.

[0112] In the embodiments of the present disclosure, based on installing a GPU acceleration card inside the device and deploying an inference server, it is possible to determine the target shared memory based on the inference process and the service process through the shared memory pool, and then store the data corresponding to the inference request in the target shared memory, and store the inference result obtained from the inference request in the target shared memory. Through cross-process inference within the container and data transmission based on shared memory, it is possible to reduce the performance overhead caused by data copying, establish inference capabilities locally, reduce the time-consuming of network requests, improve the data transmission efficiency during the inference request process, avoid timeout problems, and thus improve the inference efficiency, inference performance, and accuracy.

[0113] In some embodiments, there is also provided an inference device, which is configured with an inference server and a GPU in a device for inference operations. The inference device includes:

[0114] A shared memory determination module, configured to start a service process and start the inference process represented by the inference server, and determine the target shared memory from the shared resource pool inside the device when the inference server has inference capabilities;

[0115] An inference request transmission module, configured to construct an inference request, obtain data corresponding to the inference request based on a service process, and store the data corresponding to the inference request in the target shared memory;

[0116] An inference result saving module, configured to call an inference process corresponding to a GPU to obtain data corresponding to the inference request from the target shared memory for performing an inference operation to obtain an inference result, and store the inference result in the target shared memory.

[0117] In an exemplary embodiment of the present disclosure, determining the target shared memory from a shared memory pool inside the slave device includes: obtaining a memory identifier from the shared memory pool inside the slave device, and determining the target shared memory according to the memory identifier; the memory identifier is determined according to positions corresponding to shared memories of multiple sub-blocks obtained by splitting the shared memory; wherein, the application process of the shared memory pool includes: after starting an inference process, applying for a shared memory pool according to a configuration file of a model applied by the inference request.

[0118] In an exemplary embodiment of the present disclosure, applying for a shared memory pool according to a configuration file of a model applied by the inference request includes: determining a memory size of the shared memory in the shared memory pool according to the volume of a single request of the model, the maximum number of concurrent requests supported by the model, and a reserved buffer space, and applying for the shared memory pool according to the memory size.

[0119] In an exemplary embodiment of the present disclosure, starting the inference process represented by the inference server includes: after successfully starting the inference process by calling a start script of the inference process, writing a process identifier of the inference process into a specified file, and configuring a state of the inference process to a first state; listening to a semaphore of the inference process through a process listening function, and configuring the state of the inference process to a second state according to the semaphore or the specified file; when the state of the inference process is configured to the second state, restarting the inference process.

[0120] In an exemplary embodiment of the present disclosure, configuring the state of the inference process to the second state according to the semaphore or the specified file includes: if the semaphore terminates or the process identifier in the specified file does not exist, configuring the state of the inference process to the second state.

[0121] In an exemplary embodiment of the present disclosure, before starting the inference process, the method further includes: when it is detected that there is an un-terminated inference process on the device, terminating the un-terminated inference process, and clearing the shared memory allocated on the device.

[0122] In an exemplary embodiment of the present disclosure, the inference process corresponding to the GPU calls to obtain the data corresponding to the inference request from the target shared memory to perform an inference operation to obtain an inference result, including: calling the inference process to obtain the data corresponding to the inference request from the target shared memory; monitoring the state of the inference process, and based on the data corresponding to the inference request, performing an inference operation when the state of the inference process is the third state.

[0123] In an exemplary embodiment of the present disclosure, the monitoring of the state of the inference process and performing an inference operation when the state of the inference process is the third state includes: accessing the interface of the inference process through a timer and determining the interface access status; if the interface access status is an interface access exception, performing a restart operation on the inference process; changing the state of the inference process to the third state and recording the restart time; spin-waiting for the inference process to restart, and when the time difference between the current time and the restart time is greater than the time threshold, performing a re-restart operation on the inference process until the inference process restarts successfully; when the inference process restarts successfully or the interface access status is normal interface access, performing a loading operation on the model to perform an inference operation.

[0124] In an exemplary embodiment of the present disclosure, the method further includes: determining, through an inference process control module, an interface for the inference request to interact with the inference process, determining a model usage method, and defining a model processing method; wherein, the determining, through the inference process control module, an interface for the inference request to interact with the inference process, determining a model usage method, and defining a model processing method includes: defining, through an interface layer, an interface for the inference request to interact with the inference process; determining, through an abstraction layer, one or more of a model loading method, an unloading method, and an inference method; and determining, based on an execution layer, a model acquisition location and a processing method for input parameters in the inference request.

[0125] In an exemplary embodiment of the present disclosure, the method further includes: monitoring the inference performance of the inference request, the system state of the GPU, the capacity of the shared memory pool, and abnormal situations; wherein, the monitoring of the inference performance of the inference request, the system state of the GPU, the capacity of the shared memory pool, and abnormal situations includes: monitoring at least one of the queries per second rate and the inference speed of the inference request; monitoring one or more of the usage rate, utilization rate, power consumption, and main frequency of the GPU; monitoring whether the capacity of the shared memory pool exceeds a water level line; and monitoring one or more of process exceptions, inference exceptions, shared memory capacity exceptions, machine performance exceptions, and process state exceptions.

[0126] It should be noted that the specific details of each part in the above live interaction device and live interaction system have been described in detail in the corresponding method embodiments. For the details not disclosed, reference can be made to the method embodiments, so they will not be elaborated here.

[0127] The exemplary embodiments of the present disclosure also provide a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, it implements the above method.

[0128] In one embodiment, the computer program product can be a tangible product containing a computer program, such as a computer-readable storage medium storing the computer program. The readable storage medium can be a storage medium based on signals such as electricity, magnetism, light, electromagnetic, infrared, etc., including but not limited to: random access memory (RAM), read-only memory (ROM), magnetic tape, floppy disk, flash memory (Flash), hard disk drive (HDD), solid state drive (SSD), etc. Exemplarily, the computer program product can be implemented as a non-volatile storage medium storing the computer program, such as read-only memory, Nand Flash, etc.

[0129] In one embodiment, the computer program product can be an intangible product containing a computer program. Exemplarily, the computer program product can be implemented as a virtual digital product, such as an executable file storing the computer program, an installation package and other digital files.

[0130] The code of the computer program can be written in one or more programming languages. Programming languages such as C language, Java, C++, etc. The program code can be executed entirely on the user's computing device, or partially on the user's computing device, or executed as an independent software package, or partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user's computing device through any type of network, such as a local area network (LAN), a wide area network (WAN), etc., or can be connected to an external computing device (for example, through an Internet connection provided by an operator).

[0131] The computer program can be carried or transmitted by signals such as electricity, magnetism, light, electromagnetic, infrared, etc. The electronic device can convert the signal carrying the computer program into a digital signal and then run the computer program. When the computer program runs on the electronic device, its code is used to cause the electronic device to execute (more specifically, can cause the processor of the electronic device to execute) the method steps of various exemplary embodiments of the present disclosure, such as the above method can be executed.

[0132] Exemplary embodiments of the present disclosure also provide an electronic device, such as the above terminal device or server. The electronic device may include a processor and a memory. The memory stores executable instructions of the processor, such as a computer program. The processor executes the method steps of various exemplary embodiments of the present disclosure by executing the executable instructions. In addition, the electronic device may further include a display for displaying a graphical user interface.

[0133] Next, the electronic device will be exemplarily described in the form of a general computing device. It should be understood that the electronic device here is only an example and should not limit the functions and usage scope of the embodiments of the present disclosure.

[0134] The electronic device may include: a processor, a memory, a bus, an I / O (input / output) interface, a network adapter, and a display.

[0135] The memory may include volatile memory, such as RAM and cache units, and may also include non-volatile memory, such as ROM. The memory may further include one or more program modules, and such program modules include but are not limited to: an operating system, one or more application programs, other program modules, and program data. Implementations of a network environment may be included in each or some combination of these examples. For example, the program modules may include each module in the above device.

[0136] The processor may include one or more processing units. For example, the processor may include an AP (Application Processor), a modem processor, a GPU (Graphics Processing Unit), an ISP (Image Signal Processor), a controller, an encoder, a decoder, a DSP (Digital Signal Processor), a baseband processor, and / or an NPU (Neural-Network Processing Unit), etc.

[0137] The processor can be used to execute the executable instructions stored in the memory, such as executing the above method.

[0138] The bus is used to implement connections between different components of the electronic device and may include a data bus, an address bus, and a control bus.

[0139] The electronic device can communicate with one or more external devices (such as a keyboard, a mouse, an external controller, etc.) through the I / O interface.

[0140] An electronic device can communicate with one or more networks through a network adapter. For example, the network adapter can provide mobile communication solutions such as 3G / 4G / 5G, or wireless communication solutions such as wireless local area network, Bluetooth, near field communication, etc. The network adapter can communicate with other modules of the electronic device through a bus.

[0141] The electronic device can display a graphical user interface through a display.

[0142] In addition, other hardware and / or software modules can be set in the electronic device, including but not limited to: microcode, device driver, redundant processor, external disk drive array, RAID system, tape drive, and data backup storage system, etc.

[0143] As can be seen from the above, the technical solution of the present disclosure can be implemented as a method, a device, a system, a computer program product, a storage medium, an electronic device, etc. Those skilled in the art can understand that various aspects of the present disclosure can be specifically implemented in the following forms, namely: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or an implementation combining hardware and software aspects, such as can be respectively referred to as "circuit", "module" or "system".

[0144] It should be understood that the present disclosure is not limited to the specific method steps or structures already described and shown in the drawings, and various modifications and changes can be made without departing from its scope. Based on the specific implementation manners provided by the present disclosure, those skilled in the art will easily think of other implementation manners. Therefore, the specific implementation manners provided by the present disclosure are only exemplary, and the scope and spirit of the present disclosure are pointed out by the claims, and should cover any variations, uses or adaptations of the present disclosure, which follow the general principles of the present disclosure and include well-known common general knowledge or conventional technical means in the technical field not disclosed by the present disclosure.

[0145] Through the description of the above embodiments, those skilled in the art can easily understand that the exemplary embodiments described here can be implemented by software, or by a combination of software and necessary hardware. Therefore, the technical solution according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (which can be a personal computer, a server, a terminal device, or a network device, etc.) to execute the method according to the embodiments of the present disclosure.

[0146] In addition, the above-mentioned drawings are only schematic illustrations of the processes included in the method according to the exemplary embodiments of the present disclosure, rather than for limiting purposes. It is easy to understand that the processes shown in the above-mentioned drawings do not indicate or limit the chronological order of these processes. Additionally, it is also easy to understand that these processes can be executed synchronously or asynchronously in, for example, multiple modules.

[0147] It should be noted that although several modules or units of the device for action execution are mentioned in the above detailed description, such a division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of two or more of the above-mentioned modules or units can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.

[0148] Those skilled in the art will readily conceive of other embodiments of the present disclosure after considering the specification and practicing the content disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include well-known knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and embodiments are only regarded as exemplary, and the true scope and spirit of the present disclosure are pointed out by the claims.

[0149] It should be understood that the present disclosure is not limited to the exact structures already described and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.

Claims

1. A reasoning method, characterized in that: The device applied to the inference operation is configured with an inference server and a GPU, and the inference method includes: Starting a business process and starting an inference process represented by the inference server, and determining a target shared memory from a shared resource pool inside a device when the inference server has inference capability; Construct an inference request, obtain data corresponding to the inference request based on the business process, and store the data corresponding to the inference request to the target shared memory; The inference process corresponding to the GPU is called to obtain data corresponding to the inference request from the target shared memory to perform an inference operation to obtain an inference result, and the inference result is stored in the target shared memory.

2. The inference method according to claim 1, characterized in that: Determining the target shared memory from the shared memory pool inside the device includes: Obtaining a memory identifier from a shared memory pool inside a device, and determining the target shared memory according to the memory identifier; the memory identifier is determined according to positions corresponding to the shared memories of the multiple sub-blocks obtained by splitting the shared memory; The application process of the shared memory pool includes: After starting the inference process, apply for a shared memory pool based on the configuration file of the model applied to the inference request.

3. The inference method according to claim 2, characterized in that: The applying for a shared memory pool according to the configuration file of the model applied by the inference request includes: The memory size of the shared memory in the shared memory pool is determined according to the volume of a single request of the model, the maximum number of concurrent requests supported by the model, and the reserved buffer space, and the shared memory pool is applied for according to the memory size.

4. The inference method according to claim 1, characterized in that: The starting of the inference process represented by the inference server includes: Calling a startup script of an inference process to start the inference process, and after the inference process is successfully started, writing a process identifier of the inference process into a designated file, and configuring a state of the inference process to be a first state; Monitoring the semaphore of the reasoning process through a process monitoring function, and configuring the state of the reasoning process to a second state according to the semaphore or a specified file; When the state of the reasoning process is configured as the second state, restarting the reasoning process.

5. The inference method according to claim 4, characterized in that: The configuring the state of the reasoning process to the second state according to the semaphore or the specified file includes: If the semaphore is terminated or the process identifier in the designated file does not exist, the state of the reasoning process is configured to be the second state.

6. The inference method according to claim 1, characterized in that: Before starting the reasoning process, the method further includes: When it is detected that there is an unterminated inference process on the device, the unterminated inference process is terminated, and the shared memory allocated on the device is cleared.

7. The inference method according to claim 1, characterized in that: The inference process corresponding to the calling GPU obtains data corresponding to the inference request from the target shared memory to perform an inference operation to obtain an inference result, including: Calling the inference process to obtain data corresponding to the inference request from the target shared memory; The state of the reasoning process is monitored, and when the state of the reasoning process is a third state, a reasoning operation is performed based on the data corresponding to the reasoning request.

8. The inference method according to claim 7, characterized in that: The monitoring of the state of the reasoning process and performing the reasoning operation when the state of the reasoning process is the third state includes: Access the interface of the reasoning process through the timer and determine the interface access status; If the interface access status is interface access abnormality, executing a restart operation of the reasoning process; Changing the state of the reasoning process to a third state and recording a restart time; Spin and wait for the reasoning process to restart. When the time difference between the current time and the restart time is greater than the time threshold, restart the reasoning process again until the reasoning process is successfully restarted. When the inference process is successfully restarted or the interface access status is that the interface access is normal, the model is loaded to perform an inference operation.

9. The inference method according to claim 1, characterized in that: The method further comprises: Determine the interface for interaction between the inference request and the inference process, determine the model usage method, and define the model processing method through the inference process control module; The method of determining the interface for interaction between the inference request and the inference process, determining the model usage method, and defining the model processing method through the inference process control module includes: Define an interface for the inference request to interact with the inference process through an interface layer; Determine one or more of a loading method, an unloading method, and an inference method of the model through an abstraction layer; Determines where the model is obtained and how input parameters in inference requests are handled based on the execution layer.

10. The inference method according to claim 1, characterized in that: The method further comprises: Monitoring the inference performance of the inference request, the system status of the GPU, the capacity of the shared memory pool, and abnormal conditions; The monitoring of the inference performance of the inference request, the system status of the GPU, the capacity of the shared memory pool, and abnormal conditions includes: monitoring at least one of a query rate per second and an inference speed of inference requests; Monitoring one or more of the usage rate, utilization rate, power consumption, and main frequency of the GPU; Monitoring whether the capacity of the shared memory pool exceeds a watermark; One or more of process anomalies, inference anomalies, shared memory capacity anomalies, machine performance anomalies, and process status anomalies are monitored.