Request processing method and device, electronic equipment and storage medium

By using the first and second processing modes to process stability and high-performance requests respectively in the large model, the problem of low computing resource utilization is solved, and efficient and accurate request processing is achieved.

CN120386633APending Publication Date: 2025-07-29KUNLUNXIN TECHNOLOGY (BEIJING) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510518242.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

When large models deal with stability and high-performance inference requests, the prior art is difficult to efficiently utilize computing resources, resulting in performance degradation and increased costs.

Method used

The first processing mode and the second processing mode are respectively used to process the pending requests. The first mode ensures the same result, and the second mode introduces randomness to obtain different results, and efficient processing of stability and high-performance requests is achieved through different configurations of the hardware device.

Benefits of technology

It improves the utilization rate of computing resources, achieves the accuracy of different types of requests, and ensures the consistency of user experience and computing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120386633A_ABST
    Figure CN120386633A_ABST
Patent Text Reader

Abstract

The invention provides a request processing method, and relates to the technical field of artificial intelligence, in particular to the technical field of large models and the technical field of distributed computing. According to the specific implementation scheme, a plurality of to-be-processed requests are processed according to at least one processing mode, respective processing results of the plurality of to-be-processed requests are obtained, and the at least one processing mode comprises a first processing mode and a second processing mode; the first processing mode is a mode for obtaining the same processing result under the condition that the to-be-processed request is processed for multiple times, and the second processing mode is a mode for obtaining different processing results under the condition that the to-be-processed request is processed for multiple times; and outputting respective processing results of the plurality of to-be-processed requests. The invention further provides a request processing device, electronic equipment and a storage medium.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, and particularly to the fields of large model technology and distributed computing technology. More specifically, the present disclosure provides a request processing method, apparatus, electronic device, and storage medium. Background Art

[0002] With the development of artificial intelligence technology, the applications of large models are constantly increasing. A large model can be a large language model (LLM). In the inference process of a large model, randomness factors can be introduced to generate rich and diverse inference results. Summary of the Invention

[0003] The present disclosure provides a request processing method, apparatus, device, and storage medium.

[0004] According to one aspect of the present disclosure, there is provided a request processing method, the method comprising: processing a plurality of requests to be processed according to at least one processing mode to obtain respective processing results of the plurality of requests to be processed, wherein the at least one processing mode includes a first processing mode and a second processing mode, the first processing mode being a mode of obtaining the same processing result when the request to be processed is processed multiple times, and the second processing mode being a mode of obtaining different processing results when the request to be processed is processed multiple times; outputting the respective processing results of the plurality of requests to be processed.

[0005] According to another aspect of the present disclosure, there is provided a request processing apparatus, the apparatus comprising: a processing module, configured to process a plurality of requests to be processed according to at least one processing mode to obtain respective processing results of the plurality of requests to be processed, wherein the at least one processing mode includes a first processing mode and a second processing mode, the first processing mode being a mode of obtaining the same processing result when the request to be processed is processed multiple times, and the second processing mode being a mode of obtaining different processing results when the request to be processed is processed multiple times; and an output module, configured to output the respective processing results of the plurality of requests to be processed.

[0006] According to another aspect of the present disclosure, there is provided an electronic device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method provided by the present disclosure.

[0007] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions for causing a computer to execute the method provided by the present disclosure.

[0008] According to another aspect of the present disclosure, there is provided a computer program product including a computer program which, when executed by a processor, implements the method provided according to the present disclosure.

[0009] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understandable through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:

[0011] Figure 1 is a schematic diagram of an exemplary system architecture to which the request processing method and apparatus according to an embodiment of the present disclosure can be applied;

[0012] Figure 2 is a flowchart of a request processing method according to an embodiment of the present disclosure;

[0013] Figure 3 is a schematic diagram of a request processing method according to an embodiment of the present disclosure;

[0014] Figure 4 is a block diagram of a request processing apparatus according to an embodiment of the present disclosure; and

[0015] Figure 5 is a block diagram of an electronic device to which the request processing method according to an embodiment of the present disclosure can be applied. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0016] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to assist in understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, descriptions of well-known functions and structures are omitted below for clarity and conciseness.

[0017] Taking the generative large model as an example, the generative large model can adopt various sampling methods to achieve rich and diverse inference results. The various sampling methods can include nucleus sampling (top-p). However, in scenarios such as mathematical calculations, in the case of performing multiple inferences on a calculation request, the multiple inference results should be the same. Such a request that needs to obtain multiple exactly consistent inference results can be a stability inference request.

[0018] To process stability inference requests, the inference engine can adopt a stability processing mode for inference. However, the stability processing mode is not suitable for processing non-stability inference requests that require the introduction of randomness. In the case where there are few stability inference requests, adopting the stability processing mode will lead to a decline in the performance of the large model and an increase in inference costs. It can be understood that non-stability inference requests include nucleus sampling parameters, which require the introduction of random factors during the inference process and require high computing resources. Therefore, non-stability inference requests can also be referred to as high-performance inference requests, and non-stability inference server clusters can also be referred to as high-performance inference server clusters.

[0019] Therefore, in order to make full use of the computing resources of the server cluster, the present disclosure provides a request processing method, which will be described below.

[0020] Figure 1 Schematically shows an exemplary system architecture to which the request processing method and apparatus according to an embodiment of the present disclosure can be applied.

[0021] It should be noted that Figure 1 The illustration is only an example of the system architecture to which the embodiments of the present disclosure can be applied, to help those skilled in the art understand the technical content of the present disclosure, but it does not mean that the embodiments of the present disclosure cannot be used in other devices, systems, environments or scenarios.

[0022] As Figure 1 shown, the system architecture 100 according to this embodiment may include terminal devices 1011,... terminal devices 1012, a network 102, and a server cluster 103. The network 102 is used to provide a medium for communication links between the terminal devices 1011,... terminal devices 1012 and the server cluster 103. The network 102 can also be used to provide a medium for communication links within the server cluster 103. The network 102 can include various connection types, such as wired and / or wireless communication links, etc.

[0023] Multiple users can use multiple terminal devices to interact with the server cluster 103 through the network 102 to receive or send messages, etc. For example, the terminal device 1011 can send an inference request to the server cluster 103 through the network 102. The terminal device 1012 can send an inference request to the server cluster 103 through the network 102.

[0024] On at least one of the terminal devices 1011,... terminal devices 1012, various communication client applications can be installed, such as knowledge reading applications, web browser applications, search applications, instant messaging tools, email clients, and / or social platform software, etc. (only as examples). The terminal device can be various electronic devices with a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop portable computers, and desktop computers, etc.

[0025] The server cluster 103 can be a server that provides various services. For example, an inference engine can be deployed in the server cluster 103 to process inference requests provided by users.

[0026] The server cluster 103 can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services ("Virtual Private Server", or simply "VPS" for short). The server can also be a server of a distributed system.

[0027] The request processing method can be applied to the server cluster 103. The server cluster 103 includes multiple server nodes. The multiple server nodes can include server node 1031, server node 1032, server node 1033, and server node 1034. Multiple requests can be processed respectively by using the server cluster 103.

[0028] It should be understood that Figure 1 the numbers of the terminal devices, the network, and the server nodes in

[0029] are merely illustrative. According to the implementation requirements, there can be any number of terminal devices, networks, and server nodes.

[0030] Figure 2 is a flowchart of a request processing method according to an embodiment of the present disclosure.

[0031] As Figure 2 shown, the method 200 can include operation S210 to operation S220.

[0032] In operation S210, process multiple requests to be processed according to at least one processing mode, and obtain the processing results of each of the multiple requests to be processed.

[0033] In the embodiments of the present disclosure, the request to be processed can be an inference request provided by a user via the above terminal device. For example, the user can issue an inference request through a prompt text based on natural language.

[0034] In the embodiments of the present disclosure, at least one processing mode may include a first processing mode and a second processing mode. The first processing mode may be a mode in which the same processing result is obtained when a to-be-processed request is processed multiple times. The second processing mode may be a mode in which different processing results are obtained when a to-be-processed request is processed multiple times. The above server cluster 103 may process multiple to-be-processed requests respectively based on different processing modes. For example, based on the first processing mode, the above stability inference request may be processed. For another example, based on the second processing mode, the above high-performance inference request may be processed.

[0035] In operation S220, the processing results of multiple to-be-processed requests are output.

[0036] For example, after the server cluster 103 processes multiple to-be-processed requests based on multiple processing modes, the multiple processing results may be returned to multiple terminal devices via the above network 102.

[0037] Through the embodiments of the present disclosure, the stability inference request and the high-performance inference request can be processed based on different processing modes, the efficient processing of the high-performance inference request can be achieved, the processing result of the high-performance inference request with higher accuracy and more precision can be obtained, and the accurate processing of the stability inference request can also be achieved. The processing precision of different types of requests can be improved while the utilization rate of computing resources is increased.

[0038] It can be understood that the above description of the present disclosure is given by taking the to-be-processed request as an inference request as an example. However, the present disclosure is not limited thereto, and the to-be-processed request may also be a training request.

[0039] It can be understood that the method of the present disclosure has been described above, and the method of the present disclosure will be further described below.

[0040] In some embodiments, in some implementation manners of the above operation S210, processing multiple to-be-processed requests according to at least one processing mode to obtain the processing results of the multiple to-be-processed requests respectively includes: using multiple hardware devices to process the multiple to-be-processed requests according to at least one processing mode to obtain the processing results of the multiple to-be-processed requests respectively. For example, the hardware device may be the above server node. The processing mode of the hardware device may be the first processing mode or the second processing mode.

[0041] In some embodiments, multiple hardware devices may be deployed with the same inference engine or the same training engine. For example, before starting inference, the hardware device may be set to a first processing mode or a second processing mode. Through the embodiments of the present disclosure, in the case of deploying the same inference engine to multiple hardware devices, multiple hardware devices in different processing modes can be used to process stable inference requests and high-performance inference requests, which can fully improve the utilization rate of hardware resources. It can be understood that the descriptions of stable training requests and high-performance training requests are similar to those of stable inference requests and high-performance inference requests, and the present disclosure will not elaborate herein.

[0042] It can be understood that the above describes multiple hardware devices of the present disclosure, and the following will further describe the first processing mode and the second processing mode of the present disclosure.

[0043] Figure 3 It is a schematic diagram of a request processing method according to an embodiment of the present disclosure.

[0044] As Figure 3 shown, via network 302, multiple requests of the same input batch can be obtained as multiple requests to be processed. Next, the multiple requests to be processed can be input into the inference engine 30 deployed on multiple hardware devices. Next, the hardware device for processing each request to be processed can be determined.

[0045] In operation S311, it is determined whether the random processing parameter of the request to be processed is a preset value.

[0046] For example, the random processing parameter may be a nucleus sampling parameter. The preset value may be 0. It can be understood that the random processing parameter can represent the degree of randomness. In other embodiments, the random processing parameter may also be a temperature parameter.

[0047] In some embodiments, in response to determining that the random sampling parameter of the request to be processed is a preset value, the request to be processed can be used as a first request to be processed. Next, after using at least one request to be processed with a preset random processing parameter among the multiple requests to be processed as at least one first request to be processed, operation S312 is executed.

[0048] In operation S312, at least one first request to be processed is provided to at least one first hardware device.

[0049] In some embodiments, the first hardware device may be a hardware device in the first processing mode. The number of first hardware devices may be different from the number of first requests to be processed. For example, the above server nodes 1031 and 1032 may be in the first processing mode and serve as the first hardware devices.

[0050] In operation S313, at least one first processing result is obtained by processing at least one first request to be processed according to a first processing mode using at least one first hardware device.

[0051] For example, the first request to be processed may be "Please calculate the algebraic sum of 1 to 100". According to the first processing mode, 1, 2,..., 100 can be added in sequence to obtain the first processing result (5050). It can be understood that in the case of sequential addition, almost no randomness is introduced, and stable inference can be achieved. The above description of the stable inference request also applies to the first request for inference to be processed, and details are not repeated here. Through the embodiments of the present disclosure, based on the random sampling parameter, it can be quickly determined whether the request to be processed is a stable processing request, and the request can be efficiently provided to the hardware device for inference, improving the request processing efficiency.

[0052] It can be understood that the present disclosure has been described above in conjunction with the first processing mode. The following will be described in conjunction with the second processing mode.

[0053] In some embodiments, in response to determining that the random sampling parameter of the request to be processed is not a preset value, the request to be processed can be used as a second request to be processed. Next, after using at least one request to be processed with a non-preset random processing parameter among multiple requests to be processed as at least one second request to be processed, operation S314 is executed.

[0054] In operation S314, at least one second request to be processed is provided to at least one second hardware device.

[0055] In some embodiments, the second hardware device may be a hardware device in the second processing mode. The number of second hardware devices may be different from the number of second requests to be processed. For example, the above server nodes 1033 and 1034 may be in the second processing mode and serve as second hardware devices.

[0056] In operation S315, at least one second processing result is obtained by processing at least one second request to be processed according to the second processing mode using at least one second hardware device.

[0057] For example, a request to be processed may be "Write a poem related to the moon". The request to be processed may have a nucleus sampling parameter, and the nucleus sampling parameter is not 0 and is greater than 0.8, indicating that after inferring multiple initial tokens, multiple tokens with a cumulative probability exceeding 0.8 are used as target tokens to generate a second processing result. Each token corresponds to a probability. The probability may be a value greater than 0 and less than 1. It can be understood that when the nucleus sampling parameter is greater than 0, randomness can be introduced, and multiple different results can be obtained by inferring the same request to be processed multiple times.

[0058] It can be understood that some ways of obtaining multiple processing results have been described above, and some ways of outputting multiple processing results will be described below.

[0059] In operation S320, at least one first processing result and at least one second processing result are output in the same output batch corresponding to the input batch.

[0060] For example, the first processing result and the second processing result are output in the same batch. Thus, different users can obtain results based on the same or similar latency, which helps to improve the experience of the user group.

[0061] It can be understood that the method of the present disclosure has been further described above, and the first hardware device and the second hardware device of the present disclosure will be further described below.

[0062] In some embodiments, according to the number of first pending requests of multiple previous batches corresponding to the input batch, the average value of the numbers is determined. According to the average value of the numbers, the target number of the first hardware device is determined. For example, if the current number of the first hardware device is less than the target number, the processing mode of one or more second hardware devices can be switched to the first processing mode to increase the number of the first hardware device. For another example, if the current number of the first hardware device is greater than the target number, the processing mode of one or more first hardware devices can be switched to the second processing mode to increase the number of the second hardware device. Thus, the hardware resources of multiple hardware devices deploying the same inference engine can be fully utilized to effectively process different types of pending requests.

[0063] It can be understood that the method of the present disclosure has been described above, and the device of the present disclosure will be described below.

[0064] Figure 4 It is a schematic block diagram of a request processing device according to an embodiment of the present disclosure.

[0065] As Figure 4 shown, the device 400 may include a processing module 410 and an output module 420.

[0066] The processing module 410 is configured to process multiple pending requests according to at least one processing mode to obtain processing results of the multiple pending requests respectively.

[0067] In the embodiments of the present disclosure, at least one processing mode includes a first processing mode and a second processing mode. The first processing mode is a mode in which the same processing result is obtained when the pending request is processed multiple times, and the second processing mode is a mode in which different processing results are obtained when the pending request is processed multiple times.

[0068] An output module 420 for outputting the processing results of multiple requests to be processed respectively.

[0069] In some embodiments, the processing module includes: a processing sub-module for processing multiple requests to be processed according to at least one processing mode by using multiple hardware devices to obtain the processing results of the multiple requests to be processed respectively. The same inference engine or the same training engine is deployed on the multiple hardware devices.

[0070] In some embodiments, the multiple processing results include at least one first processing result obtained by a first processing mode, the multiple hardware devices include at least one first hardware device in the first processing mode, and the processing sub-module includes: a first providing unit for providing the requests to be processed with at least one random processing parameter being a preset value as at least one first request to be processed to at least one first hardware device. A first processing unit for processing at least one first request to be processed by using at least one first hardware device according to the first processing mode to obtain at least one first processing result.

[0071] In some embodiments, the multiple processing results include at least one second processing result obtained by a second processing mode, the multiple hardware devices include at least one second hardware device in the second processing mode. The processing sub-module includes: a second providing unit for providing the requests to be processed with at least one random processing parameter not being a preset value as at least one second request to be processed to at least one second hardware device. A second processing unit for processing at least one request to be processed by using at least one second hardware device according to the second processing mode to obtain at least one processing result.

[0072] In some embodiments, the multiple requests to be processed are multiple requests of the same input batch, and the multiple processing results include at least one first processing result obtained by a first processing mode and at least one second processing result obtained by a second processing mode. The output module is further configured to output at least one first processing result and at least one second processing result in the same output batch corresponding to the input batch.

[0073] In the technical solution of the present disclosure, the processing of the collection, storage, use, processing, transmission, provision, and disclosure of the user's personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0074] According to the embodiments of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0075] In some embodiments, the electronic device provided by the present disclosure may include: at least one processor; and a memory communicatively connected to the at least one processor. The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the above-mentioned method 200.

[0076] In some embodiments, the present disclosure provides a non-transitory computer-readable storage medium storing computer instructions. The computer instructions are used to cause a computer to execute the above-mentioned method 200.

[0077] In some embodiments, the present disclosure provides a computer program product, including a computer program, which implements the above-mentioned method 200 when executed by a processor. The following will be further described in conjunction with Figure 5 for further illustration.

[0078] Figure 5 FIG. shows a schematic block diagram of an exemplary electronic device 500 that may be used to implement embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, for example, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as, for example, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0079] As Figure 5 shown, the device 500 includes a computing unit 501, which may execute various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 502 or a computer program loaded from a storage unit 508 into a random access memory (RAM) 503. In the RAM 503, various programs and data required for the operation of the device 500 may also be stored. The computing unit 501, the ROM 502, and the RAM 503 are connected to each other via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.

[0080] Multiple components in device 500 are connected to I / O interface 505, including: an input unit 506, such as a keyboard, a mouse, etc.; an output unit 507, such as various types of displays, speakers, etc.; a storage unit 508, such as a disk, an optical disc, etc.; and a communication unit 509, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 509 allows device 500 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0081] The computing unit 501 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 501 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 501 executes the various methods and processes described above, such as the request processing method. For example, in some embodiments, the request processing method can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as the storage unit 508. In some embodiments, part or all of the computer program can be loaded and / or installed onto device 500 via the ROM 502 and / or the communication unit 509. When the computer program is loaded into the RAM 503 and executed by the computing unit 501, one or more steps of the request processing method described above can be executed. Alternatively, in other embodiments, the computing unit 501 can be configured to execute the request processing method in any other suitable manner (e.g., by means of firmware).

[0082] The various embodiments of the systems and techniques described above in this document may be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard parts (ASSPs), system on chip (SOC) systems, complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which may be a special-purpose or general-purpose programmable processor that receives data and instructions from, and transmits data and instructions to, a storage system, at least one input device, and at least one output device.

[0083] The program code for implementing the methods of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus, such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The program code may execute entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine or entirely on the remote machine or server.

[0084] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory, a read-only memory, an erasable programmable read-only memory (EPROM) or flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0085] To provide for interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a cathode ray tube (CRT) monitor or a liquid crystal display (LCD)) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic, speech, or tactile input).

[0086] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.

[0087] A computer system may include a client and a server. The client and the server are generally far from each other and typically interact via a communication network. The relationship between the client and the server is generated by computer programs that run on respective computers and have a client-server relationship with each other.

[0088] It should be understood that various forms of the processes shown above may be used, with steps reordered, added, or deleted. For example, the steps recited in this disclosure may be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this is not limited herein.

[0089] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent replacements, and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.

Claims

1. A request processing method, comprising: processing a plurality of requests to be processed according to at least one processing mode to obtain respective processing results of the plurality of requests to be processed, wherein at least one of the processing modes includes a first processing mode and a second processing mode, the first processing mode being a mode in which the same processing result is obtained when the requests to be processed are processed multiple times, and the second processing mode being a mode in which different processing results are obtained when the requests to be processed are processed multiple times; outputting the respective processing results of the plurality of requests to be processed.

2. The method according to claim 1, wherein, The processing a plurality of requests to be processed according to at least one processing mode to obtain respective processing results of the plurality of requests to be processed includes: using a plurality of hardware devices to process the plurality of requests to be processed according to at least one of the processing modes to obtain respective processing results of the plurality of requests to be processed, wherein the plurality of hardware devices are deployed with the same inference engine or the same training engine.

3. The method according to claim 2, wherein, The plurality of processing results include at least one first processing result obtained by the first processing mode, and the plurality of hardware devices include at least one first hardware device in the first processing mode, The using a plurality of hardware devices to process the plurality of requests to be processed according to at least one of the processing modes to obtain respective processing results of the plurality of requests to be processed includes: providing the requests to be processed with at least one random processing parameter being a preset value as at least one first request to be processed to at least one of the first hardware devices; using at least one of the first hardware devices to process at least one of the first requests to be processed according to the first processing mode to obtain at least one first processing result.

4. The method according to claim 2, wherein The plurality of processing results include at least one second processing result obtained by the second processing mode, and the plurality of hardware devices include at least one second hardware device in the second processing mode, The using a plurality of hardware devices to process the plurality of requests to be processed according to at least one of the processing modes to obtain respective processing results of the plurality of requests to be processed includes: providing the requests to be processed with at least one random processing parameter not being a preset value as at least one second request to be processed to at least one second hardware device; using at least one of the second hardware devices to process at least one of the requests to be processed according to the second processing mode to obtain at least one of the processing results.

5. The method according to claim 1, wherein The plurality of requests to be processed are a plurality of requests in the same input batch, and the plurality of processing results include at least one first processing result obtained by the first processing mode and at least one second processing result obtained by the second processing mode, The outputting the respective processing results of the plurality of requests to be processed includes: outputting at least one of the first processing results and at least one of the second processing results in the same output batch corresponding to the input batch.

6. A request processing apparatus, comprising: A processing module, configured to process multiple requests to be processed according to at least one processing mode, and obtain the processing results of the multiple requests to be processed respectively, where at least one of the processing modes includes a first processing mode and a second processing mode, the first processing mode is a mode in which the same processing result is obtained when the requests to be processed are processed multiple times, and the second processing mode is a mode in which different processing results are obtained when the requests to be processed are processed multiple times; An output module, configured to output the processing results of the multiple requests to be processed respectively.

7. The apparatus according to claim 6, wherein, The processing module includes: A processing sub-module, configured to use multiple hardware devices to process the multiple requests to be processed according to at least one of the processing modes, and obtain the processing results of the multiple requests to be processed respectively, where the same inference engine or the same training engine is deployed on the multiple hardware devices.

8. An electronic device, comprising: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 1 to 5.

9. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 5.

10. A computer program product, comprising a computer program which, when executed by a processor, implements the method according to any one of claims 1 to 5.