Task Execution Method, Device, Electronic Device, and Storage Medium

By dynamically adjusting hardware resources to adapt to the accuracy overflow problem of large models, the problem of model debugging is solved, the hardware resource utilization rate and model stability are improved, mixed accuracy calculation is realized, and application scenarios are expanded.

CN117271113BActive Publication Date: 2025-07-11BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311101089.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-29
Publication Date
2025-07-11
Estimated Expiration
2043-08-29

AI Technical Summary

Technical Problem

During the calculation of large-scale model, the accuracy overflow problem makes it difficult to debug the model, and it is difficult to determine that the accuracy overflow caused by multiplication or addition operations, affecting the task execution efficiency.

Method used

By determining whether the simulation output data of the target sub-model is within the numerical range of the first hardware resource, if not within the range, dynamically adjust it to perform tasks with higher accuracy second hardware resources to ensure that the output data is within the appropriate numerical range. Using the combination of the simulation model and different hardware resources, the hardware resources are dynamically configured to meet different accuracy requirements.

Benefits of technology

It improves the utilization rate of hardware resources, ensures the stable operation of the model, expands the application scenarios of the model, realizes mixed accuracy calculation, and reduces resource overhead and accuracy overflow risks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117271113B_ABST
    Figure CN117271113B_ABST
Patent Text Reader

Abstract

The present disclosure provides a task execution method, which relates to the field of artificial intelligence technology, particularly to the fields of large model technology and deep learning technology. The specific implementation solution is as follows: determining the simulation output data of a target sub-model according to the input data of the target sub-model in a deep learning model, where the target sub-model includes at least one target operator, and the target task corresponding to the target sub-model is executed by a first hardware resource; in response to determining that the simulation output data is not within the first numerical range corresponding to the first hardware resource, determining a second hardware resource according to the simulation output data, where the second numerical range corresponding to the second hardware resource is greater than the first numerical range; and using the second hardware resource to execute the target task to process the input data to obtain the target output data of the target sub-model. The present disclosure also provides a task execution device, an electronic device, and a storage medium.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, and more particularly to the fields of large model technology and deep learning technology. More specifically, the present disclosure provides a task execution method, apparatus, electronic device, and storage medium. Background Art

[0002] With the development of artificial intelligence technology, large models are increasingly widely used in the field of deep learning. Large models can be trained or inferred using data with different precisions. Summary of the Invention

[0003] The present disclosure provides a task execution method, apparatus, device, and storage medium.

[0004] According to one aspect of the present disclosure, there is provided a task execution method, the method comprising: determining simulation output data of a target sub-model according to input data of the target sub-model in a deep learning model, wherein the target sub-model comprises at least one target operator, and a target task corresponding to the target sub-model is executed by a first hardware resource; in response to determining that the simulation output data is not within a first numerical range corresponding to the first hardware resource, determining a second hardware resource according to the simulation output data, wherein a second numerical range corresponding to the second hardware resource is greater than the first numerical range; and using the second hardware resource to execute the target task to process the input data to obtain target output data of the target sub-model.

[0005] According to another aspect of the present disclosure, there is provided a task execution apparatus, the apparatus comprising: a first determination module, configured to determine simulation output data of a target sub-model according to input data of the target sub-model in a deep learning model, wherein the target sub-model comprises at least one target operator, and a target task corresponding to the target sub-model is executed by a first hardware resource; a second determination module, configured to, in response to determining that the simulation output data is not within a first numerical range corresponding to the first hardware resource, determine a second hardware resource according to the simulation output data, wherein a second numerical range corresponding to the second hardware resource is greater than the first numerical range; and a first execution module, configured to use the second hardware resource to execute the target task to process the input data to obtain target output data of the target sub-model.

[0006] According to another aspect of the present disclosure, there is provided an electronic device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method provided by the present disclosure.

[0007] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions for causing a computer to execute the method provided according to the present disclosure.

[0008] According to another aspect of the present disclosure, there is provided a computer program product including a computer program which, when executed by a processor, implements the method provided according to the present disclosure.

[0009] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understandable through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:

[0011] Figure 1 is a flowchart of a task execution method according to an embodiment of the present disclosure;

[0012] Figure 2 is a schematic diagram of a deep learning model according to an embodiment of the present disclosure;

[0013] Figure 3 is a block diagram of a task execution device according to an embodiment of the present disclosure; and

[0014] Figure 4 is a block diagram of an electronic device to which the task execution method can be applied according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0015] The following makes an explanation of exemplary embodiments of the present disclosure in conjunction with the drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0016] With the increase in the depth and breadth of the model, great challenges are faced in the accuracy verification and accuracy alignment of large models. When an accuracy overflow problem occurs during the calculation process of a large model, the debugging of the model will become extremely difficult.

[0017] In some embodiments, after the operator calculation of the model is completed, the accuracy of the output result of the operator can be detected to determine whether the output result exceeds the numerical range of the corresponding data type. For example, a Scale operator includes a multiplication operation and an addition operation. The data type corresponding to the Scale operator can be 16-bit floating point number (FP16). The accuracy of the output result of the Scale operator can be detected. If the detection result indicates that the accuracy of the output result is not within the numerical range corresponding to the 16-bit floating point number, it can be determined that an accuracy overflow problem has occurred, and it is difficult to determine whether the accuracy overflow is caused by the multiplication operation or the addition operation, thereby making it difficult to efficiently optimize the execution of related tasks.

[0018] To efficiently execute tasks related to the model, the present disclosure provides a task execution method, which will be further described below.

[0019] Figure 1 It is a flowchart of a task execution method according to an embodiment of the present disclosure.

[0020] As Figure 1 shown, the method 100 may include operation S110 to operation S130.

[0021] In operation S110, according to the input data of the target sub-model in the deep learning model, the simulated output data of the target sub-model is determined.

[0022] In the embodiments of the present disclosure, the target sub-model may include at least one target operator. For example, the target operator may be an operator that may have an accuracy overflow problem. The target operator may include operators such as cumulative multiplication and accumulation. The operators such as cumulative multiplication and accumulation with a dependency relationship may be used as the target sub-model. It can be understood that if the output of the first operator is used as the input of the second operator, the first operator and the second operator have a dependency relationship.

[0023] In the embodiments of the present disclosure, the target task corresponding to the target sub-model can be executed by the first hardware resource. For example, the first hardware resource may include at least one processor core and the corresponding storage space.

[0024] In the embodiments of the present disclosure, various methods can be used to determine the simulated output data of the target sub-model. For example, the historical input data and historical output data of the target sub-model can be obtained. If the difference between the input data of the target sub-model and a piece of historical input data is small, the historical output data corresponding to the historical input data can be used as the simulated output data. It can be understood that before processing the input data of the target sub-model, the previous tasks of the target task can be executed by different hardware resources to process multiple historical input data and obtain multiple historical output data.

[0025] In operation S120, in response to determining that the simulation output data is not within the first numerical range corresponding to the first hardware resource, the second hardware resource is determined according to the simulation output data.

[0026] In the embodiments of the present disclosure, a hardware resource corresponds to a numerical range. For example, the numerical range corresponding to 16-bit floating-point numbers can be used as the numerical range corresponding to the hardware resource. The numerical range corresponding to 32-bit floating-point numbers (FP32) can also be used as the numerical range corresponding to the hardware resource. The hardware resources required for a hardware device to process 32-bit floating-point numbers can be greater than the hardware resources required for the hardware device to process 16-bit floating-point numbers.

[0027] In the embodiments of the present disclosure, the second numerical range corresponding to the second hardware resource is greater than the first numerical range. For example, the first hardware resource can correspond to the numerical range of 16-bit floating-point numbers. The second hardware resource can correspond to the numerical range of 32-bit floating-point numbers.

[0028] In the embodiments of the present disclosure, the second hardware resource can be determined according to the numerical range in which the simulation output data is located.

[0029] In operation S130, the target task is executed using the second hardware resource to process the input data, and the target output data of the target sub-model is obtained.

[0030] In the embodiments of the present disclosure, the input data can be input into the target sub-model to obtain the target output data. For example, the number of processor cores corresponding to the second hardware resource can be greater than the number of processor cores corresponding to the first hardware resource.

[0031] Through the embodiments of the present disclosure, the hardware resources required by the target sub-model can be dynamically configured. Thus, the hardware resources of the hardware device can be fully utilized, and the utilization rate of the hardware resources can be improved while ensuring the stable operation of the model.

[0032] In addition, through the embodiments of the present disclosure, the computing requirements of different precisions can also be met, the application scenarios of the model can be expanded, and it is also helpful to achieve mixed-precision computing.

[0033] It can be understood that the method of the present disclosure has been described above, and the deep learning model and the hardware device of the present disclosure will be further described below.

[0034] In some embodiments, the hardware device can be a hardware device cluster.

[0035] In some embodiments, the deep learning model can include multiple sub-models. The following will be combined with Figure 2 for further description.

[0036] Figure 2 is a schematic diagram of a deep learning model according to an embodiment of the present disclosure.

[0037] As Figure 2 shown, the deep learning model 200 may include a first sub-model 210, a second sub-model 220, and a target sub-model 230.

[0038] In the embodiments of the present disclosure, a hardware device may be divided into multiple hardware channels. The multiple hardware channels correspond to multiple hardware resources. For example, there may be a precision overflow problem with the target operator of the target sub-model. As Figure 2 shown, the target sub-model 230 may include a cumulative multiplication operator 231 and an accumulation operator 232. Also, for example, the hardware channels may include dynamic hardware channels. The target task corresponding to the target sub-model 230 may be run by the hardware resources corresponding to the dynamic hardware channels. The dynamic hardware channels may correspond to multiple hardware resources. The multiple hardware resources may correspond to multiple precisions. The multiple hardware resources may include a first hardware resource. A second hardware resource may also be determined from the multiple hardware resources.

[0039] For example, the first sub-model 210 may be run by a fourth hardware resource, and the second sub-model 220 may be run by a fifth hardware resource. It is difficult for the operators in the first sub-model 210 or the second sub-model 220 to have a precision overflow problem. The first sub-model 210 may include an addition operator 211 and an addition operator 212. The second sub-model 220 may include a multiplication operator 221 and a multiplication operator 222. During the running process of the deep learning model, the fourth hardware resource and the fifth hardware resource may be continuously used to run the first sub-model and the second sub-model respectively.

[0040] It can be understood that the hardware device and the deep learning model of the present disclosure have been described above, and the target task of the present disclosure will be further described below.

[0041] In some embodiments, the target task may include one of a forward calculation task and a backward calculation task. Correspondingly, the input data may include one of forward input data and backward input data. For example, when the input data is forward input data, the output data may be a forward calculation result. Also, for example, when the input data is backward input data, the output data may be gradient data.

[0042] It can be understood that the target task of the present disclosure has been described above, and the simulation model of the present disclosure will be described below.

[0043] In some embodiments, determining the simulation output data of the target sub-model includes: processing the input data using the simulation model corresponding to the target sub-model to obtain the simulation output data.

[0044] In the embodiments of the present disclosure, the simulation model is determined based on multiple historical input data and multiple historical output data of the target sub-model. For example, multiple historical input data of the dynamic hardware channel and the corresponding multiple historical output data can be obtained. According to the multiple historical input data and the multiple historical output data, linear regression is performed to obtain the simulation model. It can be understood that the simulation model can also be obtained according to other models. For example, the historical input data is used as the training sample, and the historical output data is used as the label to train a fully connected network (FCN) as the simulation model. It can also be understood that when the data types of the input data are the same, the hardware resources required by the fully connected network or the linear regression model can be less than the hardware resources required by the multiple target operators.

[0045] In the embodiments of the present disclosure, the simulation task corresponding to the simulation model can be executed using the preset hardware resources. For example, the preset hardware resources can be determined according to the historical output data.

[0046] It can be understood that the simulation model of the present disclosure has been described above, and the simulation output data of the present disclosure will be further described below.

[0047] In some embodiments, at least one target operator can be N target operators. N can be an integer greater than 1. The simulation model can include N simulation sub-models, and each simulation sub-model can correspond to one target operator. As Figure 2 shown, the N target operators can include the above-mentioned multiplication operator 231 and the above-mentioned accumulation operator 232. The N simulation sub-models can include the simulation sub-model corresponding to the multiplication operator 231 and the simulation sub-model corresponding to the accumulation operator 232.

[0048] In some embodiments, determining the simulation output data of the target sub-model can include: using the N simulation sub-models to process the input data to obtain N simulation output sub-data corresponding to the N target operators as the simulation output data.

[0049] In the embodiments of the present disclosure, using the N simulation sub-models to process the input data to obtain N simulation output sub-data corresponding to the N target operators includes: using the first simulation sub-model to process the input data to obtain the first simulation output sub-data. For example, the input data can be processed using the simulation sub-model corresponding to the multiplication operator 231 to obtain the first simulation output sub-data.

[0050] In the embodiments of the present disclosure, processing the input data by using N simulation sub-models to obtain N pieces of simulation output sub-data corresponding to N target operators may further include: processing the (n-1)th piece of simulation output sub-data by using the nth simulation sub-model to obtain the nth piece of simulation output sub-data. n may be an integer greater than 1 and less than or equal to N. Taking n = 2 as an example, the simulation sub-model corresponding to the above-mentioned accumulation operator 232 may process the first piece of simulation output sub-data to obtain the second piece of simulation output sub-data. The first piece of simulation output sub-data and the second piece of simulation output sub-data may be used as simulation output data.

[0051] In the embodiments of the present disclosure, it may be determined whether N pieces of simulation output sub-data are within the first numerical range corresponding to the first hardware resource. For example, taking the numerical range of 16-bit floating-point numbers corresponding to the first hardware resource as an example, if the first piece of simulation output sub-data exceeds the numerical range of 16-bit floating-point numbers, the second hardware resource may be determined from multiple hardware resources corresponding to the dynamic hardware channels according to the first piece of simulation output sub-data. For another example, if it is determined that any piece of simulation output sub-data is within the first numerical range, the first hardware resource may be used to execute the target task. Thereby, the resource overhead can be reduced and the utilization rate of the hardware resource can be improved.

[0052] It can be understood that after determining multiple pieces of simulation output sub-data above, it is determined whether the simulation output sub-data is within the first numerical range. However, the present disclosure is not limited thereto, and it may also be determined whether the simulation output sub-data is within the first numerical range after determining one piece of simulation output sub-data.

[0053] In the embodiments of the present disclosure, in response to determining that the simulation output data is not within the first numerical range corresponding to the first hardware resource, determining the second hardware resource according to the simulation output data includes: in response to determining that the first piece of simulation output sub-data is not within the first numerical range, determining the second hardware resource according to the first piece of simulation output sub-data. For example, after obtaining the first piece of simulation output sub-data, it may be determined whether the first piece of simulation output sub-data is within the first numerical range. If the first piece of simulation output sub-data is not within the first numerical range, the second hardware resource may be determined according to the first piece of simulation output sub-data. It is also possible to stop processing by using the subsequent simulation sub-models to save computing resources and reduce the resource overhead.

[0054] In the embodiments of the present disclosure, in response to determining that the simulation output data is not within the first numerical range corresponding to the first hardware resource, the second hardware resource is determined according to the simulation output data, including: in response to determining that the nth simulation output sub-data is not within the first numerical range, the second hardware resource is determined according to the nth simulation output sub-data. For example, in the case where the first n - 1 simulation output sub-data are within the first numerical range, after obtaining the nth simulation output sub-data, it can be determined whether the nth simulation output sub-data is within the first numerical range. If the nth simulation output sub-data is not within the first numerical range, the second hardware resource can be determined according to the nth simulation output sub-data. When n is less than N, the processing using the subsequent simulation sub-models can also be stopped to save computing resources and reduce resource overhead.

[0055] It can be understood that some methods for determining the second hardware resource are described above. In the embodiments of the present disclosure, in order to obtain the simulation output data quickly and at low cost, the input data is processed using the simulation model. However, since the calculation methods of the simulation model and the target sub-model are different, there may be a large difference between the simulation output data and the real output data, which will be further described below.

[0056] In some embodiments, the above method may further include: in response to determining that the simulation output data is within the first numerical range, the first hardware resource is used to execute the target task to process the input data and obtain the first output data. For example, if the simulation output data is within the first numerical range, the first hardware resource can be used to execute the target task to obtain the first output data. It can be understood that in this case, the first output data can also be used as the target output data. Next, it can be determined whether the difference between the first output data and the simulation output data is greater than or equal to a preset difference threshold.

[0057] In some embodiments, the above method may further include: in response to determining that the difference between the first output data and the simulation output data is greater than or equal to the preset difference threshold, the first hardware resource is adjusted to the third hardware resource. The third numerical range corresponding to the third hardware resource is greater than the first numerical range. For example, if the difference is greater than or equal to the preset difference threshold, it can be determined that the error of the simulation model is large. In order to stably execute the subsequent tasks of the target task, a higher-precision data type corresponding hardware resource can be used to execute the subsequent tasks. Through the embodiments of the present disclosure, the hardware resource can be adjusted in a timely manner, the risk of precision overflow can be reduced, and dynamic precision management can be achieved to adapt to various types of computing requirements.

[0058] For another example, if the difference is less than the preset difference threshold, it can be determined that the error of the simulation model is small, and the first hardware resource can be used to execute the subsequent tasks.

[0059] Figure 3It is a block diagram of a task execution device according to an embodiment of the present disclosure.

[0060] As Figure 3 shown, the device 300 may include a first determination module 310, a second determination module 320, and a first execution module 330.

[0061] The first determination module 310 is configured to determine simulation output data of the target sub-model according to the input data of the target sub-model in the deep learning model. The target sub-model includes at least one target operator. The target task corresponding to the target sub-model is executed by the first hardware resource.

[0062] The second determination module 320 is configured to, in response to determining that the simulation output data is not within the first numerical range corresponding to the first hardware resource, determine a second hardware resource according to the simulation output data. The second numerical range corresponding to the second hardware resource is greater than the first numerical range.

[0063] The first execution module 330 is configured to execute the target task by using the second hardware resource to process the input data to obtain the target output data of the target sub-model.

[0064] In some embodiments, the first determination module includes: a processing sub-module configured to process the input data by using a simulation model corresponding to the target sub-model to obtain simulation output data.

[0065] In some embodiments, the simulation model is determined according to multiple historical input data and multiple historical output data of the target sub-model.

[0066] In some embodiments, there are N target operators, where N is an integer greater than 1. The simulation model includes N simulation sub-models, and each simulation sub-model corresponds to one target operator.

[0067] In some embodiments, the processing sub-module includes: a processing unit configured to process the input data by using the N simulation sub-models to obtain N simulation output sub-data corresponding to the N target operators as the simulation output data.

[0068] In some embodiments, the processing unit includes: a first processing sub-unit configured to process the input data by using the first simulation sub-model to obtain the first simulation output sub-data. A second processing sub-unit configured to process the (n - 1)th simulation output sub-data by using the nth simulation sub-model to obtain the nth simulation output sub-data. n is an integer greater than 1 and less than or equal to N.

[0069] In some embodiments, the second determination module is further configured to: in response to determining that the nth simulation output sub-data is not within the first numerical range, determine a second hardware resource according to the nth simulation output sub-data.

[0070] In some embodiments, the apparatus 300 further includes: a second execution module, configured to execute a target task by using first hardware resources in response to determining that the simulation output data is within a first numerical range, so as to process input data and obtain first output data.

[0071] In some embodiments, the apparatus 300 further includes: an adjustment module, configured to adjust the first hardware resources to third hardware resources in response to determining that the difference between the first output data and the simulation output data is greater than or equal to a preset difference threshold. A third numerical range corresponding to the third hardware resources is greater than the first numerical range.

[0072] In some embodiments, the input data includes one of forward input data and reverse input data.

[0073] In the technical solution of the present disclosure, the processing of collection, storage, use, processing, transmission, provision, and disclosure of user personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0074] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0075] Figure 4 FIG. shows a schematic block diagram of an exemplary electronic device 400 that can be used to implement embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processing, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0076] As Figure 4 shown, the device 400 includes a computing unit 401, which can execute various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 402 or a computer program loaded from a storage unit 408 into a random access memory (RAM) 403. In the RAM 403, various programs and data required for the operation of the device 400 can also be stored. The computing unit 401, the ROM 402, and the RAM 403 are connected to each other through a bus 404. An input / output (I / O) interface 405 is also connected to the bus 404.

[0077] Multiple components in device 400 are connected to I / O interface 405, including: input unit 406, such as a keyboard, mouse, etc.; output unit 407, such as various types of displays, speakers, etc.; storage unit 408, such as a disk, optical disc, etc.; and communication unit 409, such as a network card, modem, wireless communication transceiver, etc. Communication unit 409 allows device 400 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunications networks.

[0078] Computing unit 401 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of computing unit 401 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Computing unit 401 executes the various methods and processes described above, such as the task execution method. For example, in some embodiments, the task execution method can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as storage unit 408. In some embodiments, part or all of the computer program can be loaded and / or installed onto device 400 via ROM 402 and / or communication unit 409. When the computer program is loaded into RAM 403 and executed by computing unit 401, one or more steps of the task execution method described above can be performed. Alternatively, in other embodiments, computing unit 401 can be configured to execute the task execution method in any other suitable way (e.g., by means of firmware).

[0079] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGA), application-specific integrated circuits (ASIC), application-specific standard products (ASSP), system-on-chip systems (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a dedicated or general-purpose programmable processor, and can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0080] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing devices, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program codes can be executed entirely on the machine, partially on the machine, executed partially on the machine as an independent software package and partially on a remote machine, or executed entirely on a remote machine or server.

[0081] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0082] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) monitor or an LCD (liquid crystal display)) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and the input received from the user can be in any form (including acoustic input, voice input, or tactile input).

[0083] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), and the Internet.

[0084] A computer system can include a client and a server. The client and the server are generally far apart from each other and typically interact through a communication network. The client-server relationship is created by computer programs that run on the respective computers and have a client-server relationship with each other.

[0085] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this is not limited herein.

[0086] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.

Claims

1. A task execution method, comprising: Processing input data of a target sub - model by using a simulation model corresponding to the target sub - model in a deep - learning model to obtain simulation output data of the target sub - model, wherein the target sub - model includes at least one target operator, a target task corresponding to the target sub - model is executed by a first hardware resource, the first hardware resource includes at least one processor core and a storage space corresponding to at least one of the processor cores, the simulation model is a linear regression model or a fully - connected network, the linear regression model is obtained by performing linear regression based on a plurality of historical input data and a plurality of historical output data, and the fully - connected network is trained by using the historical input data as training samples and the historical output data as labels; In response to determining that the simulation output data is not within a first numerical range corresponding to the first hardware resource, determining a second hardware resource according to the simulation output data, wherein a second numerical range corresponding to the second hardware resource is greater than the first numerical range, the number of processor cores corresponding to the second hardware resource is greater than the number of processor cores corresponding to the first hardware resource, and the first numerical range and the second numerical range are numerical ranges corresponding to floating - point numbers; and Executing the target task by using the second hardware resource to process the input data to obtain target output data of the target sub - model.

2. The task execution method according to claim 1, wherein At least one of the target operators is N, where N is an integer greater than 1, the simulation model includes N simulation sub - models, and each simulation sub - model corresponds to one of the target operators. The processing the input data by using the simulation model corresponding to the target sub - model to obtain the simulation output data includes: Processing the input data by using N simulation sub - models to obtain N simulation output sub - data corresponding to the N target operators as the simulation output data.

3. The task execution method according to claim 2, wherein, The processing the input data by using N simulation sub - models to obtain N simulation output sub - data corresponding to the N target operators includes: Processing the input data by using the first simulation sub - model to obtain the first simulation output sub - data; Processing the (n - 1) - th simulation output sub - data by using the n - th simulation sub - model to obtain the n - th simulation output sub - data, where n is an integer greater than 1 and less than or equal to N.

4. The task execution method according to claim 3, wherein, The in response to determining that the simulation output data is not within the first numerical range corresponding to the first hardware resource, determining the second hardware resource according to the simulation output data includes: In the case where the first (n - 1) simulation output sub - data are within the first numerical range, in response to determining that the n - th simulation output sub - data is not within the first numerical range, determining the second hardware resource according to the n - th simulation output sub - data.

5. The task execution method according to claim 1, further comprising: In response to determining that the simulation output data is within the first numerical range, executing the target task by using the first hardware resource to process the input data to obtain first output data.

6. The task execution method according to claim 5, further comprising: In response to determining that the difference between the first output data and the simulation output data is greater than or equal to a preset difference threshold, adjust the first hardware resource to a third hardware resource, where a third numerical range corresponding to the third hardware resource is greater than the first numerical range.

7. The task execution method according to claim 1, wherein The input data includes one of forward input data and reverse input data.

8. A task execution device, comprising: A processing sub-module, configured to process the input data of the target sub-model by using a simulation model corresponding to the target sub-model in a deep learning model to obtain the simulation output data of the target sub-model, where the target sub-model includes at least one target operator, the target task corresponding to the target sub-model is executed by a first hardware resource, the first hardware resource includes at least one processor core and a storage space corresponding to at least one of the processor cores, the simulation model is a linear regression model or a fully connected network, the linear regression model is obtained by performing linear regression based on a plurality of historical input data and a plurality of historical output data, and the fully connected network is trained by using the historical input data as training samples and the historical output data as labels; A second determination module, configured to, in response to determining that the simulation output data is not within a first numerical range corresponding to the first hardware resource, determine a second hardware resource according to the simulation output data, where a second numerical range corresponding to the second hardware resource is greater than the first numerical range, the number of processor cores corresponding to the second hardware resource is greater than the number of processor cores corresponding to the first hardware resource, and the first numerical range and the second numerical range are numerical ranges corresponding to floating-point numbers; and A first execution module, configured to execute the target task by using the second hardware resource to process the input data to obtain the target output data of the target sub-model.

9. The task execution device according to claim 8, wherein, At least one of the target operators is N, where N is an integer greater than 1, the simulation model includes N simulation sub-models, and the simulation sub-model corresponds to one of the target operators. The processing sub-module includes: A processing unit, configured to process the input data by using the N simulation sub-models to obtain N simulation output sub-data corresponding to the N target operators as the simulation output data.

10. The task execution device according to claim 9, wherein, The processing unit includes: A first processing sub-unit, configured to process the input data by using the first simulation sub-model to obtain the first simulation output sub-data; A second processing sub-unit, configured to process the (n - 1)th simulation output sub-data by using the nth simulation sub-model to obtain the nth simulation output sub-data, where n is an integer greater than 1 and less than or equal to N.

11. The task execution device according to claim 10, wherein, The second determination module is further configured to: In the case where the first (n - 1) simulation output sub-data is within the first numerical range, in response to determining that the nth simulation output sub-data is not within the first numerical range, determine the second hardware resource according to the nth simulation output sub-data.

12. The task execution device according to claim 8, further comprising: A second execution module, configured to execute the target task by using the first hardware resource in response to determining that the simulation output data is within the first numerical range, so as to process the input data and obtain first output data.

13. The task execution device according to claim 12, further comprising: An adjustment module, configured to adjust the first hardware resource to a third hardware resource in response to determining that the difference between the first output data and the simulation output data is greater than or equal to a preset difference threshold, where a third numerical range corresponding to the third hardware resource is greater than the first numerical range.

14. An electronic device, comprising: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor, so that the at least one processor can execute the method according to any one of claims 1 to 7.

15. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 7.

16. A computer program product, comprising a computer program, where the computer program implements the method according to any one of claims 1 to 7 when being executed by a processor.