Task processing method and device and electronic equipment

By obtaining device status information, utilizing the computing power allocation decision model and value estimation network optimization model parameters, computing power is allocated to target devices, solving the problem of insufficient device computing power, improving resource utilization and task completion rate, and reducing device power consumption.

CN120670160APending Publication Date: 2025-09-19GUANGZHOU AUTOMOBILE GROUP CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510765245.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-09
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

The devices lack computing power when performing intelligent tasks, resulting in high costs and waste of resources.

Method used

By obtaining the status information of server-associated devices, the computing power allocation decision model is used to predict the candidate computing power allocation results, and computing power is allocated to the target device to perform intelligent tasks. The model parameters are optimized in combination with the value estimation network to maximize the cumulative reward and achieve reasonable resource scheduling.

Benefits of technology

It improves the resource utilization of the server and the intelligent task completion rate of the equipment, reduces the power consumption of the equipment, improves the intelligent experience of the equipment, and solves the problem of insufficient computing power of old equipment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120670160A_ABST
    Figure CN120670160A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a task processing method and device and electronic equipment, and the method comprises the steps: obtaining the state information of a plurality of pieces of equipment associated with a server, each piece of equipment having an intelligent task, and the state information of the equipment comprises the usable computing power of the equipment and the computing power required by the intelligent task in the equipment; obtaining a candidate computing power distribution result of the intelligent task in each device based on the residual computing power of the server and the state information of the plurality of devices; if the candidate computing power distribution result corresponding to the equipment indicates that the computing power is distributed to the intelligent task in the equipment and the available computing power of the equipment is smaller than the computing power required by the intelligent task of the equipment, acquiring the equipment as target equipment; and executing the intelligent task in the target equipment through the computing power allocated to the target equipment by the server. Through the method, the equipment can reasonably utilize the computing power of the server, so that the execution success rate and the execution efficiency of the intelligent task in the equipment are improved while the computing power utilization rate of the server is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and more specifically, to a task processing method, device, and electronic device. Background Art

[0002] With the advancement of intelligent technology, more and more devices are now required to perform intelligent tasks. Take autonomous driving technology as an example. It uses artificial intelligence, sensors, computer vision, and advanced algorithms to enable cars to navigate and drive autonomously without human drivers, providing safe and efficient transportation solutions. This is driving the development of modern cars towards a smarter and more user-friendly direction.

[0003] However, due to the different configurations of vehicles, allocating higher computing power to all vehicles will lead to higher costs and waste of computing power. Therefore, how to solve the problem of insufficient computing power when the equipment performs tasks is a technical problem that needs to be solved urgently. Summary of the Invention

[0004] In view of this, the embodiments of the present application propose a task processing method, device and electronic device, which can enable the device to reasonably utilize the computing power of the server, so as to improve the computing power utilization of the server while improving the execution success rate and execution efficiency of intelligent tasks in the device, effectively alleviating the problem of insufficient computing power when the device performs intelligent tasks.

[0005] In the first aspect, an embodiment of the present application provides a task processing method, which includes: obtaining status information of each of multiple devices associated with a server, each of the devices having an intelligent task, and the status information of the device including the available computing power of the device and the computing power required for the intelligent task in the device; based on the remaining computing power of the server and the status information of the multiple devices, obtaining candidate computing power allocation results for the intelligent tasks in each of the devices; if the candidate computing power allocation result corresponding to the device indicates that computing power is allocated to the intelligent task in the device, and the available computing power of the device is less than the computing power required for the intelligent task of the device, obtaining the device as the target device; and executing the intelligent task in the target device through the computing power allocated by the server to the target device.

[0006] In a second aspect, an embodiment of the present application provides a task processing device, which includes: an information acquisition module for acquiring status information of each of multiple devices associated with a server, each of the devices having an intelligent task, and the status information of the device including the available computing power of the device and the computing power required for the intelligent task in the device; a computing power allocation module for obtaining candidate computing power allocation results for the intelligent tasks in each device based on the remaining computing power of the server and the status information of the multiple devices; a device determination module for acquiring the device as a target device when the candidate computing power allocation result corresponding to the device indicates that computing power is allocated to the intelligent task in the device, and the available computing power of the device is less than the computing power required for the intelligent task of the device; and a task execution module for executing the intelligent task of the target device by using the computing power allocated to the target device.

[0007] In one possible implementation, the computing power allocation module is further used to predict the candidate computing power allocation results of the intelligent tasks in each device based on the status information of the multiple devices using the computing power allocation decision model. The task processing device also includes a sample acquisition module, a value estimation module, and a parameter adjustment module. The sample acquisition module is used to obtain a training sample set, wherein the training sample set includes multiple sample data, and the sample data includes the sample remaining computing power of the server and the sample status information of multiple sample devices associated with the server; the computing power allocation module prediction module is also used to predict based on the sample data using the computing power allocation decision model to obtain the predicted computing power allocation results of the intelligent tasks in each sample device; the value estimation module is used to predict based on the sample data and the predicted computing power allocation results of the intelligent tasks in each device using the value estimation network to obtain the predicted cumulative reward; the parameter adjustment module is used to adjust the model parameters of the computing power allocation decision model and the value estimation network based on the predicted cumulative reward to maximize the cumulative reward and obtain the trained computing power allocation decision model.

[0008] In one embodiment, each of the sample data also includes a reference computing power allocation result and a reference reward cumulative reward of the intelligent tasks in each sample device; the sample acquisition module is also used to adjust the model parameters of the computing power allocation decision model and the value estimation network based on the predicted cumulative reward to maximize the cumulative reward and obtain the trained computing power allocation decision model; obtain the reference cumulative reward generated by allocating computing power to the intelligent tasks of each sample device according to the reference computing power allocation result; the parameter adjustment module is also used to soft-update the model parameters of the computing power allocation decision model based on the predicted cumulative reward through a gradient descent algorithm; and soft-update the model parameters of the value estimation network based on the strategy loss entropy and the predicted cumulative reward through a gradient descent algorithm, the strategy loss entropy is determined based on the predicted computing power allocation result and the reference computing power allocation result, the predicted cumulative reward is positively correlated with the intelligent task completion rate, the rationality of resource utilization and the reference strategy entropy, respectively, the reference strategy entropy is determined based on the dynamic coefficient and the strategy loss, and the dynamic coefficient decays with the number of training iterations of the computing power allocation decision model.

[0009] In one possible implementation, the intelligent task completion rate and strategy loss are determined based on the predicted computing power allocation result, and the rationality of the resource utilization is determined based on the remaining computing power of the server, the available computing power of each device, and the computing power allocation result.

[0010] In one possible implementation, the computing power allocation module is further used to determine whether the server supports allocating computing power to at least one device based on the remaining computing power of the server; if supported, the computing power allocation decision model is used to predict the candidate computing power allocation results of the intelligent tasks in each device based on the remaining computing power of the server and the status information of the multiple devices.

[0011] In one possible implementation, the computing power allocation module is further used to determine the computing power to be allocated based on the remaining computing power of the target device and the computing power required for the intelligent task in the target device; and to determine the computing power allocated by the server to the target device based on the computing power to be allocated, wherein the computing power allocated by the server to the target device is not less than the computing power to be allocated.

[0012] In one possible implementation, the status information of the device further includes the execution time of the smart task of the device and the remaining power of the device.

[0013] In one possible implementation, the device includes a vehicle, and the intelligent task includes an intelligent driving task.

[0014] In a third aspect, an embodiment of the present application further provides an electronic device comprising: one or more processors; a memory; computer-readable instructions are stored on the memory, and when the computer-readable instructions are executed by the processor, the video text recognition method as described above is implemented.

[0015] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium having computer-readable instructions stored thereon. When the computer-readable instructions are executed by a processor, the method described above is implemented.

[0016] In a fifth aspect, embodiments of the present application provide a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device retrieves the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the above-described method.

[0017] The embodiment of the present application provides a task processing method, device and electronic device. The method includes: obtaining the status information of each of the multiple devices associated with the server, each of the devices has an intelligent task, and the status information of the device includes the available computing power of the device and the computing power required for the intelligent task in the device; based on the remaining computing power of the server and the status information of the multiple devices, obtaining the candidate computing power allocation result of the intelligent task in each of the devices; if the candidate computing power allocation result corresponding to the device indicates that computing power is allocated to the intelligent task in the device, and the available computing power of the device is less than the computing power required for the intelligent task of the device, obtaining the device as the target device; executing the intelligent task in the target device through the computing power allocated by the server to the target device. By adopting the above method, the server's resources can be allocated to the target device for use, thereby effectively improving the rationality of the server's resource utilization. Each target device can complete its intelligent task based on the allocated resources, thereby improving the completion rate of the intelligent tasks in each device associated with the server. In addition, since the server provides computing resources for the target device, a large amount of computing work of the target device is completed on the server, so the power consumption of the target device can also be effectively reduced. Ultimately, it is possible to dispatch resources in real time based on the current status of the device, meet device needs, enhance the owner's intelligent experience, reduce the computing power cost of the device, and solve the problem of insufficient computing power of old equipment. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.

[0019] Figure 1 A flowchart of a task processing method provided in an embodiment of the present application is shown; Figure 2 Shown Figure 1 Flow diagram of step S120; Figure 3 Another flowchart of a task processing method provided by an embodiment of the present application is shown; Figure 4 Another flowchart of a task processing method provided by an embodiment of the present application is shown; Figure 5 A schematic diagram showing the training and prediction stages of a task processing method used in an embodiment of the present application is shown; Figure 6 An application scenario diagram of a task processing method provided by an embodiment of the present application is shown; Figure 7 A schematic diagram showing another application scenario of a task processing method proposed in an embodiment of the present application is shown; Figure 8 is another flowchart illustrating a task processing method according to a specific embodiment of the present application; Figure 9 is a connection block diagram of a task processing device according to a specific embodiment of the present application; Figure 10 A schematic structural diagram of an electronic device suitable for implementing the embodiments of the present application is shown. DETAILED DESCRIPTION

[0020] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this application will be thorough and complete and will fully convey the concepts of the example embodiments to those skilled in the art.

[0021] In addition, described feature, structure or characteristic can be combined in one or more embodiments in any suitable manner.In the following description, many specific details are provided so as to provide a full understanding of the embodiments of the present application. However, it will be appreciated by those skilled in the art that the technical scheme of the present application can be put into practice without one or more of the specific details, or other methods, components, devices, steps etc. can be adopted. In other cases, known methods, devices, implementations or operations are not shown or described in detail to avoid blurring the various aspects of the application.

[0022] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically separate entities. That is, these functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.

[0023] The flowcharts shown in the accompanying drawings are for illustrative purposes only and do not necessarily include all contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps may be decomposed, while others may be combined or partially combined. Therefore, the actual execution order may vary depending on the actual situation.

[0024] It should be noted that the term "plurality" used in this document refers to two or more. "And / or" describes a relationship between associated objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A alone, A and B together, or B alone. The character " / " generally indicates an "or" relationship between the associated objects.

[0025] Figure 1 The task processing method of the present application is specifically shown. The method can be applied to an electronic device, which can be a server (such as a vehicle backend server or a device management server). The method includes: Step S110: Obtain status information of each of the multiple devices associated with the server.

[0026] The status information of the device includes the available computing power of the device and the computing power required for intelligent tasks in the device.

[0027] Each of the devices has an intelligent task, and the status information of the device includes the available computing power of the device and the computing power required for the intelligent task in the device. The server can be composed of a cluster of physical servers in a large-scale data center, equipped with high-performance CPUs, GPUs, storage, and network devices. Its resources can be abstracted into shared computing power through virtualization technology and dynamically allocated to devices.

[0028] Multiple devices can be of the same type (for example, multiple devices are vehicles or terminal devices, etc.) or of different types. You can set them up according to actual needs.

[0029] The available computing power of a device can be either the current remaining computing power or the computing power available for running smart tasks. A device can have one or more smart tasks. If there are multiple smart tasks, the computing power required for each smart task may vary. Smart tasks in different devices may also vary.

[0030] In one possible implementation, the device includes a vehicle, and the intelligent tasks include intelligent driving tasks (automatic cruise tasks), cockpit dialogue tasks, and timed intelligent tasks.

[0031] Device status information may also include the device's remaining battery life, the execution time of smart tasks in the device, and the expected operating time of the device based on the current remaining battery life (or, if the device is a vehicle, the expected mileage of the vehicle based on the current remaining battery life). The above device status information is for illustrative purposes only and may include more information, which should be configured based on actual circumstances.

[0032] To obtain status information for multiple devices associated with a server, the server can periodically send query requests to the associated devices and receive status information returned by the devices in response to the query requests. Alternatively, the server can send status information to the associated devices based on status changes, either periodically or in real time. This can be configured based on actual needs.

[0033] Step S120: Based on the remaining computing power of the server and the status information of the multiple devices, obtain the candidate computing power allocation results of the intelligent tasks in each of the devices.

[0034] The candidate computing power allocation result of the intelligent task in the device is used to indicate whether to allocate computing power to the intelligent task in the device.

[0035] In one possible implementation, the above step S120 may specifically be to use the computing power allocation decision model to predict candidate computing power allocation results for the intelligent tasks in each device based on the status information of the multiple devices.

[0036] The computing power allocation decision model is obtained by iterative training based on sample data, and the sample data includes the sample remaining computing power of the server and sample status information of multiple sample devices associated with the server.

[0037] The computing power allocation decision model may be any one of a neural network model and a reinforcement learning model.

[0038] The computing power allocation decision model can extract features from the server's remaining computing power and the status information of the multiple devices, and then perform classification predictions based on the extracted features to obtain candidate computing power allocation results for the intelligent tasks in each device. The allocation results for the intelligent tasks in each device can be expressed as the probability of allocating computing power to each device. For any device, if the probability of allocating computing power to the intelligent task of the device is greater than a preset probability threshold, it can be indicated that computing power needs to be allocated to the intelligent task of the device; if the probability of allocating computing power to the intelligent task of the device is not greater than the preset probability threshold, it can be indicated that computing power does not need to be allocated to the intelligent task of the device.

[0039] Among them, when the computing power allocation model is obtained based on the training of the neural network model, the sample data also includes sample labels, and the sample labels are used to indicate the optimal allocation results of intelligent tasks in multiple sample devices; the computing power allocation model can make predictions based on the sample remaining computing power of the server and the sample status information of multiple sample devices associated with the server to obtain the predicted computing power allocation results of the intelligent tasks in each sample device; based on the predicted computing power allocation results of each intelligent task in the sample device and the sample labels, the loss is calculated to obtain the predicted loss, and the computing power allocation model is adjusted based on the predicted loss.

[0040] When the computing power allocation model is obtained based on reinforcement learning training, the computing power allocation model can be used to make predictions based on the sample remaining computing power of the server and the sample status information of multiple sample devices associated with the server to obtain the predicted computing power allocation results of the intelligent tasks in each sample device; the value estimation network can make value predictions based on the predicted computing power allocation results and the sample remaining computing power of the server and the sample status information of multiple sample devices associated with the server to obtain cumulative rewards; thereafter, the model parameters of the value estimation network and the computing power allocation model can be adjusted based on the cumulative rewards to maximize the cumulative rewards and obtain the trained computing power allocation decision model.

[0041] The training process of the computing power allocation model described above is merely illustrative, and other training methods are possible, which are not specifically limited here.

[0042] To ensure that the server can allocate computing power to its associated devices, refer to Figure 2 In one embodiment, step S120 includes: Step S122: Determine whether the server supports allocating computing power to at least one device based on the remaining computing power of the server.

[0043] Exemplarily, the remaining computing power of the server may be compared with the preset computing power. If the remaining computing power is greater than the preset computing power, it may be determined that the server is sufficient to support allocating computing power to at least one device.

[0044] The remaining computing power of the server can also be compared with the maximum computing power value (average computing power or median computing power, etc.) required for the intelligent tasks of each device. If the remaining computing power is greater than the maximum computing power value (average computing power value or median computing power value, etc.), it can be determined that the remaining computing power of the server can support the allocation of computing power to at least one device. You can also calculate the difference between the available computing power of each device and the computing power required for the intelligent tasks in the device to obtain the computing power difference corresponding to each device; if the remaining computing power of the server is greater than the computing power difference corresponding to each device, it can be determined whether the server supports allocating computing power to at least one device.

[0045] The above method of determining whether the server supports allocating computing power to at least one device is merely illustrative, and there may be more implementation methods, which are not specifically limited here.

[0046] Step S124: If supported, the computing power allocation decision model is used to predict candidate computing power allocation results for the intelligent tasks in each device based on the remaining computing power of the server and the status information of the multiple devices.

[0047] Regarding the process of using the computing power allocation decision model to predict candidate computing power allocation results for intelligent tasks on each device based on the server's remaining computing power and the status information of the multiple devices, please refer to the detailed description of the aforementioned embodiments and will not be repeated here. Step S130: If the candidate computing power allocation result corresponding to the device indicates that computing power should be allocated to the intelligent task on the device, and the available computing power of the device is less than the computing power required for the intelligent task on the device, the device is acquired as the target device.

[0048] It is worth mentioning that when the device status information also includes the remaining battery power of the device and the execution duration of the smart task, the remaining battery power of the target device should support the target device to complete its smart task. For example, the remaining battery power of the target device can be greater than a preset battery power threshold or the remaining battery power of the target device should be greater than the power required to complete its smart task.

[0049] Among them, allocating the computing power of the server to the target device can be to allocate computing power of a specified size to the target device, and the computing power of the specified size can be determined according to the computing power required for the intelligent tasks in each device (such as according to the average value, median value, maximum value or minimum value of the computing power required for the intelligent tasks in each device).

[0050] Among them, the process of allocating the server's computing power to the target device can be: allocating computing power from the server's shared resource pool, quickly binding it to the target device through a virtualization layer (such as KVM) as the computing power allocated to the target device (wherein, the allocated resources can be accompanied by a life cycle (such as a 5-minute lease) to be automatically released after the end of the life cycle), establishing a low-latency link between the target device and the computing power resources allocated to it by the server, so as to migrate part or all of the computing logic of the target device to the server through the low-latency link, so as to use the computing resources allocated to it by the server for calculation.

[0051] To ensure that the computing power is allocated reasonably so that as many intelligent tasks as possible can be executed, please refer to Figure 3 In one embodiment, the method further comprises: Step S132: Determine the computing power to be allocated according to the remaining computing power of the target device and the computing power required by the intelligent task in the target device.

[0052] For example, the computing power to be allocated may be determined as a result of subtracting the computing power required for the intelligent task in the target device from the remaining computing power of the target device.

[0053] Step S134: Allocate the computing power of the server to the target device according to the computing power to be allocated, wherein the computing power allocated to the device is not less than the computing power to be allocated.

[0054] By adopting the above steps S132-S134, computing power can be allocated to the target device according to the computing power gap of the target device, so that the computing power allocated to the target device is more reasonable, thereby helping to improve the utilization rate of the computing power. Among them, the above steps S132-S134 can be executed after executing step S130, or can be executed during the execution of step S130, that is, they can be executed by the above-mentioned computing power allocation decision model, and can be set according to actual conditions.

[0055] It is worth mentioning that when the remaining computing power of the device is sufficient to support the computing power required for its intelligent task and the remaining power of the device can support the device to run its intelligent task, the device can directly allocate its remaining computing power to its intelligent task to perform its intelligent task.

[0056] When the status information of the device includes the remaining power of the device, if the remaining power of the device is insufficient to support the device in executing its smart task, a power replenishment reminder may be generated and the smart task may no longer be executed.

[0057] Step S140: executing the intelligent task in the target device by using the computing power allocated by the server to the target device.

[0058] Specifically, the target device can upload data related to its intelligent task (such as video data, sensor data, etc.) to the server. The server then uses its allocated computing power to execute the corresponding intelligent task based on this data. It is worth noting that after the server executes the intelligent task and obtains the task execution results, it will also return the task execution results to the target device, so that the target device can perform subsequent operations based on the task execution results.

[0059] Using the method in the above embodiment, a computing power allocation decision model is used to predict candidate computing power allocation results for intelligent tasks in each device based on the remaining computing power of the server and the status information of the multiple devices. If a candidate computing power allocation result corresponding to a target device indicates that computing power should be allocated to its intelligent task, and the target device's available computing power is less than the computing power required for the intelligent task, the target device is allocated the computing power of the server so that the target device can perform its intelligent task after being allocated the computing power of the server. This can achieve the goal of allocating server resources to target devices, thereby effectively improving the rationality of server resource utilization. Each target device can complete its intelligent task based on the allocated resources, thereby improving the completion rate of intelligent tasks in each device associated with the server. In addition, because the server provides computing resources to the target device, a large amount of the target device's computing work is completed on the server, which can also effectively reduce the power consumption of the target device. Ultimately, it is possible to achieve real-time resource scheduling based on the current status of the device, meet device needs, enhance the owner's intelligent experience, reduce the computing power cost of the device, and solve the problem of insufficient computing power in old devices.

[0060] Second embodiment See also Figure 4 As shown, based on the above task processing method, the task processing method also includes: Step S160: Acquire a training sample set, where the training sample set includes a plurality of sample data.

[0061] The training sample set includes multiple sample data, and each sample data includes the sample remaining computing power of the server and sample status information of multiple sample devices associated with the server.

[0062] In some scenarios, the sample data may also include reference computing power allocation results and reference cumulative rewards.

[0063] Among them, the sample status information of the sample device may include the available computing power of the sample device and the computing power required for the intelligent tasks in the sample device; in some scenarios, the sample status information may also include the remaining power of the sample device and the execution time of the intelligent tasks in the sample device.

[0064] In one possible implementation, a training sample set can be obtained by obtaining multiple sample data sets labeled based on expert experience. The sample data includes the server's sample remaining computing power, sample status information for multiple sample devices associated with the server, and a reference computing power allocation result. The reference computing power allocation result can be the optimal computing power allocation result labeled by the experts based on their experience. The reference computing power allocation result can correspond to a reference cumulative reward, which refers to the reference cumulative reward generated by allocating computing power to the intelligent tasks of each sample device according to the reference computing power allocation result. By utilizing the above training sample set, real-time device status (computing power / electricity) and expert experience labels can be combined for training, which can accelerate the initial convergence of the model.

[0065] In another possible implementation, the training sample set can be obtained by obtaining sample remaining computing power of a server in the computing power allocation system at multiple time instants and sample status information of multiple sample devices associated with the server to obtain sample data. The sample data can also include a reference computing power allocation result, which can be a reference computing power allocation result for each device's intelligent tasks obtained by randomly allocating computing power to the intelligent tasks of multiple devices associated with the server. The reference cumulative reward refers to the reference cumulative reward generated by allocating computing power to the intelligent tasks of each sample device according to the reference computing power allocation result.

[0066] Under this embodiment, the above-mentioned training samples can be the behavior of randomly triggering the computing power allocation system to allocate computing power to at least one vehicle in the actual operating environment. Each time the behavior of allocating computing power to a device is triggered, the status information of each device when the behavior is triggered and the remaining computing power of the server are obtained, and a reference cumulative reward generated by triggering the behavior and the vehicle status and the remaining computing power of the server at the next moment are calculated as an experience, which are packaged and sent to form an experience (sample data) and stored in the buffer database.

[0067] Step S170: using the computing power allocation decision model to make predictions based on the sample data, and obtaining predicted computing power allocation results for the intelligent tasks in each sample device.

[0068] The computing power allocation decision model is used to generate specific computing power allocation actions (such as "allocate 30% computing power to device A") based on sample data (i.e., current status (computing power, device requirements, etc.)).

[0069] That is, the above-mentioned step S170 can specifically be to use the computing power allocation decision model to make a prediction based on the sample remaining computing power of the server and the sample status information of multiple sample devices associated with the server to obtain the predicted computing power allocation result of the intelligent task in each sample device.

[0070] Step S180: Utilize the value estimation network to make predictions based on the sample data and the predicted computing power allocation results of the intelligent tasks in each device to obtain a predicted cumulative reward.

[0071] The value estimation network is used to evaluate the long-term value of states or state-action pairs and guide the optimization strategy of the computing power allocation decision model.

[0072] Please refer to Figure 5 In the case where the sample data also includes reference computing power allocation results and reference cumulative rewards, in order to improve the convergence speed of the value estimation network, in one embodiment, the above-mentioned step S160 can also be to use the value estimation network to make a prediction based on the sample remaining computing power of the server and the sample status information of multiple sample devices associated with the server, the predicted computing power allocation results of the intelligent tasks in each device, the reference computing power allocation results and the reference cumulative rewards to obtain the predicted cumulative rewards.

[0073] In one possible implementation, the cumulative reward is positively correlated with the intelligent task completion rate, the rationality of resource utilization, and the reference strategy entropy, respectively. The reference strategy entropy is determined based on the dynamic coefficient and strategy loss, and the dynamic coefficient decays with the number of training iterations of the computing power allocation decision model; the intelligent task completion rate and strategy loss are determined based on the predicted computing power allocation result, and the rationality of resource utilization is determined based on the computing power of the server, the computing power of each device, and the computing power allocation result.

[0074] The intelligent task completion rate can be determined based on the number of intelligent tasks requiring computing power allocation indicated in the predicted computing power allocation result and the total number of intelligent tasks. For example, it can be the ratio between the number of intelligent tasks requiring computing power allocation indicated in the predicted computing power allocation result and the total number of intelligent tasks. The reasonableness of resource utilization can be calculated based on the remaining computing power utilization of the server and the reference computing power utilization. For example, the remaining computing power utilization can be calculated based on the computing power of the server, the computing power of each device, and the computing power allocation result. Subsequently, the reasonableness of resource utilization is calculated by subtracting the absolute value of the difference between the ratio of the actual utilization to the target utilization and the target value (e.g., 1) from the specified value (e.g., 1) to obtain the reasonableness of resource utilization.

[0075] The cumulative reward mentioned above may be at least one of the reference cumulative rewards to predict the cumulative rewards.

[0076] Among them, the cumulative reward can be expressed as , is the discount factor, which is a value between 0 and 1. The cumulative reward is calculated as the cumulative discounted reward generated from time t to one degree in the future, and T represents the end time.

[0077] For example, in this embodiment, at any time t, according to the current scheduling model strategy , generate scheduling actions (whether to allocate computing power to each sample device, that is, refer to the computing power allocation result) ; Afterwards, the scheduling action can be executed , calculate the reward value , receive the next moment state value ; ( , , , ) forms an experience and stores it in the experience pool; among them, ,in, Indicates the completion rate of intelligent tasks, Indicates the reasonableness of resource utilization. represents the strategy loss, Represents the dynamic coefficient, which decays with the number of training iterations of the computing power allocation decision model.

[0078] Step S190: Adjust the model parameters of the computing power allocation decision model and the value estimation network based on the predicted cumulative reward to maximize the cumulative reward and obtain a trained computing power allocation decision model.

[0079] In one embodiment, the above-mentioned step S190 may be: based on the predicted cumulative reward, the model parameters of the computing power allocation decision model are soft-updated by a gradient descent algorithm; based on the strategy loss entropy and the predicted cumulative reward, the model parameters of the value estimation network are soft-updated by a gradient descent algorithm, and the strategy loss entropy is determined based on the predicted computing power allocation result and the reference computing power allocation result.

[0080] By adopting the above steps S160-S190, it is possible to use the computing power allocation decision model to make predictions based on the sample data to obtain the predicted computing power allocation results of the intelligent tasks in each sample device; use the value estimation network to make predictions based on the sample data and the predicted computing power allocation results of the intelligent tasks in each device to obtain the predicted cumulative rewards; adjust the model parameters of the computing power allocation decision model and the value estimation network based on the predicted cumulative rewards to maximize the cumulative rewards, and obtain the trained computing power allocation decision model. This can enable the trained computing power allocation model to focus on high-reward computing power allocation, and the computing power allocation model can gradually improve its stability as the model training converges.

[0081] Furthermore, when the sample data includes randomly allocating computing power to intelligent tasks in multiple devices associated with the server, the reference computing power allocation results of the intelligent tasks of each device are obtained; and the computing power of the intelligent tasks of each sample device is allocated according to the reference computing power allocation results, and the reference cumulative rewards are generated, the value estimation network is iterated in combination with the reference rewards and the reference computing power allocation results. The value estimation network can quickly learn the value estimation method, thereby effectively improving the convergence speed of the value estimation network.

[0082] Furthermore, since the cumulative reward is positively correlated with the intelligent task completion rate, the rationality of resource utilization, and the reference policy entropy, the reference policy entropy is determined based on the dynamic coefficient and policy loss, and the dynamic coefficient decays with the number of training iterations of the computing power allocation decision model; the intelligent task completion rate and policy loss are determined based on the predicted computing power allocation results, and the rationality of resource utilization is determined based on the computing power of the server, the computing power of each device, and the computing power allocation results. Therefore, in the early stages of training (when μ is high): the reference policy entropy can promote exploration and discover better allocation strategies; in the later stages of training (when μ gradually decreases), the reference policy entropy also gradually decreases, making the training stage more focused on high-reward computing power allocation, so that the value estimation network finally obtained after training needs to be optimized to achieve the maximum cumulative reward value, while meeting the task completion rate. and reasonable utilization of resources While maximizing, entropy maximization must also be satisfied. Among them, policy entropy maximization helps the computing power allocation model adopt different scheduling behaviors, traverse more different situations, and is more conducive to finding the optimal solution and avoiding falling into local optimality.

[0083] Please refer to Figure 6 and Figure 7 As shown, the task processing method is applied to the server in the computing power allocation system. The computing power allocation system also includes multiple devices associated with the server, each device corresponds to an intelligent task, such as device n corresponds to intelligent task n, where each device is a vehicle, and the computing power allocation decision model is an example of a decision model based on reinforcement learning.

[0084] The server can obtain the status information of the vehicle it is associated with, which may include the chip computing power of the vehicle. , Remaining battery power of the vehicle , computing power required for intelligent tasks and the execution time required for intelligent tasks , at this time, the vehicle status information can be identified as ={( , ), ( , ),……,( , ), ( , ),……},in,( , ) represents the remaining computing power and remaining power of vehicle 1, ( , ) represents the required computing power of intelligent task 1 in vehicle 1 and the execution time of intelligent task 1.

[0085] In the resource allocation problem of the server in the vehicle intelligent environment, the computing power allocation model deployed in the server mainly predicts whether to release computing power. The resulting behavior space consists of {all vehicle tasks release computing power}, which can be represented by a Boolean number (yes or no). The behavior space can be expressed as: ={ , , , ...} , the corresponding strategy formula , Indicates the computing power allocation status of smart task 1, which can be (yes / no).

[0086] The goal of this embodiment is to schedule the computing power of the server through the computing power allocation decision model, so as to perfectly solve the downstream task requirements by utilizing the computing power of the server. Therefore, the reward indicator can be set to include the completion rate of the downstream task. and reasonable utilization of resources .in, The higher it is, the higher the downstream task completion rate is. The higher it is, the more reasonable the resource utilization is. Generally speaking, the optimization goal of the reinforcement learning algorithm is to find an optimal strategy to maximize the cumulative reward value returned by the model. Therefore, in addition to optimizing to achieve the maximum cumulative reward value, the method of this application must also satisfy entropy maximization. Maximizing policy entropy helps the model take different scheduling behaviors, traverse more different situations, and is more conducive to finding the optimal solution and avoiding falling into local optimality. Since exploration is a dynamically changing process, when the policy entropy is a fixed value, there is a certain error in the optimal performance of the model's final convergence. Therefore, this application further optimizes the reward function and adds a dynamic coefficient , the dynamic coefficient μ and the policy loss The product of μ and μ is used as the reference policy entropy, where μ is negatively correlated with the number of iterations, that is, the dynamic coefficient decays with the number of training iterations of the computing power allocation decision model. In the early stage of the computing power allocation decision model training, the training goal tends to explore more unknown areas in the environment and find more optimal solutions and possibilities. Therefore, the reference policy entropy should reach its maximum value. At this time, the dynamic coefficient As the training process continues, the model will gradually explore all possibilities, and the training goal will change from finding more optimal solutions to determining the existing optimal solution. Therefore, the reference policy entropy should gradually decrease, and the dynamic coefficient will gradually decrease. At the end of the training, the reference policy entropy should reach a minimum value, and the dynamic coefficient is at its minimum. Therefore, the optimized reward function used in this application is: .

[0087] When training the computing power allocation decision model, a continuous action space reinforcement learning algorithm is used. This algorithm combines maximum entropy reinforcement learning with the actor-critic framework to solve the reinforcement learning problem of scheduling. The main goal of this algorithm is to learn a computing power allocation strategy that maximizes the estimated cumulative reward while maximizing the entropy of the strategy, that is, maximizing the randomness of the strategy.

[0088] Specifically, we first need to construct a policy neural network as the actor network (computing power allocation decision model) in the algorithm to generate the computing power allocation strategy (computing power allocation result); the input of the network is the remaining computing power of the server and the status information of the associated vehicles (i.e., vehicle computing power). , remaining power , computing power required for intelligent tasks , the execution time of smart tasks ), the output is the probability of each vehicle's computing power allocation behavior { , , ...}, and the behavior finally selected according to the probability {downstream task } (i.e., the probability of allocating computing power to each intelligent task, and the result of whether computing power is ultimately allocated to each intelligent task).

[0089] Next, we need to evaluate the merits of these behaviors. Specifically, we construct a Soft Q neural network as the critic network (i.e., the value prediction network). The value prediction network's input includes the input of the actor network and its output. The output of the value prediction network includes the value corresponding to each behavior in the behavior space, thereby generating a cumulative reward.

[0090] Before training the computing power allocation decision network and the value estimation network, a cache database can be established. Specifically, in an intelligent vehicle environment with N vehicles, each vehicle is equipped with the same AI framework and different computing power hardware. In this environment, after randomly allocating computing power to the intelligent tasks of multiple devices associated with the server, the reference computing power allocation results of the intelligent tasks of each device are obtained. The computing power of the intelligent tasks of each sample device is allocated according to the reference computing power allocation results, and the reference cumulative reward Rt is generated, where , the status information of each vehicle , Reference computing power allocation results , reference cumulative reward Rt, vehicle status information at the next moment The server's remaining computing power at that moment is packaged as a piece of experience (sample data) and sent to a buffer database to store the experience for replay during training. A buffer database of appropriate size must also be established to avoid slow convergence when all data is used for training, or poor training results when using less sample data.

[0091] The specific training process is to use the computing power allocation decision model to predict the computing power allocation based on the status information of each vehicle at time t in the sample data and the remaining computing power of the server, and obtain the predicted computing power allocation result. , Reference computing power allocation results , reference cumulative reward Rt, vehicle status information at the next moment The predicted cumulative rewards are then calculated based on the server's remaining computing power at that moment and the predicted computing power allocation results. Subsequently, the value estimation network and computing power allocation decision model are updated using both gradient updates and soft updates based on the predicted cumulative rewards.

[0092] It's worth noting that the aforementioned training enables two parts of model-based training: first, the model assesses the computing power of different downstream tasks; second, the model rationally allocates existing computing power resources and computing power tasks, selecting them based on factors such as vehicle computing power and duration to ensure the optimal use of computing power resources. During training, the input is data (experience) stored in a cache database. Each time, a portion of this data is extracted as a training set, and the actor and critic networks are updated. Network parameters are continuously updated to maximize cumulative rewards. Ultimately, the scheduling model's neural network gradually fits the real-world environment equations, outputting a reasonable strategy for allocating computing power.

[0093] like Figure 6As shown, in a real-world intelligent vehicle environment, when a vehicle initiates or schedules an intelligent task, it packages and uploads its vehicle status information (vehicle configuration, intelligent task information, duration, start time, remaining battery life, etc.) to a server. The server first determines whether the vehicle is qualified to execute the intelligent task. (If the vehicle's battery life is insufficient to support the intelligent task (e.g., only 20% battery life is insufficient for a 100 km high-speed cruise intelligent task), the intelligent task is terminated. If so, the server determines whether computing power can be allocated to at least one vehicle. If computing power can be allocated to at least one vehicle, steps S110-S130 are executed as described above. The scheduling model determines whether the vehicle is qualified for intelligent tasks. If the vehicle has sufficient computing power, the task is executed directly locally. If the target vehicle has insufficient computing power, the model determines the appropriate resource based on the current computing power. If computing power is sufficient, computing power is allocated to the target vehicle to execute the intelligent task. If computing power is insufficient, computing power is not allocated to the target terminal, and the intelligent task is not executed.

[0094] It should be understood that, although the steps in the flowcharts of the above embodiments are shown in sequence as indicated by the arrows, these steps are not necessarily performed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be performed in other orders. Moreover, at least a portion of the steps in the flowcharts of the above embodiments may include multiple steps or multiple stages, and these steps or stages are not necessarily performed at the same time, but can be performed at different times. The execution order of these steps or stages is not necessarily to be performed in sequence, but can be performed in turn or alternately with other steps or at least a portion of steps or stages in other steps.

[0095] See also Figure 9 As shown, an embodiment of the present application provides a task processing device 200, which includes an information acquisition module 210 for acquiring status information of each of multiple devices associated with a server, wherein the status information of the device includes the available computing power of the device and the computing power required for the intelligent task in the device; a computing power allocation module 220 for obtaining candidate computing power allocation results for the intelligent task in each device based on the remaining computing power of the server and the status information of the multiple devices; a device determination module 230 for acquiring the device as a target device when the candidate computing power allocation result corresponding to the device indicates that computing power is allocated to the intelligent task in the device, and the available computing power of the device is less than the computing power required for the intelligent task of the device; and a task execution module 240 for executing the intelligent task of the target device by using the computing power allocated to the target device.

[0096] In one possible implementation, the computing power allocation module is further used to predict the candidate computing power allocation results of the intelligent tasks in each device based on the status information of the multiple devices using the computing power allocation decision model. The task processing device also includes a sample acquisition module, a value estimation module, and a parameter adjustment module. The sample acquisition module is used to obtain a training sample set, and the training sample set includes multiple sample data, and the sample data includes the sample remaining computing power of the server and the sample status information of multiple sample devices associated with the server; the computing power allocation module prediction module is also used to use the computing power allocation decision model to make predictions based on the sample data to obtain the predicted computing power allocation results of the intelligent tasks in each sample device; the value estimation module is used to use the value estimation network to make predictions based on the sample data and the predicted computing power allocation results of the intelligent tasks in each device to obtain the predicted cumulative rewards; the parameter adjustment module is used to adjust the model parameters of the computing power allocation decision model and the value estimation network based on the predicted cumulative rewards to maximize the cumulative rewards and obtain the trained computing power allocation decision model.

[0097] In one embodiment, each of the sample data also includes a reference computing power allocation result and a reference reward cumulative reward of the intelligent tasks in each sample device; the sample acquisition module is also used to adjust the model parameters of the computing power allocation decision model and the value estimation network based on the predicted cumulative reward to maximize the cumulative reward and obtain the trained computing power allocation decision model; obtain the reference cumulative reward generated by allocating computing power to the intelligent tasks of each sample device according to the reference computing power allocation result; the parameter adjustment module is also used to soft-update the model parameters of the computing power allocation decision model based on the predicted cumulative reward through a gradient descent algorithm; and soft-update the model parameters of the value estimation network based on the strategy loss entropy and the predicted cumulative reward through a gradient descent algorithm, the strategy loss entropy is determined based on the predicted computing power allocation result and the reference computing power allocation result, the predicted cumulative reward is positively correlated with the intelligent task completion rate, the rationality of resource utilization and the reference strategy entropy, respectively, the reference strategy entropy is determined based on the dynamic coefficient and the strategy loss, and the dynamic coefficient decays with the number of training iterations of the computing power allocation decision model.

[0098] In one possible implementation, the intelligent task completion rate and strategy loss are determined based on the predicted computing power allocation result, and the rationality of the resource utilization is determined based on the remaining computing power of the server, the available computing power of each device, and the computing power allocation result.

[0099] In one possible implementation, the computing power allocation module 230 is further used to determine the computing power to be allocated based on the remaining computing power of the target device and the computing power required for the intelligent task in the target device; and allocate the computing power of the server to the target device based on the computing power to be allocated, wherein the computing power allocated by the server to the target device is not less than the computing power to be allocated.

[0100] In one possible implementation, the status information of the device further includes the execution time of the smart task of the device and the remaining power of the device.

[0101] In one possible implementation, the computing power allocation module 220 is further used to determine whether the server supports allocating computing power to at least one device based on the remaining computing power of the server; if supported, the computing power allocation decision model is used to predict the candidate computing power allocation results of the intelligent tasks in each device based on the remaining computing power of the server and the status information of the multiple devices.

[0102] In one possible implementation, the task processing device 200 also includes a computing power determination module, which is used to determine the computing power to be allocated based on the remaining computing power of the target device and the computing power required for the intelligent task in the target device; based on the computing power to be allocated, determine the computing power allocated by the server to the target device, wherein the computing power allocated by the server to the target device is not less than the computing power to be allocated.

[0103] In one possible implementation, the device includes a vehicle, and the intelligent task includes an intelligent driving task.

[0104] Each module in the above-mentioned device can be implemented in whole or in part by software, hardware, or a combination thereof. Each of the above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory in the computer device in the form of software, so that the processor can call and execute the operations corresponding to each of the above modules. It should be noted that the device embodiment in this application corresponds to the aforementioned method embodiment. The specific principles in the device embodiment can be found in the contents of the aforementioned method embodiment, which will not be repeated here.

[0105] The following will be combined Figure 10 An electronic device provided by this application is described.

[0106] See also Figure 10 Based on the method provided in the above embodiment, the embodiment of the present application also provides another electronic device 100 including a processor 102 that can execute the above method. The electronic device 100 can be a server.

[0107] The electronic device 100 further includes a memory 104 . The memory 104 stores a program capable of executing the contents of the aforementioned embodiments, and the processor 102 can execute the program stored in the memory 104 .

[0108] The processor 102 may include one or more cores for processing data and a message matrix unit. The processor 102 utilizes various interfaces and circuits to connect various components within the electronic device 100. It executes instructions, programs, code sets, or instruction sets stored in the memory 104, and accesses data stored in the memory 104 to perform various functions and process data within the electronic device 100. Optionally, the processor 102 may be implemented in hardware using at least one of a digital signal processing (DSP), a field-programmable gate array (FPGA), and a programmable logic array (PLA). The processor 102 may integrate one or a combination of a central processing unit (CPU), a graphics processing unit (GPU), and a modem. The CPU primarily handles the operating system, user interface, and application programs; the GPU is responsible for rendering and drawing display content; and the modem handles wireless communications. It is understood that the modem may not be integrated into the processor 102 and may be implemented as a separate communication chip.

[0109] In this embodiment, the processor 102 includes a main controller and a system-on-chip to implement the aforementioned method steps.

[0110] The memory 104 may include random access memory (RAM) or read-only memory (ROM). The memory 104 may be used to store instructions, programs, code, code sets, or instruction sets. The memory 104 may include a program storage area and a data storage area. The program storage area may store instructions for implementing an operating system, instructions for implementing at least one function, and instructions for implementing the various method embodiments described below. The data storage area may also store data acquired by the electronic device 100 during use.

[0111] The electronic device 100 may also include a network module and a screen. The network module is used to receive and transmit electromagnetic waves, converting them into electrical signals, thereby communicating with a communications network or other devices, such as an audio playback device. The network module may include various existing circuit components for performing these functions, such as an antenna, a radio frequency transceiver, a digital signal processor, an encryption / decryption chip, a subscriber identity module (SIM) card, memory, and the like. The network module can communicate with various networks such as the Internet, an intranet, or a wireless network, or with other devices via a wireless network. Such wireless networks may include cellular telephone networks, wireless local area networks, or metropolitan area networks. The screen can display interface content and perform data interaction, such as displaying the aforementioned interface and triggering operations via the screen.

[0112] The embodiment of the present application also provides a structural block diagram of a computer-readable storage medium. The computer-readable storage medium stores program code, which can be called by a processor to execute the method described in the above method embodiment.

[0113] The computer-readable storage medium may be an electronic memory such as flash memory, EEPROM (Electrically Erasable Programmable Read-Only Memory), EPROM, hard disk, or ROM. Alternatively, the computer-readable storage medium includes a non-transitory computer-readable storage medium. The computer-readable storage medium has storage space for program code for executing any of the method steps described above. This program code can be read from or written to one or more computer program products. The program code may be compressed, for example, in a suitable format.

[0114] The present application also provides a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the methods described in the various optional implementations described above.

[0115] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A task processing method, characterized in that: The method comprises: Obtaining status information of each of a plurality of devices associated with the server, each of the devices having an intelligent task, the status information of the device including the available computing power of the device and the computing power required for the intelligent task in the device; Based on the remaining computing power of the server and the status information of the multiple devices, obtaining candidate computing power allocation results for the intelligent tasks in each of the devices; If the candidate computing power allocation result corresponding to the device indicates that computing power is allocated to the intelligent task in the device, and the available computing power of the device is less than the computing power required by the intelligent task of the device, the device is obtained as the target device; The intelligent task in the target device is executed by using the computing power allocated by the server to the target device.

2. The method according to claim 1, characterized in that The obtaining, based on the remaining computing power of the server and the status information of the plurality of devices, candidate computing power allocation results for the intelligent tasks in each of the devices includes: A computing power allocation decision model is used to predict candidate computing power allocation results for intelligent tasks in each device based on the status information of the multiple devices, wherein the computing power allocation model is trained based on the following steps: Acquire a training sample set, where the training sample set includes a plurality of sample data, where the sample data includes a sample remaining computing power of a server and sample status information of a plurality of sample devices associated with the server; Using the computing power allocation decision model to make predictions based on the sample data, a predicted computing power allocation result for the intelligent tasks in each sample device is obtained; Using the value estimation network to make predictions based on the sample data and the predicted computing power allocation results of the intelligent tasks in each device, a prediction cumulative reward is obtained; Based on the predicted cumulative reward, the model parameters of the computing power allocation decision model and the value estimation network are adjusted to maximize the cumulative reward, thereby obtaining a trained computing power allocation decision model.

3. The method according to claim 2, characterized in that Each of the sample data also includes the reference computing power allocation results and reference reward accumulation rewards of the intelligent tasks in each sample device; The obtaining of the training sample set includes: After obtaining the random computing power allocation for the intelligent tasks of multiple devices associated with the server, the reference computing power allocation result of the intelligent tasks of each device is obtained; Obtain the reference cumulative rewards generated by allocating computing power to the intelligent tasks of each sample device according to the reference computing power allocation results; The adjusting of the model parameters of the computing power allocation decision model and the value estimation network based on the predicted cumulative reward includes: Based on the predicted cumulative reward, the model parameters of the computing power allocation decision model are soft-updated using a gradient descent algorithm; Based on the policy loss entropy and the predicted cumulative reward, the model parameters of the value estimation network are soft-updated through a gradient descent algorithm. The policy loss entropy is determined based on the predicted computing power allocation result and the reference computing power allocation result. The predicted cumulative reward is positively correlated with the intelligent task completion rate, the rationality of resource utilization, and the reference policy entropy. The reference policy entropy is determined according to the dynamic coefficient and policy loss. The dynamic coefficient decays with the number of training iterations of the computing power allocation decision model.

4. The method according to claim 3, characterized in that The intelligent task completion rate and strategy loss are determined based on the predicted computing power allocation result, and the rationality of the resource utilization is determined based on the remaining computing power of the server, the available computing power of each device, and the computing power allocation result.

5. The method according to claim 2, characterized in that The computing power allocation decision model is used to predict candidate computing power allocation results for intelligent tasks in each device based on the remaining computing power of the server and the status information of the multiple devices, including: determining, based on the remaining computing power of the server, whether the server supports allocating computing power to at least one device; If supported, the computing power allocation decision model is used to predict candidate computing power allocation results for the intelligent tasks in each device based on the remaining computing power of the server and the status information of the multiple devices.

6. The method according to claim 1, characterized in that Before executing the intelligent task in the target device using the computing power allocated by the server to the target device, the method further includes: Determining the computing power to be allocated based on the remaining computing power of the target device and the computing power required by the intelligent task in the target device; The computing power allocated by the server to the target device is determined based on the computing power to be allocated, wherein the computing power allocated by the server to the target device is not less than the computing power to be allocated.

7. The method according to claim 1, characterized in that The device status information also includes the execution time of the device's smart tasks and the remaining power of the device.

8. The method according to any one of claims 1 to 7, characterized in that The device includes a vehicle, and the intelligent task includes an intelligent driving task.

9. A task processing device, characterized in that: The device comprises: An information acquisition module is used to obtain status information of each of a plurality of devices associated with the server, each of the devices having an intelligent task, and the status information of the devices includes the available computing power of the device and the computing power required for the intelligent task in the device; A computing power allocation module, configured to obtain candidate computing power allocation results for intelligent tasks in each device based on the remaining computing power of the server and the status information of the multiple devices; a device determination module configured to, when the candidate computing power allocation result corresponding to the device indicates that computing power is allocated to the intelligent task in the device and the available computing power of the device is less than the computing power required for the intelligent task of the device, obtain the device as a target device; The task execution module is used to execute the intelligent task of the target device by using the computing power allocated to the target device.

10. An electronic device, characterized in that: include: one or more processors; Memory; One or more computer-readable instructions, wherein the one or more computer-readable instructions are stored in the memory and configured to be executed by the one or more processors, the one or more computer-readable instructions being configured to perform the method according to any one of claims 1-8.

11. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer-readable instructions, which can be called by a processor to execute the method according to any one of claims 1 to 8.