A computing power resource scheduling method and system based on the vehicle Internet of Things
By adopting a computing power resource scheduling method based on the Internet of Vehicles in smart cars, the problem that smart cars are difficult to meet the high computing power needs is solved, and a stable computing expected environment and efficient computing resource utilization are achieved.
Patent Information
- Application Number
- CN202510180507.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-19
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2045-02-19
AI Technical Summary
Smart cars are difficult to meet the future demand for high computing power, especially when fully autonomous driving is achieved. Current chip processes are limited by cost, power consumption and performance.
Through a computing power resource scheduling method based on the Internet of Vehicles, an edge-side server receives vehicle-side computing tasks, determines the target computing terminal set based on task type and resource requirements, and calculates the computing power stability value of each computing terminal by responding to the change trend of the time, and finally offloads the computing task to the computing terminal with the highest computing power stability value.
It provides a stable computing expected environment, ensures that the response time of computing tasks takes into account historical execution of computing tasks, and improves the utilization efficiency of computing resources and the timeliness and reliability of data processing.
Smart Images

Figure CN119668883B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of information technology, and particularly relates to a computing power resource scheduling method and system based on the vehicle networking. Background Art
[0002] In recent years, intelligent vehicles have gradually become a hot spot in the market development. However, with the improvement of the intelligence level of intelligent vehicles, the computing power requirements for the vehicle computing platform are higher. For example, to achieve fully autonomous driving, a computing power of about 1000 TOPS is required, which has reached the computing power of the human brain. At present, however, due to the limitations of cost, power consumption and performance in the chip process, intelligent vehicles are difficult to meet the future requirements for computing power. Summary of the Invention
[0003] The Summary of the Invention is provided to introduce a set of concepts that will be further described in the following detailed implementation manners, so as to overcome one or more problems described above. The Summary of the Invention is not intended to identify the key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter.
[0004] According to an embodiment of the present invention, a computing power resource scheduling method based on the vehicle networking includes:
[0005] A first server receives a first computing task sent by a vehicle terminal;
[0006] Obtain a first task type and a first resource requirement of the computing task according to the first computing task;
[0007] Determine a target computing terminal set that meets the first computing task according to the first resource requirement;
[0008] According to the first task type, count the response duration when each computing terminal in the target computing terminal set executes the computing task, and determine the duration change trend when the computing terminal executes the computing task;
[0009] Determine a first computing power stability value of each computing terminal in the target computing terminal set based on the duration change trend;
[0010] Sort the computing terminals in the target computing terminal set based on the first computing power stability value, and determine the computing terminal with the highest computing power stability value as the target server for offloading the first computing task;
[0011] The first server is arranged on the edge side.
[0012] According to an embodiment of the present invention, the first resource requirement includes model parameters used by the first computing task, GPU resource requirement and tolerance duration;
[0013] The obtaining process of the target computing terminal set includes:
[0014] Determine a set of available first computing terminals according to the GPU resource requirements;
[0015] Determine the first weights of the computing terminals in the set of first computing terminals based on the network communication overhead;
[0016] Determine the second weights of the computing terminals in the set of first computing terminals according to the model parameters and historical computing tasks;
[0017] Determine the third weights based on at least the first weights and the second weights;
[0018] Sort the computing terminals in the set of first computing terminals in descending order according to the magnitudes of the third weights, and determine several computing terminals ranked at the front as the target set of computing terminals.
[0019] According to an embodiment of the present invention, determining the duration change trend when a computing terminal executes a computing task includes:
[0020] Obtain the computing task execution records of each computing terminal after the first moment respectively, and the available resource information of the computing terminal when the computing task starts to be executed;
[0021] Determine the first resource feature according to the resource requirement information of the computing task, the available resource information, and the response duration corresponding to the computing task;
[0022] Sort the first resource features corresponding to each computing task according to the timestamp information when the computing task starts;
[0023] Determine the duration change trend corresponding to the task type to which the computing tasks already executed by the computing terminal belong according to the sorted first resource features.
[0024] According to an embodiment of the present invention, the response duration corresponding to the computing task is the time interval from when the computing task starts to be offloaded from the first server to when the computing task is completed, or the response duration corresponding to the computing task is the time interval from when the computing task is started on the computing terminal to when the computing task is completed.
[0025] According to an embodiment of the present invention, the computing task corresponding to the computing task execution record has the same computing task type as the first computing task, or the difference between the computing task corresponding to the computing task execution record and the first resource requirement is less than a first threshold.
[0026] According to an embodiment of the present invention, the first computing power stability value is determined based on the difference value between the first response duration and the second response duration when the computing terminal executes a computing task, or based on the ratio of the first response duration to the second response duration when the computing terminal executes a computing task. The first response duration is the average response duration of several computing tasks finally executed by the computing terminal, and the second response duration is the average response duration when the computing terminal executes several computing tasks before the second time window.
[0027] According to an embodiment of the present invention, before determining the first computing power stability value, it is determined whether the duration change trend contains outliers;
[0028] In response to the existence of outliers, it is determined whether the time point corresponding to the outliers is greater than the first time threshold from the current moment;
[0029] In response to the time point corresponding to the outliers being lower than the first time threshold from the current moment, the corresponding computing terminal is removed from the target computing terminal set.
[0030] According to an embodiment of the present invention, the first computing power stability value is the reduction value of the second response duration relative to the first response duration, or the reduction ratio of the second response duration relative to the first response duration.
[0031] According to an embodiment of the present invention, the first resource requirement includes model parameters used by the first computing task, GPU resource requirements, and tolerance duration;
[0032] The process of obtaining the target computing terminal set includes:
[0033] Determine the available first computing terminal set according to the computing resource requirements;
[0034] Sort the first computing terminal set in ascending order according to the proportion of available resources, and determine several computing terminals with a low available resource ratio as the target computing terminal set.
[0035] According to an embodiment of the present invention, a computing power resource scheduling system based on vehicle-to-everything includes:
[0036] A task receiving unit, disposed on the edge side, for receiving the first computing task sent by the vehicle terminal;
[0037] A target computing terminal set obtaining unit, configured to obtain the first task type and the first resource requirement of the computing task according to the first computing task, and determine a target computing terminal set that meets the first computing task according to the first resource requirement;
[0038] The computing power stability value acquisition unit is used to count the response time of each computing terminal when executing computing tasks in the target computing terminal set according to the first task type, determine the duration change trend of the computing tasks executed by the computing terminals, and determine the first computing power stability value of each computing terminal in the target computing terminal set based on the duration change trend;
[0039] The target computing power resource selection unit sorts the computing terminals in the target computing terminal set based on the first computing power stability value, and determines the computing terminal with the highest computing power stability value as the target server for offloading the first computing task.
[0040] Through the present invention, a stable computing prediction environment can be provided, that is, the response time of the computing process takes into account the computing tasks executed in the past.
[0041] The present invention content is provided to introduce a series of concepts that will be further described in the following specific embodiments in a simplified form. The present invention content is not intended to identify the key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter. In addition, note that the present invention is not limited to the specific embodiments described in the specific embodiments and / or other parts of this document. These embodiments presented herein are for illustrative purposes only. Based on the teachings contained herein, other embodiments will be apparent to those skilled in the relevant art(s). BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 It is a schematic flow chart of a computing power resource scheduling method based on the Internet of Vehicles in an embodiment of the present invention;
[0043] Figure 2 It is a schematic structural diagram of a computing power resource scheduling system based on the Internet of Vehicles in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0044] In order to enable those skilled in the art of this technology to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the protection scope of the present invention.
[0045] In the description of the present invention, the terms "first", "second", etc. in the claims and the above-mentioned drawings are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order different from those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily limit to those clearly listed steps or units, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices. The present invention is applicable to a general vehicle-road-cloud-network structure. For a better understanding of the solution of the present invention, the implementation of some calculation processes is explained as follows.
[0046] The vehicle is connected to the server through a wireless network. Here, the server can be a cloud server or a server on the edge side; among them, the server on the edge side wirelessly covers the vehicle through the corresponding network base station; correspondingly, the edge network bandwidth used here is the network bandwidth between the network base station corresponding to the target edge server and the cloud server; the cloud network bandwidth is the network bandwidth between the cloud server and the external network of the network cluster.
[0047] The cloud server can include multiple GPU computing power servers, so that computing tasks can be executed in the cloud. Although the cloud is used to describe the computing power resources here, some of the resources used in the calculation are on the edge side. Since the allocation process of the computing resources is invisible to the vehicle side, in the present invention, only the edge server set on the edge side for communicating with the vehicle is considered. This edge server is connected to the computing power resources set in the cloud or on the edge side to perform offloading of the computing power resources.
[0048] In implementation, the edge server can obtain the running load parameters of each computing power node in the network cluster and the edge network bandwidth, as well as the cloud network bandwidth of the computing power nodes in the network cluster. The acquisition of these parameters provides necessary data support for the subsequent selection of computing nodes. And each computing power resource obtains the resource occupancy of resources such as CPU, memory, GPU, and / or disk through a resource monitoring program to obtain the proportion of system available resources in one or more dimensions. Through this information, the system can accurately understand the current load situation and network transmission capacity of each target edge server, providing data support for the subsequent selection of computing nodes. The resource proportion of each computing node can be saved to a location accessible by the edge side server, such as in a cache database, and the available computing nodes are configured by screening the types and available resources.
[0049] In most cases, the server with the most available resources is determined as the target server, which is most suitable for performing computational tasks under the current load and network bandwidth conditions, so as to ensure that the computational tasks are assigned to the server with the lightest load, thereby improving the utilization efficiency of computing resources and avoiding server overload. However, in some scenarios, it is found that when the same model executes the same type of computational task, the response time will jitter, which may be caused by the occupation of the computing resources of the host or network jitter. In some cases, it is also found that there is a certain correlation between this process and the computational task, and the above factors need to be considered so that the computational tasks are always unloaded and executed as expected, that is, from the vehicle side, the trend of the response time of the computational task execution should be a basically stable curve. Based on this, the solution of the present invention is provided as follows.
[0050] According to an embodiment of the present invention, a computing power resource scheduling method based on the vehicle network includes:
[0051] The first server receives a first computational task sent by the vehicle side;
[0052] Obtain the first task type and the first resource requirement of the computational task according to the first computational task;
[0053] Determine a set of target computing terminals that meet the first computational task according to the first resource requirement;
[0054] According to the first task type, count the response time when each computing terminal executes the computational task in the set of target computing terminals, and determine the duration change trend when the computing terminal executes the computational task;
[0055] Determine the first computing power stability value of each computing terminal in the set of target computing terminals based on the duration change trend;
[0056] Sort the computing terminals in the set of target computing terminals based on the first computing power stability value, and determine the computing terminal with the highest computing power stability value as the target server for unloading the first computational task;
[0057] The first server is arranged on the edge side.
[0058] Wherein, after passing authentication, the vehicle side is connected to the first server and sends the computational task to the first server; the first server is arranged on the edge side to minimize the occupation of delay and bandwidth;
[0059] The first server receives the first computing task sent by the vehicle terminal. This computing task can include information such as the input value of the computing task, the resource requirements of the computing task, the model name, and the model size. Among them, the input value of the computing task can be data obtained by collection, such as raw data obtained by sensors, image information, etc. After these information are sent to the first server, they are cached. After determining the terminal to be offloaded later, the corresponding data are sent to the offloading terminal that actually executes the computing task. After data processing, the data are used as the input of the model and the data processing result is obtained. This data processing result is sent to the second server, and after that, it is sent to the vehicle terminal by the first server to complete a computing task.
[0060] To complete the computing task, the first server first needs to determine the type of the computing task and determine the computing nodes with stable performance according to the type of the computing task. Here, the first task type and the first resource requirements of the computing task are obtained according to the first computing task. Among them, the type of the computing task is used to indicate the division of the performance of specific computing nodes when allocating resources, such as image recognition, video decoding, or complex algorithms. In some cases, it can rely on the data source to distinguish the computing task. For example, tags such as cameras, lidar, and millimeter-wave radars can be used as auxiliary tags for the computing task type, while path planning, solving the optimal value, risk assessment, target classification, target detection, semantic segmentation, and instance segmentation can be used as tags at the application level. Based on the above tags, the classification of the computing task can be realized. The requirements for computing resources are configured according to the computing task. For example, the operation of the model depends on the minimum memory size or the video memory resources of the GPU. When the computing resources do not meet the resource requirements, there will be situations where the execution of the computing task times out or the result is untrustworthy.
[0061] After that, the set of target computing terminals that meet the first computing task is determined according to the first resource requirements. When screening, factors such as bandwidth or network latency can also be considered to avoid nodes with insufficient network resources being used as the nodes that actually execute the computing task.
[0062] After that, according to the first task type, the response duration of each computing terminal when executing the computing task is counted in the set of target computing terminals, and the duration change trend corresponding to the task type to which the computing task already executed by the computing terminal belongs is determined. This process determines the stability of a computing node executing a computing task based on the response stability trend of the computing terminal for one or more computing tasks. Regarding the duration change trend, a first time window can be determined, and the state of the system is sampled at a speed of, for example, once every 2 seconds according to the first time window to determine the state of each computing node at each moment. Taking an example to illustrate, the state is defined as , where is the state of the system, is the idle ratio of the CPU, is the number of free memories, is the video memory in the idle state. Other parameters that can be configured to define the system state can also include the read / write capabilities of the hard disk and network quality information. Obviously, the system state at this time does not reflect the association with the executed task. At this time, the information related to the task is added to the system state, that is is changed to , where is the execution task when the corresponding response duration; since when no task is executed, the parameters do not have reference significance. Therefore, when calculating the trend of the duration change, the system state parameters used are composed of the available resources of the system and the response duration when calculating the corresponding calculation task. Since in most cases, it is not a continuous execution of one type of task. For example, within a time window, due to the execution of a calculation task by the first computing terminal, when the first server allocates a second calculation task, since the first computing terminal is occupied, the second computing terminal actually executes the calculation task. At this time, for a specific computing node, corresponding to multiple types of calculation tasks, multiple calculation tasks and corresponding state sequences can be obtained; then select one or more of the state sequences and determine whether the change trend of the response duration is increasing or decreasing to determine whether the expected response time is within the expected range; The present invention is at least based on the following assumption, that is, when the duration of the task finally executed by a computing node is , when it executes the same task, the corresponding task duration is at least ; Obviously, if the actual execution duration is lower than when the same task is newly executed, it is obviously what is expected; and when it is higher than , it meets the expectation of the present invention, that is, the computing node with a task duration lower than will have a higher score and will be selected. Further, since the tasks actually executed by a computing node will change, then regardless of whether the types of tasks it executes change, a series of task change trends obtained based on multiple task types can be used to determine whether the computing performance of the computing node is deteriorating; thereby performing performance evaluation of the computing node based on this; It should be noted that the response duration can be the time interval from unloading the calculation task to completing the calculation task, or the interval from determining the unloading node of the calculation task to the computing node completing the calculation task. When using the latter, obviously the network overhead has been taken into account;
[0063] After that, the first computing power stability value of each computing terminal in the target computing terminal set can be determined based on the duration change trend; the computing power stability value can be defined in various forms. For example, it can be defined as the percentage change in response time when executing the same type of task. For example, when the response time increase ratio for the last 5 executed tasks and the historical execution of the same type of tasks is , it is defined as If is greater than 1, it means the duration decreases and the system performance does not affect the execution of the computing task; if it is a value less than 1, it means the system performance has declined;
[0064] Furthermore, it can also be the difference between a preset waiting duration and the increased value of the response time when executing the same type of task. Obviously, according to this definition, if the response time increases, the stability value obtained by this algorithm is lower.
[0065] In an embodiment of the present invention, according to the first task type, the response duration of each computing terminal when executing computing tasks is statistically counted in the target computing terminal set, so as to determine the duration change trend corresponding to the task type to which the executed computing tasks of the computing terminal belong. This is applicable in most cases because most of the same type of tasks are executed on terminals with matching computing power.
[0066] It should be understood that the corresponding parameters can adopt inconsistent definitions. For example, directly using the increase ratio or increased value of the response time as the first computing power stability value parameter. In this case, a smaller value can be taken as a better benchmark;
[0067] After that, based on the first computing power stability value, the computing terminals in the target computing terminal set are sorted, and the computing terminal with the highest computing power stability value is determined as the target server for offloading the first computing task.
[0068] When processing the scheduling of computing power in the manner of the present invention, it can make it possible to obtain relatively stable nodes as the nodes for offloading actual computing tasks when meeting resource requirements, so that the data processing is more timely and reliable.
[0069] According to an embodiment of the present invention, the first resource requirement includes model parameters used by the first computing task, GPU resource requirements, and tolerance duration;
[0070] The process of obtaining the target computing terminal set includes:
[0071] Determine the available first computing terminal set according to the GPU resource requirements;
[0072] Based on the network communication overhead, determine the first weight of the computing terminals in the first computing terminal set;
[0073] Determine the second weight of the computing terminals in the first computing terminal set according to the model parameters and historical computing tasks;
[0074] Determine the third weight based at least on the first weight and the second weight;
[0075] Sort the computing terminals in the first computing terminal set in descending order according to the magnitude of the third weight, and determine several computing terminals ranked at the front as the target computing terminal set.
[0076] The model parameters used in the first computing task can be the name, size (0.5B, 7B or 32B), precision (such as int8 or FP16), GPU resource requirements (such as 3GB, 6GB, etc.) and tolerance duration (such as 100ms, 10ms) of the general model; among them, the video memory requirement of the GPU has a greater impact in most cases (when the computing nodes update the hardware devices in a timely manner), because most of the subsequent text determines the computing resources through the video memory; since the network communication overhead needs to be considered during the computing process, the network communication overhead here can be obtained through measurement, such as sending execution data packets and tracking to obtain. On the basis of what has been described above, the available first computing terminal set can be determined according to the GPU resource requirements, which meets the requirements of in-vehicle computing tasks, and the matching computing terminals can be determined according to one or more screening items. At this time, the obtained computing terminals have consistent weights and can be directly returned.
[0077] In the present invention, further determine the first weight of the computing terminals in the first computing terminal set based on the network communication overhead. Nodes with lower latency have higher weights; in an example of the present invention, the first weight is defined as , where is the latency;
[0078] Determine the second weight of the computing terminals in the first computing terminal set according to the model parameters and historical computing tasks. Here, nodes that meet the computing requirements and have a higher resource idle ratio have higher weights, and computing nodes with a lower response duration during the execution of historical tasks have higher weights; in an embodiment of the present invention, the second weight is the idle resource ratio of the computing nodes that meet the first computing resource requirements, and the idle resource ratio is the product of the GPU video memory idle ratio, the CPU idle ratio, the memory idle ratio and the average response duration of several recently executed computing tasks;
[0079] Determine the third weight based at least on the first weight and the second weight; here, the product of the first weight and the second weight is selected as the final weight, and the computing terminals are sorted; in some embodiments, other factors can also be considered, such as determining a weight based on the execution records of the first task type, so as to reduce the loading duration when there is a cache in the computing node.
[0080] According to an embodiment of the present invention, determining the duration change trend when a computing terminal executes a computing task includes:
[0081] Obtain the computing task execution records of each computing terminal after the first moment respectively, and the available resource information of the computing terminal when the computing task starts to execute;
[0082] Determine the first resource feature according to the resource requirement information of the computing task, the available resource information, and the response duration corresponding to the computing task;
[0083] Sort the first resource features corresponding to each computing task according to the timestamp information when the computing task starts;
[0084] Determine the duration change trend corresponding to the task type to which the computing tasks already executed by the computing terminal belong according to the sorted first resource features.
[0085] Wherein, the first moment can be a preset duration such as 1h, and it should not be too high to avoid deviation from the actual system situation. Obtain the system state, resource information, and response duration at the start of each computing task within this time period, and construct a vector list based on this, and sort it according to the task start time corresponding to the vector list, so as to obtain a list. Based on this, the duration change trend can be analyzed. In a more specific implementation, select the last executed computing task as the computing benchmark, determine the task type and the corresponding duration change trend to which the corresponding computing task belongs, and use it as the basis for whether the computing performance deteriorates.
[0086] According to an embodiment of the present invention, a plurality of computing terminals are configured to execute tasks of a preset type, and the specific task type is considered when determining the duration change trend when the computing terminal executes a computing task; at this time, the corresponding steps are:
[0087] Obtain the computing task execution records of each computing terminal after the first moment respectively, and the available resource information of the computing terminal when the computing task starts to execute, and the computing tasks included in the computing task execution records have the first task type;
[0088] Determine the first resource feature according to the resource requirement information of the computing task, the available resource information, and the response duration corresponding to the computing task;
[0089] Sort the first resource features corresponding to each computing task according to the timestamp information when the computing task starts;
[0090] Determine the duration change trend corresponding to the task type to which the computing tasks already executed by the computing terminal belong according to the sorted first resource features.
[0091] According to an embodiment of the present invention, the response duration corresponding to the computing task is the time interval from when the computing task starts to be offloaded from the first server until the computing task is completed, or the response duration corresponding to the computing task is the time interval from when the computing task is started on the computing terminal until the computing task is completed.
[0092] Two different types of time intervals can be used to define the response duration. In the case of good network, the duration from the end of offloading to the completion of the task can be used; while in the case of high requirements for response latency, the duration from the start of offloading to the completion of the task can be selected.
[0093] According to an embodiment of the present invention, the computing task corresponding to the computing task execution record has the same computing task type as the first computing task, or the difference between the computing task corresponding to the computing task execution record and the first resource requirement is less than the first threshold.
[0094] When the computing task corresponding to the computing task execution record has the same computing task type as the first computing task, the nodes screened based on the duration change trend have executed approximate tasks or processed computing tasks with approximate resource consumption. When they execute the computing task again, their expected values can generally be more consistent with the response duration when actually executing the computing task; however, since the computing tasks are not always continuous, at this time, considering the demand for computing resources, if the required resources are approximate, the deviation of the duration to complete the computing task should also be within an acceptable range. At this time, select the computing node whose difference between the computing task corresponding to the computing task execution record and the first resource requirement is less than the first threshold to provide an approximate computing expectation.
[0095] According to an embodiment of the present invention, the first computing power stability value is determined based on the difference value between the first response duration and the second response duration when the computing terminal executes the computing task, or is determined based on the ratio of the first response duration and the second response duration when the computing terminal executes the computing task. The first response duration is the average value of the response durations of several computing tasks finally executed by the computing terminal, and the second response duration is the average value of the response durations when the computing terminal executes several computing tasks before the second time window.
[0096] The first computing power stability value is mainly the difference between the response values of several recently executed computing tasks and the computing tasks executed in a period before the current moment, and in this way, it is determined whether the performance deteriorates when the computing task is executed.
[0097] According to an embodiment of the present invention, before determining the first computing power stability value, it is determined whether the duration change trend contains outliers;
[0098] In response to the existence of outliers, it is determined whether the time point corresponding to the outliers is greater than the first time threshold from the current moment;
[0099] In response to the time point corresponding to the outlier being lower than the first time threshold from the current moment, the corresponding computing terminal is removed from the target computing terminal set.
[0100] Some points can be determined as outliers through the Local Outlier Factor (LOF) algorithm; here, it is premised that in most cases, the system works normally, and when an anomaly occurs, the computing node will be processed and resume normal operation. Therefore, in most cases, outliers are points whose response duration exceeds the expected value or the mean. By determining whether the outlier is the most recently executed computing task, the system state can be determined whether it is abnormal. If it is abnormal, it can be removed from the target computing terminal set. The corresponding computing node can be executed in other computing tasks or other computing power scheduling strategies.
[0101] According to an embodiment of the present invention, the first computing power stability value is the reduction value of the second response duration relative to the first response duration, or the reduction ratio of the second response duration relative to the first response duration.
[0102] In this way, the reliability of the computing power stability value can be reliably quantified. When it is positive, it represents system optimization, and when it is negative, it represents system degradation.
[0103] According to an embodiment of the present invention, the first resource requirement includes model parameters used by the first computing task, GPU resource requirements, and tolerance duration;
[0104] The process of obtaining the target computing terminal set includes:
[0105] Determine the available first computing terminal set according to the computing resource requirements;
[0106] Sort the first computing terminal set in ascending order according to the proportion of available resources, and determine several computing terminals with a low available resource ratio as the target computing terminal set.
[0107] In this way, nodes with a higher load ratio can be selected, and the available resources in the next cycle can be maximized.
[0108] According to an embodiment of the present invention, there is also provided a computing power resource scheduling system based on the vehicle-to-everything network. The system specifically includes:
[0109] A task receiving unit, disposed on the edge side, for receiving the first computing task sent by the vehicle terminal;
[0110] A target computing terminal set obtaining unit, for obtaining the first task type and the first resource requirement of the computing task according to the first computing task, and determining the target computing terminal set that meets the first computing task according to the first resource requirement;
[0111] A computing power stability value acquisition unit, configured to count the response time of each computing terminal when executing a computing task within a target computing terminal set according to a first task type, determine the duration change trend of the computing task executed by the computing terminal, and determine the first computing power stability value of each computing terminal in the target computing terminal set based on the duration change trend;
[0112] A target computing power resource selection unit, ranks the computing terminals in the target computing terminal set based on the first computing power stability value, and determines the computing terminal with the highest computing power stability value as the target server for offloading the first computing task.
[0113] Those of ordinary skill in the art can realize that the modules and algorithm steps described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.
[0114] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described devices and equipment can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.
[0115] In the embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division, and there may be other division methods in actual implementation. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection can be through some interfaces, and the indirect coupling or communication connection of the devices or modules can be in an electrical, mechanical or other forms.
[0116] The modules described as separate components may or may not be physically separated, and the components displayed as modules may or may not be physical modules, that is, they may be located in one place, or may be distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of the embodiments of the present invention.
[0117] In addition, the various functional modules in the embodiments of the present invention can be integrated into one processing module, or each module can exist physically alone, or two or more modules can be integrated into one module.
[0118] When the above-mentioned functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method for sending / receiving energy-saving signals in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, ROM, RAM, magnetic disks, or optical discs that can store program codes.
[0119] The above description is only for the preferred embodiments of the present application and the explanation of the applied technical principles. Those skilled in the art should understand that the scope of the invention involved in the present application is not limited to the technical solution formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the inventive concept. For example, the technical solutions formed by mutually replacing the above features with the (but not limited to) technical features with similar functions disclosed in the present application.
[0120] It should be understood that the magnitudes of the sequence numbers of the steps in the content of the present invention and the embodiments do not absolutely mean the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention. For the purpose of illustration and description, the foregoing description of the implementation of the present disclosure has been given. The foregoing description is not exhaustive nor is it intended to limit the present disclosure to the exact form disclosed. According to the above teachings, various modifications and variations are possible, or various modifications and variations may be obtained from the practice of the present disclosure. These embodiments are selected and described to illustrate the principles of the present disclosure and its practical applications, so that those skilled in the art can utilize the present disclosure in various embodiments and various modifications suitable for the specific purposes conceived.
Claims
1. A computing resource scheduling method based on Internet of Vehicles, characterized in that: include: The first server receives a first computing task sent by the vehicle end; Acquire a first task type and a first resource requirement of the first computing task according to the first computing task; Determine a target computing terminal set that meets the first computing task according to the first resource requirement; According to the first task type, counting the response time of each computing terminal in the target computing terminal set when executing the computing task, and determining the change trend of the response time of each computing terminal when executing the computing task; Determine a first computing power stability value of each computing terminal in the target computing terminal set based on a response time change trend; Sort computing terminals in the target computing terminal set based on the first computing power stability value, and determine the computing terminal with the highest computing power stability value as the target server for offloading the first computing task; The first server is arranged at the edge side; The process of determining the change trend of the response time when each computing terminal performs a computing task includes: Respectively obtain the computing task execution record of each computing terminal after the first moment, and the available resource information of the computing terminal when the computing task execution starts; Determine a first resource feature according to resource requirement information of the computing task, available resource information, and a response time corresponding to the computing task; Sort the first resource features corresponding to each computing task according to the timestamp information when the computing task starts; Determine, according to the sorted first resource feature, a change trend of the response time corresponding to the task type to which the computing task executed by the computing terminal belongs; The first computing power stability value is an increase ratio or increase value of a response time when each computing terminal performs a computing task.
2. A computing resource scheduling method based on Internet of Vehicles as claimed in claim 1, characterized in that: The first resource requirement includes model parameters, GPU resource requirements and tolerance duration used by the first computing task; The process of acquiring the target computing terminal set includes: Determine a first available computing terminal set according to GPU resource requirements; Determining a first weight of a computing terminal within a first computing terminal set based on network communication overhead; Determine a second weight of a computing terminal in the first computing terminal set according to the model parameters and the historical computing tasks; determining a third weight based at least on the first weight and the second weight; Sort the computing terminals in the first computing terminal set in descending order according to the size of the third weight, and determine a number of computing terminals ranked first as the target computing terminal set; The second weight is the idle resource ratio of the computing node that meets the first computing resource requirement, and the idle resource ratio is the product of the GPU video memory idle ratio, the CPU idle ratio, the memory idle ratio and the average response time of several recently executed computing tasks.
3. The method for scheduling computing resources based on the Internet of Vehicles according to claim 1, characterized in that: The response duration corresponding to the computing task is the time interval from when the computing task is unloaded from the first server to when the computing task is completed, or the response duration corresponding to the computing task is the time interval from when the computing task is started at the computing terminal to when the computing task is completed.
4. The method for scheduling computing resources based on the Internet of Vehicles according to claim 1, characterized in that: The computing task corresponding to the computing task execution record has a computing task type consistent with the first computing task, or the difference between the computing task corresponding to the computing task execution record and the first resource requirement is less than a first threshold.
5. The method for scheduling computing resources based on the Internet of Vehicles according to claim 1, characterized in that: The first computing power stability value is determined based on the difference between the first response time and the second response time when the computing terminal performs a computing task, or based on the ratio of the first response time and the second response time when the computing terminal performs a computing task, the first response time is the average response time of several computing tasks last executed by the computing terminal, and the second response time is the average response time when the computing terminal performs several computing tasks before the second time window.
6. A method for scheduling computing resources based on Internet of Vehicles as claimed in claim 5, characterized in that: Before determining the first computing power stability value, determine whether the response time change trend contains outliers; In response to the existence of the outlier, determining whether the time point corresponding to the outlier is greater than a first time threshold from the current time; In response to the time point corresponding to the outlier being less than a first time threshold from the current moment, the corresponding computing terminal is removed from the target computing terminal set.
7. A method for scheduling computing resources based on Internet of Vehicles as claimed in claim 5, characterized in that: The first computing power stability value is a reduction value of the second response time relative to the first response time, or a reduction ratio of the second response time relative to the first response time.
8. The method for scheduling computing resources based on the Internet of Vehicles according to claim 1, characterized in that: The first resource requirement includes model parameters, GPU resource requirements and tolerance duration used by the first computing task; The process of acquiring the target computing terminal set includes: Determining a first available computing terminal set according to the first resource requirement; The first computing terminal set is sorted from small to large according to the proportion of available resources, and a number of computing terminals with low available resource ratios are determined as the target computing terminal set.
9. A computing resource scheduling system based on Internet of Vehicles, characterized in that: include: A task receiving unit, arranged at the edge side, for receiving a first computing task sent by the vehicle end; a target computing terminal set acquisition unit, configured to acquire a first task type and a first resource requirement of the first computing task according to the first computing task, and determine a target computing terminal set that satisfies the first computing task according to the first resource requirement; A computing power stability value acquisition unit, used to count the response time of each computing terminal when executing a computing task in the target computing terminal set according to the first task type, determine the response time change trend of each computing terminal when executing the computing task, and determine the first computing power stability value of each computing terminal in the target computing terminal set based on the response time change trend; a target computing power resource selection unit, which sorts computing terminals in the target computing terminal set based on the first computing power stability value, and determines the computing terminal with the highest computing power stability value as the target server for offloading the first computing task; The process of determining the change trend of the response time when each computing terminal performs a computing task includes: Respectively obtain the computing task execution record of each computing terminal after the first moment, and the available resource information of the computing terminal when the computing task execution starts; Determine a first resource feature according to resource requirement information of the computing task, available resource information, and a response time corresponding to the computing task; Sort the first resource features corresponding to each computing task according to the timestamp information when the computing task starts; Determine, according to the sorted first resource feature, a change trend of the response time corresponding to the task type to which the computing task executed by the computing terminal belongs; The first computing power stability value is an increase ratio or increase value of a response time when each computing terminal performs a computing task.
Citation Information
Patent Citations
Computing power resource measurement method based on deep reinforcement learning
CN115168027A
Task allocation method, system and device and nonvolatile storage medium
CN116775304A