Delay prediction method and device, electronic device, and storage medium

By combining Gaussian sampling and multi-layer perceptron convolution operations, the limitations of hierarchical modeling and uniform sampling methods are overcome, accurate prediction of multi-scale parallel network delays is achieved, and the accuracy of delay prediction and resource utilization efficiency are improved.

CN114897126BActive Publication Date: 2025-10-03GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210535322.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-17
Publication Date
2025-10-03
Estimated Expiration
2042-05-17

AI Technical Summary

Technical Problem

In the existing technology, the hierarchical modeling method cannot accurately fit the runtime delay of multi-scale parallel networks, and the uniform subnet sampling method cannot accurately obtain the subnet structure, resulting in inaccurate delay prediction and waste of computing resources.

Method used

The mixed Gaussian sampling method is used to sample the target model to obtain multiple subnetwork structures, and convolution operations are performed through a multi-layer perceptron to determine the prediction results of the delay information.

Benefits of technology

The accuracy and efficiency of delay information are improved, the comprehensiveness and application scope of subnet network structure are increased, and the consumption of computing resources is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114897126B_ABST
    Figure CN114897126B_ABST
Patent Text Reader

Abstract

The disclosed embodiments relate to a method and apparatus for delay prediction, an electronic device, and a storage medium, and relate to the field of computer technology. The delay prediction method includes: performing mixed Gaussian sampling on a target model corresponding to a target operation to obtain multiple subnet network structures of the target model; performing a convolution operation on the multiple subnet network structures to perform delay prediction, determining a predicted result of the target model's delay information, and performing the target operation on a processing object based on the predicted result. The technical solutions in the disclosed embodiments can improve the prediction results of the target model's delay information and achieve universality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technology, and in particular to a delay prediction method, a delay prediction device, an electronic device, and a computer-readable storage medium. Background Art

[0002] In order to accurately evaluate the performance of a machine learning model, its delay parameters can be predicted.

[0003] In related technologies, hierarchical modeling is generally used to fit the overall model's operational latency. However, this approach cannot accurately fit the operational latency of multi-scale parallel networks. Furthermore, because hierarchical modeling requires layer-by-layer sampling and measurement, it is only applicable to simple latency modeling scenarios and has certain limitations. Furthermore, uniform subnet sampling cannot accurately capture the subnet structure, resulting in low sampling efficiency and a significant waste of computing resources.

[0004] It should be noted that the information disclosed in the above background technology section is only used to enhance the understanding of the background of the present disclosure, and therefore may include information that does not constitute prior art known to ordinary technicians in the field. Summary of the Invention

[0005] The purpose of the present disclosure is to provide a method and apparatus for delay prediction, an electronic device, and a storage medium, thereby overcoming, at least to a certain extent, the problem of inaccurate delay prediction caused by limitations and defects of related technologies.

[0006] Other features and advantages of the present disclosure will become apparent from the following detailed description, or may be learned in part by practice of the present disclosure.

[0007] According to a first aspect of the present disclosure, a delay prediction method is provided, comprising: performing mixed Gaussian sampling on a target model corresponding to a target operation to obtain multiple subnet network structures of the target model; performing a convolution operation on the multiple subnet network structures to perform delay prediction, determining a prediction result of the delay information of the target model, and performing the target operation on the object to be processed based on the prediction result.

[0008] According to a second aspect of the present disclosure, a delay prediction device is provided, including: a subnet acquisition module, used to perform mixed Gaussian sampling on a target model corresponding to a target operation to obtain multiple subnet network structures of the target model; a delay information prediction module, used to perform a convolution operation on the multiple subnet network structures to perform delay prediction, determine a prediction result of the delay information of the target model, and perform the target operation on the object to be processed based on the prediction result.

[0009] According to a third aspect of the present disclosure, an electronic device is provided, comprising: a processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to execute the delay prediction method of the first aspect and its possible implementation method by executing the executable instructions.

[0010] According to a fourth aspect of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the delay prediction method of the first aspect and its possible implementation methods are implemented.

[0011] In the delay prediction method, delay prediction device, electronic device, and computer-readable storage medium provided in the embodiments of the present disclosure, on the one hand, by performing mixed Gaussian sampling on the target model corresponding to the target operation, the distribution range of the subnet network structure can be increased, thereby avoiding the limitations of the distribution range caused by the sampling method in the related art, improving the comprehensiveness and accuracy of the subnet network structure, and saving computing resources. On the other hand, by performing convolution operations on multiple subnet network structures to determine the prediction results of the delay information of each subnet network structure, and then obtaining the prediction results of the delay information of the entire target model, the distribution range of the delay information can be increased, the accuracy and effectiveness of the delay information of the predicted target model can be improved, and end-side delay prediction can be achieved, avoiding limitations, and improving convenience and application scope.

[0012] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] The accompanying drawings are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the specification, are used to explain the principles of the present disclosure. Obviously, the drawings described below are only some embodiments of the present disclosure, and those skilled in the art can derive other drawings based on these drawings without inventive effort.

[0014] Figure 1 A schematic diagram showing a system architecture to which the delay prediction method according to an embodiment of the present disclosure can be applied.

[0015] Figure 2 A schematic diagram schematically illustrates a delay prediction method in an embodiment of the present disclosure.

[0016] Figure 3 The structural diagram of the target model in the embodiment of the present disclosure is schematically shown.

[0017] Figure 4 The following schematically illustrates a flow chart of obtaining a subnet network structure in an embodiment of the present disclosure.

[0018] Figure 5 A schematic diagram of probability density functions of different sampling methods in an embodiment of the present disclosure is schematically shown.

[0019] Figure 6 The following schematically illustrates the flow chart of sampling in an embodiment of the present disclosure.

[0020] Figure 7 The following schematically illustrates a flow chart of determining prediction results in an embodiment of the present disclosure.

[0021] Figure 8 A comparison diagram schematically illustrates the time delay in an embodiment of the present disclosure.

[0022] Figure 9 The distribution diagram of the predicted delay in the embodiment of the present disclosure is schematically shown.

[0023] Figure 10 A block diagram of a delay prediction device in an embodiment of the present disclosure is schematically shown.

[0024] Figure 11 A block diagram schematically illustrates an electronic device in an embodiment of the present disclosure. DETAILED DESCRIPTION

[0025] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be implemented in a variety of forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that the present disclosure will be more comprehensive and complete and will fully convey the concepts of the example embodiments to those skilled in the art. The described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments. In the following description, many specific details are provided to provide a full understanding of the embodiments of the present disclosure. However, those skilled in the art will appreciate that the technical solutions of the present disclosure may be practiced while omitting one or more of the specific details, or that other methods, components, devices, steps, etc. may be employed. In other cases, well-known technical solutions are not shown or described in detail to avoid obscuring various aspects of the present disclosure.

[0026] In addition, the accompanying drawings are merely schematic illustrations of the present disclosure and are not necessarily drawn to scale. Identical reference numerals in the figures denote identical or similar parts, and thus repetitive descriptions thereof will be omitted. Some of the block diagrams shown in the accompanying drawings are functional entities that do not necessarily correspond to physically or logically separate entities. These functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.

[0027] In related technologies, hierarchical modeling can be used to fit the overall latency. However, compared to the serial architecture commonly used in classification tasks, hierarchical modeling cannot accurately fit the runtime latency of multi-scale parallel networks. Hierarchical modeling requires layer-by-layer sampling and measurement, making it applicable only to simple scenarios for latency modeling. Uniform subnetwork sampling cannot sample subnetwork structures with both high and low resource consumption within a limited number of samples, potentially resulting in ineffective fitting.

[0028] To solve the technical problems in the related art, the present disclosure provides a method for predicting time delays, which can be applied to scenarios where target tasks are to be achieved. The target tasks can be various types of classification tasks, such as image detection tasks, speech recognition tasks, and the like.

[0029] Figure 1 A schematic diagram shows a system architecture to which the delay prediction method and apparatus according to the embodiments of the present disclosure can be applied.

[0030] like Figure 1 As shown, the system architecture 100 may include a client 101 and a server 102. The client 101 may be an intelligent device, such as a smart phone, a computer, a tablet computer, a smart speaker or other intelligent device. The client 101 obtains the object to be processed and sends the object to be processed to the server 102, so that the server 102 processes the object to be processed according to the target model corresponding to the target operation. The object to be processed may include, for example, an image to be processed, a voice to be processed, a text to be processed, etc., which may be determined specifically according to the type of target task represented by the target operation. The server 102 may be a background system that provides delay prediction related services in the embodiment of the present disclosure, and may include a portable computer, a desktop computer, a smart phone or other electronic device with computing functions or a cluster formed by multiple electronic devices, which is used to process the object to be processed sent by the client and the target model corresponding to the target task. In addition, the client may also not need to send the object to be processed to the server, but only perform delay prediction through the client itself, and based on the prediction result of the delay information, perform target operations such as classification, segmentation, and detection on the object to be processed through the target model.

[0031] This delay prediction method can be applied to the application scenario of delay prediction for the target model corresponding to the target operation. Figure 1 As shown in , the client samples the target model represented by the search space associated with the target operation to obtain multiple subnet network structures of the target model represented by the search space. Furthermore, the multiple subnet network structures can be input into a multilayer perceptron for convolution. The multilayer perceptron performs latency prediction on the multiple subnet network structures, and determines a prediction result for the latency information of the target model including the multiple subnet network structures.

[0032] The server 102 may be the same as the client 101 , that is, both the client 101 and the server 102 are smart devices, such as smart phones.

[0033] It should be noted that the latency prediction method provided in the embodiments of the present disclosure can be executed by the server 102. Accordingly, the latency prediction method can be set in the server 102 through a program or other means. The latency prediction method provided in the embodiments of the present disclosure can also be executed by the client 101. Accordingly, the latency prediction method can be set in the client 101 through a program or other means. In the embodiments of the present disclosure, the latency prediction method is described as being executed by the client.

[0034] Next, take the client as the execution subject, refer to Figure 2 The delay prediction method in the embodiment of the present disclosure is described in detail.

[0035] In step S210, mixed Gaussian sampling is performed on the target model corresponding to the target operation to obtain multiple subnet network structures of the target model.

[0036] In the embodiments of the present disclosure, the target operation can be a target task, and the target task can be various types of tasks, such as classification tasks, detection tasks, and segmentation tasks, etc., which can be determined according to the actual application scenario and actual needs. The target model can be the model used by the target operation, and can be any type of machine learning model or deep learning model. For example, the target model can be a convolutional neural network, a recursive neural network, a generative adversarial network, etc., which are not specifically limited here.

[0037] In some embodiments, a target model, or search space, can be determined. For example, the E2NAS model (E2NAS architecture) can be used as the target model and search space. E2NAS is a hardware-constrained solution called Mixer-Hard-aware NAS (MHANAS for short). It combines GDAS, FairDarts, and FBNet to achieve greater efficiency and speed. This search space not only improves hardware-based NAS but also improves hardware-free NAS searches, mitigating crashes.

[0038] Figure 3 The network structure diagram of the E2NAS model is shown schematically in Figure 3As shown in , the network structure of the E2NAS model mainly includes four stages, from the first stage to the fourth stage. Each of the four stages contains multiple basic models. The first stage only contains basic models, and the second stage to the fourth stage all include basic models and efficient fusion models. The number of basic models in each stage is twice the stage value, where the stage value is used to indicate the stage. For example, the first stage includes 2 basic models, the second stage includes 4 basic models, the third stage includes 6 basic models, the fourth stage includes 8 basic models, and so on. It should be noted that, from the first stage to the fourth stage, the number of network layers of the basic model of each stage increases successively. For example, the number of network layers of the basic model in the first stage is 1 layer, the number of network layers of the basic model in the second stage is 2 layers, the number of network layers of the basic model in the third stage is 3 layers, and the number of network layers of the basic model in the fourth stage is 4 layers. Among them, in the second stage to the fourth stage, the basic model is connected to the efficient fusion model, and they are connected in sequence according to the basic model and the efficient fusion model. Reference Figure 3 As shown in , the basic model can include: an input layer, multiple downsampling modules, and an output layer.

[0039] like Figure 3 As shown in , the input of the current stage is the output of the previous stage, and there are two inputs with the same information in the input of the current stage. That is, the output of the previous stage serves as both inputs of the current stage, and the other inputs correspond one-to-one with the remaining outputs of the previous stage. The current stage can be any stage in the network structure, such as the second, third, and fourth stages.

[0040] Continue to refer Figure 3 As shown in , the E2NAS network structure mainly includes the following parts: input layer 301, feature extraction layer 302, first stage stage1 to fourth stage stage4 (303-306), task head 307, and output layer 308. Among them, the feature extraction layer 302 is used to extract feature vectors. The task head can be used to represent the prediction head of the target task, such as the head of the detection task, the head of the segmentation task, or the head of any task, etc. Among them, the first stage includes a basic model 309, and the remaining stages include a basic model and an efficient fusion model 310. The basic model can include multiple downsampling layers, and the efficient fusion model is used to fuse the output of the basic model.

[0041] The subnet network structure refers to the network structure formed by the samples obtained by sampling the target model, which can be a part of the target model. The sizes of multiple subnet networks can be the same or different. Based on the sampling method, multiple subnet network structures can be the same or different, and multiple subnet network structures can be combined into a complete target model. Multiple subnet networks can be trained on a training data set, and their accuracy on a verification data set can be optimized as a goal. Generally speaking, a search algorithm can be used to obtain the optimal subnet network structure for the optimization target. In the embodiment of the present disclosure, multiple subnet network structures refer to networks before optimization, that is, multiple subnet network structures are not optimal subnet network structures. For example, if the target model A includes 3 subnets, then subnet network structure 1, subnet network structure 2, and subnet network structure 3 can constitute the target model A.

[0042] In some embodiments, mixed Gaussian sampling may be performed on the target model based on multiple operators of the target model to obtain multiple subnet network structures of the target model. Figure 4 The flowchart for obtaining the subnet network structure is schematically shown in FIG. Figure 4 As shown in , it mainly includes the following steps:

[0043] In step S410, multiple candidate operators are selected from multiple operators corresponding to the target model;

[0044] In step S420, mixed Gaussian sampling is performed on the target model based on the multiple candidate operators to obtain the multiple subnet network structures.

[0045] In embodiments of the present disclosure, a target model may include multiple operators, which can be used to represent all types of operators in the target model, such as convolution operators, pooling operators, or other types of operators. To improve accuracy, the multiple operators of the target model may be screened to determine multiple candidate operators from the multiple operators. The candidate operators may be a subset of the multiple operators. The candidate operators may be of the same type, such as all convolution operators, but may be of the same or different sizes. The sizes of the multiple candidate operators are specifically determined based on the model parameters of the target model. The model parameters of the target model may include, but are not limited to, the convolution kernel size, block depth, and width of the target model network to select the candidate operators. Convolution kernels of different sizes represent different fields of view, allowing for both global and local information. The larger the convolution kernel, the more types of candidate operators there are. The candidate operators may be convolution operators, such as any one or more combinations of 3×3 convolution operators, 5×5 convolution operators, or 7×7 convolution operators.

[0046] Each candidate operator corresponds to a number of convolution kernels. When the convolution kernel is 3×3, the number of convolution kernels (midchannel) can be selected from three options, such as 1x, 2x, or 4x. When the convolution kernel is 5×5, the number of convolution kernels can also be selected from three options, such as 1x, 2x, or 4x. In other words, the size of the convolution kernel is selectable, as is the number of convolution kernels. When obtaining multiple candidate operators through the search module, you can also choose whether to connect a pooling module after the search module. This is determined by the search strategy. The search strategy determines how to accurately find the optimal network structure parameter configuration. The search strategy can be an iterative process, such as random search, Bayesian optimization, evolutionary algorithm, reinforcement learning, or gradient-based algorithm. If the search strategy requires pooling, the pooling module is connected after the search module; if not, it is not connected. When connecting the pooling module, you can also choose a different number of convolution kernels. The specific method is the same as the above steps and will not be further described here.

[0047] After obtaining multiple candidate operators, the target model can be subjected to mixed Gaussian sampling through multiple candidate operators to obtain multiple subnet network structures. According to the central limit theorem, operators in a search space are randomly sampled, and the operators obtained each time are random. Assume that the structure of the target model is n layers, and each layer randomly selects one operator from k operators. Ignoring the case of an empty set, the subnet network structure is composed of n layers, that is, n random variables, and the delay of the subnet network structure is the sum of n random variables. When n approaches positive infinity, the limiting distribution is Gaussian distribution, that is, the data pair (network structure, delay) conforms to a single Gaussian distribution, such as Figure 5 The following diagram shows the probability density functions of different sampling methods. To achieve a wider distribution for client-side tasks, a subnet structure with low latency can be fitted. However, related technologies cannot control the variance of the Gaussian distribution, so they cannot fit subnet structures with low latency, resulting in certain limitations.

[0048] In order to solve the above problems, in an embodiment of the present disclosure, the target model can be sampled by a mixed Gaussian sampling method to obtain multiple subnet network structures of the target model. Figure 6 The flow chart for sampling is shown schematically in FIG. Figure 6 As shown in FIG, the method mainly includes the following steps S610 and S620:

[0049] In step S610, for each layer of the target model, a candidate operator is randomly selected from the multiple candidate operators and filtered to obtain multiple initial subnet network structures; the multiple initial subnet network structures satisfy Gaussian mixture distribution;

[0050] In step S620, the multiple initial subnet network structures are randomly sampled to obtain the multiple subnet network structures.

[0051] In the disclosed embodiment, the target model has an n-layer network structure. Assuming that one operator is randomly sampled from the candidate operators for each layer of the target model, a random sampling result, i.e., n random sampling results, can be obtained. The random sampling results can further be filtered. The filtering process can include removing empty sets to minimize the impact of empty sets on accuracy. Through random sampling and removing empty sets, multiple initial subnetwork network structures of the target model can be obtained.

[0052] Furthermore, each of the multiple initial subnet network structures can be randomly sampled again until a subnet network structure that meets the sampling criteria is obtained. Meeting the sampling criteria can mean that the distribution of the latency information of the sampled subnet network structure meets a preset distribution, or that the number of sampled subnet network structures meets a preset number. The preset distribution can be distribution in a desired area, such as in an area with low latency or in an area with high latency. For example, if the requirement is to distribute in an area with low latency, and the latency information of the sampled subnet network structure is actually more distributed in lightweight areas with low latency, then the distribution is determined to meet the preset distribution. The preset number can be set based on actual needs, for example, to 1000 or 2000, etc. Here, a preset number of 1000 is used as an example. When the preset number is 1000, assuming the target model has 32 layers, the number of candidate operators k for each layer is 5, and one operator is randomly sampled from the candidate operators for each layer, there are a total of 31 initial subnet network structures. Each initial subnet network structure can be further randomly sampled again to obtain a subnet of the initial subnet network structure, thereby obtaining multiple subnet network structures. When randomly sampling each initial subnet network structure again, 32-33 samples can be collected for each initial subnet network structure. The number of samples for each initial subnet network structure can be different and can be adjusted according to actual needs, for example, to 32 or 33, as long as all initial subnet network structures are sampled to obtain a preset number of 1,000 subnet network structures, for example, until 1,000 subnet network structures are sampled.

[0053] For example, the number of multiple subnet network structures can be expressed as formula (1):

[0054]

[0055] The delay information of multiple subnet network structures satisfies the Gaussian mixture distribution. Since the delay information of each subnet network structure satisfies a single Gaussian distribution, the Gaussian mixture distribution is used to fuse multiple single Gaussian distributions. Figure 5 As shown in , since the delay information of each subnet network structure can be a single Gaussian distribution, in the process of sampling to obtain multiple subnet network structures, the single Gaussian distribution of the delay information of each subnet network structure is fused to obtain a Gaussian mixture distribution, making the model more complex and generating diverse samples. The Gaussian mixture function used to represent the Gaussian mixture distribution is: Among them, μ k represents the mean value of the delay information, σ k is the variance of the delay information. When sampling multiple subnet network structures with a Gaussian mixture distribution, since the mean and variance of the delay information of each subnet network structure can be perceived, it is easier to control the distribution of the sampled data so that the distribution meets actual needs and avoids limitations. Based on this, using a mixed Gaussian function for sampling can obtain a wider distribution than uniform sampling directly in the search space, increasing the distribution range of the samples, and being able to adjust the distribution range of the sampled data in a timely and convenient manner, thereby improving the accuracy and comprehensiveness of the sample distribution and avoiding the limitation of the delay data in related technologies that can only cover part of the area.

[0056] Continue to refer Figure 5 As shown in , since the delay information of each subnet network structure can be a single Gaussian distribution 501, but the distribution range of the Gaussian distribution is small, it cannot fit the area with small delay. In order to obtain the distribution state of the target area, the target area can be sampled. The target area can be a first area and a second area. The first area can be, for example, an area covered by 20-40, and the second area can be, for example, a range covered by 80-100. For the first area, random sampling can be performed near the peak of the first area to obtain the result 502 of the first area; for the second area, random sampling can be performed near the peak of the second area to obtain the result 503 of the second area. Furthermore, the results of the first area and the results of the second area can be merged to obtain a merged result 504, and the merged result satisfies the Gaussian mixture distribution.

[0057] For example, assuming the target model is a 32-layer model and the number of candidate operators per layer is k = 5, there are a total of 31 initial subnet network structures. Each initial subnet network structure is further randomly sampled, and each initial subnet network structure has a single Gaussian distribution. For example, 32-33 samples are collected for each initial subnet network structure until a preset number of subnet network structures are collected, for example, a total of 1000 subnet network structures are collected.

[0058] In the embodiment of the present disclosure, by performing mixed Gaussian sampling on the target model and sensing the variance and mean of the delay information of each subnet network structure, the distribution range of the delay information of the subnet network structure can be increased, the application scope can be increased, and the accuracy of the obtained subnet network structure can be improved.

[0059] Continue to refer Figure 2 As shown in , in step S220, the multiple sub-network structures are subjected to convolution operation to perform delay prediction, the prediction result of the delay information of the target model is determined, and the target operation is performed on the object to be processed based on the prediction result.

[0060] In the embodiment of the present disclosure, in order to avoid the problems in the related art, the delay of multiple subnet network structures can be predicted by MLP (Multi-Layer Perceptrons), the prediction results of the delay information of each subnet network structure can be obtained, and the prediction results of the delay information of each subnet network structure can be merged to obtain the prediction results of the delay information of the entire target model.

[0061] Multi-layer perceptron (MLP) is used for feature fusion. For example, the multi-layer perceptron can include two fully connected layers and an activation layer. Among them, the core operation of the fully connected layer is still the convolution operation, that is, the matrix-vector product. The fully connected layer is equivalent to a feature space transformation, which integrates all features to obtain global feature information. The fully connected layer can also be regarded as an extreme case of a convolution layer, where the convolution kernel size is the input matrix size, so the height and width of the output matrix are both 1. The activation layer can be a tanh function. The activation function can increase the nonlinearity of the model, and the tanh function can increase the convergence speed. In addition, the activation layer can also be other functions, which are not specifically limited here.

[0062] After obtaining multiple subnet network structures, the multiple subnet network structures can be input into a multilayer perceptron (MLP), so that the MLP performs convolution operations on the multiple subnet network structures through fully connected layers and activation layers, thereby obtaining a prediction result of the delay information. For example, the prediction result of the delay information can be determined by combining the network parameters of the multilayer perceptron and the attribute parameters of each of the subnet network structures. The network parameters of the multilayer perceptron can be weight parameters of the multilayer perceptron, and the attribute parameters of the subnet network structure can be the variance and mean of the delay information of the subnet network structure. The variance and mean are used to represent the degree of distribution or dispersion of the delay information.

[0063] Figure 7 A flowchart for determining the prediction result is schematically shown, referring to Figure 7 As shown in , it mainly includes the following steps:

[0064] In step S710, the network parameters of the multilayer perceptron are multiplied by the variance of the delay information of each subnet network structure to obtain a processing result;

[0065] In step S720, the processing result and the average value of the delay information of each sub-network structure are added to determine a prediction result.

[0066] In the disclosed embodiments, the variance and mean of the delay information for each subnet network structure can be obtained, and then a prediction result can be obtained by combining the network parameters of the multilayer perceptron with the variance and mean of the delay information of each subnet network structure. Furthermore, the network parameters of the multilayer perceptron can be multiplied by the variance of the delay information of each subnet network structure, and the product obtained as a processing result. The processing result can then be added to the mean of the delay information of each subnet network structure to determine a prediction result for the delay information of each subnet network structure.

[0067] For example, the prediction result of the delay information can be calculated by a multi-layer perceptron:

[0068] MLP model = nn.Sequential(

[0069] nn.Linear(self.in_features,self.mid_features),

[0070] nn.Tanh(),

[0071] nn.Linear(self.mid_features,self.out_features),

[0072] latency=(MLP model*std)+mean)

[0073] Based on the multilayer perceptron and the mean and variance of each subnet's network structure, a delay prediction result for each subnet can be calculated. Because each subnet's network structure is different, the variance and mean of the delay information for each subnet also vary. Therefore, the delay prediction results for each subnet can be the same or different, depending on the variance and mean of the subnet's network structure.

[0074] Furthermore, since multiple subnet network structures can be combined into a complete target model, the delay information of multiple subnet network structures can also be combined into the delay information of the target model. Based on this, the prediction results of the delay information of each subnet network structure contained in the target model can be merged and fused to obtain the prediction results of the delay information of the target model. Exemplarily, the prediction results of the delay information of each subnet network structure can be added to obtain the prediction results of the delay information of the entire target model. For example, if the subnet network structure of the target model includes 1000, the prediction results of the delay information of 1000 subnet network structures can be added to obtain the prediction results of the delay information of the entire target model. Exemplarily, the prediction result 1, prediction result 2 to prediction result 1000 are added to obtain the prediction results of the delay information of the entire target model.

[0075] It should be noted that before inputting the subnet network structure into the multilayer perceptron, the subnet network structure may be converted into a preset data format. The preset data format may be a tensor data format to improve data processing efficiency.

[0076] After obtaining the target model's prediction results for latency information, a target operation can be performed on the object to be processed based on the prediction results. The object to be processed can be an image, speech, or text, and the specific operation type can be determined based on the actual application scenario. Based on this, the target model, which has determined the latency information prediction results, can be used to perform the target operation on the object to achieve the corresponding function. For example, the object to be processed can be classified, detected, or segmented.

[0077] In the embodiment of the present disclosure, a multi-layer perceptron (MLP) is used to calculate the prediction result of the delay information of the target model, and a plurality of subnet network structures are obtained by sampling the target model less, and then the multi-layer perceptron is used to fit the delay information of the entire target model, thereby reducing computing resources and time consumption and improving the prediction efficiency of the delay information of the target model. Moreover, by using a multi-layer perceptron to calculate the prediction result of the delay information of the target model, it is possible to realize the prediction of all types of target models on the end side, thereby improving universality and convenience. In addition, since the distribution state of the delay information of the subnet network structure can be controlled by the variance and mean of the delay information of the subnet network structure, the accuracy and rationality of the predicted delay information are improved, and the prediction effect is improved.

[0078] In the embodiment of the present disclosure, different modeling methods can be used for comparison. For example, 2500 subnet network structures can be sampled and the runtime delay of the subnet network structure can be tested in the terminal. Among them, 1000 subnet network structures are used to implement MLP fitting, and 1500 subnet network structures are used to verify the fitting effect of MLP modeling and hierarchical modeling. Figure 8 The comparison of actual and predicted latencies shown in Figure 2 shows that hierarchical modeling exhibits significant deviations in high-latency scenarios. This demonstrates that uniform subnetwork sampling cannot capture subnetwork structures with both high and low latencies within a limited number of samples, resulting in ineffective MLP fitting and low accuracy.

[0079] For example, by using a mixed Gaussian sampling method, it is possible to sample subnetwork structures with higher and lower latency within the same number of samples, thereby obtaining a wider latency distribution and increasing the distribution range of latency information. Specifically, E2NAS is used as the architecture as the target model and search space. The search module can independently select multiple operators from two different operators, 3×3Conv and 5×5Conv, as candidate operators, and each candidate operator can select three numbers of convolution kernels. All search modules can be followed by the choice of whether to connect to the pooling module Pooling Transition, and there are three numbers of convolution kernels to choose from. Using Gaussian mixed sampling and uniform sampling, 1,000 subnetwork network structures were sampled respectively, and the network runtime latency was tested on the same type of mobile phones.

[0080] Figure 9 The figure schematically shows the distribution of the predicted delay of the subnet network structure obtained using different sampling methods. Figure 9 As shown in Figure A, 1000 subnet network structures are obtained to predict the delay information. The prediction results of the delay information obtained by the mixed Gaussian sampling method cover an area between [0,150], and the prediction results of the delay information obtained by the uniform subnet sampling method cover an area between [30,80]. By comparison, it can be seen that the prediction results of the delay information obtained by the mixed Gaussian sampling method cover a larger area, which covers more areas with small delays, and areas with small delays are the areas needed. And by Figure 9 As can be seen from Figure B, the delay distribution obtained by the mixed Gaussian sampling method does not fit the uniform distribution.

[0081] In summary, the technical solution provided in the embodiments of the present disclosure and its sampling method can be applied to all tasks that use partial sampled data to simulate large amounts of data behavior, including but not limited to all structural search tasks NAS. By using a multi-layer perceptron to predict the delay of the target model, the prediction efficiency and convenience of the model delay are improved, and the versatility of the delay prediction is improved. By sampling the target model through mixed Gaussian sampling to obtain multiple subnet network structures, the sampling range can be increased, and the comprehensiveness and accuracy can be improved. There is no need to sample layer by layer, so it can be applied to various types of application scenarios, which increases the scope of application. Furthermore, it avoids the problem that the uniform subnet sampling method cannot sample samples with a wider distribution range within a limited number of samples, increases the range of the subnet network structure, improves accuracy, and reduces computing resources.

[0082] The present disclosure provides a delay prediction device, referring to Figure 10 As shown in , the delay prediction device 1000 may include:

[0083] The subnet acquisition module 1001 is used to perform mixed Gaussian sampling on the target model corresponding to the target operation to obtain multiple subnet network structures of the target model;

[0084] The delay information prediction module 1002 is used to perform a convolution operation on the multiple sub-network structures to perform delay prediction, determine the prediction result of the delay information of the target model, and perform the target operation on the object to be processed based on the prediction result.

[0085] In an exemplary embodiment of the present disclosure, the subnet acquisition module includes: a candidate operator determination module, which is used to select multiple candidate operators from multiple operators corresponding to the target model; and a subnet structure determination module, which is used to perform mixed Gaussian sampling on the target model based on the multiple candidate operators to obtain the multiple subnet network structures.

[0086] In an exemplary embodiment of the present disclosure, the candidate operator determination module includes: an operator selection module, configured to determine the plurality of candidate operators according to model parameters of the target model.

[0087] In an exemplary embodiment of the present disclosure, the subnet structure determination module includes: an initial acquisition module, which is used to randomly extract a candidate operator from the multiple candidate operators for each layer of the target model and perform filtering processing to obtain multiple initial subnet network structures; the multiple initial subnet network structures constitute a Gaussian mixture distribution; and a random sampling module, which is used to randomly sample the multiple initial subnet network structures respectively to obtain the multiple subnet network structures.

[0088] In an exemplary embodiment of the present disclosure, the delay information prediction module includes: a prediction control module, which is used to perform convolution operations on the multiple subnet network structures based on a multi-layer perceptron to obtain prediction results of the delay information of the multiple subnet network structures; and a fusion module, which is used to fuse the prediction results of the delay information of the multiple subnet network structures to obtain the prediction results of the delay information of the target model.

[0089] In an exemplary embodiment of the present disclosure, the prediction control module includes: a parameter prediction module, which is used to determine the prediction result of the delay information by combining the network parameters of the multilayer perceptron and the attribute parameters of each subnet network structure.

[0090] In an exemplary embodiment of the present disclosure, the parameter prediction module includes: a first processing module, which is used to multiply the network parameters of the multilayer perceptron with the variance of each subnet network structure to obtain a processing result; and a second processing module, which is used to add the processing result and the mean of each subnet network structure to determine the prediction result.

[0091] It should be noted that the specific details of each module in the above-mentioned delay prediction device have been described in detail in the corresponding delay prediction method, and therefore will not be repeated here.

[0092] Figure 11 Schematic diagram of an electronic device suitable for implementing an exemplary embodiment of the present disclosure is shown. The terminal of the present disclosure can be configured as follows Figure 11 The form of the electronic device shown, however, needs to be explained. Figure 11 The electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.

[0093] The electronic device of the present disclosure includes at least a processor and a memory, where the memory is used to store one or more programs. When the one or more programs are executed by the processor, the processor can implement the method of the exemplary embodiment of the present disclosure.

[0094] Specifically, such as Figure 11As shown, the electronic device 1100 may include: a processor 1110, an internal memory 1121, an external memory interface 1122, a Universal Serial Bus (USB) interface 1130, a charging management module 1140, a power management module 1141, a battery 1142, an antenna 1, an antenna 2, a mobile communication module 1150, a wireless communication module 1160, an audio module 1170, a speaker 1171, a receiver 1172, a microphone 1173, an earphone interface 1174, a sensor module 1180, a display screen 1190, a camera module 1191, an indicator 1192, a motor 1193, a button 1194 and a subscriber identification module (Subscriber Identification Module, SIM) card interface 1195, etc. The sensor module 1180 may include a depth sensor, a pressure sensor, a gyroscope sensor, an air pressure sensor, a magnetic sensor, an acceleration sensor, a distance sensor, a proximity light sensor, a fingerprint sensor, a temperature sensor, a touch sensor, an ambient light sensor, and a bone conduction sensor, etc.

[0095] It should be understood that the structures illustrated in the embodiments of the present application do not constitute a specific limitation on the electronic device 1100. In other embodiments of the present application, the electronic device 1100 may include more or fewer components than shown, or may combine or separate certain components, or arrange the components differently. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0096] The processor 1110 may include one or more processing units, for example: the processor 1110 may include an application processor, a modem processor, a graphics processor, an image signal processor, a controller, a video codec, a digital signal processor, a baseband processor and / or a neural network processor (Neural-Network Processing Unit, NPU), etc. Among them, different processing units can be independent devices or integrated into one or more processors. In addition, a memory can be provided in the processor 1110 for storing instructions and data. The delay prediction method in this exemplary embodiment can be executed by an application processor, a graphics processor or an image signal processor. When the method involves processing related to a neural network, it can be executed by an NPU.

[0097] The internal memory 1121 can be used to store computer executable program code, which includes instructions. The internal memory 1121 can include a program storage area and a data storage area. The external memory interface 1122 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 1100.

[0098] The communication functions of mobile terminal 1100 are implemented through a mobile communication module, antenna 1, a wireless communication module, antenna 2, a modem processor, and a baseband processor. Antenna 1 and antenna 2 are used to transmit and receive electromagnetic wave signals. The mobile communication module can provide 2G, 3G, 4G, and 5G mobile communication solutions for mobile terminal 1100. The wireless communication module can provide wireless communication solutions such as wireless LAN, Bluetooth, and near-field communication for mobile terminal 200.

[0099] The display module is used to implement display functions, such as displaying user interfaces, images, and videos. The camera module is used to implement shooting functions, such as capturing images and videos. The audio module is used to implement audio functions, such as playing audio and capturing voice. The power module is used to implement power management functions, such as charging the battery, powering the device, and monitoring the battery status.

[0100] The present application also provides a computer-readable storage medium, which may be included in the electronic device described in the above embodiment; or may exist independently without being assembled into the electronic device.

[0101] Computer-readable storage media can be, for example, but not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination thereof. More specific examples of computer-readable storage media can include, but are not limited to, an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present disclosure, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or device.

[0102] Computer-readable storage media can transmit, propagate, or transfer programs for use by or in conjunction with an instruction execution system, apparatus, or device. Program code contained on a computer-readable storage medium can be transmitted using any suitable medium, including but not limited to wireless, wireline, optical cable, RF, or any suitable combination thereof.

[0103] The computer-readable storage medium carries one or more programs. When the one or more programs are executed by an electronic device, the electronic device implements the method described in the following embodiments.

[0104] Through the description of the above embodiments, it is easy for those skilled in the art to understand that the example embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solution according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, and includes several instructions to enable a computing device (which can be a personal computer, a server, a terminal device, or a network device, etc.) to execute the method according to the embodiments of the present disclosure.

[0105] Furthermore, the figures above are merely illustrative of the processes included in the methods according to exemplary embodiments of the present disclosure and are not intended to be limiting. It is readily understood that the processes illustrated in the figures above do not indicate or limit the temporal order of these processes. Furthermore, it is readily understood that these processes may be executed synchronously or asynchronously, for example, in multiple modules.

[0106] It should be noted that although several modules or units of the device for action execution are mentioned in the detailed description above, this division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of two or more modules or units described above can be concretized in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided into multiple modules or units to be concretized.

[0107] Those skilled in the art will readily conceive of other embodiments of the present disclosure after considering the specification and practicing what is disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary technical means in the art that are not disclosed in the present disclosure. The description and examples are to be regarded as exemplary only, and the true scope and spirit of the present disclosure are indicated by the claims. It should be understood that the present disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and that various modifications and changes can be made without departing from its scope. The scope of the present disclosure is limited only by the appended claims.

Claims

1. A delay prediction method, characterized in that: include: Performing mixed Gaussian sampling on the target model corresponding to the target operation to obtain multiple subnet network structures of the target model; Performing convolution operations on the multiple subnetwork structures to perform delay prediction, determining a prediction result of the delay information of the target model, and performing the target operation on the object to be processed through the target model corresponding to the target operation based on the prediction result; the object to be processed includes an image to be processed, a voice to be processed, or a text to be processed; The target model corresponding to the target operation is subjected to mixed Gaussian sampling to obtain multiple subnet network structures of the target model, including: Selecting multiple candidate operators from multiple operators corresponding to the target model; For each layer of the target model, randomly extract a candidate operator from the multiple candidate operators and perform filtering processing to obtain multiple initial subnet network structures; the multiple initial subnet network structures form a Gaussian mixture distribution; Random sampling is performed on the multiple initial subnet network structures to obtain the multiple subnet network structures.

2. The delay prediction method according to claim 1, wherein: The selecting a plurality of candidate operators from a plurality of operators corresponding to the target model includes: The plurality of candidate operators are determined according to the model parameters of the target model.

3. The delay prediction method according to claim 1, wherein: The performing a convolution operation on the multiple sub-network structures to perform delay prediction and determining a prediction result of the delay information of the target model includes: Performing a convolution operation on the multiple subnet network structures based on a multi-layer perceptron to obtain prediction results of the delay information of the multiple subnet network structures; The prediction results of the time delay information of the multiple subnet network structures are integrated to obtain the prediction result of the time delay information of the target model.

4. The delay prediction method according to claim 3, wherein: The performing a convolution operation on the multiple sub-network structures based on a multi-layer perceptron to obtain prediction results of the delay information of the multiple sub-network structures includes: Determine the prediction results of the delay information of the multiple subnet network structures by combining the network parameters of the multilayer perceptron and the attribute parameters of each subnet network structure.

5. The delay prediction method according to claim 4, characterized in that: The step of determining the prediction results of the delay information of the plurality of subnet network structures by combining the network parameters of the multilayer perceptron and the attribute parameters of each subnet network structure includes: Performing a multiplication operation on the network parameters of the multilayer perceptron and the variance of each subnet network structure to obtain a processing result; An addition operation is performed on the processing result and the average value of each subnet network structure to determine a prediction result of the delay information.

6. A delay prediction device, applied to a server, characterized in that: include: A subnet acquisition module is used to perform mixed Gaussian sampling on the target model corresponding to the target operation to obtain multiple subnet network structures of the target model; a delay information prediction module, configured to perform a convolution operation on the multiple subnetwork structures to perform delay prediction, determine a prediction result of the delay information of the target model, and based on the prediction result, perform the target operation on the object to be processed through the target model corresponding to the target operation; the object to be processed includes an image to be processed, a voice to be processed, or a text to be processed; The target model corresponding to the target operation is subjected to mixed Gaussian sampling to obtain multiple subnet network structures of the target model, including: Selecting multiple candidate operators from multiple operators corresponding to the target model; For each layer of the target model, randomly extract a candidate operator from the multiple candidate operators and perform filtering processing to obtain multiple initial subnet network structures; the multiple initial subnet network structures form a Gaussian mixture distribution; Random sampling is performed on the multiple initial subnet network structures to obtain the multiple subnet network structures.

7. An electronic device, characterized in that: include: processor; as well as a memory for storing executable instructions of the processor; The processor is configured to execute the delay prediction method according to any one of claims 1 to 5 by executing the executable instructions.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the delay prediction method according to any one of claims 1 to 5 is implemented.

Citation Information

Patent Citations

  • Constrained optimization estimation method for network link delay

    CN106452976A

  • Method and system for target acquisition and tracking using marine radar

    KR1020170031829A