Object processing methods and apparatus, electronic devices, storage media
By distinguishing between non-zero and zero operators in the search space, a new search strategy is adopted to obtain the target network structure, which solves the problems of network search crashes and wasted computing resources in the prior art, and achieves efficient and accurate network optimization.
Patent Information
- Application Number
- CN202210535958.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-17
- Publication Date
- 2025-12-02
- Estimated Expiration
- 2042-05-17
AI Technical Summary
Existing network search methods based on differentiable search structures are prone to crashing, leading to limitations in network structure optimization. Network search results based on hardware constraints have poor accuracy and high computational resource consumption.
By distinguishing between non-zero and zero operators in the search space, a new search strategy is adopted to select candidate operators in each layer of the network, obtain the target network structure, and combine the multilayer perceptron to predict the latency value, thereby optimizing the network search process.
It improves the accuracy and efficiency of web searches, avoids crashes, enhances the stability and interpretability of results, and reduces computational resource consumption.
Smart Images

Figure CN114723021B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of computer technology, and more specifically, to an object processing method, an object processing apparatus, an electronic device, and a computer-readable storage medium. Background Technology
[0002] In order for neural networks to generalize without overfitting the training dataset, it is crucial to find the optimal structure through network search.
[0003] In related technologies, network searches are generally performed using search methods based on differentiable search structures or network searches based on hardware constraints.
[0004] Among the methods described above, search methods based on differentiable search structures are prone to crashes during optimization, which may limit the network structure and render the search results invalid. Furthermore, hardware-constrained network search methods yield inaccurate results, lack interpretability, and require significant computational resources.
[0005] It should be noted that the information disclosed in the background section above is only used to enhance the understanding of the background of this disclosure, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention
[0006] The purpose of this disclosure is to provide an object processing method and apparatus, electronic device, and storage medium, thereby overcoming, at least to some extent, the problem of inaccurate network structure caused by the limitations and defects of related technologies.
[0007] Other features and advantages of this disclosure will become apparent from the following detailed description, or may be learned in part from practice of this disclosure.
[0008] According to a first aspect of this disclosure, an object processing method is provided, comprising: obtaining a search space; selecting candidate operators from a plurality of operators of each layer of the network in the search space; performing a network search on each layer of the network in the search space according to the candidate operators to obtain a target network structure; wherein the plurality of operators include non-zero operators and / or zero operators; and performing processing operations on the object to be processed based on the target network structure to obtain a processing result.
[0009] According to a second aspect of this disclosure, an object processing apparatus is provided, comprising: a search space determination module for acquiring a search space; a network search module for selecting candidate operators from a plurality of operators in each layer of the network in the search space, and performing a network search on each layer of the network in the search space according to the candidate operators to obtain a target network structure; wherein the plurality of operators includes non-zero operators and / or zero operators; and an operation execution module for performing processing operations on the object to be processed based on the target network structure to obtain a processing result.
[0010] According to a third aspect of this disclosure, an electronic device is provided, comprising: a processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to perform the object processing method of the first aspect and possible implementations thereof by executing the executable instructions.
[0011] According to a fourth aspect of this disclosure, a computer-readable storage medium is provided, on which a computer program is stored, which, when executed by a processor, implements the object processing method of the first aspect described above and its possible implementations.
[0012] The object processing method, object processing apparatus, electronic device, and computer-readable storage medium provided in this disclosure, on the one hand, determine candidate operators for each layer of the network by using non-zero operators and / or zero operators in the search space, and perform network search on each layer of the network based on the candidate operators to obtain the target network structure. Because the operator types are distinguished, the candidate operators for each layer of the network can be determined quickly and accurately, thereby improving the accuracy of the network search for each layer and increasing the efficiency and accuracy of the network search. On the other hand, obtaining the target network structure by performing network search on each layer of the network using the candidate operators of each layer avoids the crash problems caused in related technologies, improves stability and reliability, and enhances the rationality and interpretability of the results. It also avoids the limitations caused by crash problems in related technologies, improves the effectiveness of the network search, and reduces computational resources.
[0013] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description
[0014] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure. It is obvious that the drawings described below are merely some embodiments of this disclosure, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort.
[0015] Figure 1A schematic diagram of a system architecture to which the object processing method of the present disclosure embodiments can be applied is shown.
[0016] Figure 2 The diagram illustrates an object processing method according to an embodiment of the present disclosure.
[0017] Figure 3 The diagram illustrates the structure of the search space in an embodiment of this disclosure.
[0018] Figure 4 The schematic diagram illustrates the process of determining the selection probability of an operator in an embodiment of this disclosure.
[0019] Figure 5 The schematic diagram illustrates a process flow for data processing in an embodiment of this disclosure.
[0020] Figure 6 The schematic diagram illustrates the flowchart for determining the output result in an embodiment of this disclosure.
[0021] Figure 7 The diagram illustrates the first comparison results of different search methods in an embodiment of this disclosure.
[0022] Figure 8 This diagram illustrates a second comparison of results for different search methods in an embodiment of the present disclosure.
[0023] Figure 9 A block diagram of an object processing apparatus according to an embodiment of the present disclosure is shown schematically.
[0024] Figure 10 A block diagram of an electronic device according to an embodiment of the present disclosure is shown schematically. Detailed Implementation
[0025] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided to make this disclosure more comprehensive and complete, and to fully convey the concept of the example embodiments to those skilled in the art. The described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a full understanding of embodiments of this disclosure. However, those skilled in the art will recognize that the technical solutions of this disclosure can be practiced with one or more of the specific details omitted, or other methods, components, apparatus, steps, etc., can be employed. In other instances, well-known technical solutions are not shown or described in detail to avoid obscuring various aspects of this disclosure.
[0026] Furthermore, the accompanying drawings are merely illustrative of this disclosure and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and therefore repeated descriptions of them will be omitted. Some block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.
[0027] Search methods in related technologies are prone to crashing during optimization. This often leads to the selection of numerous parameter-free operations in the early stages of the search, such as identity activation and pooling, causing the network structure to get trapped in local solutions. This phenomenon is exacerbated by introducing hardware resource constraints, resulting in invalid search results and affecting effectiveness and accuracy. While some search methods use single-path selection during forward computation, significantly reducing training complexity, this can also lead to model crashes. Another search method can mitigate the problem of excessive crashes during the search process. However, this method cannot effectively constrain hardware resources during optimization because its differentiable search decouples each path using activation functions but lacks gradient truncation, allowing multiple paths to be selected simultaneously. It cannot accurately calculate the hardware constraints of each training path, such as network latency. Other hardware-constrained network searches generate multiple paths in each training iteration, making it impossible to precisely know the specific paths. Therefore, all losses are calculated, resulting in unreasonable and uninterpretable results.
[0028] To address the technical problems in related technologies, this disclosure provides an object processing method that can be applied to scenarios where a network search is performed on a search space to obtain a target structure network, and processing operations are performed on the object to be processed based on the target network structure to achieve the target task. The target task can be various types of classification tasks, such as image detection tasks, speech recognition tasks, etc.
[0029] Figure 1 A schematic diagram of a system architecture for object processing methods and apparatus to which embodiments of the present disclosure can be applied is shown.
[0030] like Figure 1As shown, the system architecture 100 may include a client 101 and a server 102. The client 101 can be a smart device, such as a smartphone, computer, tablet, or smart speaker. The server 102 can be the same as the client 101, meaning both are smart devices, such as smartphones. Alternatively, the server 102 may differ from the client 101; for example, it can be a backend system providing object processing-related services in this embodiment, and may include a single electronic device or a cluster of multiple electronic devices with computing capabilities, such as a portable computer, desktop computer, or smartphone, for processing the search space and the objects to be processed.
[0031] Client 101 acquires the object to be processed and sends it to server 102. Server 102 performs a network search in the search space according to a search strategy to obtain the target network structure, and then performs processing operations on the acquired object to be processed based on the target network structure. The object to be processed may include, for example, an image to be processed, speech to be processed, text to be processed, etc., depending on the type of processing operation. Alternatively, the client may not need to send the search space and the object to be processed to the server, but can simply perform a network search based on the search space and object processing on its own, obtain the target network structure through the search strategy, and perform processing operations such as classification, segmentation, and detection on the object to be processed based on the target network structure.
[0032] It should be noted that the object processing method provided in this embodiment can be executed by the server 102. Accordingly, the object processing method can be set in the server 102 through a program or other means. The object processing method provided in this embodiment can also be executed by the client 101. Accordingly, the object processing method can be set in the client 101 through a program or other means, thereby realizing network search on the client side. In this embodiment, the example of the object processing method being executed by the client is used for explanation.
[0033] This object processing method can be applied to web search scenarios. The explanation will focus on an example where the object processing method is executed by the client. (Reference) Figure 1 As shown, the client performs a network search on the model represented by the search space to obtain the optimal network structure of the search space as the target network structure. The client then performs processing operations on the object to be processed through the target network structure to obtain the processing results such as classification, segmentation, and detection corresponding to the processing operations.
[0034] Next, taking the example of the object processing method being executed by the client, refer to... Figure 2 The object processing methods in the embodiments of this disclosure will be described in detail.
[0035] In step S210, the search space is obtained.
[0036] This disclosure can be applied to web search scenarios. Web search is used to determine the optimal network structure in a search space represented by neural networks, and its essence lies in the process of automatically performing tasks to discover more complex architectures. Specifically, search strategies can be used to test and evaluate a large number of architectures in the search space, and the architecture that best satisfies the given problem objective is selected by maximizing the fitness function.
[0037] In the process of network search, the search space can be defined first, and then candidate network structures can be determined from the search space through search strategies. The candidate network structures are evaluated, and the next round of search is carried out based on the feedback to determine the optimal network structure.
[0038] Based on this, in this embodiment, the search space can first be determined. The search space can be a hardware-constrained scheme, MHANAS (Mixer-Hard-aware NAS), which combines features of GDAS (Gradient-based search using Differentiable Architecture Sampler), differentiable search Fair Darts, and FBNet, making it more efficient and faster. This search space not only improves hardware-constrained network search but also improves unconstrained network search, mitigating crashes. Because this scheme is more efficient and faster, it can also be called E2NAS.
[0039] Figure 3 The diagram illustrates the network structure of the search space. (See reference...) Figure 3 As shown, the search space, or E2NAS network structure, mainly comprises four stages: Stage 1 to Stage 4. Each stage contains multiple basic models. Stage 1 contains only basic models, while Stages 2 through 4 each include both basic models and efficient fusion models. The number of basic models in each stage is twice the stage number, where the stage number indicates the current stage. For example, Stage 1 includes 2 basic models, Stage 2 includes 4, Stage 3 includes 6, Stage 4 includes 8, and so on. It should be noted that the number of network layers in the basic models increases sequentially from Stage 1 to Stage 4. For example, Stage 1 has 1 network layer, Stage 2 has 2, Stage 3 has 3, and Stage 4 has 4. In Stages 2 through 4, basic models are connected to efficient fusion models, and this connection occurs sequentially, with basic models followed by efficient fusion models. (Reference) Figure 3As shown, the basic model may include: an input layer, multiple downsampling modules, and an output layer.
[0040] like Figure 3 As shown, the input of the current stage is the output of the previous stage, and two inputs in the current stage have the same information, meaning the output of the previous stage serves as both inputs to the current stage. The other inputs correspond one-to-one with the remaining outputs of the previous stage. The current stage can be any stage in the network structure, such as any of the second, third, and fourth stages.
[0041] Continue to refer to Figure 3 As shown, the E2NAS network structure mainly includes the following parts: input layer 301, feature extraction layer 302, first stage 1 to fourth stage 4 (303-306), task head 307, and output layer 308. The feature extraction layer 302 is used to extract feature vectors. The task head can represent the prediction head of the target task, such as the head of a detection task, a segmentation task, or an arbitrary task. The first stage includes a basic model 309, and the remaining stages, represented by the second to fourth stages, include a basic model and an efficient fusion model 310. The basic model can include multiple downsampling layers, and the efficient fusion model is used to fuse the outputs of the basic model.
[0042] It should be noted that the search space can also be any type of neural network model, as long as it can perform network search operations; no special restrictions are imposed here.
[0043] In step S220, candidate operators are selected from multiple operators in each layer of the network in the search space, and network search is performed on each layer of the network in the search space according to the candidate operators to obtain the target network structure; wherein, the multiple operators include non-zero operators and / or zero operators.
[0044] In this embodiment of the disclosure, after determining the search space, the network search problem can be modeled, and its network search objective is set to minimize the loss value corresponding to the weights of the searched target network structure, so as to obtain the target network structure. The objective of the network model search is to find the optimal network structure (target network structure) that minimizes the loss value of the weights corresponding to the target network structure. Based on this, the network search problem can be a constrained optimization problem, specifically modeled as shown in formula (1):
[0045]
[0046] Where α is the network structure parameter searched in the search space, and the network structure parameter can be used to determine the network structure; W is the convolution weight, t is the time delay budget, G is the network structure in the search space, and L is the loss value of the weight W(G(α)) corresponding to the network structure parameter α. Lat(·) represents the end-to-end time delay value (the time delay value obtained by the end measurement). Applying the Lagrange method, the above formula (1) for calculating the loss value of the weight corresponding to the network structure during the network search process can be re-expressed as formula (2):
[0047]
[0048] Here, λ is a hyperparameter dependent on t. To stabilize the search process, t can be restricted to the log space, thereby guiding the global optimization objective. Based on this, equation (2) can also be expressed as equation (3):
[0049]
[0050] It should be noted that "Lat" refers to the latency value obtained from testing the network structure on the client side, and this latency value can be modeled using a Multi-Layer Perceptron (MLP). For example, the latency information of the network structure in the search space can be predicted using an MLP to obtain the corresponding latency value. The MLP is used for feature fusion and can include two fully connected layers and an activation layer. The activation layer can be a tanh function. The network structure in the search space can be input into the MLP, which then performs convolution operations on the network structure through the fully connected layers and the activation layer to obtain the latency value of the network structure. For example, the latency value can be determined by combining the network parameters of the MLP and the attribute parameters of the network structure. The network parameters of the MLP can be the weight parameters of the MLP, and the attribute parameters of the network structure can be the variance and mean of the latency information of the network structure, for example, the variance and mean of the latency information corresponding to the subnets of the network structure. Using the MLP, the latency value of the network structure in the search network can be accurately obtained.
[0051] After modeling the network search problem, a search strategy can be determined. To alleviate the long-standing crash problem in resource-constrained DNAS, this disclosure proposes a novel search strategy, the inverter. The inverter explicitly separates non-zero operations and zero operations in the search space, and then performs a network search in the search space based on the non-zero and zero operators.
[0052] In some embodiments, each layer of the search space includes multiple operators. These operators can represent all types of operators in the search space, such as convolution operators, pooling operators, or other types of operators. Convolution operators can also include various types, such as 3×3 convolution operators, 5×5 convolution operators, etc. To improve accuracy, the multiple operators in each layer of the search space can be filtered to select multiple candidate operators for each layer. The selected candidate operators can be combined to form paths for each layer. For example, the paths for a certain layer in the search space might be 3×3 convolution operators and 5×5 convolution operators. For each layer of the search space, there may be multiple operators, which can include non-zero operators and zero operators. Non-zero operators can include, but are not limited to, 3×3 convolution operators, 5×5 convolution operators, 7×7 convolution operators, and the Identity operator, etc. The number of non-zero operators can be multiple, such as 10 or other values. Furthermore, the types and number of operators contained in each layer of the search space can be the same or different, depending on the specific structure of each layer and the search network. The number of candidate operators can be determined according to actual needs; for example, there can be one or more.
[0053] For example, candidate operators for each layer of the search space can be determined based on the types of operators contained in each layer. To accurately select candidate operators, during forward propagation, the selection probability of each operator can be determined based on its type, and the corresponding candidate operators can be further determined from all operators contained in each layer of the search space based on the selection probabilities. Since the types of operators are different, the selection probability for each type of operator is calculated differently to improve accuracy and specificity.
[0054] Figure 4 The flowchart illustrating the determination of the selection probability of the operator is shown in the figure. (See reference) Figure 4 As shown, the main steps include:
[0055] In step S410, it is determined whether the operator is a zero operator; if yes, proceed to step S420; if no, proceed to step S430.
[0056] In this step, each operator in each layer of the search space can be evaluated to determine the type of each operator.
[0057] In step S420, if the operator is of type zero, the selection probability of the non-zero operator is determined according to the magnitude of the first parameter of the search space.
[0058] In this step, the first parameter is used to represent the index value of the zero operator in the network structure parameters, which can be used as follows: The selection probability of an operator can also be called the operator's weight. The selection probability of a zero operator can be specifically calculated using formula (4):
[0059]
[0060] To reduce time consumption and memory usage, Gumbel-softmax and discretization strategies are used to transform formula (4). For the zero operator, the selection probability of the zero operator can be calculated based on whether the magnitude of the first parameter satisfies the numerical condition. For example, if the first parameter is greater than the numerical threshold, it can be considered that the first parameter satisfies the numerical condition. The numerical threshold can be 0.5 or other values, which are determined according to actual needs. Specifically, when the first parameter is greater than the numerical threshold of 0.5, the selection probability of the zero operator can be the first probability, which can be 1; when the first parameter is less than 0.5 or equal to 0.5, the selection probability of the zero operator can be the second probability, which can be 0. The selection probability of the zero operator can be as shown in formula (5):
[0061]
[0062] In step S430, if the operator is a non-zero operator, the selection probability of the non-zero operator is determined according to the second parameter of the search space.
[0063] In this step, the second parameter represents the index value of the non-zero operators in the network structure parameters. The second parameter can be determined based on the ratio of the network structure parameters associated with zero operators to those associated with non-zero operators, and can be expressed as follows: The selection probability of a non-zero operator can be determined based on the second parameter, specifically calculated according to formula (6):
[0064]
[0065] Furthermore, to save computation time and memory usage, the selection probability of the non-zero operator can be calculated according to formula (7):
[0066]
[0067] Here, g is a variable extracted from Gumbel(0,1). In backpropagation, to achieve stable training, the selection probability of the zero operator is estimated as shown in Equation (5), while the selection probability of the non-zero operator is approximated as shown in Equation (8):
[0068]
[0069] In this embodiment, when the operator is a zero operator, its selection probability can be calculated using the sigmoid function; when the operator is a non-zero operator, its selection probability can be calculated using the softmax function. It should be noted that the sum of the selection probabilities of all operators contained in each layer of the search space is 1. By distinguishing the types of operators and using different parameters to calculate the selection probability of operators corresponding to different types, the probability of each operator can be determined more accurately, avoiding the influence of other types of operators and improving reliability and specificity.
[0070] After determining the selection probability of each operator, candidate operators can be selected from the multiple operators contained in each layer of the network based on the required number of candidate operators and their selection probabilities. For example, candidate operators can be selected based on the required number of candidate operators and whether their selection probabilities meet preset conditions. Operators with selection probabilities that meet preset conditions can be identified as candidate operators. The criteria for meeting the preset conditions differ depending on the required number of candidate operators. For example, when there is only one candidate operator, meeting the preset conditions could mean that the operator has the highest selection probability, or that the operator's selection probability appears most frequently, etc., depending on the actual needs, and is not specifically limited here. When there are multiple candidate operators, meeting the preset conditions could mean that the selection probabilities are arranged in descending order within the top N positions, depending on the actual needs.
[0071] For example, for a certain network layer, if the selection probabilities of the 3×3 convolution operator and the zero operator both meet the preset conditions, the 3×3 convolution operator and the zero operator can be identified as candidate operators for that network layer. It should be noted that the above steps can be repeated until candidate operators are selected for each network layer in the search space.
[0072] After selecting candidate operators for each layer of the network, the path for each layer can be determined based on the relationships between the candidate operators; that is, the path can be determined according to the candidate operators to be executed. The relationships can be used to represent the execution order of the candidate operators. Furthermore, the paths of each layer can be combined to determine the target network structure obtained by performing a network search on the search space. The target network structure can be a part of the search space, and it can be updated in real time according to the search strategy until the target network structure with the best performance is obtained.
[0073] It should be added that after the target network structure is found, the performance parameters of the target network structure can be evaluated, and the next round of network search can be performed based on the evaluation results to update the target network structure.
[0074] While determining the path and target network structure for each layer, input data can be fed into each layer of the search space. The input data is then processed by the candidate operators selected from each layer to obtain the output results of each layer. Figure 5 The flowchart illustrating the data processing is shown in the image. Figure 5 As shown, input data 501 is input into each layer of the network, and the input data is processed by the candidate operators of each layer of the network 502 to obtain the output result 503 of the input data in each layer of the network.
[0075] Figure 6 The diagram illustrates the specific flowchart for determining the output result. (See reference...) Figure 6 As shown, the main steps include:
[0076] In step S610, if the candidate operator is a non-zero operator, the corresponding operator operation is performed on the input data through the candidate operator to determine the first operation result.
[0077] In step S620, based on the first operation result and the weights of the non-zero operators, the output result of the path of the input data through each layer of the network in the search space is determined.
[0078] In this embodiment, a first operation result can be obtained by performing corresponding operator operations on the input data using non-zero operators. The first operation result of the input data with respect to the non-zero operators and the weights of the non-zero operators are then multiplied, and combined with the second operation result of the input data with respect to the zero operator, to determine the output result of each layer of the network in the search space. For example, the first operation result can be multiplied with the weights of the non-zero operators, and then added to the second operation result. The first operation result can be determined based on the type of non-zero operators among the candidate operators, such as a convolution result or other results; the second operation result is the result corresponding to the zero operator, such as 0. In related technologies, softmax is used to approximate the output in order to make the search space continuous, but this method may lead to severe crashes. To avoid crashes, in this embodiment, the output result of each layer of the network in the search space can be calculated using formula (9):
[0079]
[0080] Where X represents the input data and O represents the set of operators.
[0081] For example, if a candidate operator selected by a certain layer of the search space is a 3×3 convolution operator and a zero operator, then the input data X needs to undergo a convolution operation using the 3×3 convolution operator to obtain the first operation result. This result is then multiplied by the weights of the non-zero operators (i.e., the case where the zero operator is removed from the operator) and added back to the non-zero operator weights (the weights of the non-zero operators are determined based on the difference between the weights of 1 and the zero operator), thus determining the output result corresponding to the path of the input data through that layer of the network. By distinguishing operator types and combining non-zero and zero operators, the output result of each layer of the search space is adjusted and controlled through the weights of non-zero and zero operators, avoiding inaccuracies due to limitations and improving the accuracy and comprehensiveness of the output results. Repeating this process yields the output result after the input data passes through the entire target network structure. Furthermore, the performance parameters of the target network structure can be determined through the output results to evaluate the target network structure.
[0082] Continue to refer to Figure 2 As shown, in step S230, based on the target network structure, processing operations are performed on the object to be processed to obtain the processing result.
[0083] In this embodiment of the disclosure, after obtaining the target network structure that satisfies the optimization objective by performing a network search on the search space, the object to be processed can be processed by performing the corresponding operator operation on all the candidate operators contained in each layer of the target network structure, thereby completing the processing operation corresponding to the target task and obtaining the corresponding processing result.
[0084] For example, after obtaining the target network structure of the search space, processing operations can be performed on the object to be processed based on the target network structure. The object to be processed can be an image, speech, text, etc., and the specific object to be processed can be determined according to the type of processing operation and the actual application scenario. Based on this, the optimal target network structure obtained through network search can be used to perform processing operations on the object to be processed, realizing the function corresponding to the processing operation. The processing operation can be an operation under the target task scenario, specifically determined according to the type of target task, and the processing operation can include the operator operations corresponding to all candidate operators contained in the target network structure. The target task can be various types of tasks, such as classification tasks, detection tasks, and segmentation tasks, etc., and can be determined according to the actual application scenario and actual needs. Based on this, the processing operation can include, but is not limited to, classification operations, detection operations, and segmentation operations, so that the object to be processed can be classified, detected, or segmented based on the target network structure, etc., without specific limitations here.
[0085] For example, the object to be processed can be subjected to the corresponding operator operation by each candidate operator in the target network structure. According to formula (9), the first operation result is obtained by performing the corresponding operator operation on the object to be processed through the non-zero operator. The first operation result and the weight of the non-zero operator are multiplied, and then multiplied with the weight of the zero operator. Then, the result of the second operation on the object to be processed for the zero operator is added to calculate the output result of the object to be processed. The processing operation corresponding to the target task is then performed on the object to be processed based on the output result. In this way, the output result can be determined by combining the types of all candidate operators and their corresponding weights, which improves the accuracy and comprehensiveness of the output result.
[0086] In this embodiment, by determining the non-zero and / or zero operators of each layer of the network in the search space, selecting candidate operators for each layer from the non-zero and / or zero operators, and performing a network search on each layer of the network in the search space based on the candidate operators, the target network structure is obtained. Because the operator types are distinguished, the candidate operators for each layer of the network can be determined quickly and accurately, thereby improving the accuracy of the search for each layer of the network and increasing the efficiency and accuracy of the network search. Furthermore, because a precise network search is performed on each layer of the network, the accuracy of the target network structure is improved, and the rationality and interpretability of the results are enhanced. Obtaining the target network structure by performing a network search on each layer of the network using candidate operators avoids the crash problems caused by related technologies, improving stability and reliability. At the same time, it avoids the limitations caused by crash problems in related technologies, improves the effectiveness of the network search, and reduces computational resources.
[0087] In this embodiment, target network structures obtained by different network search strategies can be compared to verify the target network structure. Within the same search space, different search strategies (MHANAS, GDAS, and Fair DARTS) can be used to perform network searches, obtaining the target network structure corresponding to each strategy. For example, the performance parameters of the target network structures corresponding to different search strategies can be compared using multiple evaluation metrics. These metrics may include, but are not limited to, human pose estimation metrics and network latency values, and may also include other evaluation metrics, which are not specifically limited here. Furthermore, the performance parameters of the obtained target network structures can be compared based on multiple different levels of hardware resources. Different levels of hardware resources can be represented by penalty terms. Performance parameters, for example, can be the accuracy of the target network structure.
[0088] First, the MHANAS and GDAS search methods are compared. A fair search space is designed using DW-HRNet as the architecture, and all network transition modules are fixed. The search module can independently select one candidate operator from four different operators: 3x3 Conv, 5x5 Conv, 7x7 Conv, and the activation operator Identity. The total search space size is 4. 30 .
[0089] The GDAS search method, even without any hardware resource constraints, suffers from an overabundance of identities due to crashes. Introducing hardware resource constraints during the search process only exacerbates this problem, ultimately rendering the search results invalid.
[0090] Table 1
[0091]
[0092] As shown in Table 1, the search strategy in this embodiment, under different levels of hardware resource constraints, yields target network structures with higher accuracy than those obtained by GDAS search. (Reference) Figure 7 As shown in the figure, the accuracy of the human pose recognition index and latency value of the target network structure obtained by the search strategy provided in this embodiment is higher than that of the network structure obtained by the GDAS search method.
[0093] Next, the network structures obtained by using the MHANAS search method and the Fair DARTS search method are compared. A fair search space was designed using DW-HRNet as the architecture, and all network transition modules were fixed. The search module can independently select multiple candidate operators from three different operators: 3x3Conv, 5x5Conv, and 7x7Conv. The total search space size is (2... 3 30.
[0094] Fair DARTS cannot introduce hardware resource constraints during the search process; it can only limit the computational load of the network structure after the search is complete by using a threshold. This disclosure provides three different thresholds to obtain network structures with different computational loads. While Fair DARTS effectively improves the crash phenomenon, it cannot introduce hardware resource constraints. As shown in Table 2, the search strategy provided in this disclosure has higher accuracy than the network structures obtained by the Fair DARTS search method at the same latency, especially in scenarios with lower latency, where its accuracy is even higher. Figure 8 As shown in the image.
[0095] Table 2
[0096]
[0097] In summary, the technical solution provided in this disclosure obtains the target network structure by determining the non-zero and / or zero operators of each layer of the network in the search space and performing a network search on each layer of the network based on candidate operators. Because the operator types are distinguished, candidate operators for each layer of the network can be quickly and accurately determined, thereby improving the accuracy of the search for each layer and enhancing the efficiency and accuracy of the network search. Furthermore, the precise network search for each layer improves the accuracy of the target network structure and the reasonableness of the results. Obtaining the target network structure by performing a network search on each layer using candidate operators proposes a new approach to network model search algorithms, effectively solving / alleviating problems encountered in hardware-constrained network model search, such as the crash problem, and providing greater interpretability. Simultaneously, it avoids the limitations caused by the crash problem in related technologies, improving the effectiveness of the network search and reducing computational resources.
[0098] This disclosure provides an object processing apparatus, with reference to... Figure 9 As shown, the object processing device 900 may include:
[0099] Search space determination module 901 is used to obtain the search space;
[0100] The network search module 902 is used to select candidate operators from multiple operators of each layer of the network in the search space, and perform network search on each layer of the network in the search space according to the candidate operators to obtain the target network structure; wherein, the candidate operators include non-zero operators and / or zero operators;
[0101] The operation execution module 903 is used to perform processing operations on the object to be processed based on the target network structure to obtain the processing result.
[0102] In one exemplary embodiment of this disclosure, the network search module includes: a candidate operator determination module, configured to determine candidate operators for each layer of the network in the search space based on the type of operators contained in each layer of the network.
[0103] In one exemplary embodiment of this disclosure, the candidate operator determination module includes: a probability calculation module, configured to determine the selection probability of each operator according to the type of the operator, and determine the candidate operator from all operators according to the selection probability.
[0104] In one exemplary embodiment of this disclosure, the probability calculation module includes: a first calculation module, configured to determine the selection probability of the zero operator based on the magnitude of a first parameter of the search space if the operator is of type zero; and a second calculation module, configured to determine the selection probability of the non-zero operator based on a second parameter of the search space if the operator is of type non-zero.
[0105] In one exemplary embodiment of this disclosure, the network search module includes: a path determination module, configured to determine the path of each layer of the search space based on the candidate operators of each layer of the search space; and a network structure determination module, configured to combine the paths of each layer of the network to obtain the target network structure.
[0106] In one exemplary embodiment of this disclosure, the apparatus further includes: a non-zero operator processing module, configured to perform corresponding operations on the input data through the non-zero operator if the candidate operator is a non-zero operator, and determine a first operation result; and an output result determination module, configured to determine the output result of the path of each layer of the network in the search space based on the first operation result and the weight of the non-zero operator.
[0107] In one exemplary embodiment of this disclosure, the output result determination module is configured to: perform a multiplication operation on the operation result and the weight of the non-zero operator, and combine the second operation result of the input data for the zero operator to determine the output result of each layer of the network in the search space.
[0108] It should be noted that the specific details of each module in the above object processing device have been described in detail in the corresponding object processing methods, so they will not be repeated here.
[0109] Exemplary embodiments of this disclosure also provide an electronic device. This electronic device may be the terminal 101 or server 102 in the above embodiments. Generally, the electronic device may include a processor and a memory, the memory being used to store executable instructions of the processor, the processor being configured to perform the above-described object processing method by executing the executable instructions.
[0110] The following is based on Figure 10 Taking a mobile terminal 1000 as an example, the construction of this electronic device will be described by way of example. Those skilled in the art will understand that, apart from components specifically designed for mobile purposes, Figure 10 The structure can also be applied to fixed types of equipment.
[0111] like Figure 10As shown, the mobile terminal 1000 may specifically include: a processor 1001, a memory 1002, a bus 1003, a mobile communication module 1004, an antenna 1, a wireless communication module 1005, an antenna 2, a display screen 1006, a camera module 1007, an audio module 1008, a power module 1009, and a sensor module 1010.
[0112] Processor 1001 may include one or more processing units, such as an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, an encoder, a decoder, a digital signal processor (DSP), a baseband processor, and / or a neural network processing unit (NPU). The image denoising method in this exemplary embodiment can be executed by an AP, GPU, or DSP. When the method involves neural network-related processing, it can be executed by an NPU. For example, the NPU can load neural network parameters and execute neural network-related algorithm instructions.
[0113] An encoder encodes (compresses) images or videos to reduce data size for easier storage or transmission. A decoder decodes (decompresses) the encoded data to restore the original image or video data. The mobile terminal 1000 can support one or more encoders and decoders, such as image formats like JPEG (Joint Photographic Experts Group), PNG (Portable Network Graphics), and BMP (Bitmap), and video formats like MPEG (Moving Picture Experts Group) 1, MPEG10, H.1063, H.1064, and HEVC (High Efficiency Video Coding).
[0114] The processor 1001 can be connected to the memory 1002 or other components via the bus 1003.
[0115] The memory 1002 can be used to store computer executable program code, which includes instructions. The processor 1001 executes various functional applications and data processing of the mobile terminal 1000 by running the instructions stored in the memory 1002. The memory 1002 can also store application data, such as images, videos, and other files.
[0116] The communication function of mobile terminal 1000 can be implemented through mobile communication module 1004, antenna 1, wireless communication module 1005, antenna 2, modem processor, and baseband processor. Antenna 1 and antenna 2 are used to transmit and receive electromagnetic wave signals. Mobile communication module 1004 can provide 3G, 4G, 5G and other mobile communication solutions for mobile terminal 1000. Wireless communication module 1005 can provide wireless communication solutions such as wireless LAN, Bluetooth, and near-field communication for mobile terminal 1000.
[0117] The display screen 1006 is used to implement display functions, such as displaying the user interface, images, and videos. The camera module 1007 is used to implement shooting functions, such as capturing images and videos. The audio module 1008 is used to implement audio functions, such as playing audio and capturing voice. The power module 1009 is used to implement power management functions, such as charging the battery, supplying power to the device, and monitoring battery status. The sensor module 1010 may include one or more sensors to implement corresponding sensing and detection functions. For example, the sensor module 1010 may include an inertial sensor, which is used to detect the motion posture of the mobile terminal 1000 and output inertial sensing data.
[0118] This application also provides a computer-readable storage medium, which may be included in the electronic device described in the above embodiments; or it may exist independently and not assembled into the electronic device.
[0119] Computer-readable storage media can be, for example—but not limited to—electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatuses, or devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0120] A computer-readable storage medium can be sent, propagated, or transmitted for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable storage medium can be transmitted using any suitable medium, including but not limited to: wireless, wireline, optical fiber, RF, etc., or any suitable combination thereof.
[0121] A computer-readable storage medium carries one or more programs that, when executed by an electronic device, cause the electronic device to perform the methods described in the following embodiments.
[0122] From the above description of the embodiments, those skilled in the art will readily understand that the exemplary embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solutions according to the embodiments of this disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, external hard drive, etc.) or on a network, including several instructions to cause a computing device (such as a personal computer, server, terminal device, or network device, etc.) to execute the methods according to the embodiments of this disclosure.
[0123] Furthermore, the above figures are merely illustrative of the processes included in the method according to exemplary embodiments of this disclosure and are not intended to be limiting. It is readily understood that the processes shown in the above figures do not indicate or limit the temporal order of these processes. Additionally, it is readily understood that these processes may be executed synchronously or asynchronously, for example, in multiple modules.
[0124] It should be noted that although several modules or units for the device used to perform actions have been mentioned in the detailed description above, this division is not mandatory. In fact, according to embodiments of this disclosure, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.
[0125] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the disclosure herein. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and embodiments are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the claims. It should be understood that this disclosure is not limited to the precise structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.
Claims
1. An object processing method, characterized in that, Applied to classification tasks, including: Obtain search space; Based on the type of operators contained in each layer of the network in the search space, candidate operators for each layer of the network are determined from multiple operators in the search space, and network search is performed on each layer of the network in the search space based on the candidate operators to obtain the target network structure; wherein, the multiple operators include non-zero operators and zero operators; Based on the target network structure, processing operations are performed on the object to be processed to obtain the processing result; the object to be processed includes an image to be processed, speech to be processed, or text to be processed. The step of performing a network search on each layer of the search space based on the candidate operators to obtain the target network structure includes: Based on the candidate operators of each layer of the search space, determine the path of each layer of the search space; The paths of each network layer are combined to obtain the target network structure.
2. The object processing method according to claim 1, characterized in that, The step of determining candidate operators for each layer of the network in the search space based on the type of operators contained in each layer of the network includes: If the operator is a zero operator, the selection probability of the zero operator is determined based on the magnitude of the first parameter of the search space; the first parameter represents the index value of the zero operator in the network structure parameters; the first parameter is expressed as... , These are the network structure parameters retrieved from the search space; If the operator is a non-zero operator, the selection probability of the non-zero operator is determined according to the second parameter of the search space; the second parameter represents the index value of the non-zero operator in the network structure parameters, and is determined based on the ratio of the network structure parameters associated with zero operators to the network structure parameters associated with non-zero operators. The second parameter is expressed as follows: ; The candidate operator is determined from all operators based on the selection probability.
3. The object processing method according to claim 1, characterized in that, The method further includes: If the candidate operator is a non-zero operator, the corresponding operator operation is performed on the input data through the non-zero operator to determine the first operation result; Based on the result of the first operation and the weights of the non-zero operators, the output results of the paths of each layer of the network in the search space are determined.
4. The object processing method according to claim 3, characterized in that, The step of determining the output of the path for each layer of the network in the search space based on the result of the first operation and the weights of the non-zero operators includes: The first operation result and the weights of the non-zero operators are multiplied together, and the second operation result of the zero operator is combined with the input data to determine the output result of each layer of the network in the search space.
5. An object processing apparatus, applied to a server, characterized in that, Applied to classification tasks, including: The search space determination module is used to obtain the search space. A network search module is used to determine candidate operators for each layer of the network from multiple operators in the search space based on the type of operators contained in each layer of the network, and to perform a network search on each layer of the network in the search space based on the candidate operators to obtain the target network structure; wherein, the multiple operators include non-zero operators and zero operators; An operation execution module is used to perform processing operations on the object to be processed based on the target network structure to obtain processing results; the object to be processed includes an image to be processed, speech to be processed, or text to be processed. The step of performing a network search on each layer of the search space based on the candidate operators to obtain the target network structure includes: Based on the candidate operators of each layer of the search space, determine the path of each layer of the search space; The paths of each network layer are combined to obtain the target network structure.
6. An electronic device, characterized in that, include: processor; as well as Memory for storing the executable instructions of the processor; The processor is configured to execute the object processing method of any one of claims 1-4 by executing the executable instructions.
7. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the object processing method according to any one of claims 1-4.
Citation Information
Patent Citations
Network structure searching method, device and equipment and computer storage medium
CN114266350A
Time delay prediction method and device, electronic equipment and storage medium
CN114897126A
Neural Architecture Search Method, Image Processing Method And Apparatus, And Storage Medium
US20220215227A1