Joint search method and device for neural network and hardware

Through the joint search method of neural network and hardware, combined with gradient computing and heterogeneous multi-core architecture, the problem of insufficient accuracy of computing resources and hardware performance in neural network architecture search in the prior art is solved, and more efficient calculations and more accurate search results are achieved.

CN120163196AActive Publication Date: 2025-06-17SUZHOU INST FOR ADVANCED STUDY USTC +1
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510637864.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-19
Publication Date
2025-06-17
Estimated Expiration
2045-05-19

AI Technical Summary

Technical Problem

The existing neural network architecture search technology has defects in accuracy evaluation and hardware performance search, resulting in large overhead of computing resources, long calculation time, and insufficient hardware performance accuracy.

Method used

A joint search method for neural networks and hardware is proposed. By sampling the neural network search space based on preset conditions, using computing units to perform gradient calculation tasks, combining heterogeneous multi-core architecture to simulate hardware performance, and determining the adaptability value of the neural network to be evaluated.

Benefits of technology

It realizes the reduction in computing overhead and improvement in computing speed of service equipment such as computers, and at the same time improves the search accuracy of the target neural network, avoiding the problem of failure of the single-core architecture to fully consider network layer differences and computing preferences.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120163196A_ABST
    Figure CN120163196A_ABST
Patent Text Reader

Abstract

The invention provides a joint search method and device for a neural network and hardware, is applied to the technical field of computers, and is used for solving the problem that defects exist in the aspects of precision evaluation and hardware performance search in the prior art. The method comprises the following steps: sampling a neural network search space corresponding to a current round of iteration to obtain G to-be-evaluated neural networks corresponding to the current round of iteration, and determining a precision score value of the gth to-be-evaluated neural network; performing hardware performance simulation on the hardware configuration parameters of the to-be-evaluated neural networks by using the heterogeneous multi-core architecture, and determining a hardware performance value of the gth to-be-evaluated neural network; according to the ratio of the precision score value to the hardware performance value of the gth to-be-evaluated neural network, determining the adaptability value of the gth to-be-evaluated neural network, and determining a candidate neural network corresponding to the current round of iteration in the G to-be-evaluated neural networks; and determining a target neural network suitable for executing the target processing task from the multiple rounds of respective corresponding candidate neural networks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and more particularly to a method and device for jointly searching a neural network and hardware. Background Art

[0002] While neural networks perform excellently in various fields, they bring an increase in computational intensity and design difficulty. For this reason, the optimization work combining neural network architecture search and hardware accelerators has become a research hotspot.

[0003] The neural network architecture search technology in the related art has defects in accuracy evaluation and hardware performance search. Specifically, in terms of accuracy evaluation, the existing evaluation methods rely on over-parameterized supernet training, resulting in huge computational resource overhead and computational time overhead problems in service devices such as computers; in terms of hardware performance search, the methods in the related art perform hardware performance search based on a single-core architecture, failing to fully consider the differences and computational preferences of network layers, resulting in insufficient hardware performance accuracy obtained by the search and low computational speed of service devices such as computers. Summary of the Invention

[0004] In view of the above problems, the present invention provides a method and device for jointly searching a neural network and hardware.

[0005] According to a first aspect of the present invention, there is provided a method for jointly searching a neural network and hardware, applied to a computing device, including: sampling a neural network search space corresponding to the current round of iteration in a storage unit based on preset neural network search conditions to obtain G neural networks to be evaluated corresponding to the current round of iteration; based on the g-th neural network to be evaluated among the G neural networks to be evaluated corresponding to the current round of iteration, using a computing unit to execute a gradient calculation task to obtain a set of target gradient values of the g-th neural network to be evaluated, G≥1, 1≤g≤G; determining an accuracy score value of the g-th neural network to be evaluated according to the set of target gradient values of the g-th neural network to be evaluated; using a heterogeneous multi-core architecture to perform hardware performance simulation on hardware configuration parameters of the neural network to be evaluated to determine a hardware performance value of the g-th neural network to be evaluated; determining an adaptability value of the g-th neural network to be evaluated according to a ratio between the accuracy score value and the hardware performance value of the g-th neural network to be evaluated, and obtaining adaptability values of the G neural networks to be evaluated; determining the neural network to be evaluated corresponding to the highest adaptability value among the adaptability values of the G neural networks to be evaluated as the candidate neural network corresponding to the current round of iteration; determining a target neural network suitable for performing a target processing task from the candidate neural networks corresponding to multiple rounds.

[0006] According to an embodiment of the present invention, the target processing task includes at least one of the following: a target detection task, a text processing task, and an image classification task.

[0007] According to an embodiment of the present invention, for the g-th neural network to be evaluated among the G neural networks to be evaluated corresponding to the current round of iteration, determining the set of target gradient values of the g-th neural network to be evaluated includes: determining M batches of training label data input to the g-th neural network to be evaluated, where M≥1; for the m-th batch of training label data among the M batches, determining the loss value of the m-th batch of the g-th neural network to be evaluated according to the m-th batch of training label data of the g-th neural network to be evaluated, where 1≤m≤M; respectively calculating the derivatives of the loss value of the m-th batch and the respective weight parameters of the N g-th neural networks to be evaluated to obtain N gradient values of the m-th batch, obtaining M×N gradient values of the M batches, where N≥1; screening the M×N gradient values to obtain the set of target gradient values of the g-th neural network to be evaluated.

[0008] According to an embodiment of the present invention, screening the M×N gradient values to obtain the set of target gradient values of the g-th neural network to be evaluated includes: for the n-th weight parameter among the N weight parameters of the g-th neural network to be evaluated, dividing the M gradient values corresponding to the n-th weight parameter into multiple gradient intervals to obtain a plurality of gradient intervals, where 1≤n≤N; according to the density of each gradient interval, selecting the gradient interval with the largest density among the plurality of gradient intervals as the n-th target gradient interval corresponding to the n-th weight parameter to obtain N target gradient intervals; according to the n-th target gradient interval, determining the n-th set of target gradient values corresponding to the n-th weight parameter to obtain N sets of target gradient values; using the N sets of target gradient values as the set of target gradient values of the g-th neural network to be evaluated.

[0009] According to an embodiment of the present invention, determining the accuracy score value of the g-th neural network to be evaluated according to the set of target gradient values of the g-th neural network to be evaluated includes: for the n-th weight parameter of the g-th neural network to be evaluated, determining the gradient average value of the n-th set of target gradient values; determining the gradient variance value of the n-th set of target gradient values; according to the ratio between the gradient average value and the gradient variance value, determining the n-th initial accuracy score value to obtain N initial accuracy score values; according to the N initial accuracy score values, determining the accuracy score value of the g-th neural network to be evaluated.

[0010] According to an embodiment of the present invention, the heterogeneous multi-core architecture includes multiple hardware cores; using the heterogeneous multi-core architecture to perform hardware performance simulation on the hardware configuration parameters of the neural network to be evaluated, and determining the hardware performance value of the g-th neural network to be evaluated, including: mapping the g-th neural network to be evaluated to multiple hardware cores to obtain multiple target hardware cores; determining the target number of computing units and target bandwidth of each target hardware core according to the ratio of the computing volume requirement and the broadband requirement of each target hardware core; optimizing the data flow parameters of each target hardware core to obtain the target data flow parameters of each target hardware core; using the performance simulation framework to perform hardware performance simulation on the target number of computing units, target bandwidth and target data flow parameters of each target hardware core, and determining the hardware performance value of the g-th neural network to be evaluated.

[0011] According to an embodiment of the present invention, the g-th neural network to be evaluated includes multiple operators; mapping the g-th neural network to be evaluated to multiple hardware cores to obtain multiple target hardware cores, including: clustering all operators according to the dimension parameters of each operator to obtain the category of each operator; mapping the operators with the same category to the same hardware core to obtain multiple target hardware cores.

[0012] According to an embodiment of the present invention, optimizing the data flow parameters of each target hardware core to obtain the target data flow parameters of each target hardware core includes: for the data flow parameters of any target hardware core, iteratively optimizing the data flow parameters to obtain candidate data flow parameters; determining the candidate data flow parameters that meet the preset iteration conditions of the data flow parameters as the data flow parameters.

[0013] According to an embodiment of the present invention, before optimizing the data flow parameters of each target hardware core, the method further includes: pruning the data flow parameters of any target hardware core to obtain the pruned data flow parameters.

[0014] The second aspect of the present invention provides a joint search device for a neural network and hardware, including: a storage unit for storing a neural network search space; a computing unit configured to: sample the neural network search space corresponding to the current iteration in the storage unit based on a preset neural network search condition to obtain G neural networks to be evaluated corresponding to the current iteration; perform a gradient calculation task based on the g-th neural network to be evaluated among the G neural networks to be evaluated corresponding to the current iteration to obtain a set of target gradient values of the g-th neural network to be evaluated, where G≥1 and 1≤g≤G; determine an accuracy score value of the g-th neural network to be evaluated according to the set of target gradient values of the g-th neural network to be evaluated; a heterogeneous multi-core architecture configured to: perform a hardware performance simulation on the hardware configuration parameters of the neural network to be evaluated to determine the hardware performance value of the g-th neural network to be evaluated; wherein, the computing unit is further configured to: determine the adaptability value of the g-th neural network to be evaluated according to the ratio between the accuracy score value and the hardware performance value of the g-th neural network to be evaluated, and obtain the adaptability values of the G neural networks to be evaluated; determine the neural network to be evaluated corresponding to the highest adaptability value among the adaptability values of the G neural networks to be evaluated as the candidate neural network corresponding to the current iteration; and determine a target neural network applicable to perform a target processing task from the candidate neural networks corresponding to multiple respective iterations.

[0015] The third aspect of the present invention provides an electronic device, including: one or more processors; a memory for storing one or more computer programs, wherein the above one or more processors execute the above one or more computer programs to implement the steps of the above method.

[0016] The fourth aspect of the present invention further provides a computer-readable storage medium, on which a computer program or instruction is stored, and when the above computer program or instruction is executed by a processor, the steps of the above method are implemented.

[0017] The fifth aspect of the present invention further provides a computer program product, including a computer program or instruction, and when the above computer program or instruction is executed by a processor, the steps of the above method are implemented.

[0018] According to an embodiment of the present invention, based on a preset neural network search condition, sample the neural network search space corresponding to the current round of iteration in the storage unit to obtain G neural networks to be evaluated corresponding to the current round of iteration; based on the g-th neural network to be evaluated among the G neural networks to be evaluated corresponding to the current round of iteration, use the computing unit to execute a gradient calculation task to obtain a set of target gradient values of the g-th neural network to be evaluated; according to the set of target gradient values of the g-th neural network to be evaluated, determine the accuracy score value of the g-th neural network to be evaluated; use the heterogeneous multi-core architecture to perform hardware performance simulation on the hardware configuration parameters of the neural network to be evaluated, and determine the hardware performance value of the g-th neural network to be evaluated; according to the ratio between the accuracy score value and the hardware performance value of the g-th neural network to be evaluated, determine the adaptability value of the g-th neural network to be evaluated, and obtain the adaptability values of the G neural networks to be evaluated; determine the neural network to be evaluated corresponding to the highest adaptability value among the adaptability values of the G neural networks to be evaluated as the candidate neural network corresponding to the current round of iteration; determine the target neural network applicable to execute the target processing task from the candidate neural networks corresponding to multiple rounds. Since the gradient calculation task is executed by the computing unit, the relationship between the gradient information and the model convergence of the neural network to be evaluated is fully utilized when determining the accuracy score value corresponding to the neural network to be evaluated, avoiding the design of a heavy and complex super network, realizing the reduction of the computing overhead of service devices such as computers, and improving the search accuracy of the target neural network while ensuring the improvement of the computing speed of service devices such as computers; in addition, since the hardware performance of the neural network to be evaluated is simulated by relying on the heterogeneous multi-core architecture, the problem that the single-core architecture fails to fully consider the differences and computing preferences of network layers is avoided, so that service devices such as computers do not need to reconfigure the hardware performance evaluation architecture for different neural networks to be evaluated, ensuring the stable operation of service devices such as computers and reducing the storage space occupation of service devices such as computers. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Through the following description of the embodiments of the present invention with reference to the drawings, the above content and other objects, features and advantages of the present invention will become clearer. In the drawings:

[0020] Figure 1 Schematically shows an application scenario diagram of a joint search method and device for a neural network and hardware according to an embodiment of the present invention.

[0021] Figure 2 Schematically shows a flowchart of a joint search method for a neural network and hardware according to an embodiment of the present invention.

[0022] Figure 3 Schematically shows a schematic diagram of a heterogeneous multi-core architecture according to an embodiment of the present invention.

[0023] Figure 4A schematic diagram showing data flow parameters according to an embodiment of the present invention is shown.

[0024] Figure 5 A schematic diagram showing a joint search architecture of a neural network and hardware according to an embodiment of the present invention is shown.

[0025] Figure 6 A schematic diagram showing test results of a hardware searcher according to an embodiment of the present invention is shown.

[0026] Figure 7 A block diagram showing the structure of a joint search device of a neural network and hardware according to an embodiment of the present invention is shown.

[0027] Figure 8 A block diagram showing an electronic device suitable for implementing a joint search method of a neural network and hardware according to an embodiment of the present invention is shown. Detailed implementation manners

[0028] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present invention. In the following detailed description, for the sake of explanation, many specific details are set forth to provide a comprehensive understanding of the embodiments of the present invention. However, it is obvious that one or more embodiments can also be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessarily confusing the concepts of the present invention.

[0029] The terms used herein are merely for describing specific embodiments and are not intended to limit the present invention. The terms "including", "comprising" and the like used herein indicate the presence of the described features, steps, operations and / or components, but do not exclude the presence or addition of one or more other features, steps, operations or components.

[0030] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.

[0031] In the case of using expressions such as "at least one of A, B, and C", generally, it should be interpreted according to the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include, but is not limited to, a system having only A, only B, only C, having A and B, having A and C, having B and C, and / or having A, B, and C).

[0032] In the joint search of related technologies, the acquisition of accuracy information is mostly based on training methods and predictor-based methods. After generating different neural networks, each neural network needs to be fully trained to evaluate its accuracy. By integrating all subnetworks in the design space into an over-parameterized supernet, the supernet only needs to be trained once to avoid the process of repeatedly training each subnetwork. Specifically, the subnetwork can inherit weights from the pre-trained supernet, so as to quickly obtain accuracy information. On the other hand, the predictor-based method completes the prediction of model accuracy by learning the accuracy performance of the network structure. Therefore, the above method usually needs to collect a batch of actual training accuracy data of neural networks to form a data set containing network structures and their corresponding accuracy. Based on this data set, a deep learning model is trained to predict the accuracy that different network structures may achieve. However, the above-mentioned neural network ranking method that relies on training faces huge computing resource overhead and time overhead. Even the one-time training method based on the supernet still requires more than 100 GPU hours of time overhead. In addition, when the search space changes, it is necessary to re-design the supernet or prepare for predictor training, which faces a large reconstruction cost and limited versatility.

[0033] To obtain hardware performance, hardware search-based work needs to search for the most suitable hardware configuration parameters for a certain designed network structure. The hardware information of the neural network performing reasoning under this parameter is used as hardware feedback. Since the hardware design is not determined during the joint search process, the determination of the hardware design parameters can be used as a sub-search problem. For the hardware search space of existing joint search work, the parallel size of the hardware computing array, the data reuse size and the classic data flow selection are explored based on the single-core accelerator architecture. However, the above method performs hardware structure search based on the single-core architecture and classic data flow, which fails to fully consider the differences and computing preferences of the network layer, resulting in insufficient hardware adaptability obtained by the search.

[0034] In view of this, an embodiment of the present invention provides a joint search method for a neural network and hardware, which is applied to a computing device and includes: sampling a neural network search space corresponding to the current round of iteration in a storage unit based on a preset neural network search condition to obtain G neural networks to be evaluated corresponding to the current round of iteration; based on the g-th neural network to be evaluated among the G neural networks to be evaluated corresponding to the current round of iteration, using a computing unit to execute a gradient calculation task to obtain a set of target gradient values of the g-th neural network to be evaluated, where G≥1 and 1≤g≤G; determining an accuracy score value of the g-th neural network to be evaluated according to the set of target gradient values of the g-th neural network to be evaluated; using a heterogeneous multi-core architecture to perform hardware performance simulation on hardware configuration parameters of the neural network to be evaluated to determine a hardware performance value of the g-th neural network to be evaluated; determining an adaptability value of the g-th neural network to be evaluated according to a ratio between the accuracy score value and the hardware performance value of the g-th neural network to be evaluated, and obtaining adaptability values of the G neural networks to be evaluated; determining the neural network to be evaluated corresponding to the highest adaptability value among the adaptability values of the G neural networks to be evaluated as the candidate neural network corresponding to the current round of iteration; and determining a target neural network applicable to execute a target processing task from the candidate neural networks corresponding to multiple rounds respectively.

[0035] Figure 1 FIG. schematically shows an application scenario diagram of a joint search method and apparatus for a neural network and hardware according to an embodiment of the present invention.

[0036] As Figure 1 shown, the application scenario according to this embodiment may include a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a server 105. The network 104 is used to provide a medium for a communication link between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.

[0037] Users may use the first terminal device 101, the second terminal device 102, and the third terminal device 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various communication client applications may be installed on the first terminal device 101, the second terminal device 102, and the third terminal device 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (only for example).

[0038] The first terminal device 101, the second terminal device 102, and the third terminal device 103 may be various electronic devices having a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop portable computers, and desktop computers, etc.

[0039] Server 105 may be a server that provides various services. For example, it can be a background management server (only for example) that supports the websites browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103. The background management server can analyze and process data such as user requests received, and feedback the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal devices.

[0040] It should be noted that the neural network and hardware joint search method provided by the embodiments of the present invention can generally be executed by server 105. Correspondingly, the neural network and hardware joint search device provided by the embodiments of the present invention can generally be set in server 105. The neural network and hardware joint search method provided by the embodiments of the present invention can also be executed by a server or a server cluster different from server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or server 105. Correspondingly, the neural network and hardware joint search device provided by the embodiments of the present invention can also be set in a server or a server cluster different from server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or server 105.

[0041] It should be understood that Figure 1 the numbers of terminal devices, networks, and servers in

[0042] are merely illustrative. According to the implementation requirements, there can be any number of terminal devices, networks, and servers. Figure 1 The following will be based on Figures 2 to 6 the described scenario to describe in detail the neural network and hardware joint search method of the embodiments of the invention through

[0043] Figure 2 FIG. schematically shows a flowchart of the neural network and hardware joint search method according to an embodiment of the present invention.

[0044] As Figure 2 shown, the neural network and hardware joint search method of this embodiment includes operations S210 to S270.

[0045] In operation S210, based on a preset neural network search condition, sample the neural network search space corresponding to the current round of iteration in the storage unit to obtain G neural networks to be evaluated corresponding to the current round of iteration.

[0046] According to an embodiment of the present invention, the preset neural network search condition may be the user's demand preference for the neural network. For example, the accuracy requirement that the obtained neural network needs to meet or the performance adaptation preference with the accelerator hardware, etc.

[0047] According to an embodiment of the present invention, the storage unit may be a storage hardware such as a solid-state drive or a mechanical hard drive, and is used to store the neural network search space. Among them, the neural network search space may be a neural network template and a series of operators with variable types and sizes. The above neural network template has variable network layers and fixed network layers. The position of the preset variable network layer in the template can be filled with operators to form numerous possible neural network architectures. Among them, the neural network search space may be a search space dominated by convolutional operators to form a Convolutional Neural Network (CNN) system search space, or may be a search space dominated by attention calculations to form a Transformer system search space. Based on the preset neural network search condition, sample the operators in the neural network search space corresponding to the current round of iteration, screen the operators that can adapt to the hardware accelerator, and fill the sampled operators into the neural network template to obtain G neural networks to be evaluated corresponding to the current round of iteration.

[0048] According to an embodiment of the present invention, in the joint search method of the neural network and the hardware, multiple rounds of iterative optimization are performed. Among them, the operations in each round are the same. For any round, it is necessary to sample the neural network search space corresponding to the current round of iteration based on the preset neural network search condition to obtain G neural networks to be evaluated corresponding to the current round of iteration, where the preset neural network search condition does not need to be input repeatedly.

[0049] In operation S220, based on the g-th neural network to be evaluated among the G neural networks to be evaluated corresponding to the current round of iteration, use the computing unit to execute the gradient calculation task to obtain the target gradient value set of the g-th neural network to be evaluated, where G≥1 and 1≤g≤G.

[0050] According to an embodiment of the present invention, the computing unit may include at least one of a Central Processing Unit (CPU), a Graphics Processing Unit (GPU), a Tensor Processing Unit (TPU), a Field-Programmable Gate Array (FPGA), etc., depending on the complexity of the related tasks, performance requirements, and application scenarios.

[0051] According to an embodiment of the present invention, the gradient calculation task may be to calculate the gradient of the neural network to be evaluated based on the training data related to the neural network to be evaluated.

[0052] According to an embodiment of the present invention, the gradient of the neural network is obtained through backpropagation calculation and is used for updating the weight parameters of the neural network to minimize the loss function. The higher the consistency of the gradient of the neural network to be evaluated, the smaller the probability of oscillation in the update of the relevant parameters of the neural network to be evaluated, and the faster it converges.

[0053] According to an embodiment of the present invention, the present invention uses several mini-batch training data to obtain a set of target gradient values of the neural network to be evaluated. Based on the consistency of the set of target gradient values in different batches, it can represent the probability of oscillation in the update of the relevant parameters of the neural network to be evaluated.

[0054] According to an embodiment of the present invention, the several mini-batch training data used by each neural network to be evaluated in each round are the same.

[0055] In operation S230, according to the set of target gradient values of the g-th neural network to be evaluated, determine the accuracy score value of the g-th neural network to be evaluated.

[0056] According to an embodiment of the present invention, for the relevant parameters of the neural network to be evaluated, the gradient consistency is evaluated, making full use of the relationship between the gradient information and the model convergence of the neural network to obtain the accuracy score value of the neural network to be evaluated. While improving the accuracy evaluation efficiency, it gets rid of the dependence on the design of a heavy and complex super network, adapts to different evaluation scenarios, reduces the difficulty of accuracy scoring, and thus reduces the system overhead and improves the system processing efficiency.

[0057] In operation S240, use the heterogeneous multi-core architecture to perform hardware performance simulation on the hardware configuration parameters of the neural network to be evaluated, and determine the hardware performance value of the g-th neural network to be evaluated.

[0058] According to an embodiment of the present invention, the hardware configuration parameters may be data flow parameters, the number of processing elements, and bandwidth limitations.

[0059] According to an embodiment of the present invention, the heterogeneous multi-core architecture may be a processor architecture with multiple hardware cores. The hardware configuration parameters of each hardware core are different. Multiple hardware cores can perform hardware performance simulation on the hardware configuration parameters of the neural network to be evaluated in parallel, and can provide a comprehensive hardware performance evaluation for the neural network to be evaluated, so as to select the most suitable neural network architecture for the hardware accelerator in the subsequent iterative search process.

[0060] In operation S250, according to the ratio between the accuracy score value and the hardware performance value of the g-th neural network to be evaluated, determine the adaptability value of the g-th neural network to be evaluated, and obtain the adaptability values of G neural networks to be evaluated.

[0061] According to an embodiment of the present invention, the adaptability value of the neural network to be evaluated, combined with the accuracy score value and the hardware performance value of the neural network to be evaluated, can reflect the performance differences of different neural networks to be evaluated.

[0062] According to an embodiment of the present invention, in each round, the above operations are respectively performed on G neural networks to be evaluated to obtain their respective adaptability values.

[0063] In operation S260, determine the neural network to be evaluated corresponding to the highest adaptability value among the adaptability values of G neural networks to be evaluated as the candidate neural network corresponding to the current round of iteration.

[0064] According to an embodiment of the present invention, by screening the neural network to be evaluated corresponding to the highest adaptability value among the adaptability values of G neural networks to be evaluated, the optimal neural network to be evaluated for the current round is determined.

[0065] In operation S270, determine the target neural network applicable to execute the target processing task from the candidate neural networks corresponding to multiple rounds.

[0066] According to an embodiment of the present invention, the target processing task includes at least one of the following: target detection task, text processing task, and image classification task. For example, when the target processing task is an image classification task, the obtained target neural network can be capable of automatically learning the key features in the image, having sufficient neural network depth and complexity, etc., so that the target neural network can be applicable to the corresponding computer device (such as a computer device like a hardware accelerator) to execute the image classification task.

[0067] Among them, applying the target neural network to the target processing task can be: for the selected target processing task, determine the input-output requirements associated with the target processing task, and based on the input-output requirements, select the corresponding target neural network, so as to be able to complete the processing of the input data based on the target neural network to complete the target processing task. For example, for an image classification task, determine the required input image data and the input-output format of the image; based on the required input image data and the input-output format of the image, determine the required target neural network; subsequently, use the target neural network to process the input image data to complete the classification task of the image.

[0068] According to an embodiment of the present invention, in the joint search method of a neural network and hardware, the candidate neural network obtained in the current iteration round is used as an optimization condition to optimize the neural network search space corresponding to the next round, that is, the G neural networks to be evaluated obtained in the next round are different from the G neural networks to be evaluated in the current round. In each round, according to the respective G neural networks to be evaluated, the respective candidate neural networks are screened out. Among the candidate neural networks of each round, the candidate neural network that meets the predetermined iteration condition is determined as the target neural network.

[0069] According to an embodiment of the present invention, the predetermined iteration condition may be to complete the iteration of a preset number of rounds, compare the fitness values of the candidate neural networks of each round, and use the candidate neural network corresponding to the highest fitness value as the target neural network. For example, it is preset that the joint search method of the neural network and hardware needs to complete 3 rounds of iteration. In the first round, candidate neural network A is obtained, and the fitness value of candidate neural network A is 1; in the second round, the neural network search space is optimized according to candidate neural network A, and candidate neural network B is obtained, and the fitness value of candidate neural network B is 1.5; in the third round, the neural network search space is optimized according to candidate neural network B, and candidate neural network C is obtained, and the fitness value of candidate neural network C is 2. At this point, 3 rounds of iteration are completed. Among them, the fitness value of candidate neural network C obtained in the third round is the highest, and candidate neural network C is determined as the target neural network.

[0070] According to an embodiment of the present invention, the predetermined iteration condition may also be to stop the iteration when the fitness value of the candidate neural network screened out from the G neural networks to be evaluated in a certain round reaches a preset threshold, and determine the candidate neural network as the target neural network. For example, the preset threshold is set to 2, and the joint search method of the neural network and hardware is executed. In the first round, candidate neural network A is obtained, and the fitness value of candidate neural network A is 1; in the second round, the neural network search space is optimized according to candidate neural network A, and candidate neural network B is obtained, and the fitness value of candidate neural network B is 1.5; in the third round, the neural network search space is optimized according to candidate neural network B, and candidate neural network C is obtained, and the fitness value of candidate neural network C is 2. At this point, the fitness value of candidate neural network C reaches the preset threshold, the iteration is stopped, and candidate neural network C is determined as the target neural network.

[0071] According to an embodiment of the present invention, based on a preset neural network search condition, sample the neural network search space corresponding to the current round of iteration in the storage unit to obtain G neural networks to be evaluated corresponding to the current round of iteration; based on the g-th neural network to be evaluated among the G neural networks to be evaluated corresponding to the current round of iteration, use the computing unit to execute a gradient calculation task to obtain a set of target gradient values of the g-th neural network to be evaluated; according to the set of target gradient values of the g-th neural network to be evaluated, determine the accuracy score value of the g-th neural network to be evaluated; use a heterogeneous multi-core architecture to perform hardware performance simulation on the hardware configuration parameters of the neural network to be evaluated, and determine the hardware performance value of the g-th neural network to be evaluated; according to the ratio between the accuracy score value and the hardware performance value of the g-th neural network to be evaluated, determine the adaptability value of the g-th neural network to be evaluated, and obtain the adaptability values of the G neural networks to be evaluated; determine the neural network to be evaluated corresponding to the highest adaptability value among the adaptability values of the G neural networks to be evaluated as the candidate neural network corresponding to the current round of iteration; determine the target neural network applicable to execute the target processing task from the candidate neural networks corresponding to multiple rounds. Since the computing unit is used to execute the gradient calculation task, the relationship between the gradient information and the model convergence of the neural network to be evaluated is fully utilized when determining the accuracy score value corresponding to the neural network to be evaluated, avoiding the design of a heavy and complex super network, realizing the reduction of the computing overhead of service devices such as computers, and improving the search accuracy of the target neural network while ensuring the improvement of the computing speed of service devices such as computers; in addition, since the heterogeneous multi-core architecture is relied on to simulate the hardware performance of the neural network to be evaluated, the problem that the single-core architecture fails to fully consider the differences and computing preferences of network layers is avoided, so that service devices such as computers do not need to reconfigure the hardware performance evaluation architecture for different neural networks to be evaluated, ensuring the stable operation of service devices such as computers and reducing the storage space occupation of service devices such as computers.

[0072] According to an embodiment of the present invention, for the g-th neural network to be evaluated among the G neural networks to be evaluated corresponding to the current round of iteration, determining the set of target gradient values of the g-th neural network to be evaluated includes: determining M batches of training label data input to the g-th neural network to be evaluated, M≥1; for the m-th batch of training label data among the M batches, according to the m-th batch of training label data of the g-th neural network to be evaluated, determine the loss value of the m-th batch of the g-th neural network to be evaluated, 1≤m≤M; calculate the derivatives of the loss value of the m-th batch and the weight parameters of the N g-th neural networks to be evaluated respectively to obtain N gradient values of the m-th batch, and obtain M×N gradient values of the M batches, N≥1; screen the M×N gradient values to obtain the set of target gradient values of the g-th neural network to be evaluated.

[0073] According to an embodiment of the present invention, the training labeled data may include training data and labels corresponding to the training data, wherein the training labeled data may be M batches of mini-batch data.

[0074] According to an embodiment of the present invention, for the m-th batch of training labeled data among the M batches, the training data is input into the neural network to be evaluated, and the result output by the neural network to be evaluated is used to calculate the loss with the label corresponding to the training data, so as to obtain the loss value of the m-th batch of the g-th neural network to be evaluated.

[0075] According to an embodiment of the present invention, there are N weight parameters in the neural network to be evaluated. For any batch of training labeled data, each of the N weight parameters has a corresponding gradient value. Specifically, the determination of the gradient value is shown in formula (1). The loss value of the m-th batch and the derivatives of each of the N weight parameters of the g-th neural network to be evaluated are calculated respectively to obtain N gradient values of the m-th batch.

[0076] (1);

[0077] Wherein, represents the n-th weight parameter among the N weight parameters, represents the m-th batch of training labeled data, wherein, , the m-th batch of training labeled data includes b training data x and b labels y, represents the i-th training data among the b training data x, represents the i-th label corresponding to the i-th training data among the b labels y, represents calculating the loss, represents taking the partial derivative, represents the N gradient values of the m-th batch, represents the neural network to be evaluated when inputting the i-th training data, represents the weight parameters of the neural network to be evaluated.

[0078] According to an embodiment of the present invention, based on the N gradient values of the m-th batch obtained above, M×N gradient values of M batches for the g-th neural network to be evaluated can be obtained. Since too large or too small gradients are not conducive to the convergence of the neural network to be evaluated, blindly using all information may lead to evaluation bias. Therefore, the M×N gradient values are screened, and the too large or too small gradient values are defined as gradient noise, and a set of target gradient values for the g-th neural network to be evaluated is selected.

[0079] According to an embodiment of the present invention, screening M×N gradient values to obtain a set of target gradient values of the g-th neural network to be evaluated, including: for the n-th weight parameter among the N weight parameters of the g-th neural network to be evaluated, partitioning the M gradient values corresponding to the n-th weight parameter to obtain a plurality of gradient intervals, where 1≤n≤N; selecting, according to the density of each gradient interval, the gradient interval with the largest density among the plurality of gradient intervals as the n-th target gradient interval corresponding to the n-th weight parameter, to obtain N target gradient intervals; determining, according to the n-th target gradient interval, a set of n-th target gradient values corresponding to the n-th weight parameter, to obtain N sets of target gradient values; and using the N sets of target gradient values as the set of target gradient values of the g-th neural network to be evaluated.

[0080] According to an embodiment of the present invention, for the n-th weight parameter among the N weight parameters of the g-th neural network to be evaluated, partitioning the M gradient values corresponding to the n-th weight parameter. For example, the first gradient interval is 0≤gradient value<1, the second gradient interval is 1≤gradient value<2, and the third gradient interval is 2≤gradient value<3.

[0081] According to an embodiment of the present invention, the density of a gradient interval can be the number of gradient values in each interval. Selecting, among the plurality of gradient intervals, the gradient interval with the largest density as the n-th target gradient interval corresponding to the n-th weight parameter. For example, if there are 3 gradient values in the first gradient interval, 10 gradient values in the second gradient interval, and 2 gradient values in the third gradient interval, then the second gradient interval is taken as the target gradient interval. And the 10 gradient values in this target gradient interval are used as the set of target gradient values.

[0082] According to an embodiment of the present invention, determining an accuracy score value of the g-th neural network to be evaluated according to the set of target gradient values of the g-th neural network to be evaluated, including: for the n-th weight parameter of the g-th neural network to be evaluated, determining the gradient average value of the set of n-th target gradient values; determining the gradient variance value of the set of n-th target gradient values; determining, according to the ratio between the gradient average value and the gradient variance value, the n-th initial accuracy score value, to obtain N initial accuracy score values; and determining the accuracy score value of the g-th neural network to be evaluated according to the N initial accuracy score values.

[0083] According to an embodiment of the present invention, the precision score can be obtained by using several small batches of training data to obtain gradients and comparing the consistency of the gradients obtained from different batches to represent the oscillation probability of network weight updates. Specifically, the consistency within each network layer of the neural network to be evaluated is summed, the consistency between different network layers is multiplied, and a logarithmic function is used for scaling to avoid too large precision score values. For example, the neural network to be evaluated includes 3 network layers. For the first network layer corresponding to 4 weight parameters, the ratio between the average gradient value and the gradient variance value corresponding to the first weight parameter is determined to obtain the initial precision score value of the first weight parameter corresponding to the first network layer, and then the initial precision score values of the 4 weight parameters corresponding to the first network layer are obtained respectively; after summing the initial precision score values of the 4 weight parameters corresponding to the first network layer respectively, the initial precision score value of the first network layer is obtained, and then the initial precision score values of the 3 network layers are obtained respectively. The initial precision score values of the 3 network layers are multiplied to determine the precision score value of the g-th neural network to be evaluated.

[0084] Specifically, the determination of the precision score value can be as shown in formula (2).

[0085] (2);

[0086] where D represents the number of network layers, represents all the weight parameters of the d-th network layer, represents the target gradient interval, represents the indicator function, represents the average gradient value of the n-th set of target gradient values, represents the gradient variance value of the n-th set of target gradient values, represents the precision score value of the g-th neural network to be evaluated. In addition, the sign of the gradient value is retained before the absolute value operation to capture the direction of gradient update and the oscillation of the weights, and log represents taking the logarithm.

[0087] According to an embodiment of the present invention, the higher the value of the precision score, the more stable the gradients of different batches are, and the faster the training loss of the network architecture of the neural network to be evaluated decreases. Since the determination operation of the precision score value of the present invention eliminates the training process in calculating the network ranking in the related art, avoids designing a heavy and complex super network, realizes the reduction of the computing overhead of service devices such as computers, and improves the search precision of the target neural network while ensuring the improvement of the computing speed of service devices such as computers, and reduces the burden on service devices such as computers.

[0088] According to an embodiment of the present invention, a heterogeneous multi-core architecture includes multiple hardware cores; using the heterogeneous multi-core architecture to perform hardware performance simulation on the hardware configuration parameters of a neural network to be evaluated, and determining the hardware performance value of the g-th neural network to be evaluated, including: mapping the g-th neural network to be evaluated to multiple hardware cores to obtain multiple target hardware cores; determining the target number of computing units and target bandwidth of each target hardware core according to the ratio of the computing amount requirement and bandwidth requirement of each target hardware core; optimizing the data flow parameters of each target hardware core to obtain the target data flow parameters of each target hardware core; using a performance simulation framework to perform hardware performance simulation on the target number of computing units, target bandwidth, and target data flow parameters of each target hardware core to determine the hardware performance value of the g-th neural network to be evaluated.

[0089] Figure 3 FIG. schematically shows a schematic diagram of a heterogeneous multi-core architecture according to an embodiment of the present invention.

[0090] According to an embodiment of the present invention, as Figure 3 shown, the heterogeneous multi-core architecture may be a multi-core processor architecture with multiple hardware cores, and the computing sub-units (PEs) allocation, data flow scheduling, etc. of each hardware core are all different. Among them, the data and performance calculation results required by all computing sub-units interact with caches (global cache, local cache) through a multi-level interconnection network (global interconnection, local interconnection).

[0091] According to an embodiment of the present invention, the neural network to be evaluated can be decomposed into a series of Q operators to be executed . Denote the heterogeneous multi-core architecture as , representing a heterogeneous multi-core architecture with H hardware cores.

[0092] According to an embodiment of the present invention, mapping the neural network to be evaluated to multiple hardware cores can be denoted as , then the mapping of the operator to the core can be represented as an assignment problem. That is, find a mapping that satisfies the injective condition from the set to . Specifically, it can be to use the heterogeneous multi-core architecture to simulate a hardware accelerator, map the neural network to be evaluated to multiple hardware cores of the heterogeneous multi-core architecture as the load of the hardware accelerator.

[0093] According to an embodiment of the present invention, the ratio of the computing amount requirement and bandwidth requirement of each target hardware core can be to determine the ratio of the computing amount requirement and bandwidth requirement of the operator for each target hardware core according to the preset hardware constraint requirements, so as to allocate the computing amount requirement and bandwidth requirement of each target hardware core to obtain the target number of computing units and target bandwidth of each target hardware core.

[0094] According to an embodiment of the present invention, the data flow parameters describe the scheduling of operators between the computing units and memory of the target hardware core, and further reflect the interconnection structure between the computing units. Specifically, it can be the six-dimensional loop order, parallel dimension, shard size, and parallelism included in the one-dimensional array of operators for each target hardware core. Among them, in the two-dimensional array, these parameters will be repeated at a higher level, and the data flow parameter space of the target hardware core shows exponential growth.

[0095] According to an embodiment of the present invention, optimizing the data flow parameters of each target hardware core can be based on a genetic algorithm for optimization, selecting the optimal data flow configuration for the operators of the target hardware core, and obtaining the target data flow parameters of each target hardware core.

[0096] According to an embodiment of the present invention, each target hardware core can be represented as , representing the target data flow parameters, the number of target computing units, and the target bandwidth of the h-th target hardware core. When each target hardware core is determined, the above-mentioned target data flow parameters, the number of target computing units, and the target bandwidth are sent into a mature performance simulation framework, such as Maestro or Timegoop, then the hardware performance value of the neural network to be evaluated can be obtained to represent the power consumption, area, latency, etc. of the target hardware core.

[0097] According to an embodiment of the present invention, the determination operation of the hardware performance value of the present invention covers data scheduling, data parallelism, and the hardware interconnection structure, realizes the adaptation to different neural networks and different operators. Under the condition of similar hardware performance evaluation time, the hardware performance output by the heterogeneous multi-core architecture is improved in different applications, and the improvement of the energy-delay product can reach up to 235.8 times, thereby making the operation of service devices such as computers stable and reducing the computing time of service devices such as computers.

[0098] According to an embodiment of the present invention, the g-th neural network to be evaluated includes multiple operators; mapping the g-th neural network to multiple hardware cores to obtain multiple target hardware cores, including: clustering all operators according to the dimension parameters of each operator to obtain the category of each operator; mapping the operators with the same category to the same hardware core to obtain multiple target hardware cores.

[0099] According to an embodiment of the present invention, the dimension parameters of each operator include six dimension parameters, and K, C, R, S, X, and Y are used to represent the number of output channels, the number of input channels, the height of the convolution kernel, the width of the convolution kernel, the height of the input feature map, and the width of the input feature map respectively. Among them, for the matrix multiplication operator, it is regarded as equivalent to the convolution calculation with a convolution kernel size of 1, thereby covering the main operators of convolutional neural networks and neural network loads of the Transformer architecture class.

[0100] According to an embodiment of the present invention, the mapping only needs to concern the similarity of the computing characteristics between each operator, and uniformly map similar operators to the same hardware core. Specifically, the DB-scan clustering algorithm is adopted, based on the above six-dimensional parameters, and operator clustering is performed in a high-dimensional space. Among them, the proximity of the values on different dimensional parameters means the similarity preference in the data reuse pattern and parallel dimension. The above clustering algorithm divides the categories of operators by the distance between neighboring points, and can automatically determine the number of operator categories according to the situation of the clustering points. Among them, the similarity distance between different operators is calculated using the Euclidean distance.

[0101] According to an embodiment of the present invention, the heterogeneous multi-core architecture can design a common parallel scheme for operators with similar computing characteristics, which has stronger adaptability compared with the prior art where all operators share the same set of hardware performance calculation methods. Therefore, there is no need to construct different architectures for different operators, which can reduce the storage space occupation of service devices such as computers.

[0102] According to an embodiment of the present invention, the data flow parameters of each target hardware core are optimized to obtain the target data flow parameters of each target hardware core, including: for the data flow parameters of any target hardware core, iteratively optimize the data flow parameters to obtain candidate data flow parameters; determine the candidate data flow parameters that meet the preset iteration conditions of the data flow parameters as the target data flow parameters.

[0103] According to an embodiment of the present invention, the data flow parameters are iteratively optimized based on the genetic algorithm. Each iteration of the data flow parameters serves as each generation of individuals in the genetic algorithm, that is, the candidate data flow parameters.

[0104] Figure 4 Schematically shows a schematic diagram of the data flow parameters according to an embodiment of the present invention.

[0105] According to an embodiment of the present invention, taking one-dimensional data flow parameters as an example, the data flow parameters include the order, parallel dimension, parallelism, and shard size of different dimensions, and can be represented by a 2x7 genome. Among them, as Figure 4 shown, based on the load description of the neural network to be evaluated corresponding to each operator of the neural network to be evaluated, map the operator to the hardware core, and encode the obtained target hardware core to obtain the data flow parameters. Among them, the first row is the parallel dimension (corresponding to Figure 4 the first column "K" and the eighth column "C" in Figure 4In the second to seventh columns “S - Y” and the ninth to fourteenth columns “S - R”), the calculation order of the above six - dimensional parameters of the operator corresponding to this calculation order, and the second row shows the corresponding parallelism and shard size. Further, for a two - dimensional array, it can be regarded as the superposition of multiple one - dimensional arrays. Therefore, the encoding can be represented by a 2×14 genome. At this time, columns 8 - 14 represent one - dimensional arrays, and columns 1 - 7 represent the scheduling and parallelization on top of the one - dimensional arrays.

[0106] According to an embodiment of the present invention, based on the characteristics of the genetic algorithm, candidate data - flow parameters with high data - flow parameter fitness will be selected to generate new candidate data - flow parameters. Specifically, new candidate data - flow parameters will be generated from the existing candidate data - flow parameters. During the optimization process, mutation for cross - mutation update of a single candidate data - flow parameter can be swapping the loop order, modifying the parallel dimension, mutating the parallelism and shard size. Generating new data - flow parameters for multiple data - flow parameters in each generation can also be the cross - over between two - dimensional and one - dimensional genes, cross - over of parallel dimensions, etc.

[0107] According to an embodiment of the present invention, the iterative convergence period of the genetic algorithm has a certain degree of randomness, which varies according to the cross - mutation probability and the operator. To ensure the quality of iterative optimization, the preset iterative condition can be to set the number of iterations and select the optimal candidate data - flow parameter as the target data - flow parameter. Among them, for the case where the iterative convergence is relatively fast, an early - stopping mechanism is adopted to reduce the search time. When the candidate data - flow parameters have not been updated for three consecutive iterations, the convergence condition is reached, and the iterative optimization will be terminated in advance.

[0108] According to an embodiment of the present invention, before optimizing the data - flow parameters of each target hardware core, the method further includes: pruning the data - flow parameters of any target hardware core to obtain the pruned data - flow parameters.

[0109] According to an embodiment of the present invention, to reduce the computational burden of hardware performance, pruning of data - flow parameters is performed respectively based on values and based on policies. Pruning based on values means that in the exploration of a two - dimensional array, the exploration values of the shard size and parallelism are constrained to be multiples of 4, which will significantly reduce the hardware - performance search points and conform to the experience of hardware design. Pruning based on policies selects representative operators for data - flow search based on the load situation borne by the hardware core.

[0110] According to an embodiment of the present invention, when pruning is not performed, the data flow parameters of a hardware core need to perform hardware performance calculations for all operators. This repeated process is time-consuming, and at the same time, the evaluation time of the evaluation tool at the back end of the computer system will linearly increase with the number of operators. In addition, since the operators mapped to the same hardware core are the clustering results in the mapping stage, there is redundancy in the operation of performing hardware performance calculations for all operators. Pruning based on a policy searches for the optimal data flow parameter configuration for the operators and applies this data flow parameter to all operators on this hardware core.

[0111] According to an embodiment of the present invention, pruning the data flow parameters for any target hardware core avoids data redundancy and improves the computing speed of service devices such as computers.

[0112] Figure 5 Schematically shows a schematic diagram of the joint search architecture of a neural network and hardware according to an embodiment of the present invention.

[0113] According to an embodiment of the present invention, the above-mentioned joint search method for a neural network and hardware of the present invention can be implemented by a neural network joint search architecture as shown in Figure 5 The neural network joint search architecture includes a neural network search space, an accuracy evaluator, and a hardware searcher.

[0114] Input the preset neural network search conditions, training label data, and hardware constraint requirements into the neural network joint search architecture. The neural network search space corresponding to the current round of iteration is sampled based on the preset neural network search conditions, and a neural network population including G neural networks to be evaluated is output; the accuracy evaluator determines the target gradient value set of the g-th neural network to be evaluated among the G neural networks to be evaluated corresponding to the current round of iteration, and determines the accuracy score value of the g-th neural network to be evaluated according to the target gradient value set of the g-th neural network to be evaluated; the hardware searcher uses a heterogeneous multi-core architecture to perform hardware performance simulation on the hardware configuration parameters of the neural network to be evaluated, and determines the hardware performance value of the g-th neural network to be evaluated; subsequently, the neural network joint search architecture determines the fitness value of the g-th neural network to be evaluated according to the ratio between the accuracy score value and the hardware performance value of the g-th neural network to be evaluated, determines the neural network to be evaluated corresponding to the highest fitness value among the G fitness values of the neural networks to be evaluated as the candidate neural network corresponding to the current round of iteration, and optimizes the sampling distribution of the neural network search space for the next round based on the candidate neural network corresponding to the current round of iteration, and determines the target neural network that meets the predetermined iteration conditions and outputs it.

[0115] According to an embodiment of the present invention, integrating the designs of the above precision evaluator and hardware searcher, the present invention has the characteristics of being fast and efficient in the scenario of joint search. Among them, based on the ImageNet (ImageNet Large Scale Visual Recognition Challenge) dataset and the CIFAR-10 (Canadian Institute For Advanced Research) dataset, the neural network joint search architecture is evaluated for convolutional neural network and vision transformer scenarios. The experimental results show that the neural network joint search architecture of the present invention efficiently explores the neural network search space. Compared with the existing search methods, the present invention saves 12.96 times the energy consumption delay product, and at the same time improves the search efficiency by 48 times. In addition, compared with the prior art, the neural network searched by the present invention can achieve a 56% reduction in the energy consumption delay product while maintaining a similar network accuracy, and the search time is reduced from 420 hours to about 3 hours.

[0116] According to an embodiment of the present invention, the following test operations are performed on the above precision evaluator. The precision evaluator ranks the precision score values of neural networks for convolutional neural network and vision transformer scenarios respectively, and verifies the reliability of the ranking based on datasets of various scales. As shown in Table 1, the benchmark tool (NASBench-201) for CNN (Convolutional Neural Network) comes from the invention datasets (CIFAR10, CIFAR100, ImgNet16), and the benchmark tool (AutoFormer) for ViT (Vision Transformer) comes from the sampling of existing supernets, with a sampling scale of 1000 random neural networks (ImageNet-1k), and their true test precisions are obtained. The evaluation metrics for reliability use the Spearman coefficient (S) and the Kendall coefficient (K) to verify the consistency between the predicted precision ranking given by the precision evaluator of the present invention and the actual precision ranking. Both ranges are [-1, 1], and the larger the coefficient, the higher the correlation, that is, the more reliable the ranking.

[0117]

[0118] The results show that the prediction accuracy rankings and actual accuracies given by the accuracy evaluator of the present invention maintain good consistency compared with the evaluators of related technologies (ZenScore, ZiCo, GradSign, #Param, #FLOP, Snip, DSS, Param, FLOPs). For example, the Spearman coefficient (S) and Kendall coefficient (K) measured by the present invention reach approximately 0.64 and 0.82 respectively, indicating that it is reliable to use the accuracy evaluator of the present invention for accuracy evaluation; in addition, the accuracy evaluator of the present invention still maintains the highest consistency in the case of cross-neural network search space and cross-datasets, providing guarantee for model comparison during search.

[0119] In addition, the neural network search results after adopting the accuracy evaluator of the present invention can be shown in Table 2.

[0120]

[0121] It can be seen that compared with the evaluators of related technologies, the neural networks searched by the accuracy evaluator designed by the present invention show the highest neural network accuracy score values in various datasets. For example, in the vision transformer scenario, the neural network accuracy score value given by the accuracy evaluator of the present invention is relatively high; in addition, due to the characteristic of being free from training of the accuracy evaluator of the present invention, compared with other evaluators, the search time of the accuracy evaluator of the present invention can be reduced to 3 hours, bringing a 480x improvement in search efficiency, and thus reducing the computing time of service devices such as computers while reducing the occupancy of storage space of service devices such as computers.

[0122] According to an embodiment of the present invention, the following test operations are performed on the above-mentioned hardware searcher. Classical network models under CNN and ViT are selected, and large-scale and small-scale models are covered to apply the hardware searcher of the present invention for hardware performance search. Among them, the hardware constraints of small-scale networks (ResNet18, EfficientNetB1, ViT-Tiny, PiT-Tiny) are 96 computing units, and the hardware constraints of large-scale networks (ResNet50, VGG16, ViT-Small, PiT-Small) are 1024 hardware units. The goal of hardware performance search is to minimize the energy-delay product (EDP, Energy-Delay Product). The comparison benchmark selects single-core hardware search, naive operator partitioning method, and search without pruning, and compares the search results and search times under various strategies (single-core, stage, pruning, the present invention).

[0123] Figure 6 A schematic diagram showing the test results of the hardware searcher according to an embodiment of the present invention is schematically shown.

[0124] As Figure 6 shown, the horizontal axis in the figure represents ResNet18, ResNet50, EfficientNetB1, VGG16, Vision Transformer - Tiny, Vision Transformer - Small, Pyramid Vision Transformer - Tiny, and Pyramid Vision Transformer - Small respectively. Among them, on the left side of the vertical axis are 1E+11, 1E+12, … representing different energy - delay products, and on the right side is the search time. This figure shows the EDP and search time under different search strategies (Single - core, Stage - based, w / o purning, the hardware searcher of the present invention) in the above - mentioned small - scale network and large - scale network environments, that is, the four columns and broken lines corresponding to each network environment. The results show that the hardware searcher of the present invention significantly improves the EDP, showing an improvement of 4.3 times to 192.6 times in CNN and 2.7x to 235.8 times in ViT. This is due to the parallel adaptation of the heterogeneous multi - core structure to heterogeneous operators. Among them, the gain of the small - scale network is better than that of the large - scale network, proving the necessity of heterogeneous multi - cores in resource - constrained scenarios. In addition, search space pruning significantly reduces the search time. The results show that the search efficiency is improved by 2.2 times to 7.4 times compared with the unpruned version, but maintains similar search results, indicating that the pruning strategy of the present invention retains the excellent solution space and effectively removes redundant search points.

[0125] According to an embodiment of the present invention, the following test operations are performed on the above - mentioned neural network joint search architecture. Joint search is carried out for the CNN scenario and the ViT scenario, paying attention to hardware adaptation and hardware performance while pursuing the accuracy of the neural network. The present invention uses a training - free evaluator and a hardware searcher to search for excellent neural networks, and then trains a single neural network for application. The search results are shown in Table 3.

[0126] In the CNN scenario, the present invention explores the neural network search space based on ProxylessNAS. Compared with existing methods (DANCE, DIAN), the search results of the present invention show the characteristics of high accuracy and high hardware efficiency (EDP). The accuracy of the neural network reaches 73.96%, with an accuracy improvement of 1.76% to 3.81% while the energy - delay product is improved by 3.54x and 12.96x respectively. In addition, the present invention does not require a pre - training and data collection process, and the search time of service devices such as computers is only 3h, achieving an efficiency improvement of 48x.

[0127] For the ViT scenario, compared with existing methods (AutoF-T, AutoF-S, AutoF-B), the neural network joint search architecture of the present invention maintains similar neural network accuracy performance, with a maximum accuracy loss of 0.2%. However, in terms of hardware efficiency, the search results of the present invention have a 14% - 36% improvement, and at the same time, the search time is increased by 131 to 161 times, from 420h to about 3h, thereby being able to reduce the search time of service devices such as computers.

[0128]

[0129] Based on the above joint search method for neural networks and hardware, the present invention also provides a neural network joint search device. The following will be combined with Figure 7 to describe this device in detail.

[0130] Figure 7 The structural block diagram of the neural network joint search device according to an embodiment of the present invention is schematically shown.

[0131] As Figure 7 shown, the neural network joint search device 700 of this embodiment includes a storage unit 710, a computing unit 720, and a heterogeneous multi-core architecture 730.

[0132] The storage unit 710 is used to store the neural network search space.

[0133] The computing unit 720 is configured to: based on preset neural network search conditions, sample the neural network search space corresponding to the current round of iteration in the storage unit 710 to obtain G neural networks to be evaluated corresponding to the current round of iteration; based on the g-th neural network to be evaluated among the G neural networks to be evaluated corresponding to the current round of iteration, perform a gradient calculation task to obtain a set of target gradient values of the g-th neural network to be evaluated, where G≥1 and 1≤g≤G; determine the accuracy score value of the g-th neural network to be evaluated according to the set of target gradient values of the g-th neural network to be evaluated.

[0134] The heterogeneous multi-core architecture 730 is configured to: perform hardware performance simulation on the hardware configuration parameters of the neural network to be evaluated, and determine the hardware performance value of the g-th neural network to be evaluated.

[0135] Among them, the computing unit 720 is further configured to: determine the adaptability value of the g-th neural network to be evaluated according to the ratio between the accuracy score value and the hardware performance value of the g-th neural network to be evaluated, and obtain the adaptability values of the G neural networks to be evaluated; determine the neural network to be evaluated corresponding to the highest adaptability value among the adaptability values of the G neural networks to be evaluated as the candidate neural network corresponding to the current round of iteration; determine the target neural network applicable to perform the target processing task from the candidate neural networks corresponding to multiple rounds.

[0136] According to an embodiment of the present invention, based on a preset neural network search condition, sample the neural network search space corresponding to the current round of iteration in the storage unit 710 to obtain G neural networks to be evaluated corresponding to the current round of iteration; based on the g-th neural network to be evaluated among the G neural networks to be evaluated corresponding to the current round of iteration, use the computing unit 720 to execute a gradient calculation task to obtain a set of target gradient values of the g-th neural network to be evaluated; according to the set of target gradient values of the g-th neural network to be evaluated, determine the accuracy score value of the g-th neural network to be evaluated; use the heterogeneous multi-core architecture 730 to perform hardware performance simulation on the hardware configuration parameters of the neural network to be evaluated, and determine the hardware performance value of the g-th neural network to be evaluated; according to the ratio between the accuracy score value and the hardware performance value of the g-th neural network to be evaluated, determine the adaptability value of the g-th neural network to be evaluated, and obtain the adaptability values of the G neural networks to be evaluated; determine the neural network to be evaluated corresponding to the highest adaptability value among the adaptability values of the G neural networks to be evaluated as the candidate neural network corresponding to the current round of iteration; determine the target neural network applicable to execute the target processing task from the candidate neural networks corresponding to multiple rounds. Since the computing unit 720 is used to execute the gradient calculation task, the relationship between the gradient information and the model convergence of the neural network to be evaluated is fully utilized when determining the accuracy score value corresponding to the neural network to be evaluated, avoiding the design of a heavy and complex super network, realizing the reduction of the computing overhead of service devices such as computers, and improving the search accuracy of the target neural network while ensuring the improvement of the computing speed of service devices such as computers; in addition, since the heterogeneous multi-core architecture 730 is relied on to simulate the hardware performance of the neural network to be evaluated, the problem that the single-core architecture fails to fully consider the differences and computing preferences of network layers is avoided, so that service devices such as computers do not need to reconfigure the hardware performance evaluation architecture for different neural networks to be evaluated, ensuring the stable operation of service devices such as computers and reducing the storage space occupied by service devices such as computers.

[0137] According to an embodiment of the present invention, any plurality of modules among the storage unit 710, the computing unit 720, and the heterogeneous multi-core architecture 730 may be combined and implemented in one module, or any one of them may be split into multiple modules. Alternatively, at least part of the functions of one or more of these modules may be combined with at least part of the functions of other modules and implemented in one module. According to an embodiment of the present invention, at least one of the storage unit 710, the computing unit 720, and the heterogeneous multi-core architecture 730 may be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on chip, a system on substrate, a system on package, an application specific integrated circuit (ASIC), or any other reasonable way of integrating or packaging circuits, etc., implemented by hardware or firmware, or implemented in any one of the three implementation manners of software, hardware, and firmware, or in an appropriate combination of any several of them. Alternatively, at least one of the storage unit 710, the computing unit 720, and the heterogeneous multi-core architecture 730 may be at least partially implemented as a computer program module, which can execute corresponding functions when the computer program module is run.

[0138] According to an embodiment of the present invention, the computing unit 720 is further configured to:

[0139] Determine M batches of training label data for inputting into the g-th neural network to be evaluated, where M≥1.

[0140] For the m-th batch of training label data among the M batches, according to the m-th batch of training label data of the g-th neural network to be evaluated, determine the loss value of the m-th batch of the g-th neural network to be evaluated, where 1≤m≤M.

[0141] Calculate the derivatives of the loss value of the m-th batch and the weight parameters of the N g-th neural networks to be evaluated respectively, to obtain N gradient values of the m-th batch, and obtain M×N gradient values of the M batches, where N≥1.

[0142] Screen the M×N gradient values to obtain the set of target gradient values of the g-th neural network to be evaluated.

[0143] According to an embodiment of the present invention, the computing unit 720 is further configured to:

[0144] For the n-th weight parameter among the N weight parameters of the g-th neural network to be evaluated, divide the M gradient values corresponding to the n-th weight parameter into multiple gradient intervals, where 1≤n≤N.

[0145] According to the density of each gradient interval, select the gradient interval with the largest density among the multiple gradient intervals as the n-th target gradient interval corresponding to the n-th weight parameter, to obtain N target gradient intervals.

[0146] Determine the nth set of target gradient values corresponding to the nth weight parameter according to the nth target gradient interval, and obtain N sets of target gradient values.

[0147] Use the N sets of target gradient values as the set of target gradient values of the gth neural network to be evaluated.

[0148] According to an embodiment of the present invention, the calculation unit 720 is further configured to:

[0149] For the nth weight parameter of the gth neural network to be evaluated, determine the average gradient value of the nth set of target gradient values.

[0150] Determine the gradient variance value of the nth set of target gradient values.

[0151] Determine the nth initial accuracy score value according to the ratio between the average gradient value and the gradient variance value, and obtain N initial accuracy score values.

[0152] Determine the accuracy score value of the gth neural network to be evaluated according to the N initial accuracy score values.

[0153] According to an embodiment of the present invention, the heterogeneous multi-core architecture 730 is further configured to:

[0154] Map the gth neural network to be evaluated to multiple hardware cores to obtain multiple target hardware cores.

[0155] Determine the target number of computing units and the target bandwidth of each target hardware core according to the ratio of the computing power requirement and the broadband requirement of each target hardware core.

[0156] Optimize the data flow parameters of each of the above-mentioned target hardware cores to obtain the target data flow parameters of each target hardware core.

[0157] Use the performance simulation framework to perform hardware performance simulation on the target number of computing units, the target bandwidth, and the target data flow parameters of each target hardware core, and determine the hardware performance value of the gth neural network to be evaluated.

[0158] According to an embodiment of the present invention, the heterogeneous multi-core architecture 730 is further configured to:

[0159] Cluster all operators according to the dimension parameters of each operator to obtain the category of each operator.

[0160] Map the operators with the same category to the same hardware core to obtain multiple target hardware cores.

[0161] According to an embodiment of the present invention, the heterogeneous multi-core architecture 730 is further configured to:

[0162] Iteratively optimize the data flow parameters for any target hardware core to obtain candidate data flow parameters.

[0163] Determine the candidate data flow parameters that meet the preset iteration conditions of the data flow parameters as the data flow parameters.

[0164] According to an embodiment of the present invention, the heterogeneous multi-core architecture 730 is further configured to:

[0165] Prune the data flow parameters for any target hardware core to obtain the pruned data flow parameters.

[0166] Figure 8 Schematically shows a block diagram of an electronic device suitable for implementing the joint search method of neural network and hardware according to an embodiment of the present invention.

[0167] As Figure 8 shown, the electronic device according to an embodiment of the present invention includes a processor 801, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 802 or the program loaded from the storage section 808 into the random access memory (RAM) 803. The processor 801 may include, for example, a general microprocessor (such as a CPU), an instruction set processor, and / or a related chipset, and / or a dedicated microprocessor (such as an application specific integrated circuit (ASIC)), etc. The processor 801 may also include on-board memory for caching purposes. The processor 801 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present invention.

[0168] In the RAM 803, various programs and data required for the operation of the electronic device are stored. The processor 801, the ROM 802, and the RAM 803 are connected to each other through a bus 804. The processor 801 performs various operations of the method flow according to an embodiment of the present invention by executing the programs in the ROM 802 and / or the RAM 803. It should be noted that the program may also be stored in one or more memories other than the ROM 802 and the RAM 803. The processor 801 may also perform various operations of the method flow according to an embodiment of the present invention by executing the programs stored in the one or more memories.

[0169] According to an embodiment of the present invention, the electronic device may further include an input / output (I / O) interface 805, and the input / output (I / O) interface 805 is also connected to the bus 804. The electronic device 800 may further include one or more of the following components connected to the input / output (I / O) interface 805: an input portion 806 including a keyboard, a mouse, etc.; an output portion 807 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage portion 808 including a hard disk, etc.; and a communication portion 809 including a network interface card such as a LAN card, a modem, etc. The communication portion 809 performs communication processing via a network such as the Internet. The drive 810 is also connected to the input / output (I / O) interface 805 as needed. A removable medium 811, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 810 as needed so that a computer program read from it can be installed into the storage portion 808 as needed.

[0170] The present invention also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or may exist separately without being assembled into the device / apparatus / system. The above computer-readable storage medium carries one or more programs, and when the above one or more programs are executed, the method according to the embodiments of the present invention is implemented.

[0171] According to an embodiment of the present invention, the computer-readable storage medium may be a non-volatile computer-readable storage medium, for example, it may include but is not limited to: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, the computer-readable storage medium may be any tangible medium that contains or stores a program, and this program can be used by or combined with an instruction execution system, device, or apparatus. For example, according to an embodiment of the present invention, the computer-readable storage medium may include the ROM 802 and / or the RAM 803 described above and / or one or more memories other than the ROM 802 and the RAM 803.

[0172] An embodiment of the present invention further includes a computer program product, which includes a computer program, and the computer program includes program code for executing the method shown in the flowchart. When the computer program product runs in a computer system, the program code is used to cause the computer system to implement the joint search method of the neural network and the hardware provided by the embodiments of the present invention.

[0173] When the computer program is executed by the processor 801, the above functions defined in the system / apparatus of the embodiments of the present invention are executed. According to the embodiments of the present invention, the systems, apparatuses, modules, units, etc. described above can be implemented by computer program modules.

[0174] In one embodiment, the computer program can rely on tangible storage media such as optical storage devices, magnetic storage devices, etc. In another embodiment, the computer program can also be transmitted and distributed in the form of signals on a network medium, and be downloaded and installed through the communication part 809, and / or be installed from the removable medium 811. The program code included in the computer program can be transmitted by any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.

[0175] In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 809, and / or be installed from the removable medium 811. When the computer program is executed by the processor 801, the above functions defined in the system of the embodiments of the present invention are executed. According to the embodiments of the present invention, the systems, devices, apparatuses, modules, units, etc. described above can be implemented by computer program modules.

[0176] According to the embodiments of the present invention, the program code for executing the computer program provided by the embodiments of the present invention can be written in any combination of one or more programming languages. Specifically, these computing programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include but are not limited to, such as Java, C++, python, the "C" language, or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (for example, by using an Internet service provider to connect through the Internet).

[0177] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that, in some alternative implementations, the functions noted in the blocks may occur in a different order than that noted in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and combinations of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or by a combination of dedicated hardware and computer instructions.

[0178] Those skilled in the art will appreciate that the features described in the various embodiments of the present invention can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in the present invention. In particular, without departing from the spirit and teachings of the present invention, the features described in the various embodiments of the present invention can be combined and / or combined in various ways. All such combinations and / or combinations fall within the scope of the present invention.

[0179] The embodiments of the present invention have been described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present invention. Although the embodiments have been described separately above, this does not mean that the measures in each embodiment cannot be used advantageously in combination. Without departing from the scope of the present invention, those skilled in the art can make various substitutions and modifications, and all such substitutions and modifications should fall within the scope of the present invention.

Claims

1. A joint search method of neural network and hardware, characterized in that: Applied to a computing device, the method comprises: Based on the preset neural network search conditions, the neural network search space corresponding to the current round iteration in the storage unit is sampled to obtain G neural networks to be evaluated corresponding to the current round iteration; Based on the g-th neural network to be evaluated among the G neural networks to be evaluated corresponding to the current round iteration, using the computing unit to perform the gradient computing task, to obtain a target gradient value set of the g-th neural network to be evaluated, G≥1, 1≤g≤G; Determining the accuracy score value of the g-th neural network to be evaluated according to the target gradient value set of the g-th neural network to be evaluated; Performing hardware performance simulation on the hardware configuration parameters of the neural network to be evaluated by using a heterogeneous multi-core architecture to determine the hardware performance value of the g-th neural network to be evaluated; Determine the adaptability value of the g-th neural network to be evaluated according to the ratio between the accuracy score value of the g-th neural network to be evaluated and the hardware performance value, and obtain the adaptability values ​​of the G neural networks to be evaluated; Determine the neural network to be evaluated corresponding to the highest fitness value among the fitness values ​​of the G neural networks to be evaluated as the candidate neural network corresponding to the current round iteration; A target neural network suitable for performing the target processing task is determined from multiple rounds of corresponding candidate neural networks.

2. The method according to claim 1, characterized in that: The target processing task includes at least one of the following: Object detection tasks, text processing tasks, and image classification tasks.

3. The method according to claim 1, characterized in that: The step of performing a gradient calculation task based on the g-th neural network to be evaluated among the G neural networks to be evaluated corresponding to the current iteration using a computing unit to obtain a target gradient value set of the g-th neural network to be evaluated includes: Determine M batches of training label data to be input into the g-th neural network to be evaluated, M≥1; For the mth batch of training label data in the M batches, determine the loss value of the mth batch of the gth neural network to be evaluated according to the mth batch of training label data of the gth neural network to be evaluated, 1≤m≤M; Calculate the loss value of the mth batch and the derivatives of the weight parameters of the N gth neural networks to be evaluated respectively, obtain N gradient values ​​of the mth batch, and obtain M×N gradient values ​​of the M batches, where N≥1; The M×N gradient values ​​are screened to obtain a target gradient value set of the g-th neural network to be evaluated.

4. The method according to claim 3, characterized in that: The screening of the M×N gradient values ​​to obtain the target gradient value set of the g-th neural network to be evaluated includes: For an nth weight parameter among the N weight parameters of the gth neural network to be evaluated, dividing the M gradient values ​​corresponding to the nth weight parameter into intervals to obtain a plurality of gradient intervals, 1≤n≤N; According to the density of each of the gradient intervals, a gradient interval with the largest density is selected from the plurality of gradient intervals as the nth target gradient interval corresponding to the nth weight parameter, to obtain N target gradient intervals; Determining, according to the nth target gradient interval, an nth target gradient value set corresponding to the nth weight parameter, to obtain N target gradient value sets; The N target gradient value sets are used as the target gradient value sets of the g-th neural network to be evaluated.

5. The method according to claim 4, characterized in that The step of determining the accuracy score value of the g-th neural network to be evaluated according to the target gradient value set of the g-th neural network to be evaluated comprises: For the nth weight parameter of the gth neural network to be evaluated, determining the gradient average of the nth target gradient value set; Determining a gradient variance value of the nth target gradient value set; Determine the nth initial accuracy score value according to the ratio between the gradient average value and the gradient variance value, and obtain N initial accuracy score values; According to the N initial accuracy score values, determine the accuracy score value of the g-th neural network to be evaluated.

6. The method according to claim 1, characterized in that The heterogeneous multi-core architecture includes multiple hardware cores; The using of a heterogeneous multi-core architecture to perform hardware performance simulation on the hardware configuration parameters of the neural network to be evaluated to determine the hardware performance value of the g-th neural network to be evaluated includes: Mapping the g-th neural network to be evaluated to the plurality of hardware cores to obtain a plurality of target hardware cores; Determining the target number of computing units and the target bandwidth of each target hardware core according to the ratio of the computing demand and the bandwidth demand of each target hardware core; Optimizing the data flow parameters of each of the target hardware cores to obtain target data flow parameters of each of the target hardware cores; Using a performance simulation framework, hardware performance simulation is performed on the target number of computing units, the target bandwidth, and the target data flow parameters of each target hardware core to determine the hardware performance value of the g-th neural network to be evaluated.

7. The method according to claim 6, characterized in that The g-th neural network to be evaluated includes multiple operators; Mapping the g-th neural network to be evaluated to the plurality of hardware cores to obtain a plurality of target hardware cores comprises: Clustering all the operators according to the dimension parameters of each operator to obtain a category of each operator; The operators of the same category are mapped to the same hardware core to obtain a plurality of target hardware cores.

8. The method according to claim 6, characterized in that The step of optimizing the data flow parameters of each target hardware core to obtain the target data flow parameters of each target hardware core includes: For any data flow parameter of the target hardware core, iteratively optimize the data flow parameter to obtain candidate data flow parameters; A candidate data stream parameter that meets a preset iteration condition of the data stream parameter is determined as the data stream parameter.

9. The method according to claim 6, characterized in that Before optimizing the data flow parameters of each of the target hardware cores, the method further includes: Pruning is performed on the data flow parameters of any of the target hardware cores to obtain pruned data flow parameters.

10. A joint search device of neural network and hardware, characterized in that: include: A storage unit for storing the neural network search space; Computing unit, configured as: Based on the preset neural network search condition, sampling the neural network search space corresponding to the current round iteration in the storage unit to obtain G neural networks to be evaluated corresponding to the current round iteration; Based on the g-th neural network to be evaluated among the G neural networks to be evaluated corresponding to the current round iteration, perform a gradient calculation task to obtain a target gradient value set of the g-th neural network to be evaluated, G≥1, 1≤g≤G; Determining the accuracy score value of the g-th neural network to be evaluated according to the target gradient value set of the g-th neural network to be evaluated; Heterogeneous multi-core architecture, configured as: Performing hardware performance simulation on the hardware configuration parameters of the neural network to be evaluated, and determining the hardware performance value of the g-th neural network to be evaluated; Wherein, the computing unit is further configured to: determine the adaptability value of the g-th neural network to be evaluated according to the ratio between the accuracy score value of the g-th neural network to be evaluated and the hardware performance value, and obtain the adaptability values ​​of the G neural networks to be evaluated; determine the neural network to be evaluated corresponding to the highest adaptability value among the adaptability values ​​of the G neural networks to be evaluated as the candidate neural network corresponding to the current round iteration; determine the target neural network suitable for executing the target processing task from the candidate neural networks corresponding to each of the multiple rounds.

Citation Information

Patent Citations

  • Adaptive search method and device for neural network

    CN113128678A

  • Neural network architecture searching method and device, equipment and medium

    CN113361680A

  • Software and hardware joint search method, device and equipment oriented to storage and calculation integrated architecture

    CN115293341A

  • System for universal hardware-neural network architecture search

    CN115906962A

  • Joint search method, apparatus and device for CNN model and accelerator, and medium

    CN118095364A