Architecture determination method and device
By constructing the architecture determination model of covariance function and target neural network, adjusting the weight parameters to determine the distribution of the objective function, the problem of inaccurate determination of CGRA architecture in the existing technology is solved, and more accurate and efficient architecture selection is achieved.
Patent Information
- Application Number
- CN202410227187.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-02-29
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2044-02-29
AI Technical Summary
When using coarse-grained reconfigurable architecture (CGRA) to process data calculation tasks, it is difficult to accurately determine the CGRA that best performs in processing the current task, and it is easy to fall into the local optimal solution.
By building an architectural determination model that includes covariance functions and target neural network, the weight parameters of the target neural network are adjusted using architecture sample data and real indicator data, and then more accurate target function distribution and target architecture data are determined.
Improves the accuracy of architecture determination, ensures a more targeted coarse-grained reconfigurable architecture selection for the current task, reduces training time and improves model accuracy.
Smart Images

Figure CN118093549B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technology, and in particular to an architecture determination method, apparatus, device, medium, and program product. Background Art
[0002] As the demand for data computing continues to grow, the requirements for processors are also getting higher and higher. The use of coarse-grained reconfigurable architecture (CGRA) arrays has become a trend to relieve processor computing pressure and accelerate computing tasks. When using CGRA arrays to process data computing tasks, it is usually necessary to reconstruct the CGRA to determine the CGRA that is most effective in processing the current task.
[0003] However, since CGRA is composed of multiple components, each with its own organizational structure, CGRA has a huge architectural space. It is difficult to find the CGRA that best processes the current task from the architectural space. Related technologies usually use heuristic algorithms based on simulated annealing, evolutionary algorithms, and other algorithms to determine the CGRA that best processes the current task. However, the above algorithms still have the technical problem of easily falling into local optimal solutions, resulting in inaccurate CGRAs. Summary of the invention
[0004] In view of the above problems, the present disclosure provides an architecture determination method, apparatus, device, medium and program product for improving the accuracy of architecture determination.
[0005] According to one aspect of the present disclosure, a method for determining an architecture is provided, comprising:
[0006] In response to having received a computing task, inputting architecture sample data into an architecture determination model to obtain test index data of the architecture sample data, wherein the architecture determination model includes a covariance function and a target neural network, the output of the target neural network is the input of the covariance function, and the architecture sample data is data for characterizing a coarse-grained reconfigurable architecture; based on the test index data and the real index data of the architecture sample data corresponding to the computing task, adjusting the weight parameters in the target neural network to obtain an adjusted architecture determination model; based on the adjusted architecture determination model, determining the target function distribution; based on a plurality of target index data obtained by respectively inputting a plurality of preset architecture data into the target function distribution, determining the target architecture data; and when it is determined that the target architecture data meets the preset conditions, determining the target coarse-grained reconfigurable architecture for processing the computing task based on the target architecture data.
[0007] Another aspect of the present disclosure provides an architecture determination device, including:
[0008] A data input module is used to input the architecture sample data into the architecture determination model in response to the received computing task, so as to obtain the test index data of the architecture sample data, wherein the architecture determination model includes a covariance function and a target neural network, the output of the target neural network is the input of the covariance function, and the architecture sample data is the data for characterizing the coarse-grained reconfigurable architecture; a model adjustment module is used to adjust the weight parameters in the target neural network based on the test index data and the real index data of the architecture sample data corresponding to the computing task, so as to obtain the adjusted architecture determination model; a distribution determination module is used to determine the target function distribution based on the adjusted architecture determination model; a data determination module is used to determine the target architecture data based on the multiple target index data obtained by respectively inputting multiple preset architecture data into the target function distribution; an architecture determination module is used to determine the target coarse-grained reconfigurable architecture for processing the computing task based on the target architecture data when it is determined that the target architecture data meets the preset conditions.
[0009] Another aspect of the present disclosure provides an electronic device, including: one or more processors; and a memory for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors execute the above-mentioned architecture determination method.
[0010] Another aspect of the present disclosure further provides a computer-readable storage medium having executable instructions stored thereon, which, when executed by a processor, causes the processor to execute the above-mentioned architecture determination method.
[0011] Another aspect of the present disclosure further provides a computer program product, including a computer program, which implements the above architecture determination method when executed by a processor.
[0012] According to the architecture determination method provided by the present disclosure, by inputting architecture sample data into an architecture determination model including a covariance function and a target neural network, test index data is obtained, and based on the test index data and the real index data of the architecture sample data corresponding to the computing task, the weight parameters of the target neural network are adjusted, thereby obtaining a more accurate architecture determination model, and the target function distribution is determined based on the adjusted architecture determination model, thereby obtaining a coarse-grained reconfigurable architecture with the best effect in processing the computing task. Since the architecture determination model is obtained through the covariance function and the target neural network, and the output of the target neural network is the input of the covariance function, the input data is adjusted through the target neural network and then input into the covariance function, so that the process of adjusting the weight parameters in the target neural network can replace the process of adjusting the hyperparameters in the covariance function, thereby making the training of the architecture determination model faster, and making the trained architecture determination model more accurate, and since the architecture determination model is trained using the real index data of the architecture sample data corresponding to the computing task, the architecture determination model is more targeted to the computing task, and the target architecture data determined based on the architecture determination model is also more accurate. Therefore, the technical problem of being unable to accurately determine a CGRA that is effective in processing the current task is at least partially solved, and the technical effect of determining a more accurate coarse-grained reconfigurable architecture for the current task is achieved. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] The above contents and other purposes, features and advantages of the present disclosure will become more apparent through the following description of the embodiments of the present disclosure with reference to the accompanying drawings, in which:
[0014] Figure 1 A schematic diagram showing an application scenario of the architecture determination method, apparatus, device, medium, and program product according to an embodiment of the present disclosure;
[0015] Figure 2 A flowchart of a method for determining an architecture according to an embodiment of the present disclosure is schematically shown;
[0016] Figure 3 A flowchart for determining target distribution according to an embodiment of the present disclosure is schematically shown;
[0017] Figure 4 A schematic diagram of an initial coarse-grained reconfigurable architecture array according to an embodiment of the present disclosure is schematically shown;
[0018] Figure 5 A flowchart of a method for determining an architecture according to another embodiment of the present disclosure is schematically shown;
[0019] Figure 6 A structural block diagram schematically shows a device for determining an architecture according to an embodiment of the present disclosure; and
[0020] Figure 7 A block diagram of an electronic device suitable for implementing the architecture determination method according to an embodiment of the present disclosure is schematically shown. DETAILED DESCRIPTION
[0021] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the present disclosure. In the following detailed description, for ease of explanation, many specific details are set forth to provide a comprehensive understanding of the embodiments of the present disclosure. However, it is apparent that one or more embodiments may also be implemented without these specific details. In addition, in the following description, descriptions of known structures and technologies are omitted to avoid unnecessary confusion of the concepts of the present disclosure.
[0022] The terms used herein are only for describing specific embodiments and are not intended to limit the present disclosure. The terms "comprise", "include", etc. used herein indicate the existence of the features, steps, operations and / or components, but do not exclude the existence or addition of one or more other features, steps, operations or components.
[0023] All terms (including technical and scientific terms) used herein have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification, and should not be interpreted in an idealized or overly rigid manner.
[0024] When using expressions such as "at least one of A, B, and C, etc.", they should generally be interpreted according to the meaning of the expression commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include but is not limited to a system having A alone, B alone, C alone, A and B, A and C, B and C, and / or A, B, C, etc.).
[0025] During the research process, it was found that with the rapid growth of demand for artificial intelligence applications, the requirements for computing, storage, and data exchange have also shown a huge growth. This has caused traditional processors to face huge performance and energy challenges. Recently, due to its computational flexibility and reconfigurability, coarse-grained reconfigurable architecture arrays have become an inevitable trend to relieve processor computing pressure and accelerate computing tasks.
[0026] The CGRA microarchitecture consists of multiple components, each with a diverse organizational structure and allowing for reconfiguration at a higher granularity. Under a specific technology flow, different CGRA microarchitectures exhibit different characteristics such as area, performance, and power consumption. Therefore, finding a microarchitecture design that balances area and performance becomes very complicated. First, the entire design space is very large, and its size expands exponentially as the number of components considered increases. Second, for each CGRA microarchitecture design with a specific benchmark, we need to use commercial EDA tools, such as for detailed simulation, to obtain important indicators such as area and time, which requires a lot of running time and computing resources.
[0027] However, designing a microarchitecture that can balance performance and maintain an appropriate area under limited resources is a challenging task for researchers. Progress in this field depends not only on a deep understanding of the design space, but also requires highly optimized simulation and evaluation processes to effectively explore the optimal microarchitecture design.
[0028] Generally speaking, Design Space Exploration (DSE) can help to individually balance performance, area, and other indicators in the design, turning the traditional architecture design problem into an optimization problem of a black box function. However, the huge design space associated with CGRA makes DSE very time-consuming and expensive. In order to find design solutions, several optimization algorithms have been developed, including heuristic algorithms based on simulated annealing, particle swarm optimization algorithms (PSO), and evolutionary algorithms. Although these algorithms reduce the number of hardware simulations, most of them are prone to falling into local optimal solutions and have relatively low coverage. At the same time, in CGRA design, how to perform DSE more effectively to find a design that meets performance requirements while maintaining an appropriate area under limited resources is a challenging problem.
[0029] Currently, when many architectures conduct architectural space exploration, the performance indicators of the architecture are obtained through corresponding performance models or power consumption models, etc. However, the architectural indicators obtained in this way are not accurate enough, which is very inaccurate for the subsequent use of optimization algorithms to solve black box functions, and will affect the quality of the final solution.
[0030] In view of this, an embodiment of the present disclosure provides an architecture determination method, in response to having received a computing task, inputting architecture sample data into an architecture determination model to obtain test indicator data of the architecture sample data, wherein the architecture determination model includes a covariance function and a target neural network, the output of the target neural network is the input of the covariance function, and the architecture sample data is data for characterizing a coarse-grained reconfigurable architecture; based on the test indicator data and the real indicator data of the architecture sample data corresponding to the computing task, adjusting the weight parameters in the target neural network to obtain an adjusted architecture determination model; based on the adjusted architecture determination model, determining the target function distribution; based on the multiple target indicator data obtained by respectively inputting multiple preset architecture data into the target function distribution, determining the target architecture data; when it is determined that the target architecture data meets the preset conditions, determining the coarse-grained reconfigurable architecture for processing the computing task based on the target architecture data.
[0031] Figure 1 The application scenario diagram of the architecture determination method, apparatus, device, medium and program product according to the embodiments of the present disclosure is schematically shown.
[0032] like Figure 1 As shown, the application scenario 100 according to this embodiment may include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used to provide a medium for a communication link between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired, wireless communication links or optical fiber cables, etc.
[0033] Users can use terminal devices 101, 102, 103 to interact with server 105 through network 104 to receive or send messages, etc. Various communication client applications can be installed on terminal devices 101, 102, 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (only examples).
[0034] The terminal devices 101 , 102 , and 103 may be various electronic devices having a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop computers, and desktop computers.
[0035] The server 105 may be a server that provides various services, such as a background management server (only an example) that provides support for websites browsed by users using the terminal devices 101, 102, and 103. The background management server may analyze and process the received data such as user requests, and feed back the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal device.
[0036] It should be noted that the architecture determination method provided in the embodiment of the present disclosure can generally be executed by the server 105. Accordingly, the architecture determination device provided in the embodiment of the present disclosure can generally be set in the server 105. The architecture determination method provided in the embodiment of the present disclosure can also be executed by a server or server cluster that is different from the server 105 and can communicate with the terminal devices 101, 102, 103 and / or the server 105. Accordingly, the architecture determination device provided in the embodiment of the present disclosure can also be set in a server or server cluster that is different from the server 105 and can communicate with the terminal devices 101, 102, 103 and / or the server 105.
[0037] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is only for illustration. Any number of terminal devices, networks and servers may be provided according to implementation requirements.
[0038] The following will be based on Figure 1 The scene described by Figure 2 to Figure 5 The architecture determination method of the disclosed embodiment is described in detail.
[0039] Figure 2 The flowchart of the architecture determination method according to the embodiment of the present disclosure is schematically shown.
[0040] like Figure 2 As shown, the method includes operations S210 to S250.
[0041] In operation S210, in response to receiving a computing task, the architecture sample data is input into the architecture determination model to obtain test indicator data of the architecture sample data, wherein the architecture determination model includes a covariance function and a target neural network, the output of the target neural network is the input of the covariance function, and the architecture sample data is data for characterizing the coarse-grained reconfigurable architecture.
[0042] In operation S220, based on the test indicator data and the real indicator data of the architecture sample data corresponding to the computing task, the weight parameters in the target neural network are adjusted to obtain an adjusted architecture determination model.
[0043] In operation S230 , the target function distribution is determined based on the adjusted architecture determination model.
[0044] In operation S240 , target architecture data is determined based on a plurality of target indicator data obtained by respectively inputting a plurality of preset architecture data into the target function distribution.
[0045] In operation S250 , when it is determined that the target architecture data satisfies a preset condition, a target coarse-grained reconfigurable architecture for processing the computing task is determined based on the target architecture data.
[0046] According to the embodiments of the present disclosure, there is no limitation on the computing task, and it can be any task that needs to be processed by a processor, such as a convolution task, a loop task, and the like.
[0047] According to an embodiment of the present disclosure, the architecture sample data may be data that can characterize a coarse-grained reconfigurable architecture, such as matrix data, vector data, and the like.
[0048] According to an embodiment of the present disclosure, the architecture determination model may be any model that can determine the indicator data of any coarse-grained reconfigurable architecture, such as a Gaussian model.
[0049] According to the embodiments of the present disclosure, there is no limitation on the target neural network, which may be a fully connected neural network, a convolutional neural network or other neural network.
[0050] According to an embodiment of the present disclosure, the covariance function may be a black box function.
[0051] According to the embodiments of the present disclosure, the number of layers of the target neural network is not limited and may be three layers, five layers, etc.
[0052] According to the embodiments of the present disclosure, there is no limitation on the presentation form of the target coarse-grained reconfigurable architecture, and it can be a variety of forms such as pictures, topologies, data flows, etc. that can represent the real coarse-grained reconfigurable architecture.
[0053] According to an embodiment of the present disclosure, different coarse-grained reconfigurable architectures can be obtained by reconstructing the coarse-grained reconfigurable architecture. The coarse-grained reconfigurable architecture includes computing nodes (Process Element, PE), and each computing node includes computing units (Function Unit, FU) such as: Mul, Mac, Add, Shift, Max, Phi, Br, etc.
[0054] According to an embodiment of the present disclosure, the reconstruction of the coarse-grained reconfigurable architecture can be achieved by reconstructing the number and type of computing units included in each computing node in the coarse-grained reconfigurable architecture, which can be specifically characterized as adjusting the data representing the computing units in the initial architecture data.
[0055] According to an embodiment of the present disclosure, the coarse-grained reconfigurable architecture may also be reconfigured by modifying the connection relationship between the plurality of computing nodes.
[0056] According to an embodiment of the present disclosure, different coarse-grained reconfigurable architectures have different indicator data values, and the indicator data may include an architecture area value, a task start interval value, power consumption, and the like.
[0057] According to an embodiment of the present disclosure, the real indicator data may be real indicator data related to the computing task obtained by scheduling the computing task to a coarse-grained reconfigurable architecture corresponding to the architecture sample data or by simulating based on the architecture sample data.
[0058] According to an embodiment of the present disclosure, based on the test index data and the real index data corresponding to the computing task, the loss value of the test index data can be determined, and the weight parameters of the target neural network can be iteratively adjusted through the loss value until the loss value is less than the preset loss value, so that a trained architecture determination model can be obtained, so that more accurate architecture index data can be obtained through the architecture determination model.
[0059] According to the embodiments of the present disclosure, since the architecture determines that the model includes a covariance function and a target neural network and the output of the target neural network is the input of the covariance function, the hyperparameters in the covariance function are fixed to specific values, and the weight parameters in the target neural network are trained so that the input data, after passing through the target neural network, is data that has been adjusted by the target neural network, and then the data is input into the covariance function, which is equivalent to realizing the training of the hyperparameters in the covariance function, avoiding the problem of high training difficulty and high consumption of computing resources caused by the need to train multiple hyperparameters separately due to the presence of multiple hyperparameters in the covariance function, and achieving beneficial effects such as reducing training time and increasing training accuracy.
[0060] According to an embodiment of the present disclosure, by including a model determined by a covariance function and a target neural network architecture, multi-objective output can be achieved, such as: outputting two indicators of an architecture area value and a task start interval value.
[0061] According to an embodiment of the present disclosure, specifically, the covariance function may be as shown in the following formula (1), and the covariance function after inputting the target neural network output data may be as shown in the following formula (2).
[0062] k(x,y)=k θ (x i ,y j ); (1)
[0063] k(x,y)=k w,θ (ψ(x i ,w),ψ(y j ,w));(2)
[0064] Among them, x is the architecture data, y is the test index value, i is the i-th architecture data, j is the j-th test index value, θ is the hyperparameter, ψ represents the nonlinear transformation using the target activation function, and w represents the weight in the target neural network.
[0065] According to the embodiments of the present disclosure, there is no limitation on the target activation function, such as tanh, relu, sigmoid, etc.
[0066] According to an embodiment of the present disclosure, a regularization operation may be used after each adjustment of the weight parameters in the target neural network, thereby preventing the target neural network from overfitting even when there is less architecture sample data.
[0067] According to an embodiment of the present disclosure, the adjusted architecture determination model is a trained architecture determination model.
[0068] According to an embodiment of the present disclosure, since the architecture determination model can be a Gaussian model, one or more objective function distributions can be determined by the architecture determination model to determine the target architecture data corresponding to the coarse-grained reconfigurable architecture used to process the computing task, and the process of determining the target architecture data based on the Gaussian model can be performed through a Bayesian optimization algorithm.
[0069] According to an embodiment of the present disclosure, the target architecture data can be simulated to determine whether the target architecture data meets preset conditions. If it meets the preset conditions, the target coarse-grained reconfigurable architecture corresponding to the target architecture data can be considered to be an architecture suitable for processing the computing task.
[0070] According to the architecture determination method provided by the present disclosure, by inputting architecture sample data into an architecture determination model including a covariance function and a target neural network, test index data is obtained, and based on the test index data and the real index data of the architecture sample data corresponding to the computing task, the weight parameters of the target neural network are adjusted, thereby obtaining a more accurate architecture determination model, and the target function distribution is determined based on the adjusted architecture determination model, thereby obtaining a coarse-grained reconfigurable architecture with the best effect in processing the computing task. Since the architecture determination model is obtained by the target neural network obtained by the covariance function and the target neural network, and the output of the target neural network is the input of the covariance function, the input data is adjusted by the target neural network and then input into the covariance function, so that the process of adjusting the weight parameters in the target neural network can replace the process of adjusting the hyperparameters in the covariance function, thereby making the training of the architecture determination model faster and more accurate, and since the architecture determination model is trained by using the real index data of the architecture sample data corresponding to the computing task, the architecture determination model is more targeted to the computing task, and the target architecture data determined based on the architecture determination model is more accurate. Therefore, the technical problem of being unable to accurately determine the CGRA that best processes the current task is at least partially solved, and the technical effect of determining a more accurate coarse-grained reconfigurable architecture for the current task is achieved.
[0071] According to an embodiment of the present disclosure, based on test indicator data and real indicator data of architecture sample data corresponding to a computing task, adjusting weight parameters in a target neural network to obtain an adjusted architecture determination model may include the following operations.
[0072] Based on the test indicator data and the real indicator data of the architecture sample data corresponding to the computing task, the loss value of the architecture determination model is determined; and the weight parameter is adjusted based on the loss value to obtain the adjusted architecture determination model.
[0073] According to an embodiment of the present disclosure, a loss value can be obtained by inputting test index data and real index data into a loss function. The loss function is not limited and may be a 0-1 loss function, an absolute value loss function, a log loss function, and the like.
[0074] According to an embodiment of the present disclosure, the weight parameters in the target neural network are iteratively adjusted by the loss value until the loss value between the output test index data and the real index data is less than the preset loss value.
[0075] According to an embodiment of the present disclosure, by adjusting only the weight parameters in the target neural network, the training speed and training accuracy of the architecture determination model can be accelerated.
[0076] Figure 3 The flowchart of determining target distribution according to an embodiment of the present disclosure is schematically shown.
[0077] like Figure 3 As shown, the adjusted architecture determination model includes an adjusted Gaussian model; the adjusted Gaussian model includes a mean parameter, and the flowchart for determining the target distribution includes operations S310 to S330.
[0078] In operation S310, a target neural network and a covariance function in the model are determined based on the adjusted architecture, and a target covariance function is determined.
[0079] In operation S320 , a joint posterior probability distribution of the Gaussian model is determined based on the mean parameter and the target covariance function.
[0080] In operation S330, the joint posterior probability distribution is sampled to obtain a target function distribution.
[0081] According to an embodiment of the present disclosure, the target covariance function may be as shown in formula (2).
[0082] According to an embodiment of the present disclosure, the architecture determination model may be a Gaussian model, and the Gaussian model may be characterized by a mean function and a covariance function.
[0083] According to an embodiment of the present disclosure, the mean function generally adopts a constant, which may be 0, etc. The mean function may be shown in the following formula (3).
[0084] m(x)=μ 0 ; (3)
[0085] Among them, μ 0 is the mean parameter.
[0086] According to an embodiment of the present disclosure, the target covariance function can be determined by adjusting the input variables in the covariance function through the weight parameters in the target neural network in the adjusted architecture determination model. For example, if the input variable of the covariance function is x, then the input variable of the target covariance function is wx, where w is the weight parameter.
[0087] According to the embodiments of the present disclosure, the method of sampling the joint posterior probability distribution is not limited, and it can be Monte Carlo sampling, random sampling, etc.
[0088] According to an embodiment of the present disclosure, the target indicator data includes an architecture area value and a task start interval value; based on multiple target indicator data obtained by respectively inputting multiple preset architecture data into the objective function distribution, determining the target architecture data may include the following operations.
[0089] Input multiple preset architecture data into the objective function distribution respectively to obtain multiple target indicator data; for each target indicator data, perform target calculation on the architecture area value and the task start interval value to obtain test evaluation values corresponding to each of the multiple target indicator data; sort the test evaluation values corresponding to each of the multiple target indicator data to obtain sorting results; and determine the target architecture data based on the sorting results.
[0090] According to an embodiment of the present disclosure, the task start interval (Initial Interval, II) can be the time interval between consecutive iterations when a computing task is processed using a preset target coarse-grained reconfigurable architecture corresponding to the preset architecture data. For example, when computing task a is processed using the preset target coarse-grained reconfigurable architecture A, the interval between consecutive iterations is the task start interval.
[0091] According to an embodiment of the present disclosure, the objective function distribution may be one or more. For each objective function distribution, by respectively inputting a plurality of preset architecture data into the objective function distribution, target indicator data corresponding one to one with the plurality of preset architecture data may be obtained.
[0092] According to an embodiment of the present disclosure, the target indicator data may include an architecture area value and a task start interval value. By performing target calculation on the target indicator data, a test evaluation value corresponding to each preset architecture data can be obtained, wherein the target calculation is not limited and may be a multiplication calculation, addition calculation, weighted average, or the like that can unify the data of multiple characterization indicators included in the target indicator data.
[0093] According to an embodiment of the present disclosure, the architecture area value and the task start interval value are subjected to target calculation to obtain test evaluation values corresponding to each of the plurality of target indicator data, so that the indicator values of the preset architecture data can be characterized more conveniently and intuitively.
[0094] According to an embodiment of the present disclosure, by sorting the results, the preset architecture data with the largest test evaluation value is used as the target architecture data.
[0095] According to an embodiment of the present disclosure, in the case where there are multiple objective function distributions, target architecture data corresponding one-to-one to the multiple objective function distributions will be obtained.
[0096] According to an embodiment of the present disclosure, by sampling the joint posterior probability distribution function of the architecture determination model and determining multiple objective function distributions, and determining multiple target architecture data through multiple objective function distributions, a faster convergence speed can be obtained, making the target architecture data more and more accurate. At the same time, multiple target architecture data can be determined in parallel, which speeds up the determination of the target architecture data.
[0097] According to an embodiment of the present disclosure, the method may further include the following operations.
[0098] Acquire initial architecture data, wherein the initial architecture data is data characterizing an initial coarse-grained reconfigurable architecture, the initial coarse-grained reconfigurable architecture includes multiple computing nodes, and the computing nodes include computing units; determine target data characterizing the computing units from the initial architecture data; and adjust the target data corresponding to each of the multiple computing nodes represented in the initial architecture data to obtain multiple preset architecture data.
[0099] According to an embodiment of the present disclosure, the initial architecture data may be preset data or data obtained from a database or a configuration file.
[0100] According to an embodiment of the present disclosure, the initial coarse-grained reconfigurable architecture may be a real architecture structure, composed of hardware devices such as circuits and components, or may be a picture, a topological representation, etc. of the real architecture structure.
[0101] According to an embodiment of the present disclosure, an initial coarse-grained reconfigurable architecture may be determined based on the Open-CGRA open source CGRA framework.
[0102] According to an embodiment of the present disclosure, the initial architecture data may be matrix data or vector data, etc., which can characterize the initial coarse-grained reconfigurable architecture.
[0103] According to an embodiment of the present disclosure, a computing node may be represented by a matrix row, and an element in each row represents a computing unit.
[0104] According to an embodiment of the present disclosure, the elements in each row, i.e., the target data, are adjusted multiple times, and a preset architecture data is output after each adjustment. By performing permutations and combinations until data different from the previous preset architecture data can no longer be output, a plurality of different preset architecture data are obtained, for example: the initial architecture data is A=[1, 0, 0], and the preset architecture data obtained by adjusting the target data in the initial architecture data is B=[0, 1, 0].
[0105] According to an embodiment of the present disclosure, the process of adjusting the target data is equivalent to adjusting the type and / or number of computing units for each computing node. For example, computing node B includes 7 different types of computing units such as Mul, Mac, Add, Shift, Max, Phi, and Br. By reducing the computing units, such as deleting Shift, Max, Phi, and Br, they are represented by 0 in the preset architecture data, and retaining Mul, Mac, and Add, they are represented by 1 in the preset architecture data. When there are multiple computing nodes, by making the above adjustments to each computing node respectively, multiple different preset architecture data can be obtained. Specifically, when there are 9 computing nodes and each computing node may have 7, 7^9 preset architecture data can be obtained through the above adjustment.
[0106] According to the embodiments of the present disclosure, the connection relationship between multiple computing nodes can be fixed first, such as: fixed to a King-Mesh (extra-large mesh) topology structure, and by continuously adjusting the type and / or number of computing units in each computing node, that is, adjusting the target data, different preset target coarse-grained reconfigurable architectures can be obtained, and different preset target coarse-grained reconfigurable architectures can be obtained in an exhaustive manner.
[0107] According to an embodiment of the present disclosure, multiple preset architecture data can realize the simulation of all architectural possibilities of CGRA, so that the target architecture data can be determined from a global perspective by utilizing multiple preset architecture data, that is, the coarse-grained reconfigurable architecture that is most suitable for processing computing tasks can be determined from a global perspective.
[0108] According to an embodiment of the present disclosure, there is no limitation on the method for determining multiple preset architecture data, and the data may be determined in a variety of ways, such as by adjusting the connection relationship between computing units and / or adjusting the type, quantity and / or storage capacity of computing units.
[0109] Figure 4 A schematic diagram of an initial coarse-grained reconfigurable architecture according to an embodiment of the present disclosure is schematically shown.
[0110] like Figure 4 As shown, the initial coarse-grained reconfigurable array may include components such as data storage (Data Memory), computing nodes (Process Element, PE), etc. The computing nodes may include instruction storage units (Config Mem), computing units (Function Unit), crossbar network switches (Crossbar), selection signals (Predicate) and some registers (Regs).
[0111] According to an embodiment of the present disclosure, Data Memory is a data storage component that stores data required for computing tasks.
[0112] According to an embodiment of the present disclosure, the above method may further include the following operations.
[0113] Generate a data stream based on a computing task; determine the architecture data to be simulated from a plurality of preset architecture data; perform architecture modeling on the architecture data to be simulated to obtain the modeling data to be simulated; simulate the architecture data to be simulated based on the modeling data to be simulated to obtain the architecture area value of the architecture data to be simulated; schedule the data stream to the modeling data to be simulated based on a depth-first search scheduling method to obtain the task start interval value of the architecture to be simulated; use the architecture data to be simulated as architecture sample data; obtain the real indicator data of the architecture sample data based on the task start interval value of the architecture data to be simulated and the architecture area value of the architecture data to be simulated.
[0114] According to an embodiment of the present disclosure, a data flow can be obtained by connecting each sub-step in a computing task, for example, computing task A includes sub-steps 1, 2, and 3, and 1-2-3 are connected in a processing order to obtain a data flow. The data flow can be a data flow graph.
[0115] According to an embodiment of the present disclosure, the architecture data to be simulated may be determined from a plurality of preset architecture data by random sampling.
[0116] According to the embodiments of the present disclosure, a programming language can be used to perform architecture modeling on the architecture to be simulated. Specifically, preset architecture information such as the connection relationship between computing nodes, the connection relationship between computing units, etc. can be obtained from a database, and the architecture information is combined with the architecture data to be simulated and described or modeled using a programming language, thereby obtaining simulation modeling data, wherein the programming language is not limited and can be Python, C, etc.
[0117] According to the embodiments of the present disclosure, the method of simulating the architecture data to be simulated based on the modeling data to be simulated is not limited, and the operation process, components, circuit connections, etc. of the preset target coarse-grained reconfigurable architecture can be simulated by a simulation tool.
[0118] According to an embodiment of the present disclosure, specifically, the simulation tool may be an EDA (Electronic Design Automation) tool. More specifically, the format of the modeling data to be simulated may be converted into a language recognizable by the simulation tool, and then the converted modeling data to be simulated may be simulated using the simulation tool to obtain the architecture area value.
[0119] According to an embodiment of the present disclosure, in one implementation, the programming language used is Python, and the modeling data to be simulated can be translated into SystemVerilog, a language that can be recognized by EDA tools, through the PyMTL toolkit, and then the EDA tool Design Compiler (DC) is used to perform simulation under the SMIC 28nm process library to obtain the corresponding area data.
[0120] According to an embodiment of the present disclosure, by adopting a depth-first search (DFS) scheduling method, all scheduling possibilities can be exhaustively enumerated, and all possible scheduling schemes can be traversed in a relatively short time. When scheduling the data stream to the modeling data to be simulated, the task start interval value of each scheduling scheme is recorded to find the minimum task start interval value. When determining the task start interval value of the architecture data to be simulated, the minimum task start interval value and the architecture area value can be used as real indicator data of the architecture sample data to train the architecture determination model.
[0121] According to an embodiment of the present disclosure, during the application process, a scheduling method corresponding to the minimum task startup interval may be used to schedule computing tasks to the target architecture data, thereby saving the overall processing time of the computing tasks.
[0122] According to an embodiment of the present disclosure, the method may further include the following operations.
[0123] Perform architecture modeling on the target architecture data to obtain the target modeling data; simulate the target modeling data to obtain the architecture area value of the target architecture data; schedule the computing task to the target architecture of the target modeling data to obtain the task start interval value of the target architecture data; perform target calculation on the architecture area value of the target architecture data and the task start interval value of the target architecture data to obtain a true evaluation value; determine whether the target architecture data meets the preset conditions based on the true evaluation value and the test evaluation value of the target architecture data.
[0124] According to the embodiment of the present disclosure, the target modeling data can be simulated in the same simulation manner as the simulation architecture.
[0125] According to the embodiments of the present disclosure, there is no limitation on the scheduling method for scheduling computing tasks into the target modeling data target architecture, and the depth-first search scheduling method may still be used, or other scheduling methods may be used.
[0126] According to an embodiment of the present disclosure, the method of determining whether the target architecture data meets the preset conditions based on the real evaluation value and the test evaluation value of the target architecture data is not limited and may be through the following steps.
[0127] Determine the difference between the true evaluation value and the test evaluation value, and determine whether the difference satisfies the first preset condition; if it satisfies the first preset condition, determine the current iteration number, and if the current iteration number is the first, repeat operations S210 to S240; if the current iteration number is the nth, determine whether the difference between the test evaluation value of the target architecture data obtained in the previous N iterations consecutive to the current iteration and the test evaluation value of the target architecture data obtained in the current iteration both belong to the preset range; if both belong to the preset range, determine that the target architecture data satisfies the second preset condition, wherein N and n are constants, N <n。
[0128] According to an embodiment of the present disclosure, determining whether the target architecture data satisfies a preset condition may also be determining whether the target architecture data satisfies a first preset condition, or determining whether the target architecture data satisfies a second preset condition.
[0129] According to an embodiment of the present disclosure, the preset conditions include one or more of the following: the difference between the real evaluation value and the test evaluation value falls within a preset range, and the difference between the test evaluation value and the historical test evaluation value falls within a preset range, wherein the test evaluation value is determined by the current iteration, and the historical test evaluation value is the test evaluation value corresponding to the target architecture data determined in the N iterations preceding the current iteration, and N>1.
[0130] According to an embodiment of the present disclosure, the first preset condition may be that the difference between the real evaluation value and the test evaluation value falls within a preset range, and the second preset condition may be that the difference between the test evaluation value and the historical test evaluation value falls within a preset range.
[0131] According to an embodiment of the present disclosure, the above method may further include the following operations.
[0132] When it is determined that the target architecture data does not meet the preset conditions, the weight parameters of the target neural network in the adjusted architecture determination model are readjusted based on the test index data of the target architecture data, the architecture area value of the target architecture data and the task start interval value of the target architecture data to obtain the readjusted architecture determination model.
[0133] According to an embodiment of the present disclosure, when it is determined that the target architecture data does not satisfy a preset condition, the architecture determination model can be readjusted using the target architecture data. Specifically, a loss value can be calculated using the test index data of the target architecture data and the real index data of the target architecture data, i.e., the architecture area value of the target architecture data and the task start interval value of the target architecture data. The weight parameters of the target neural network in the adjusted architecture determination model can then be readjusted using the loss value to obtain the readjusted architecture determination model.
[0134] According to an embodiment of the present disclosure, when the target architecture data does not meet the preset conditions, the adjusted architecture determination model can be adjusted, that is, retrained, through the target architecture data and the real indicator data of the target architecture data, thereby realizing the process of continuously testing and training the architecture determination model, so that a coarse-grained reconfigurable architecture suitable for processing computing tasks can be finally obtained.
[0135] Figure 5 The flowchart of the architecture determination method according to another embodiment of the present disclosure is schematically shown.
[0136] like Figure 5 As shown, the method includes operations S501 to S510.
[0137] In operation S501 , in response to receiving a computing task, a plurality of preset architecture data is determined.
[0138] In operation S502 , architecture data to be simulated is determined from a plurality of preset architecture data.
[0139] In operation S503 , the architecture data to be simulated is used as architecture sample data, and real indicator data of the architecture data to be simulated is determined.
[0140] In operation S504, the architecture sample data is input into the architecture determination model to obtain test indicator data of the architecture sample data.
[0141] In operation S505, based on the test indicator data and the real indicator data of the architecture sample data, the weight parameters in the target neural network are adjusted to obtain an adjusted architecture determination model.
[0142] In operation S506 , the target function distribution is determined based on the adjusted architecture determination model.
[0143] In operation S507 , target architecture data is determined based on a plurality of target indicator data obtained by respectively inputting a plurality of preset architecture data into the target function distribution.
[0144] In operation S508, it is determined whether the target architecture data satisfies a preset condition. If the target architecture data satisfies the preset condition, operation S509 is performed. If the target architecture data does not satisfy the preset condition, operation S510 is performed.
[0145] In operation S509 , a target coarse-grained reconfigurable architecture for processing the computing task is determined based on the target architecture data.
[0146] In operation S510, based on the test index data of the target architecture data, the architecture area value of the target architecture data, and the task start interval value of the target architecture data, the weight parameters of the target neural network in the adjusted architecture determination model are readjusted to obtain the readjusted architecture determination model.
[0147] Based on the above architecture determination method, the present disclosure also provides an architecture determination device. Figure 6 The device is described in detail.
[0148] Figure 6 The structural block diagram of the architecture determination device according to an embodiment of the present disclosure is schematically shown.
[0149] like Figure 6 As shown, the architecture determination device 600 of this embodiment includes a data input module 610 , a model adjustment module 620 , a distribution determination module 630 , a data determination module 640 and an architecture determination module 650 .
[0150] The data input module 610 is used to input the architecture sample data into the architecture determination model in response to receiving the computing task, so as to obtain the test index data of the architecture sample data, wherein the architecture determination model includes a covariance function and a target neural network, the output of the target neural network is the input of the covariance function, and the architecture sample data is the data for characterizing the coarse-grained reconfigurable architecture.
[0151] The model adjustment module 620 is used to adjust the weight parameters in the target neural network based on the test indicator data and the real indicator data of the architecture sample data corresponding to the computing task to obtain an adjusted architecture determination model.
[0152] The distribution determination module 630 is used to determine the target function distribution based on the adjusted architecture determination model.
[0153] The data determination module 640 is used to determine the target architecture data based on multiple target indicator data obtained by inputting multiple preset architecture data into the target function distribution respectively.
[0154] The architecture determination module 650 is used to determine a target coarse-grained reconfigurable architecture for processing a computing task based on the target architecture data when it is determined that the target architecture data meets a preset condition.
[0155] According to an embodiment of the present disclosure, the adjusted architecture determination model includes an adjusted Gaussian model. The adjusted Gaussian model includes a mean parameter; the distribution determination module 630 may include: a function determination submodule, a joint distribution determination submodule, and a sampling submodule.
[0156] The function determination submodule is used to determine the target neural network and covariance function in the model based on the adjusted architecture, and determine the target covariance function.
[0157] The joint distribution determination submodule is used to determine the joint posterior probability distribution of the Gaussian model based on the mean parameter and the target covariance function.
[0158] The sampling submodule is used to sample the joint posterior probability distribution to obtain the target function distribution.
[0159] According to an embodiment of the present disclosure, the target indicator data includes an architecture area value and a task start interval value; the data determination module 640 may include: an indicator data determination submodule, a test evaluation value determination submodule, a sorting submodule and an architecture data determination submodule.
[0160] The indicator data determination submodule is used to input multiple preset architecture data into the objective function distribution respectively to obtain multiple target indicator data.
[0161] The test evaluation value determination submodule is used to perform target calculation on the architecture area value and the task start interval value for each target indicator data to obtain the test evaluation values corresponding to each of the multiple target indicator data.
[0162] The sorting submodule is used to sort the test evaluation values corresponding to the multiple target indicator data to obtain the sorting results.
[0163] The architecture data determination submodule is used to determine the target architecture data based on the sorting results.
[0164] According to an embodiment of the present disclosure, the architecture determination device 600 further includes: a computing unit acquisition module, a preset architecture determination module, and a preset architecture data determination module.
[0165] The acquisition module is used to acquire initial architecture data, wherein the initial architecture data is data characterizing an initial coarse-grained reconfigurable architecture, and the initial coarse-grained reconfigurable architecture includes a plurality of computing nodes, and the computing nodes include computing units.
[0166] The target data determination module is used to determine the target data representing the computing unit from the initial architecture data.
[0167] The preset architecture determination module is used to adjust the target data corresponding to each of the multiple computing nodes represented in the initial architecture data to obtain multiple preset architecture data.
[0168] According to an embodiment of the present disclosure, the architecture determination device 600 also includes: a data flow generation module, a simulation data determination module, a first modeling module, a first simulation module, a first scheduling module, a determination module and a real indicator determination module.
[0169] The data stream generation module is used to generate data streams based on computing tasks.
[0170] The simulation data determination module is used to determine the architecture data to be simulated from a plurality of preset architecture data.
[0171] The first modeling module is used to perform architecture modeling on the architecture data to be simulated to obtain the modeling data to be simulated.
[0172] The first simulation module is used to simulate the architecture data to be simulated based on the modeling data to be simulated, and obtain the architecture area value of the architecture data to be simulated.
[0173] The first scheduling module is used to schedule the data stream to the modeling data to be simulated based on the depth-first search scheduling method, and obtain the task start interval value of the architecture to be simulated.
[0174] A determination module is used to use the architecture data to be simulated as architecture sample data;
[0175] The real indicator determination module is used to obtain the real indicator data of the architecture sample data based on the task start interval value of the architecture data to be simulated and the architecture area value of the architecture data to be simulated.
[0176] According to an embodiment of the present disclosure, the architecture determination device 600 further includes: a second modeling module, a second simulation module, a second scheduling module, a true value determination module, and a condition determination module.
[0177] The second modeling module is used to perform architecture modeling on the target architecture data to obtain target modeling data.
[0178] The second simulation module is used to simulate the target modeling data to obtain the architecture area value of the target architecture data.
[0179] The second scheduling module is used to schedule the computing tasks to the target modeling data and obtain the task start interval value of the target architecture data.
[0180] The real value determination module is used to perform target calculation on the architecture area value of the target architecture data and the task start interval value of the target architecture data to obtain a real evaluation value.
[0181] The condition determination module is used to determine whether the target architecture data meets the preset condition based on the real evaluation value and the test evaluation value of the target architecture data. According to an embodiment of the present disclosure, the architecture determination device 600 further includes: a readjustment module.
[0182] The readjustment module is used to readjust the weight parameters of the target neural network in the adjusted architecture determination model based on the test index data of the target architecture data, the architecture area value of the target architecture data and the task start interval value of the target architecture data when it is determined that the target architecture data does not meet the preset conditions, so as to obtain the readjusted architecture determination model.
[0183] According to an embodiment of the present disclosure, the model adjustment module 620 may include: a loss value determination module and a parameter adjustment module.
[0184] The loss value determination module is used to determine the loss value of the architecture determination model based on the test indicator data and the real indicator data of the architecture sample data corresponding to the computing task.
[0185] The parameter adjustment module is used to adjust the weight parameters based on the loss value to obtain an adjusted architecture determination model.
[0186] According to an embodiment of the present disclosure, any multiple modules of the data input module 610, the model adjustment module 620, the distribution determination module 630, the data determination module 640, and the architecture determination module 650 can be combined in one module for implementation, or any one of the modules can be split into multiple modules. Alternatively, at least part of the functions of one or more of these modules can be combined with at least part of the functions of other modules and implemented in one module. According to an embodiment of the present disclosure, at least one of the data input module 610, the model adjustment module 620, the distribution determination module 630, the data determination module 640, and the architecture determination module 650 can be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on a chip, a system on a substrate, a system on a package, an application specific integrated circuit (ASIC), or can be implemented by hardware or firmware such as any other reasonable way of integrating or packaging the circuit, or implemented in any one of the three implementation methods of software, hardware, and firmware, or in any appropriate combination of any of them. Alternatively, at least one of the data input module 610, the model adjustment module 620, the distribution determination module 630, the data determination module 640 and the architecture determination module 650 may be at least partially implemented as a computer program module, which may perform a corresponding function when executed.
[0187] Figure 7 A block diagram of an electronic device suitable for implementing the architecture determination method according to an embodiment of the present disclosure is schematically shown.
[0188] like Figure 7 As shown, the electronic device 700 according to an embodiment of the present disclosure includes a processor 701, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 702 or a program loaded from a storage part 708 into a random access memory (RAM) 703. The processor 701 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or a related chipset and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 701 may also include an onboard memory for caching purposes. The processor 701 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present disclosure.
[0189] In RAM 703, various programs and data required for the operation of electronic device 700 are stored. Processor 701, ROM 702 and RAM 703 are connected to each other via bus 704. Processor 701 performs various operations of the method flow according to the embodiment of the present disclosure by executing the program in ROM 702 and / or RAM 703. It should be noted that the program can also be stored in one or more memories other than ROM 702 and RAM 703. Processor 701 can also perform various operations of the method flow according to the embodiment of the present disclosure by executing the program stored in one or more memories.
[0190] According to an embodiment of the present disclosure, the electronic device 700 may further include an input / output (I / O) interface 705, which is also connected to the bus 704. The electronic device 700 may further include one or more of the following components connected to the input / output (I / O) interface 705: an input portion 706 including a keyboard, a mouse, etc.; an output portion 707 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage portion 708 including a hard disk, etc.; and a communication portion 709 including a network interface card such as a LAN card, a modem, etc. The communication portion 709 performs communication processing via a network such as the Internet. A drive 710 is also connected to the input / output (I / O) interface 705 as needed. A removable medium 711, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 710 as needed, so that a computer program read therefrom is installed into the storage portion 708 as needed.
[0191] The present disclosure also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or may exist independently without being assembled into the device / apparatus / system. The above computer-readable storage medium carries one or more programs, and when the above one or more programs are executed, the method according to the embodiment of the present disclosure is implemented.
[0192] According to an embodiment of the present disclosure, the computer-readable storage medium may be a non-volatile computer-readable storage medium, for example, may include but is not limited to: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program, which may be used by or in combination with an instruction execution system, an apparatus or a device. For example, according to an embodiment of the present disclosure, a computer-readable storage medium may include the ROM 702 and / or RAM 703 described above and / or one or more memories other than ROM 702 and RAM 703.
[0193] The embodiment of the present disclosure also includes a computer program product, which includes a computer program, and the computer program contains program code for executing the method shown in the flowchart. When the computer program product is run in a computer system, the program code is used to enable the computer system to implement the architecture determination method provided by the embodiment of the present disclosure.
[0194] The above functions defined in the system / device of the embodiment of the present disclosure are performed when the computer program is executed by the processor 701. According to the embodiment of the present disclosure, the system, device, module, unit, etc. described above can be implemented by a computer program module.
[0195] In one embodiment, the computer program may rely on tangible storage media such as optical storage devices, magnetic storage devices, etc. In another embodiment, the computer program may also be transmitted and distributed in the form of signals on a network medium, and downloaded and installed through the communication part 709, and / or installed from the removable medium 711. The program code contained in the computer program may be transmitted using any appropriate network medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.
[0196] In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 709, and / or installed from the removable medium 711. When the computer program is executed by the processor 701, the above functions defined in the system of the embodiment of the present disclosure are performed. According to the embodiment of the present disclosure, the system, device, means, module, unit, etc. described above can be implemented by a computer program module.
[0197] According to an embodiment of the present disclosure, the program code for executing the computer program provided by the embodiment of the present disclosure can be written in any combination of one or more programming languages. Specifically, these computing programs can be implemented using high-level process and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include, but are not limited to, Java, C++, python, "C" language or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, partially on the remote computing device, or entirely on the remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (for example, using an Internet service provider to connect through the Internet).
[0198] The flow charts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flow chart or block diagram can represent a module, a program segment, or a part of a code, and the above-mentioned module, program segment, or a part of a code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flow chart, and the combination of the boxes in the block diagram or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0199] It will be appreciated by those skilled in the art that the features described in the various embodiments and / or claims of the present disclosure may be combined and / or combined in a variety of ways, even if such combinations and / or combinations are not explicitly described in the present disclosure. In particular, the features described in the various embodiments and / or claims of the present disclosure may be combined and / or combined in a variety of ways without departing from the spirit and teachings of the present disclosure. All of these combinations and / or combinations fall within the scope of the present disclosure.
[0200] The embodiments of the present disclosure are described above. However, these embodiments are only for illustrative purposes and are not intended to limit the scope of the present disclosure. Although the embodiments are described above separately, this does not mean that the measures in the various embodiments cannot be used in combination to advantage. The scope of the present disclosure is defined by the attached claims and their equivalents. Without departing from the scope of the present disclosure, those skilled in the art may make a variety of substitutions and modifications, which should all fall within the scope of the present disclosure.
Claims
1. A method for determining an architecture, comprising: In response to having received a computing task, inputting architecture sample data into an architecture determination model to obtain test index data of the architecture sample data, wherein the architecture determination model includes a covariance function and a target neural network, the output of the target neural network is the input of the covariance function, the architecture sample data is data for characterizing a coarse-grained reconfigurable architecture, the coarse-grained reconfigurable architecture includes computing nodes, the computing nodes include computing units, the architecture sample data is used to characterize the type and quantity of computing units included in the computing nodes in the coarse-grained reconfigurable architecture, the architecture sample data is determined from a plurality of preset architecture data, the plurality of preset architecture data respectively correspond to different coarse-grained reconfigurable architectures, and different coarse-grained reconfigurable architectures include computing nodes with different types and quantities of computing units; Based on the test indicator data and the real indicator data of the architecture sample data corresponding to the computing task, adjusting the weight parameters in the target neural network to obtain an adjusted architecture determination model, wherein the test indicator data includes at least one of an architecture area value, a task start interval, and power consumption of the coarse-grained reconfigurable architecture; Determine the model based on the adjusted architecture and determine the target function distribution; Determining target architecture data from the plurality of preset architecture data based on a plurality of target indicator data obtained by respectively inputting the plurality of preset architecture data into the target function distribution; In a case where it is determined that the target architecture data satisfies a preset condition, a target coarse-grained reconfigurable architecture for processing the computing task is determined based on the target architecture data.
2. The method according to claim 1, wherein: The adjusted architecture determination model includes an adjusted Gaussian model; the adjusted Gaussian model includes a mean parameter; The step of determining the model based on the adjusted architecture and determining the target function distribution includes: Determine a target neural network and the covariance function in the model based on the adjusted architecture, and determine a target covariance function; Determining a joint posterior probability distribution of the Gaussian model based on the mean parameter and the target covariance function; The joint posterior probability distribution is sampled to obtain a target function distribution.
3. The method according to claim 1, wherein: The target indicator data includes an architecture area value and a task start interval value; The step of determining the target architecture data based on the multiple target indicator data obtained by respectively inputting the multiple preset architecture data into the target function distribution includes: Inputting a plurality of preset architecture data into the objective function distribution respectively to obtain a plurality of objective indicator data; For each of the target indicator data, target calculation is performed on the architecture area value and the task start interval value to obtain test evaluation values corresponding to each of the multiple target indicator data; Sorting the test evaluation values corresponding to the plurality of target indicator data to obtain a sorting result; Based on the sorting result, target architecture data is determined.
4. The method according to claim 3, further comprising: Acquire initial architecture data, wherein the initial architecture data is data characterizing an initial coarse-grained reconfigurable architecture, the initial coarse-grained reconfigurable architecture includes a plurality of computing nodes, and the computing nodes include computing units; determining target data characterizing the computing unit from the initial architecture data; The target data corresponding to each of the plurality of computing nodes represented in the initial architecture data are adjusted respectively to obtain a plurality of preset architecture data.
5. The method according to claim 4, further comprising: generating a data stream based on the computing task; Determine architecture data to be simulated from a plurality of preset architecture data; Performing architecture modeling on the architecture data to be simulated to obtain modeling data to be simulated; Simulating the architecture data to be simulated based on the modeling data to be simulated to obtain an architecture area value of the architecture data to be simulated; Based on a scheduling method of depth-first search, the data stream is scheduled to the modeling data to be simulated, and a task start interval value of the architecture to be simulated is obtained; Using the architecture data to be simulated as the architecture sample data; Based on the task start interval value of the architecture data to be simulated and the architecture area value of the architecture data to be simulated, real indicator data of the architecture sample data is obtained.
6. The method according to claim 3, further comprising: Performing architecture modeling on the target architecture data to obtain target modeling data; Simulating the target modeling data to obtain an architecture area value of the target architecture data; Scheduling the computing task to the target modeling data to obtain a task start interval value of the target architecture data; Perform the target calculation on the architecture area value of the target architecture data and the task start interval value of the target architecture data to obtain a true evaluation value; Based on the real evaluation value and the test evaluation value of the target architecture data, it is determined whether the target architecture data meets the preset condition.
7. The method according to claim 6, further comprising: When it is determined that the target architecture data does not meet the preset conditions, the weight parameters of the target neural network in the adjusted architecture determination model are readjusted based on the test index data of the target architecture data, the architecture area value of the target architecture data and the task start interval value of the target architecture data to obtain the readjusted architecture determination model.
8. The method according to claim 6, wherein: The preset conditions include one or more of the following: the difference between the actual evaluation value and the test evaluation value falls within a preset range, and the difference between the test evaluation value and the historical test evaluation value falls within the preset range, wherein the test evaluation value is determined by the current iteration, and the historical test evaluation value is the test evaluation value corresponding to the target architecture data determined in the N consecutive iterations before the current iteration, and N>1.
9. The method according to claim 1, wherein: The step of adjusting the weight parameters in the target neural network based on the test indicator data and the real indicator data of the architecture sample data corresponding to the computing task to obtain an adjusted architecture determination model includes: Determining a loss value of the architecture determination model based on the test indicator data and real indicator data of the architecture sample data corresponding to the computing task; The weight parameter is adjusted based on the loss value to obtain an adjusted architecture determination model.
10. A device for determining an architecture, comprising: A data input module, for inputting architecture sample data into an architecture determination model in response to having received a computing task, and obtaining test index data of the architecture sample data, wherein the architecture determination model includes a covariance function and a target neural network, the output of the target neural network is the input of the covariance function, the architecture sample data is data for characterizing a coarse-grained reconfigurable architecture, the coarse-grained reconfigurable architecture includes computing nodes, the computing nodes include computing units, the architecture sample data is used to characterize the type and quantity of computing units included in the computing nodes in the coarse-grained reconfigurable architecture, the architecture sample data is determined from a plurality of preset architecture data, the plurality of preset architecture data respectively correspond to different coarse-grained reconfigurable architectures, and different coarse-grained reconfigurable architectures include computing nodes with different types and quantities of computing units; a model adjustment module, configured to adjust weight parameters in the target neural network based on the test indicator data and real indicator data of the architecture sample data corresponding to the computing task, so as to obtain an adjusted architecture determination model, wherein the test indicator data includes at least one of an architecture area value, a task start interval, and a power consumption of the coarse-grained reconfigurable architecture; A distribution determination module, for determining a target function distribution based on the adjusted architecture determination model; A data determination module, configured to determine target architecture data from a plurality of preset architecture data based on a plurality of target indicator data obtained by respectively inputting a plurality of preset architecture data into the target function distribution; The architecture determination module is used to determine a target coarse-grained reconfigurable architecture for processing the computing task based on the target architecture data when it is determined that the target architecture data meets a preset condition.
Citation Information
Patent Citations
Reconfigurable computing architecture mapping method and device based on graph convolution and reinforcement learning
CN115757264A
Systems, apparatus, methods, and architectures for a neural network workflow to generate a hardware acceletator
US20200225996A1