Convolutional neural network weight reuse and power consumption collaborative optimization method and system
By transforming the convolution operation to the frequency domain and combining it with sparse Fourier transform, factorization, and machine learning, the weight reuse of convolutional neural networks is optimized, solving the problems of high computational complexity and insufficient adaptability, and achieving efficient and flexible power optimization.
Patent Information
- Application Number
- CN202510915550.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-03
- Publication Date
- 2025-11-18
AI Technical Summary
Existing methods for reusing weights in convolutional neural networks have high computational complexity, lack adaptability, and are difficult to implement efficient and flexible power optimization on resource-constrained devices.
The convolution operation is transformed from the spatial domain to the frequency domain. Fast Fourier Transform and Sparse Fourier Transform are used to reduce the amount of computation. Factorization and activation group reuse strategies are combined with machine learning algorithms to dynamically adjust the weight reuse mode. A hardware-friendly encoding scheme is designed and cross-layer weight sharing is explored and integrated into the hardware accelerator system.
It significantly reduces computational and storage requirements, improves computational efficiency and adaptability, reduces power consumption, optimizes network performance, and is suitable for resource-constrained devices.
Smart Images

Figure CN120975140A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of deep learning technology, and in particular to a method and system for co-optimizing weight reuse and power consumption in convolutional neural networks. Background Technology
[0002] Convolutional Neural Networks (CNNs) have achieved remarkable results in fields such as image recognition and natural language processing. Their weight reuse mechanism is one of the key technologies for improving computational efficiency and model performance. In recent years, with the rapid development of deep learning, weight reuse technology has continued to evolve, from simple weight sharing to complex dynamic adjustment strategies, aiming to reduce computational and storage requirements while maintaining high model accuracy.
[0003] While existing technologies have made some progress in weight reuse, several shortcomings remain. For example, traditional weight reuse methods primarily focus on the spatial domain, resulting in high computational complexity and difficulty in fully utilizing hardware resources. Furthermore, existing methods lack adaptability to input data features and network structure when dynamically adjusting weight reuse patterns, leading to limited optimization performance across different scenarios. These issues restrict the application of convolutional neural networks on resource-constrained devices, particularly in power consumption optimization, where existing technologies cannot achieve efficient and flexible weight reuse. Summary of the Invention
[0004] In view of the aforementioned existing problems, the present invention is proposed.
[0005] Therefore, this invention provides a method for co-optimizing weight reuse and power consumption in convolutional neural networks to address the problems of high computational complexity, lack of adaptability, and the need for efficient and flexible weight reuse in terms of power consumption optimization in existing convolutional neural network weight reuse methods.
[0006] To solve the above-mentioned technical problems, the present invention provides the following technical solution:
[0007] In a first aspect, the present invention provides a method for co-optimizing weight reuse and power consumption in convolutional neural networks, characterized by comprising the following steps:
[0008] The convolution operation in the convolutional neural network is transformed from the spatial domain to the frequency domain. The Fast Fourier Transform (FFT) is used to transform the input data and weights, converting the convolution operation into a dot product operation, thereby reducing the amount of computation.
[0009] In the frequency domain, the dot product in the convolution operation is factored to extract common weights to reduce the number of multiplication operations. Combined with the activation group reuse strategy, the weight repetition across filters is further utilized to optimize the calculation process.
[0010] Based on the results of dot product factorization and activation group reuse, machine learning algorithms are used to dynamically identify and adjust the weight reuse pattern. The optimal weight reuse strategy is adaptively selected according to the weight distribution of the current layer and the characteristics of the input data.
[0011] We design a hardware-friendly weight reuse encoding scheme to encode the weight repetition pattern into a form that is easy to implement in hardware. At the same time, we optimize the storage structure of weight and activation data and adopt sparse storage technology to skip the storage and calculation of zero-value weights or activation values.
[0012] Explore weight reuse strategies across different convolutional layers. By sharing weights across multiple convolutional layers, reduce the storage and computational overhead of weights, and combine them with a global power consumption optimization algorithm to optimize the power consumption performance of the entire network.
[0013] The above optimization strategies are integrated into the hardware accelerator system of convolutional neural networks to build a complete system for weight reuse and power consumption co-optimization, and a comprehensive performance evaluation is performed.
[0014] As a preferred embodiment of the convolutional neural network weight reuse and power consumption co-optimization method described in this invention, when transforming the input data and weights using Fast Fourier Transform (FFT), a sparse Fast Fourier Transform (sFFT) algorithm is employed, transforming only non-zero values to further reduce computational load. The specific steps are as follows:
[0015] The input data and weights are sparsified, and the positions of non-zero values are marked.
[0016] The sparse fast Fourier transform (sFFT) algorithm is applied, which only transforms non-zero values and skips the calculation of zero values;
[0017] The transformed non-zero values are stored in the frequency domain to prepare for subsequent dot product operations.
[0018] The frequency domain transformation of the convolution operation is completed by performing a dot product operation on the non-zero values in the frequency domain.
[0019] As a preferred embodiment of the convolutional neural network weight reuse and power consumption co-optimization method described in this invention, the following steps are employed: When factoring the dot product in the convolution operation, a block-matrix-based factorization method is used to divide the weight matrix into multiple small blocks, and factorization is performed on each small block separately to improve computational efficiency.
[0020] The weight matrix is divided into multiple small blocks, each containing a certain number of weights;
[0021] Factorize each small block, extract common weights, and reduce the number of multiplication operations;
[0022] By combining the activation group reuse strategy and utilizing the weight repetition across filters, the computation process can be further optimized.
[0023] The optimized calculation results are stored to provide input for subsequent dynamic weight reuse pattern recognition.
[0024] As a preferred embodiment of the convolutional neural network weight reuse and power consumption co-optimization method described in this invention, the step of dynamically identifying and adjusting the weight reuse pattern using machine learning algorithms employs a reinforcement learning algorithm. A reward mechanism incentivizes the system to select the optimal weight reuse strategy to adapt to different input data and network structures. The specific steps are as follows:
[0025] Initialize the reinforcement learning environment and define the state space as the weight distribution of the current layer and the features of the input data;
[0026] Define the action space as different weight reuse strategies, and set the reward function to optimize power consumption.
[0027] Through reinforcement learning algorithms, the optimal action is selected based on the current state, i.e., the weighted reuse strategy.
[0028] Based on the selected weight reuse strategy, the weight usage pattern is adjusted to provide an optimized pattern for subsequent hardware-friendly coding.
[0029] As a preferred embodiment of the convolutional neural network weight reuse and power consumption co-optimization method described in this invention, the hardware-friendly weight reuse encoding scheme is designed by employing a weight compression method based on Huffman coding to encode the weight repetition pattern into a shorter binary string, thereby reducing storage requirements. The specific steps are as follows:
[0030] Calculate the frequency of occurrence of each pattern in the statistical weighted repetition pattern.
[0031] Based on the frequency, Huffman coding is used to assign a unique binary string to each repeating pattern;
[0032] The weight repetition pattern is replaced with the corresponding binary string to achieve compressed storage of the weight;
[0033] Optimize the storage structure of weights and activation data by using sparse storage technology to skip the storage and calculation of zero-value weights or activation values.
[0034] As a preferred embodiment of the convolutional neural network weight reuse and power consumption co-optimization method described in this invention, the following steps are employed when exploring weight reuse strategies across different convolutional layers: a graph-based weight sharing method is used to model the weight sharing problem between convolutional layers as a shortest path problem in a graph. The optimal weight sharing scheme is determined by solving for the shortest path.
[0035] Construct a weight sharing graph between convolutional layers, where nodes represent convolutional layers and edges represent weight sharing relationships;
[0036] The weight sharing problem is modeled as the shortest path problem in a graph, and the path weight is defined as the power consumption optimization effect of the shared weight.
[0037] Using graph theory algorithms, such as Dijkstra's algorithm, we can find the shortest path and determine the optimal weight-sharing scheme.
[0038] Based on the shortest path results, adjust the weight sharing relationship between convolutional layers to reduce the storage and computational overhead of weights.
[0039] As a preferred embodiment of the convolutional neural network weight reuse and power consumption co-optimization method described in this invention, when integrating the optimization strategy into the hardware accelerator system of the convolutional neural network, a field-programmable gate array (FPGA) is used as the hardware accelerator. The reconfigurability of the FPGA enables flexible adjustment and efficient execution of the optimization strategy. The specific steps are as follows:
[0040] Select a suitable FPGA model and design the hardware accelerator architecture based on the optimization strategy;
[0041] The logic of the optimization strategy is mapped onto the FPGA to implement the functional modules of the hardware accelerator;
[0042] The reconfigurability of FPGAs allows for flexible adjustment of hardware accelerator configurations to adapt to different optimization strategies.
[0043] A comprehensive performance evaluation of the integrated system was conducted to verify the effectiveness of the optimization strategy and the reliability of the system.
[0044] Secondly, the present invention provides a system for co-optimizing weight reuse and power consumption in convolutional neural networks, including a frequency domain conversion module: converting the convolution operation in the convolutional neural network from the spatial domain to the frequency domain, using Fast Fourier Transform (FFT) to transform the input data and weights, converting the convolution operation into a dot product operation, thereby reducing the amount of computation;
[0045] Dot product factorization module: In the frequency domain, the dot product in the convolution operation is factored to extract common weights to reduce the number of multiplication operations, and combined with the activation group reuse strategy, the weight repetition across filters is further utilized to optimize the calculation process.
[0046] Dynamic weight reuse module: Based on the results of dot product factorization and activation group reuse, it uses machine learning algorithms to dynamically identify and adjust the weight reuse pattern, and adaptively selects the optimal weight reuse strategy according to the weight distribution of the current layer and the characteristics of the input data.
[0047] Hardware-friendly encoding module: Design a hardware-friendly weight reuse encoding scheme to encode the weight repetition pattern into a form that is easy to implement in hardware. At the same time, optimize the storage structure of weight and activation data, and adopt sparse storage technology to skip the storage and calculation of zero-value weights or activation values.
[0048] Cross-layer weight reuse module: Explores weight reuse strategies across different convolutional layers, reduces weight storage and computation overhead by sharing weights across multiple convolutional layers, and optimizes the power consumption performance of the entire network by combining global power consumption optimization algorithms;
[0049] System Integration and Evaluation Module: Integrates the above optimization strategies into the hardware accelerator system of convolutional neural networks, constructs a complete weight reuse and power consumption co-optimization system, and performs a comprehensive performance evaluation.
[0050] The beneficial effects of this invention are as follows: By transforming the convolution operation in the convolutional neural network from the spatial domain to the frequency domain, and using the Fast Fourier Transform (FFT) to transform the input data and weights, the computational load is significantly reduced. In the frequency domain, the dot product in the convolution operation is further factored to extract common weights to reduce the number of multiplication operations. Combined with the activation group reuse strategy, the weight repetition across filters is further utilized to optimize the computation process. The weight reuse pattern is dynamically identified and adjusted through machine learning algorithms. Based on the weight distribution of the current layer and the characteristics of the input data, the optimal weight reuse strategy is adaptively selected, improving the system's flexibility and adaptability. Furthermore, a hardware-friendly weight reuse encoding scheme is designed to encode the weight repetition pattern into a form easily implemented in hardware. Simultaneously, the storage structure of weights and activation data is optimized, employing sparse storage technology to skip the storage and computation of zero-value weights or activation values, further reducing storage requirements and power consumption. The invention explores weight reuse strategies across different convolutional layers. By sharing weights among multiple convolutional layers, the storage and computational overhead of weights is reduced. Combined with a global power optimization algorithm, the power performance of the entire network is optimized. Finally, the above optimization strategies were integrated into the hardware accelerator system of the convolutional neural network to build a complete system for weight reuse and power consumption co-optimization. A comprehensive performance evaluation was conducted to verify the effectiveness of the optimization strategies and the reliability of the system. Attached Figure Description
[0051] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0052] Figure 1 This is a flowchart of the method for co-optimizing weight reuse and power consumption of convolutional neural networks in Example 1. Detailed Implementation
[0053] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0054] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.
[0055] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a single or selective embodiment that is mutually exclusive with other embodiments.
[0056] Reference Figure 1 This is the first embodiment of the present invention, which provides a method for co-optimizing weight reuse and power consumption in convolutional neural networks, characterized by including the following steps:
[0057] The convolution operation in the convolutional neural network is transformed from the spatial domain to the frequency domain. The Fast Fourier Transform (FFT) is used to transform the input data and weights, converting the convolution operation into a dot product operation, thereby reducing the amount of computation.
[0058] In the frequency domain, the dot product in the convolution operation is factored to extract common weights to reduce the number of multiplication operations. Combined with the activation group reuse strategy, the weight repetition across filters is further utilized to optimize the calculation process.
[0059] Based on the results of dot product factorization and activation group reuse, machine learning algorithms are used to dynamically identify and adjust the weight reuse pattern. The optimal weight reuse strategy is adaptively selected according to the weight distribution of the current layer and the characteristics of the input data.
[0060] We design a hardware-friendly weight reuse encoding scheme to encode the weight repetition pattern into a form that is easy to implement in hardware. At the same time, we optimize the storage structure of weight and activation data and adopt sparse storage technology to skip the storage and calculation of zero-value weights or activation values.
[0061] Explore weight reuse strategies across different convolutional layers. By sharing weights across multiple convolutional layers, reduce the storage and computational overhead of weights, and combine them with a global power consumption optimization algorithm to optimize the power consumption performance of the entire network.
[0062] The above optimization strategies are integrated into the hardware accelerator system of convolutional neural networks to build a complete system for weight reuse and power consumption co-optimization, and a comprehensive performance evaluation is performed.
[0063] It should be noted that step 1 involves converting the convolution operation in the convolutional neural network from the spatial domain to the frequency domain. The input data and weights are transformed using the Fast Fourier Transform (FFT), converting the convolution operation into a dot product operation, thereby reducing the amount of computation.
[0064] First, a Fast Fourier Transform (FFT) is performed on the input data and weights in the convolutional neural network, mapping them from the spatial domain to the frequency domain. In the frequency domain, the convolution operation can be simplified to a dot product operation, which greatly reduces computational complexity. Specifically, traditional convolution operations require a large number of multiplication and addition operations, while dot product operations only require simple multiplication operations, significantly reducing the computational load.
[0065] By transforming the convolution operation to the frequency domain and leveraging the efficiency of the FFT, the computational load is significantly reduced, and computational efficiency is improved. This makes convolutional neural networks more efficient in processing large-scale data, reduces the demand for hardware resources, and also provides a foundation for subsequent optimization steps.
[0066] Step 2: In the frequency domain, factorize the dot product in the convolution operation, extract common weights to reduce the number of multiplication operations, and combine the activation group reuse strategy to further optimize the calculation process by utilizing the weight repetition across filters.
[0067] In the frequency domain, the dot product in the convolution operation is factored to extract common weights, thereby reducing the number of multiplication operations. Simultaneously, combined with an activation group reuse strategy, the weight repetition across filters is further utilized to optimize the entire computation process. Specifically, factorization breaks down the complex dot product operation into simpler sub-operations, reducing unnecessary computation.
[0068] By employing factorization and activation group reuse strategies, the computational load is further reduced, and computational efficiency is improved. This not only reduces hardware resource consumption but also improves the system's response speed, making convolutional neural networks more efficient in practical applications.
[0069] Step 3: Based on the results of dot product factorization and activation group reuse, use machine learning algorithms to dynamically identify and adjust the weight reuse pattern, and adaptively select the optimal weight reuse strategy according to the weight distribution of the current layer and the characteristics of the input data.
[0070] Machine learning algorithms, particularly reinforcement learning algorithms, are used to dynamically identify and adjust weight reuse patterns. The algorithm adaptively selects the optimal weight reuse strategy based on the weight distribution of the current layer and the characteristics of the input data. Specifically, by defining a state space and an action space and setting a reward function, the algorithm can select the optimal action, i.e., the weight reuse strategy, based on the current state.
[0071] By dynamically identifying and adjusting weight reuse patterns, the system can adaptively select the optimal weight reuse strategy, improving its flexibility and adaptability. This allows the convolutional neural network to maintain high computational performance in different scenarios and further optimizes power consumption.
[0072] Step 4: Design a hardware-friendly weight reuse encoding scheme to encode the weight repetition pattern into a form that is easy to implement in hardware. At the same time, optimize the storage structure of weight and activation data, and use sparse storage technology to skip the storage and calculation of zero-value weights or activation values.
[0073] Design a hardware-friendly weight repetition encoding scheme to encode the repetition patterns of weights into a form easily implemented in hardware. Simultaneously, optimize the storage structure of weight and activation data by employing sparse storage techniques to skip the storage and computation of zero-value weights or activation values. Specifically, calculate the frequency of each pattern by statistically analyzing the repetition patterns of weights, and assign a unique binary string to each repetition pattern using Huffman coding.
[0074] By employing a hardware-friendly weighted reusable encoding scheme and sparse storage technology, storage requirements and computational overhead are reduced. This not only improves the utilization of hardware resources but also reduces system power consumption, enabling convolutional neural networks to run efficiently on resource-constrained devices.
[0075] Step 5: Explore weight reuse strategies across different convolutional layers. By sharing weights across multiple convolutional layers, the storage and computational overhead of weights can be reduced. Combined with a global power optimization algorithm, the power performance of the entire network can be optimized.
[0076] This study explores weight reuse strategies across different convolutional layers, reducing the storage and computational overhead of weights by sharing them across multiple convolutional layers. Simultaneously, a global power optimization algorithm is incorporated to improve the overall network's power performance. Specifically, by constructing a weight sharing graph between convolutional layers, the weight sharing problem is modeled as a shortest path problem in a graph, and graph theory algorithms are used to find the optimal weight sharing scheme.
[0077] By employing a cross-layer weight reuse strategy and a global power consumption optimization algorithm, the storage and computational overhead of weights are further reduced, optimizing the overall power consumption performance of the network. This allows the convolutional neural network to significantly reduce power consumption while maintaining high computational efficiency, thereby improving the system's energy efficiency ratio.
[0078] Step 6: Integrate the above optimization strategies into the hardware accelerator system of the convolutional neural network to build a complete weight reuse and power consumption co-optimization system, and conduct a comprehensive performance evaluation.
[0079] The aforementioned optimization strategies are integrated into a hardware accelerator system for convolutional neural networks to construct a complete system for co-optimizing weight reuse and power consumption. Specifically, a suitable FPGA model is selected, the architecture of the hardware accelerator is designed according to the optimization strategies, the logic of the optimization strategies is mapped onto the FPGA, and the functional modules of the hardware accelerator are implemented.
[0080] By integrating optimization strategies into the hardware accelerator system, a complete system for co-optimizing weight reuse and power consumption was constructed, achieving efficient computation and low-power operation. Comprehensive performance evaluation verified the effectiveness of the optimization strategies and the reliability of the system, making convolutional neural networks more efficient and energy-saving in practical applications.
[0081] Specifically, when using Fast Fourier Transform (FFT) to transform the input data and weights, the Sparse Fast Fourier Transform (sFFT) algorithm is employed, transforming only non-zero values to further reduce computation. The specific steps are as follows:
[0082] The input data and weights are sparsified, and the positions of non-zero values are marked.
[0083] The sparse fast Fourier transform (sFFT) algorithm is applied, which only transforms non-zero values and skips the calculation of zero values;
[0084] The transformed non-zero values are stored in the frequency domain to prepare for subsequent dot product operations.
[0085] The frequency domain transformation of the convolution operation is completed by performing a dot product operation on the non-zero values in the frequency domain.
[0086] It should be noted that when using Fast Fourier Transform (FFT) to transform the input data and weights, a Sparse Fast Fourier Transform (sFFT) algorithm is employed, transforming only non-zero values to further reduce computation. The specific steps are as follows:
[0087] The input data and weights are sparsified, and the positions of non-zero values are marked:
[0088] First, the input data and weights in the convolutional neural network are sparsified. This process involves analyzing the input data and weight matrix, identifying non-zero values, and recording their positions. In this way, the input data and weights can be represented in sparse matrix form, preparing for the subsequent Sparse Fast Fourier Transform (sFFT) algorithm.
[0089] The sparse fast Fourier transform (sFFT) algorithm is applied, transforming only non-zero values and skipping the calculation of zero values:
[0090] Next, the sparse fast Fourier transform (sFFT) algorithm is applied to transform the sparsified input data and weights. The sFFT algorithm utilizes the sparsity of the input data and weights, performing Fourier transforms only on non-zero values and skipping the calculation of zero values. This method significantly reduces the computational cost of the Fourier transform and improves computational efficiency.
[0091] The transformed non-zero values are stored in the frequency domain to prepare for subsequent dot product operations.
[0092] The non-zero values obtained after the sFFT transformation are stored in the frequency domain. These results will serve as input for subsequent dot product operations, providing data support for the execution of convolution operations in the frequency domain.
[0093] Perform a dot product operation on the non-zero values in the frequency domain to complete the frequency domain transformation of the convolution operation:
[0094] Finally, a dot product operation is performed on the transformed non-zero values in the frequency domain to complete the frequency domain transformation of the convolution operation. In this way, the convolution operation is efficiently converted into a dot product operation, further reducing computational complexity.
[0095] Through the above steps, highly efficient optimizations of convolution operations in convolutional neural networks are achieved. First, by employing sparsification and the sparse fast Fourier transform (sFFT) algorithm, the computational cost of Fourier transforms is significantly reduced, improving computational efficiency. Second, by transforming the convolution operation from the spatial domain to the frequency domain and replacing traditional convolution calculations with dot product operations, computational complexity is further reduced. These optimizations not only improve the operational efficiency of convolutional neural networks but also reduce the demand for hardware resources, enabling convolutional neural networks to run more efficiently on resource-constrained devices.
[0096] Specifically, when factoring the dot product in the convolution operation, a block-matrix-based factorization method is used. The weight matrix is divided into multiple small blocks, and factorization is performed on each small block separately to improve computational efficiency. The specific steps are as follows:
[0097] The weight matrix is divided into multiple small blocks, each containing a certain number of weights;
[0098] Factorize each small block, extract common weights, and reduce the number of multiplication operations;
[0099] By combining the activation group reuse strategy and utilizing the weight repetition across filters, the computation process can be further optimized.
[0100] The optimized calculation results are stored to provide input for subsequent dynamic weight reuse pattern recognition.
[0101] It should be noted that when factoring the dot product in the convolution operation, a block-matrix-based factorization method is used. The weight matrix is divided into multiple small blocks, and factorization is performed on each block separately to improve computational efficiency. The specific steps are as follows:
[0102] Weight matrix partitioning: The weight matrix is divided into multiple small blocks according to certain rules, and each small block contains a certain number of weights. This partitioning method can be flexibly adjusted according to the size and shape of the weight matrix to adapt to different convolutional neural network structures.
[0103] Factorization: Factorize each small block and extract common weights. By extracting common weights, the number of multiplication operations can be reduced, thereby reducing computational complexity.
[0104] Activation group reuse strategy: By combining activation group reuse strategy with weight repetition across filters, the computation process is further optimized. This method can share computation results among multiple filters, reducing redundant computation.
[0105] Storing optimization results: The optimized calculation results are stored to provide input for subsequent dynamic weight reuse pattern recognition. This step ensures that the optimization results can be effectively utilized in subsequent steps, forming a complete optimization process.
[0106] By dividing the weight matrix into multiple small blocks and factoring them separately, efficient computation of the dot product in convolution operations is achieved, significantly reducing the number of multiplication operations. Combined with an activation group reuse strategy, the computation process is further optimized by utilizing weight repetition across filters. This block-based factorization method not only improves computational efficiency but also reduces storage requirements, as the extraction of common weights reduces the number of weights that need to be stored. Ultimately, this method provides optimized input for dynamic weight reuse pattern recognition, making the entire weight reuse and power optimization process of the convolutional neural network more efficient and flexible.
[0107] Specifically, when dynamically identifying and adjusting weight reuse patterns using machine learning algorithms, a reinforcement learning algorithm is employed. A reward mechanism incentivizes the system to select the optimal weight reuse strategy to adapt to different input data and network structures. The specific steps are as follows:
[0108] Initialize the reinforcement learning environment and define the state space as the weight distribution of the current layer and the features of the input data;
[0109] Define the action space as different weight reuse strategies, and set the reward function to optimize power consumption.
[0110] Through reinforcement learning algorithms, the optimal action is selected based on the current state, i.e., the weighted reuse strategy.
[0111] Based on the selected weight reuse strategy, the weight usage pattern is adjusted to provide an optimized pattern for subsequent hardware-friendly coding.
[0112] It should be noted that initializing the reinforcement learning environment involves defining the state space as the weight distribution of the current layer and the characteristics of the input data. This state space definition enables the system to perceive the current network state, providing a basis for subsequent weight reuse strategy selection.
[0113] Define the action space and reward function: The action space consists of different weighted reusable strategies, and the reward function is set to the power consumption optimization effect. By defining the action space and reward function, the system can clearly define the optimization objective and evaluate different strategies based on the power consumption optimization effect.
[0114] Choosing the optimal weight reuse strategy: Using reinforcement learning algorithms, the optimal action is selected based on the current state; this is known as the weight reuse strategy. Through continuous trial and error and learning, reinforcement learning algorithms can adaptively select the most suitable weight reuse strategy for the current state.
[0115] Adjusting weight usage pattern: Based on the selected weight reuse strategy, the weight usage pattern is adjusted to provide an optimized pattern for subsequent hardware-friendly coding. This adjustment process ensures more efficient use of weights and reduces redundant calculations.
[0116] Adaptability: Through reinforcement learning algorithms, the system can adaptively select the optimal weight reuse strategy based on the weight distribution of the current layer and the characteristics of the input data. This adaptability enables the system to automatically adjust the weight reuse pattern under different input data and network structures, thereby improving the system's flexibility and adaptability.
[0117] Power consumption optimization: By setting the reward function to reflect the power consumption optimization effect, the system can clearly define the optimization target and evaluate and select different strategies based on the power consumption optimization effect. This allows the system to significantly reduce power consumption while maintaining computational efficiency.
[0118] Optimize weight usage pattern: Based on the selected weight reuse strategy, adjust the weight usage pattern to provide an optimized pattern for subsequent hardware-friendly coding. This adjustment process reduces redundant calculations, improves weight utilization efficiency, and further optimizes the overall system performance.
[0119] Improve system efficiency: By dynamically adjusting the weight reuse strategy, the system can better utilize hardware resources, reduce unnecessary computing and storage overhead, and thus improve the overall efficiency of the system.
[0120] Specifically, when designing the hardware-friendly weight reuse encoding scheme, a weight compression method based on Huffman coding is adopted to encode the weight repetition pattern into a shorter binary string to reduce storage requirements. The specific steps are as follows:
[0121] Calculate the frequency of occurrence of each pattern in the statistical weighted repetition pattern.
[0122] Based on the frequency, Huffman coding is used to assign a unique binary string to each repeating pattern;
[0123] The weight repetition pattern is replaced with the corresponding binary string to achieve compressed storage of the weight;
[0124] Optimize the storage structure of weights and activation data by using sparse storage technology to skip the storage and calculation of zero-value weights or activation values.
[0125] It should be noted that the statistical weighting of repetitive patterns calculates the frequency of occurrence of each pattern.
[0126] First, the weights in the convolutional neural network are analyzed to identify recurring patterns in the weight matrix. By traversing the weight matrix and recording the frequency of each weight value, the frequency of each recurring pattern can be calculated. This process can be achieved by constructing a frequency table, where the key is the weight value and the value is the number of times that weight value appears. In this way, a comprehensive understanding of the distribution of recurring patterns in the weight matrix can be obtained.
[0127] By analyzing the recurrence patterns and frequencies of weights, a data foundation is provided for subsequent weight compression. This step enables the system to identify which weight values occur frequently, thus providing crucial information for subsequent Huffman coding and contributing to more efficient weight compression.
[0128] Based on frequency, Huffman coding is used to assign a unique binary string to each repeating pattern.
[0129] Based on the weight frequencies obtained in step 1, a Huffman tree is constructed. A Huffman tree is a binary tree structure where the weight of each node is equal to the sum of the weights of its child nodes. Using the Huffman tree, a unique binary string is assigned to each weight value, ensuring that higher-frequency weight values are assigned shorter codes, while lower-frequency weight values are assigned longer codes. This process is implemented using the Huffman coding algorithm, ensuring both efficiency and uniqueness in the encoding.
[0130] Using Huffman coding to assign binary strings to weights can significantly reduce the storage requirements for weights. By assigning short codes to high-frequency weights and long codes to low-frequency weights, storage space can be effectively reduced while maintaining the integrity of weight information.
[0131] The weights are compressed and stored by replacing the repeating pattern of the weights with the corresponding binary string.
[0132] After Huffman coding is completed, each weight value in the weight matrix is replaced with its corresponding binary string. This process involves traversing the weight matrix, finding the corresponding code for each weight value in the Huffman coding table, and replacing the original weight value with that code. In this way, the weight matrix is transformed into a more compact binary string matrix, thus achieving compressed storage of the weights.
[0133] By replacing the repeating pattern of the weights with binary strings, compressed storage of the weights is achieved. This step significantly reduces the storage space required for the weights, lowers the hardware resource requirements, and improves the system's storage efficiency.
[0134] Optimize the storage structure of weights and activation data by employing sparse storage techniques to skip the storage and calculation of zero-value weights or activation values.
[0135] After weight compression, the storage structure of weights and activation data is further optimized. Sparse storage technology is employed, storing only non-zero values and their index information, skipping the storage and calculation of zero-value weights or activation values. This process is achieved by constructing a sparse matrix, which contains only non-zero values and their corresponding row and column indices. This approach further reduces storage requirements while improving computational efficiency.
[0136] By employing sparse storage technology, the storage and computation of zero-value weights or activation values are skipped, further reducing storage requirements and computational complexity. This step not only improves the system's storage efficiency but also reduces computational resource consumption, enabling the system to have better adaptability and performance on resource-constrained devices.
[0137] Specifically, when exploring weight reuse strategies across different convolutional layers, a graph-based weight sharing method is adopted. This model the weight sharing problem between convolutional layers as a shortest path problem in a graph, and the optimal weight sharing scheme is determined by solving for the shortest path. The specific steps are as follows:
[0138] Construct a weight sharing graph between convolutional layers, where nodes represent convolutional layers and edges represent weight sharing relationships;
[0139] The weight sharing problem is modeled as the shortest path problem in a graph, and the path weight is defined as the power consumption optimization effect of the shared weight.
[0140] Using graph theory algorithms, such as Dijkstra's algorithm, we can find the shortest path and determine the optimal weight-sharing scheme.
[0141] Based on the shortest path results, adjust the weight sharing relationship between convolutional layers to reduce the storage and computational overhead of weights.
[0142] It should be noted that a weight sharing graph is constructed between convolutional layers.
[0143] First, each convolutional layer in the convolutional neural network is considered as a node in a graph. Next, the weight-sharing relationships between the convolutional layers are analyzed, representing these relationships as edges in the graph. Specifically, if there is a possibility of weight sharing between two convolutional layers, an edge is established between these two nodes, and this edge is assigned a weight value that reflects the potential effect of the shared weights on power consumption optimization.
[0144] By constructing a weight sharing graph, we achieved a visual and quantitative representation of the weight sharing relationships between convolutional layers. This provides a clear operational object and basic data structure for subsequent optimization algorithms, enabling the complex weight sharing problem to be systematically analyzed and solved in the form of graph theory.
[0145] The weight sharing problem is modeled as a shortest path problem in a graph.
[0146] In the constructed weight-sharing graph, path weights are defined as the power consumption optimization effect of shared weights. Specifically, path weights can be calculated comprehensively based on factors such as the number of shared weights, the degree of sharing, and the reduction in computational complexity. Then, the problem of finding the optimal weight-sharing scheme is transformed into the problem of finding the shortest path in the graph, that is, finding a path such that the sum of the weights of all edges on the path is minimized.
[0147] Transforming the weight sharing problem into a graph shortest path problem makes the solution clearer and more standardized. Utilizing well-established shortest path algorithms in graph theory, the optimal weight sharing scheme can be found efficiently, thus theoretically ensuring the maximization of power consumption optimization.
[0148] Solving shortest paths using graph theory algorithms
[0149] Classic graph theory algorithms, such as Dijkstra's algorithm, are used to solve the modeled shortest path problem. Dijkstra's algorithm calculates the shortest path from the starting node to all other nodes in the graph by progressively expanding the nodes. In each step, the algorithm selects the node with the currently known shortest path to expand and updates the path weights of its neighboring nodes. Ultimately, the algorithm finds the shortest path from the starting node to the target node, i.e., the optimal weight-sharing scheme.
[0150] By utilizing mature graph theory algorithms such as Dijkstra's algorithm, the optimal weight-sharing scheme can be solved efficiently and accurately. This not only ensures the reliability and accuracy of the solution process but also enables the rapid finding of the global optimum, thereby significantly improving the power consumption optimization effect of convolutional neural networks.
[0151] Adjust weight sharing relationships based on shortest path results
[0152] Based on the shortest path results obtained by Dijkstra's algorithm, the weight sharing relationship between convolutional layers is adjusted. Specifically, according to the weight sharing scheme indicated by the shortest path, the weight sharing relationship between each convolutional layer is redistributed, making the weight sharing pattern of the entire network more optimized. This process involves adjusting the weight storage and access patterns, as well as optimizing the specific implementation of convolution operations.
[0153] Adjusting weight sharing based on the shortest path result effectively reduces the storage and computational overhead of weights. By optimizing the weight sharing pattern, convolutional neural networks can significantly reduce power consumption and improve operating efficiency while maintaining high accuracy. This not only enhances the applicability of the network on resource-constrained devices but also provides strong support for the widespread application of deep learning models.
[0154] Specifically, when integrating the optimization strategy into the hardware accelerator system of the convolutional neural network, a field-programmable gate array (FPGA) is used as the hardware accelerator. Leveraging the reconfigurability of the FPGA, flexible adjustment and efficient execution of the optimization strategy are achieved. The specific steps are as follows:
[0155] Select a suitable FPGA model and design the hardware accelerator architecture based on the optimization strategy;
[0156] The logic of the optimization strategy is mapped onto the FPGA to implement the functional modules of the hardware accelerator;
[0157] The reconfigurability of FPGAs allows for flexible adjustment of hardware accelerator configurations to adapt to different optimization strategies.
[0158] A comprehensive performance evaluation of the integrated system was conducted to verify the effectiveness of the optimization strategy and the reliability of the system.
[0159] It should be noted that selecting a suitable FPGA model and designing the hardware accelerator architecture based on optimization strategies are crucial.
[0160] When integrating optimization strategies into a hardware accelerator system for convolutional neural networks, the first step is to select a suitable Field-Programmable Gate Array (FPGA) model. The selection of an FPGA must comprehensively consider factors such as its resource capacity, processing speed, power consumption, and cost to ensure it meets the hardware requirements of the convolutional neural network's weight reuse and power consumption co-optimization strategy. After selecting the FPGA model, the architecture of the hardware accelerator is designed based on the specific requirements of the optimization strategy. The architecture design must fully consider how to efficiently implement functional modules such as frequency domain transformation of convolution operations, dot product factorization, dynamic weight reuse mode adjustment, and cross-layer weight reuse on the FPGA, while also taking into account the rational allocation and utilization of hardware resources to achieve efficient execution of the optimization strategy.
[0161] By carefully selecting FPGA models and designing a hardware accelerator architecture, hardware support for convolutional neural network optimization strategies was achieved, laying the foundation for subsequent optimization strategy mapping and execution. This enables the optimization strategies to be implemented efficiently at the hardware level, thereby contributing to improved system performance and reduced power consumption.
[0162] The optimization strategy logic is mapped onto the FPGA to implement the functional modules of the hardware accelerator.
[0163] After completing the hardware accelerator architecture design, the logic of the optimization strategy is mapped in detail onto the selected FPGA. This process involves translating the specific algorithmic logic of optimization strategies, such as frequency domain transformation of convolution operations, dot product factorization, dynamic weight reuse mode adjustment, and cross-layer weight reuse, into FPGA-executable Hardware Description Language (HDL) code. Through precise logic mapping, it is ensured that each functional module can run accurately and efficiently on the FPGA, realizing the hardware acceleration function of the optimization strategy.
[0164] By successfully mapping the optimization strategy logic onto the FPGA, the functional modularity of the hardware accelerator was achieved, enabling the optimization strategy to be executed efficiently at the hardware level. This not only improves the computational efficiency of the convolutional neural network but also provides a flexible operational basis for subsequent hardware reconfigurability adjustments, further enhancing the system's adaptability and scalability.
[0165] The reconfigurability of FPGAs allows for flexible adjustments to the configuration of hardware accelerators to adapt to different optimization strategies.
[0166] Leveraging the reconfigurable nature of FPGAs, the configuration of hardware accelerators can be flexibly adjusted based on different convolutional neural network models, input data characteristics, and optimization objectives. This adjustment includes, but is not limited to, changing the connection methods of functional modules, adjusting resource allocation, and optimizing data paths, to ensure that the hardware accelerator can dynamically adapt to changes in various optimization strategies. By monitoring network operating status and power consumption in real time, rapid adjustments can be made using the reconfigurability of FPGAs, thereby achieving optimal utilization of hardware resources.
[0167] Leveraging the reconfigurability of FPGAs, hardware accelerators can flexibly adapt to different optimization strategies, enhancing the system's flexibility and adaptability. This flexibility enables the hardware accelerators to maintain high efficiency under different network structures and input data conditions, further improving the overall system performance and power consumption optimization.
[0168] A comprehensive performance evaluation of the integrated system was conducted to verify the effectiveness of the optimization strategy and the reliability of the system.
[0169] After configuring the hardware accelerator and integrating optimization strategies, a comprehensive performance evaluation of the entire system was conducted. The evaluation included, but was not limited to, key metrics such as improved computational efficiency, reduced power consumption, and maintained accuracy. The actual effectiveness of the optimization strategies was verified through comparison with traditional convolutional neural network implementations. Simultaneously, rigorous testing of the system's stability, reliability, and compatibility was performed to ensure stable and reliable operation under various operating conditions.
[0170] Comprehensive performance evaluation verified the effectiveness of the optimization strategy and the reliability of the system. This not only ensured the feasibility and advantages of the optimization strategy in practical applications but also provided a scientific basis for further optimization and improvement of the system. Performance evaluation results show that the optimization strategy significantly improved the computational efficiency and power consumption performance of the convolutional neural network while maintaining high accuracy and stability.
[0171] This embodiment also provides a system for co-optimizing weight reuse and power consumption in convolutional neural networks, including:
[0172] Frequency domain transformation module: Transforms the convolution operation in the convolutional neural network from the spatial domain to the frequency domain. It uses Fast Fourier Transform (FFT) to transform the input data and weights, converting the convolution operation into a dot product operation, thereby reducing the amount of computation.
[0173] Dot product factorization module: In the frequency domain, the dot product in the convolution operation is factored to extract common weights to reduce the number of multiplication operations, and combined with the activation group reuse strategy, the weight repetition across filters is further utilized to optimize the calculation process.
[0174] Dynamic weight reuse module: Based on the results of dot product factorization and activation group reuse, it uses machine learning algorithms to dynamically identify and adjust the weight reuse pattern, and adaptively selects the optimal weight reuse strategy according to the weight distribution of the current layer and the characteristics of the input data.
[0175] Hardware-friendly encoding module: Design a hardware-friendly weight reuse encoding scheme to encode the weight repetition pattern into a form that is easy to implement in hardware. At the same time, optimize the storage structure of weight and activation data, and adopt sparse storage technology to skip the storage and calculation of zero-value weights or activation values.
[0176] Cross-layer weight reuse module: Explores weight reuse strategies across different convolutional layers, reduces weight storage and computation overhead by sharing weights across multiple convolutional layers, and optimizes the power consumption performance of the entire network by combining global power consumption optimization algorithms;
[0177] System Integration and Evaluation Module: Integrates the above optimization strategies into the hardware accelerator system of convolutional neural networks, constructs a complete weight reuse and power consumption co-optimization system, and performs a comprehensive performance evaluation.
[0178] In summary, this invention significantly reduces computation by transforming convolution operations in convolutional neural networks from the spatial domain to the frequency domain and using Fast Fourier Transform (FFT) to transform the input data and weights. In the frequency domain, the dot product in the convolution operation is further factorized to extract common weights, reducing the number of multiplication operations. Combined with an activation group reuse strategy, the computation process is further optimized by utilizing weight repetition across filters. Machine learning algorithms dynamically identify and adjust weight reuse patterns, adaptively selecting the optimal weight reuse strategy based on the weight distribution of the current layer and the characteristics of the input data, thus improving the system's flexibility and adaptability. Furthermore, a hardware-friendly weight reuse encoding scheme is designed, encoding the weight repetition pattern into a form easily implemented in hardware. The storage structure of weights and activation data is optimized, employing sparse storage technology to skip the storage and computation of zero-value weights or activation values, further reducing storage requirements and power consumption. The invention explores weight reuse strategies across different convolutional layers, reducing weight storage and computation overhead by sharing weights across multiple convolutional layers, and optimizes the overall network's power performance by combining global power optimization algorithms. Finally, the above optimization strategies were integrated into the hardware accelerator system of the convolutional neural network to build a complete system for weight reuse and power consumption co-optimization. A comprehensive performance evaluation was conducted to verify the effectiveness of the optimization strategies and the reliability of the system.
[0179] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A method for co-optimizing weight reuse and power consumption in convolutional neural networks, characterized in that, Includes the following steps: The convolution operation in the convolutional neural network is transformed from the spatial domain to the frequency domain. The Fast Fourier Transform (FFT) is used to transform the input data and weights, converting the convolution operation into a dot product operation, thereby reducing the amount of computation. In the frequency domain, the dot product in the convolution operation is factored to extract common weights to reduce the number of multiplication operations. Combined with the activation group reuse strategy, the weight repetition across filters is further utilized to optimize the calculation process. Based on the results of dot product factorization and activation group reuse, machine learning algorithms are used to dynamically identify and adjust the weight reuse pattern. The optimal weight reuse strategy is adaptively selected according to the weight distribution of the current layer and the characteristics of the input data. We design a hardware-friendly weight reuse encoding scheme to encode the weight repetition pattern into a form that is easy to implement in hardware. At the same time, we optimize the storage structure of weight and activation data and adopt sparse storage technology to skip the storage and calculation of zero-value weights or activation values. Explore weight reuse strategies across different convolutional layers. By sharing weights across multiple convolutional layers, reduce the storage and computational overhead of weights, and combine them with a global power consumption optimization algorithm to optimize the power consumption performance of the entire network. The above optimization strategies are integrated into the hardware accelerator system of convolutional neural networks to build a complete system for weight reuse and power consumption co-optimization, and a comprehensive performance evaluation is performed.
2. The method for co-optimizing weight reuse and power consumption in convolutional neural networks as described in claim 1, characterized in that: When using Fast Fourier Transform (FFT) to transform the input data and weights, a Sparse Fast Fourier Transform (sFFT) algorithm is employed, which only transforms non-zero values to further reduce computational complexity. The specific steps are as follows: The input data and weights are sparsified, and the positions of non-zero values are marked. The sparse fast Fourier transform (sFFT) algorithm is applied, which only transforms non-zero values and skips the calculation of zero values; The transformed non-zero values are stored in the frequency domain to prepare for subsequent dot product operations. The frequency domain transformation of the convolution operation is completed by performing a dot product operation on the non-zero values in the frequency domain.
3. The method for co-optimizing weight reuse and power consumption in convolutional neural networks as described in claim 2, characterized in that: When factoring the dot product in the convolution operation, a block-matrix-based factorization method is used. The weight matrix is divided into multiple small blocks, and factorization is performed on each block separately to improve computational efficiency. The specific steps are as follows: The weight matrix is divided into multiple small blocks, each containing a certain number of weights; Factorize each small block, extract common weights, and reduce the number of multiplication operations; By combining the activation group reuse strategy and utilizing the weight repetition across filters, the computation process can be further optimized. The optimized calculation results are stored to provide input for subsequent dynamic weight reuse pattern recognition.
4. The method for co-optimizing weight reuse and power consumption in convolutional neural networks as described in claim 3, characterized in that: When dynamically identifying and adjusting weight reuse patterns using machine learning algorithms, a reinforcement learning algorithm is employed. A reward mechanism incentivizes the system to select the optimal weight reuse strategy to adapt to different input data and network structures. The specific steps are as follows: Initialize the reinforcement learning environment and define the state space as the weight distribution of the current layer and the features of the input data; Define the action space as different weight reuse strategies, and set the reward function to optimize power consumption. Through reinforcement learning algorithms, the optimal action is selected based on the current state, i.e., the weighted reuse strategy. Based on the selected weight reuse strategy, the weight usage pattern is adjusted to provide an optimized pattern for subsequent hardware-friendly coding.
5. The method for co-optimizing weight reuse and power consumption in convolutional neural networks as described in claim 4, characterized in that: When designing the hardware-friendly weight reuse encoding scheme, a weight compression method based on Huffman coding is adopted to encode the weight repetition pattern into a shorter binary string to reduce storage requirements. The specific steps are as follows: Calculate the frequency of occurrence of each pattern in the statistical weighted repetition pattern. Based on the frequency, Huffman coding is used to assign a unique binary string to each repeating pattern; The weight repetition pattern is replaced with the corresponding binary string to achieve compressed storage of the weight; Optimize the storage structure of weights and activation data by using sparse storage technology to skip the storage and calculation of zero-value weights or activation values.
6. The method for co-optimizing weight reuse and power consumption in convolutional neural networks as described in claim 5, characterized in that: When exploring weight reuse strategies across different convolutional layers, a graph-based weight sharing method is adopted. This model the weight sharing problem between convolutional layers as a shortest path problem in a graph. The optimal weight sharing scheme is determined by solving for the shortest path. The specific steps are as follows: Construct a weight sharing graph between convolutional layers, where nodes represent convolutional layers and edges represent weight sharing relationships; The weight sharing problem is modeled as the shortest path problem in a graph, and the path weight is defined as the power consumption optimization effect of the shared weight. Using graph theory algorithms, such as Dijkstra's algorithm, we can find the shortest path and determine the optimal weight-sharing scheme. Based on the shortest path results, adjust the weight sharing relationship between convolutional layers to reduce the storage and computational overhead of weights.
7. The method for co-optimizing weight reuse and power consumption in convolutional neural networks as described in claim 6, characterized in that: When integrating the optimization strategy into the hardware accelerator system of the convolutional neural network, a field-programmable gate array (FPGA) is used as the hardware accelerator. Leveraging the reconfigurability of the FPGA, flexible adjustment and efficient execution of the optimization strategy are achieved. The specific steps are as follows: Select a suitable FPGA model and design the hardware accelerator architecture based on the optimization strategy; The logic of the optimization strategy is mapped onto the FPGA to implement the functional modules of the hardware accelerator; The reconfigurability of FPGAs allows for flexible adjustment of hardware accelerator configurations to adapt to different optimization strategies. A comprehensive performance evaluation of the integrated system was conducted to verify the effectiveness of the optimization strategy and the reliability of the system.
8. A system for co-optimizing weight reuse and power consumption in convolutional neural networks, based on the method for co-optimizing weight reuse and power consumption in convolutional neural networks according to any one of claims 1 to 7, characterized in that: include, Frequency domain transformation module: Transforms the convolution operation in the convolutional neural network from the spatial domain to the frequency domain. It uses Fast Fourier Transform (FFT) to transform the input data and weights, converting the convolution operation into a dot product operation, thereby reducing the amount of computation. Dot product factorization module: In the frequency domain, the dot product in the convolution operation is factored to extract common weights to reduce the number of multiplication operations, and combined with the activation group reuse strategy, the weight repetition across filters is further utilized to optimize the calculation process. Dynamic weight reuse module: Based on the results of dot product factorization and activation group reuse, it uses machine learning algorithms to dynamically identify and adjust the weight reuse pattern, and adaptively selects the optimal weight reuse strategy according to the weight distribution of the current layer and the characteristics of the input data. Hardware-friendly encoding module: Design a hardware-friendly weight reuse encoding scheme to encode the weight repetition pattern into a form that is easy to implement in hardware. At the same time, optimize the storage structure of weight and activation data, and adopt sparse storage technology to skip the storage and calculation of zero-value weights or activation values. Cross-layer weight reuse module: Explores weight reuse strategies across different convolutional layers, reduces weight storage and computation overhead by sharing weights across multiple convolutional layers, and optimizes the power consumption performance of the entire network by combining global power consumption optimization algorithms; System Integration and Evaluation Module: Integrates the above optimization strategies into the hardware accelerator system of convolutional neural networks, constructs a complete weight reuse and power consumption co-optimization system, and performs a comprehensive performance evaluation.