A method for optimizing parallel operation control of multiple HASH algorithms in ASIC chip
By creating integrated modules in the ASIC chip and analyzing pipelines for SHA-256 processing units, combined with particle swarm optimization algorithm, the problem of insufficient utilization of processing unit resources in traditional solutions is solved, and more efficient data processing and load balancing is achieved.
Patent Information
- Application Number
- CN202411281797.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-13
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2044-09-13
AI Technical Summary
When traditional solutions deal with SHA-256 algorithms, it is difficult to dynamically adjust the resource utilization of the processing unit, resulting in insufficient resources in some stages and overload in other stages, affecting overall performance.
By creating an integration module, analyzing the pipelines of the SHA-256 processing unit, determining the mode selection matrix, scheduling different processing units and pipeline stages, adding control signals, and optimizing the task allocation scheme through the particle swarm optimization algorithm.
It improves the utilization rate of hardware resources, realizes load balancing, avoids overload or idle processing units, and improves overall processing efficiency and system real-timeness.
Smart Images

Figure CN119149202B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and in particular to a method for optimizing parallel operation control of multiple HASH algorithms in an ASIC chip. Background Art
[0002] Traditional solutions usually execute the SHA-256 algorithm based on a single processing unit, which may not be able to cope with large-scale data. Although the performance of a single processing unit may be high, the computing power of a single processing unit often becomes a bottleneck in scenarios that require processing massive amounts of data or real-time response.
[0003] In addition, traditional solutions may lack sufficient flexibility when dealing with the pipeline characteristics of the SHA-256 algorithm. Since the SHA-256 algorithm contains multiple processing stages, the computational load of each stage may be different. Traditional solutions may find it difficult to dynamically adjust according to the actual load of each stage, resulting in insufficient resource utilization of the processing unit in some stages and overload in other stages, thus affecting overall performance. Summary of the invention
[0004] The technical problem to be solved by the present invention is to provide a method for optimizing the parallel operation control of multiple HASH algorithms in an ASIC chip, which can improve the efficiency and performance of data processing.
[0005] In order to solve the above technical problems, the technical solution of the present invention is as follows:
[0006] In a first aspect, a method for optimizing parallel operation control of multiple HASH algorithms in an ASIC chip is provided, the method comprising:
[0007] Creating an integration module, the integration module is used to accommodate multiple SHA-256 processing units;
[0008] Analyze the pipeline of the SHA-256 processing unit to determine the number of pipeline stages;
[0009] According to the number of pipeline stages, a mode selection matrix is determined;
[0010] According to the mode selection matrix, different SHA-256 processing units and pipeline stages are scheduled to obtain an adjusted SHA-256 processing unit;
[0011] A control signal is added to each adjusted SHA-256 processing unit, where the control signal is used to control the start and stop of the SHA-256 processing unit operation;
[0012] Optimize the task allocation scheme between the adjusted SHA-256 processing units, using each particle to represent a specific task allocation scheme;
[0013] Define a fitness function for evaluating the performance of each task allocation scheme;
[0014] According to the fitness function, the fitness of each particle is calculated, the speed and position of the particle are updated, and the operation is iterated until the preset number of iterations is reached to obtain the final task allocation plan.
[0015] Furthermore, the pipeline of the SHA-256 processing unit is analyzed to determine the number of pipeline stages, including:
[0016] Get the circuit diagram of the SHA-256 processing unit;
[0017] Based on the circuit diagram, identify the different pipeline stages in the SHA-256 processing unit;
[0018] According to different pipeline stages in the SHA-256 processing unit, the number of pipeline stages is determined, wherein the number of pipeline stages corresponds to different stages identified in the hardware implementation, each stage is regarded as one stage of the pipeline, and the number of pipeline stages is determined by analyzing the input data and output data of each stage, as well as the dependency between the input data and the output data.
[0019] Further, based on the circuit diagram, the different pipeline stages in the SHA-256 processing unit are identified, including:
[0020] Based on the circuit diagram, identify logic units and modules, including registers, arithmetic logic units, and control units;
[0021] According to the circuit diagram and data path, the SHA-256 processing unit is divided into different functional modules, including a message filling module, an expansion module, and a compression function module;
[0022] Refine each functional module and identify the various stages that make up the pipeline;
[0023] Mark the boundaries of each pipeline stage on the circuit diagram.
[0024] Furthermore, each functional module is refined to identify the various stages that make up the pipeline, including:
[0025] According to the length of the original message in the message padding module, the message is padded according to the padding rule of SHA-256 so that its length meets specific conditions, and a 64-bit integer is appended to the end of the padded message to indicate the length of the original message, so as to obtain the processed message;
[0026] Split the processed message into fixed-size blocks, each of which is 512 bits in size;
[0027] Expand each 512-bit message block into 64 32-bit words;
[0028] Initialize a loop counter to control the number of iterations of the compression function;
[0029] According to the value of the loop counter, select the corresponding W word for processing, use the selected W word and the current working variable to calculate a series of temporary variables, update the value of the working variable according to the temporary variables and specific logical functions, increment the loop counter, and prepare for the next iteration; after completing the processing of all message blocks, use the value of the working variable as the final hash value.
[0030] Furthermore, the processed message is divided into fixed-size blocks, each block is 512 bits in size, including:
[0031] Generate an initial population. Each individual in the initial population represents a message segmentation method, including the positions of a series of segmentation points. The size of each block at the segmentation point is ≤ 512 bits.
[0032] Define an evaluation function to evaluate the performance of each split method in the SHA-256 processing process;
[0033] For each individual, the message is split according to the splitting method it represents, the split message blocks are processed with SHA-256, and the processing performance is evaluated using the evaluation function;
[0034] According to the processing performance, the corresponding individuals are determined, two individuals are randomly selected, and the positions of some of the individual division points are exchanged at a certain crossover rate to generate new individuals; the positions of the individual division points are randomly changed at a certain mutation rate to generate a new generation of populations, and iterative evolution is performed until the preset evolutionary generation is reached to generate the final division method;
[0035] The processed message is split into fixed-size blocks according to the final split method.
[0036] Furthermore, according to the processing performance, the corresponding individuals are determined, two individuals are randomly selected, and the positions of some segmentation points of the individuals are exchanged at a certain crossover rate to generate new individuals, including:
[0037] Determine the number of split points in multi-point crossover, which are used to exchange some genes of parent individuals;
[0038] Generate a random number corresponding to the number of selected split points, and for the two selected parent individuals, determine the split point positions to be exchanged according to the generated random number;
[0039] At each selected split point, the corresponding split point information of the two parent individuals is exchanged, and after multi-point crossover, two new individuals are obtained.
[0040] Further, according to the number of pipeline stages, a mode selection matrix is determined, including:
[0041] According to the number of pipeline stages, create a matrix, where the rows of the matrix represent pipeline stages and the columns represent SHA-256 processing units. Initially, each element in the matrix is set to a state value indicating unassigned or free;
[0042] According to the number of pipeline stages and the number of processing units, the pipeline stage of processing corresponding to each processing unit is determined;
[0043] For each processing unit, in its corresponding column, the row corresponding to the assigned pipeline stage is set to a state value indicating assigned or busy to obtain a mode selection matrix.
[0044] In a second aspect, a parallel operation control optimization system of multiple HASH algorithms in an ASIC chip is applied in the method described, comprising:
[0045] A creation module is used to create an integration module, and the integration module is used to accommodate multiple SHA-256 processing units;
[0046] An analysis module, for analyzing a pipeline of a SHA-256 processing unit to determine the number of pipeline stages;
[0047] A determination module, used for determining a mode selection matrix according to the number of pipeline stages;
[0048] A scheduling module, for scheduling different SHA-256 processing units and pipeline stages according to a mode selection matrix to obtain an adjusted SHA-256 processing unit;
[0049] An adding module is used to add a control signal to each adjusted SHA-256 processing unit, and the control signal is used to control the start and stop of the SHA-256 processing unit operation;
[0050] An optimization module, used to optimize the task allocation scheme between the adjusted SHA-256 processing units, using each particle to represent a specific task allocation scheme;
[0051] A definition module is used to define a fitness function for evaluating the performance of each task allocation scheme;
[0052] The calculation module is used to calculate the fitness of each particle according to the fitness function, update the speed and position of the particle, and iterate until the preset number of iterations is reached to obtain the final task allocation plan.
[0053] According to a third aspect, a computing device includes:
[0054] one or more processors;
[0055] The storage device is used to store one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the method described.
[0056] In a fourth aspect, a computer-readable storage medium stores a program, and when the program is executed by a processor, the method described is implemented.
[0057] The above scheme of the present invention includes at least the following beneficial effects.
[0058] By creating an integrated module, the work of multiple SHA-256 processing units can be managed and coordinated within a unified framework, improving the utilization of hardware resources. By analyzing the pipeline of the SHA-256 processing unit, the number of pipeline stages can be accurately determined, thereby gaining a more detailed understanding of the algorithm's execution flow and performance bottlenecks. The determination of the mode selection matrix enables different SHA-256 processing units and pipeline stages to be flexibly scheduled, helping to dynamically adjust processing strategies based on real-time workloads, thereby improving overall processing efficiency.
[0059] By scheduling different processing units and pipeline stages, reasonable resource allocation and load balancing can be achieved, overloading of certain processing units or stages can be avoided, and the stable operation and high efficiency of the entire system can be ensured. By adding control signals, the start and stop of each SHA-256 processing unit can be accurately controlled, which not only helps save unnecessary energy consumption, but also responds quickly to external requests when needed, improving the real-time performance of the system.
[0060] By optimizing the task allocation scheme, it can be ensured that each processing unit can get a reasonable and balanced amount of tasks, avoiding the situation of task accumulation or idle processing units, thereby maximizing the throughput and efficiency of the entire system. The definition of the fitness function provides an objective and quantitative standard to evaluate the performance of each task allocation scheme, which makes the optimization process more targeted and helps to quickly find the best task allocation strategy. By calculating the fitness of each particle and updating its speed and position, the optimal solution can be gradually approached. This iterative optimization method not only improves the search efficiency, but also increases the possibility of finding the global optimal solution. BRIEF DESCRIPTION OF THE DRAWINGS
[0061] Figure 1 It is a flow chart of a method for optimizing parallel operation control of multiple HASH algorithms in an ASIC chip provided by an embodiment of the present invention.
[0062] Figure 2It is a schematic diagram of a parallel operation control optimization system of multiple HASH algorithms in an ASIC chip provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0063] The exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided to enable a more thorough understanding of the present disclosure and to fully convey the scope of the present disclosure to those skilled in the art.
[0064] like Figure 1 As shown, an embodiment of the present invention provides a method for optimizing the parallel operation control of multiple HASH algorithms in an ASIC chip, the method comprising the following steps:
[0065] Step 11, creating an integration module, the integration module is used to accommodate multiple SHA-256 processing units;
[0066] Step 12, analyzing the pipeline of the SHA-256 processing unit to determine the number of pipeline stages;
[0067] Step 13, determining a mode selection matrix according to the number of pipeline stages;
[0068] Step 14, scheduling different SHA-256 processing units and pipeline stages according to the mode selection matrix to obtain an adjusted SHA-256 processing unit;
[0069] Step 15, adding a control signal to each adjusted SHA-256 processing unit, where the control signal is used to control the start and stop of the SHA-256 processing unit operation;
[0070] Step 16, optimizing the adjusted task allocation scheme between the SHA-256 processing units, using each particle to represent a specific task allocation scheme;
[0071] Step 17, defining a fitness function for evaluating the performance of each task allocation scheme;
[0072] Step 18, according to the fitness function, calculate the fitness of each particle, update the speed and position of the particle, and iterate until the preset number of iterations is reached to obtain the final task allocation solution.
[0073] In an embodiment of the present invention, step 11, the creation of the integration module provides a centralized management and operation platform for multiple SHA-256 processing units, enhances the organization of internal chip resources, facilitates unified scheduling and control, and simplifies the complexity of external interfaces and interactions. Step 12, by analyzing the pipeline, the execution details of the SHA-256 algorithm in the ASIC chip can be accurately grasped, including potential performance bottlenecks and optimization points, which helps to achieve more refined control and higher computing efficiency. Step 13, the determination of the mode selection matrix provides a basis for realizing flexible computing scheduling. It allows the working mode of the SHA-256 processing unit to be dynamically adjusted according to the pipeline level and real-time workload, thereby optimizing performance and responding to different computing requirements. Step 14, through intelligent scheduling, the collaborative work between the SHA-256 processing unit and the pipeline stage can be ensured, resource conflicts and waste can be avoided, the parallelism and efficiency of the overall operation can be improved, and the chip can process a large number of HASH computing tasks more efficiently.
[0074] Step 15. The introduction of the control signal provides a precise start-stop control mechanism for the SHA-256 processing unit, which helps to reduce power consumption, reduce unnecessary computing overhead, and achieve fast response and recovery when necessary. Step 16. By optimizing the task allocation, it can ensure that each SHA-256 processing unit has an appropriate workload, avoiding overload or idle conditions, which helps to achieve load balancing and improve the overall computing throughput and efficiency. Step 17. The definition of the fitness function provides a standard for quantitatively evaluating the performance of the task allocation scheme, which makes the optimization process more objective and measurable, and helps to quickly identify and select the best task allocation strategy. Step 18. Through the iterative optimization process based on the fitness function, the optimal task allocation scheme can be gradually approached and found. This method combines the characteristics of global search and local optimization, improves the optimization efficiency, and ensures the quality and performance of the final solution.
[0075] In another preferred embodiment of the present invention, the above step 11 may include:
[0076] Analyze the number of SHA-256 processing units that need to be accommodated, determine the interface requirements between the integrated module and other system components; set the performance indicators of the integrated module, such as data transmission rate, power consumption, etc.; design the communication bus inside the module to ensure efficient communication between each SHA-256 processing unit and with other system components. Plan the power supply and clock distribution network to ensure that all processing units can work stably, and design control logic to manage and schedule each processing unit. Use electronic design automation (EDA) tools to design circuit diagrams, perform circuit simulation, verify the correctness and performance of the design, perform physical layout on the chip, determine the location of each SHA-256 processing unit, complete wiring work, and ensure that all circuit connections are correct. Send the design data to the foundry for chip manufacturing, test the manufactured chips, and verify the functions and performance of the integrated module.
[0077] When specifically applied, this includes:
[0078] Suppose you need to design an integrated module that can accommodate 4 SHA-256 processing units. The following is a specific case description:
[0079] It is determined that it is necessary to accommodate 4 SHA-256 processing units, and the integration module is required to be able to communicate efficiently with external memory and other processing units; design a four-way parallel communication bus to ensure that each SHA-256 processing unit can communicate with the external memory independently, plan a centralized power management system to provide a stable power supply for each processing unit; design a central control unit to manage and schedule the work of the 4 SHA-256 processing units. Use EDA tools to draw circuit diagrams, including communication buses, power management systems, and central control units, and verify the correctness and performance of the design through circuit simulation to ensure that all processing units can work normally and meet performance indicators.
[0080] Rationally plan the location of the four SHA-256 processing units on the chip, as well as the layout of the communication bus, power management system, and central control unit, complete the wiring work, ensure that all circuit connections are correct, and optimize the wiring to reduce signal delays and power consumption. Send the design data to the foundry for chip manufacturing. Strictly test the manufactured chips, including functional testing and performance testing, to ensure that the integrated module meets the design requirements and can work stably.
[0081] In a preferred embodiment of the present invention, the above step 12, analyzing the pipeline of the SHA-256 processing unit to determine the number of pipeline stages, may include:
[0082] Step 121, obtaining a circuit diagram of the SHA-256 processing unit, specifically comprising: obtaining the circuit diagram of the SHA-256 processing unit from a design database, involving using a specific electronic design automation (EDA) tool to browse or export the circuit diagram; ensuring that the selected circuit diagram is the latest or specific version of the SHA-256 processing unit to reflect the current design and implementation details, and exporting the circuit diagram to a common file format (such as PDF, PNG, etc.) or directly printing it out as needed for subsequent analysis;
[0083] Step 122, identifying different pipeline stages in the SHA-256 processing unit according to the circuit diagram;
[0084] Step 123, according to different pipeline stages in the SHA-256 processing unit, determine the number of pipeline stages, wherein the number of pipeline stages corresponds to different stages identified in the hardware implementation, each stage is regarded as one stage of the pipeline, and the number of pipeline stages is determined by analyzing the input data and output data of each stage, as well as the dependency relationship between the input data and the output data, specifically including:
[0085] According to the different pipeline stages of the SHA-256 processing unit identified in step 122, list the names and main functions of all stages; ensure that each stage has a clear definition and boundaries; for each pipeline stage, list its input data and output data in detail. The input data includes the original data, control signals, the output of the previous stage, etc.; the output data includes the processed data, status flags, hash values, etc.
[0086] Analyze whether the input data of each stage depends on the output data of other stages, identify which stages are serial, that is, the next stage must wait for the completion of the previous stage before it can start, and identify which stages can be parallel, that is, their execution is not dependent on each other; divide the pipeline into several independent levels according to data dependencies; each level corresponds to one or more stages that can be executed in parallel, but there is no data dependency between these stages. If the output of one stage is a necessary input for the next stage, then these two stages should be divided into different pipeline levels.
[0087] Verify whether the divided pipeline levels are reasonable and ensure that no data dependencies are missed. If the division is found to be unreasonable, such as a certain stage is mistakenly divided into the wrong level, adjustments need to be made and the verification and adjustment process is repeated until a reasonable and efficient pipeline level division scheme is determined; record the final pipeline level division scheme in a document, including the functional description of each level, input and output data, and data dependencies.
[0088] In an embodiment of the present invention, by analyzing the circuit diagram and identifying the pipeline stages, the internal structure and function of the SHA-256 processing unit can be understood. By understanding the workload and dependencies of each pipeline stage in detail, potential performance bottlenecks and optimization points can be accurately identified. Understanding the number of pipeline stages helps to allocate resources more reasonably, avoid overloading some stages while other stages are idle, thereby maximizing the utilization of resources. By optimizing the collaborative work of the pipeline stages, resource conflicts and waiting times can be reduced, and the parallelism and efficiency of the overall operation can be improved. By finely controlling the start and stop of each pipeline stage, the working state of the processing unit can be dynamically adjusted according to actual needs, thereby reducing unnecessary power consumption.
[0089] When specifically applied, this includes:
[0090] In order to improve the security and efficiency of data processing, a SHA-256 processing unit is integrated. In order to optimize the performance of the processing unit, its pipeline is analyzed to determine the number of pipeline stages, thereby guiding the subsequent parallel operation control and resource optimization.
[0091] In step 121, the latest circuit diagram of the SHA-256 processing unit is exported from the design database using an EDA tool and saved as a PDF file.
[0092] Step 122, by analyzing the circuit diagram, the following pipeline stages in the SHA-256 processing unit are identified:
[0093] Data reception and preprocessing are responsible for receiving external data and performing necessary format conversion and filling.
[0094] Message expansion and grouping: the preprocessed data is expanded and grouped according to the requirements of the SHA-256 algorithm.
[0095] The compression function is executed, and the SHA-256 compression function is performed on each group to generate an intermediate hash value.
[0096] Hash value merging and outputting: Merge the intermediate hash values of all groups to generate and output the final 256-bit hash value.
[0097] Step 123, next, the input and output data of each stage and their dependencies are analyzed:
[0098] The data receiving and preprocessing stage mainly relies on external input data and has no internal dependencies. The input of the message expansion and grouping stage is the preprocessed data, and the output is the grouped data, which depends on the completion of the data receiving and preprocessing stage. The input of the compression function execution stage is the grouped data, and the output is the intermediate hash value, which depends on the completion of the message expansion and grouping stage. In addition, the compression function execution of each group can be carried out in parallel, but there is no dependency between each group. The input of the hash value merging and output stage is the intermediate hash value of all groups, and the output is the final hash value, which depends on the completion of all compression function execution stages.
[0099] Based on the above analysis, it is determined that the pipeline level of the SHA-256 processing unit is 4, corresponding to the 4 identified pipeline stages. This discovery provides an important basis for subsequent parallel computing control and resource optimization. For example, the overall performance can be improved by increasing the parallelism of the compression function execution stage, while ensuring reasonable resource allocation in the data reception and preprocessing, message expansion and grouping, and hash value merging and output stages to avoid performance bottlenecks.
[0100] In a preferred embodiment of the present invention, the above step 122, identifying different pipeline stages in the SHA-256 processing unit according to the circuit diagram, may include:
[0101] Step 1221, according to the circuit diagram, identifying logic units and modules, including registers, arithmetic logic units and control units, specifically including:
[0102] Take a preliminary look at the circuit diagram to familiarize yourself with its overall structure and layout; find and label all logic units in the circuit diagram, such as registers, arithmetic logic units (ALUs), and control units, which are hardware components that perform specific functions.
[0103] Register is a temporary storage unit used to store data or instructions.
[0104] Arithmetic Logic Unit, a unit used to perform arithmetic and logical operations.
[0105] A control unit is used to coordinate and manage the operations of other units.
[0106] Create a list or table recording the name, type, and location of each logical unit identified;
[0107] Step 1222, according to the circuit diagram and the data path, the SHA-256 processing unit is divided into different functional modules, including a message filling module, an expansion module and a compression function module, specifically including:
[0108] Observe the data flow path in the circuit diagram to understand how data is transmitted and processed in the processing unit. Based on the processing flow of the SHA-256 algorithm, the processing unit is divided into several main functional modules, such as message filling module, expansion module and compression function module.
[0109] A message padding module is used to pad the input message to a specified length;
[0110] The extension module is used to extend the padded message to generate additional data blocks;
[0111] The compression function module is used to execute the compression function of the SHA-256 algorithm and perform hash calculation on each data block;
[0112] Step 1223, refine each functional module and identify the various stages that constitute the pipeline;
[0113] Step 1224, marking the boundaries of each pipeline stage on the circuit diagram, specifically includes:
[0114] Using circuit diagram editing software or manual marking tools (such as colored pens or label paper), find the start and end positions of each stage on the circuit diagram based on the pipeline stage information identified in step 1223. Use the selected marking tool to clearly mark the boundaries of each pipeline stage on the circuit diagram. Different colors or symbols can be used to represent different stages.
[0115] In the embodiment of the present invention, step 1222, in this step, according to the circuit diagram and the data path, the SHA-256 processing unit is divided into different functional modules, such as a message filling module, an expansion module, and a compression function module, and these functional modules correspond to different parts of the SHA-256 algorithm. By dividing the processing unit into functional modules, analysts can better understand the function of each module and the interaction between them, and understanding the functional modules helps to identify which parts can be executed in parallel, thereby improving processing efficiency.
[0116] Step 1223, in this step, each functional module is further refined to identify the various stages that make up the pipeline. For example, in the compression function module, multiple consecutive processing steps can be identified, each corresponding to a pipeline stage. By refining the functional modules, analysts can obtain a detailed view of each stage of the pipeline. Understanding the specific operations of each pipeline stage helps to accurately locate performance bottlenecks and optimize them. Step 1224, marking the boundaries of each pipeline stage on the circuit diagram helps to visually represent the boundaries and data flow between different stages. Marking the boundaries makes the pipeline stages clearly visible on the circuit diagram.
[0117] In a preferred embodiment of the present invention, the above step 1223, which refines each functional module and identifies each stage constituting the pipeline, may include:
[0118] Step 12231, according to the length of the original message in the message filling module, according to the filling rule of SHA-256, the message is padded so that its length meets specific conditions, and a 64-bit integer is appended to the end of the padded message to indicate the length of the original message, so as to obtain a processed message, specifically including:
[0119] Get the original message to be processed and record its length (in bits); SHA-256 requires that the message length (including padding) must be a multiple of 512 bits. Therefore, it is necessary to calculate the number of bits required for padding so that the total length of the original message plus the padding reaches the closest multiple of 512 bits of the original message length and is not less than that. The length of the padding is composed of a 1 and a number of 0s, with the 1 being located at the beginning of the padding, followed by a sufficient number of 0s to reach the required length. First add a '1' bit to the end of the original message; then, add a sufficient number of '0' bits so that the total length from the beginning of the original message to the end of the padding is 64 bits less than the nearest multiple of 512 bits. Convert the recorded original message length to a 64-bit unsigned integer, which should be represented in big endian order (that is, the most significant bit is first). Append this 64-bit integer to the end of the padded message. In this way, the length of the entire message (including the original message, padding, and length addition) is a multiple of 512 bits;
[0120] Step 12232, split the processed message into fixed-size blocks, each block has a size of 512 bits;
[0121] Step 12233, expanding each 512-bit message block into 64 32-bit words, specifically includes:
[0122] The 512-bit message block is divided into 16 32-bit words, which are arranged in big-endian order (most significant byte first) and named W0 to W1. 15 Next, the SHA-256 specific algorithm is used to expand these 16 words, generating a total of 64 32-bit words, named W0 to W 63 ; For W 16 To W 63 , each word W t (t is the index of the word, ranging from 16 to 63) are calculated using the following formula:
[0123] ;
[0124] in, and It is a mixing function defined in the SHA-256 algorithm, which is used to perform bit operations and shifts on 32-bit words. and It is a mixing function defined in the SHA-256 algorithm, which is used to perform bit operations on 32-bit words. The definitions of these two functions are as follows:
[0125] For any 32-bit word W, The calculation formula is:
[0126]
[0127] in, Indicates circularly shifting W to the right Bit; Indicates logical shift W to the right bits (i.e. fill with 0 on the right). The calculation formula is:
[0128] ;
[0129] In the context of SHA-256, these operations are all performed on 32-bit unsigned integers. These operations help to mix and diffuse the data in each round of the algorithm, thus enhancing the security of the hash function.
[0130] Step 12234, initialize a loop counter to control the number of iterations of the compression function, including: In the SHA-256 algorithm, the compression function will perform 64 rounds of iterations for each 512-bit message block. Therefore, the maximum value of the loop counter should be set to 63 (starting from 0); allocate a variable in the memory as a loop counter, usually named t or counter; set the initial value of the loop counter to 0, which will be used to control the iteration process of the compression function; ensure that other required variables of the compression function (such as working variables, constants, etc.) are also correctly initialized. The loop counter will be incremented before each iteration of the compression function and will stop when it reaches 64;
[0131] Step 12235, according to the value of the loop counter, select the corresponding W word for processing, use the selected W word and the current working variable to calculate a series of temporary variables, update the value of the working variable according to the temporary variables and a specific logic function, increment the loop counter, and prepare for the next iteration; after completing the processing of all message blocks, use the value of the working variable as the final hash value, specifically including:
[0132] According to the current loop counter value t (starting from 0 and ending at 63), the 64 32-bit words W0 to W0 obtained from the previous expansion are 63 Select the corresponding W t; If this is the processing of the first message block, eight working variables a, b, c, d, e, f, g, h need to be set from the predefined initial hash value (initial value of SHA-256); if it is not the first message block, the working variable is the hash value after the previous message block is processed. Using the selected Wt and the current working variables, a series of temporary variables are calculated through the compression function defined in the SHA-256 algorithm. These temporary variables include T1 and T2, which are calculated based on a specific logic function and the current working variables.
[0133] Using the calculated temporary variables T1 and T2, and the current working variables, update the values of the working variables according to the rules defined in the SHA-256 algorithm, involving a series of additions and bit operations. Increment the value of the loop counter t by 1 to prepare for the next iteration. If the value of the loop counter t is less than 64, return to the above steps to continue processing the next W t ; If t is equal to 64, it means that all words of the current message block have been processed. If there are more message blocks to be processed, the value of the current working variable is used as the initial hash value for the next message block, and the process returns to the above steps to start processing the next message block. When all message blocks are processed, the value of the current working variable is the final hash value. This hash value is output as the result of the SHA-256 algorithm. This process describes how the compression function in the SHA-256 algorithm processes each 512-bit message block and generates the final hash value through multiple iterations. Each step strictly follows the SHA-256 specification to ensure the correctness and security of the hash result.
[0134] In the SHA-256 hash algorithm, each iteration step involves calculating temporary variables T1 and T2, and then using these variables and the current working variable to update the value of the working variable. The following is the specific implementation process:
[0135] Set the initial working variables , b, c, d, e, f, g, h are the initial hash values of SHA-256 (if it is the first message block) or the hash values of the previous message block after processing. For each message block, 64 rounds of iterations are performed, and the following operations are performed in each round of iteration:
[0136] Select W t , select the corresponding 32-bit word W according to the loop counter t (from 0 to 63) t If it is the first message block, W t is obtained directly from the original message after padding and grouping; otherwise, W t It is obtained through the message expansion algorithm.
[0137] Using Logical Functions , and the epicycle constant To calculate T1:
[0138] ;
[0139] in, is a series of bit rotations and XOR operations, is the selection function, according to The value of is used to select the output. Use the logical function and To calculate T2:
[0140] ;
[0141] in, is another series of bit rotations and XOR operations, is a majority function, according to The value of returns the bit with the highest number of occurrences. Update the working variable:
[0142] ;
[0143] Increment the loop counter:
[0144] ;
[0145] Prepare for the next iteration or complete processing:
[0146] If t<64, return to step 2 to continue iterating. If t=64, it means that all words of the current message block have been processed. If there are more message blocks to be processed, the value of the current working variable is used as the initial hash value for the next message block to be processed, and return to the above step to start processing the next message block. When all message blocks are processed, the value of the current working variable ( , b, c, d, e, f, , h) is the final hash value.
[0147] In an embodiment of the present invention, the processing of SHA-256 is decomposed, and each step provides a deeper and clearer understanding of the algorithm flow. The SHA-256 processing is divided into clear stages, each stage corresponds to a specific function, making the modular design of hardware or software simpler and more direct. This modular design improves the readability and maintainability of the code. After identifying the various stages of the pipeline, independent performance analysis and optimization can be performed for each stage. For example, the processing speed of a specific stage can be increased by parallel processing, pipeline rearrangement or hardware acceleration technology, thereby improving the overall performance. During system testing or operation, if a problem or failure occurs, the refined pipeline stage can help quickly and accurately locate the problem. The clear definition of each stage makes troubleshooting more efficient. When the SHA-256 processing unit needs to be expanded or modified, the refined pipeline structure makes these changes easier to implement, and it can be accurately known which stages need to be adjusted and how to adjust them, thereby reducing the uncertainty and risk in the modification process.
[0148] In a preferred embodiment of the present invention, the above step 12232, dividing the processed message into blocks of a fixed size, each block having a size of 512 bits, may include:
[0149] Step 122321, generating an initial population, each individual in the initial population represents a message segmentation method, including a series of segmentation point positions, and the size of each segmentation point block is ≤ 512 bits, specifically including: determining the size of the initial population, that is, how many individuals it contains; each individual represents a message segmentation method, consisting of a series of segmentation point positions; for each individual, randomly generating a segmentation point position, ensuring that the block size between each segmentation point does not exceed 512 bits, repeating the above process until a sufficient number of individuals are generated to form an initial population;
[0150] Step 122322, define an evaluation function for evaluating the performance of each segmentation method during SHA-256 processing;
[0151] Step 122323, for each individual, the message is segmented according to the segmentation method represented by the individual, SHA-256 processing is performed on the segmented message blocks, and the processing performance is evaluated using the evaluation function, specifically including: for each individual in the population, the message is segmented according to the segmentation method represented by the individual; SHA-256 hash processing is performed on each segmented message block; and the processing performance score corresponding to each individual is calculated using the previously defined evaluation function;
[0152] Step 122324, according to the processing performance, determine the corresponding individuals, randomly select two individuals, exchange some of the individual's segmentation point positions at a certain crossover rate to generate new individuals; randomly change the individual's segmentation point positions at a certain mutation rate to generate a new generation of population, iterative evolution, until the preset evolutionary generation is reached, to generate the final segmentation method, specifically including: setting a mutation rate (for example, 0.05), indicating that each segmentation point position has a 5% probability of mutating in each iteration; presetting an evolutionary generation (for example, 100 generations) as the termination condition of iterative evolution; initializing to 0 to record the current evolutionary progress; repeating the following steps until the current evolutionary generation reaches the preset evolutionary generation:
[0153] Based on the performance evaluated in the previous step, select individuals with better performance as the basis of the new generation population, and generate a part of new individuals through crossover operation (as described above); traverse each individual in the new generation population; generate a random number for each split point position of each individual; if the random number is less than the mutation rate, mutate the split point position. The mutation can be to randomly move the position of the split point, or to randomly increase or decrease a split point within the allowed range (no more than 512 bits); for each individual in the new generation population, split the message according to the split method it represents, and perform SHA-256 processing, and use the evaluation function (as described above) to evaluate the processing performance of each individual; update the current evolutionary generation: the current evolutionary generation increases by 1, and after each iteration, the best performing individual and its corresponding split method are recorded. When the iterative evolution is completed, the best performing individual and its corresponding split method are output as the final split method.
[0154] Step 122325, according to the final segmentation method, the processed message is divided into blocks of fixed size, specifically including: selecting the individual with the best performance from the evolved population, that is, the final segmentation method, dividing the original message according to the final segmentation method, and performing necessary processing on each segmented message block, such as encryption, storage or transmission.
[0155] In the embodiment of the present invention, step 122321, by generating an initial population, the algorithm can explore multiple possible message segmentation methods. The diversity of the population helps the algorithm avoid falling into a local optimal solution during the search process, thereby having a greater chance of finding a global optimal solution. Step 122322, the evaluation function provides a quantitative standard for the algorithm to compare the performance of different segmentation methods, which enables the algorithm to objectively evaluate the pros and cons of each segmentation method and select and adjust according to performance. Step 122323, by performing actual SHA-256 processing and performance evaluation on each individual, the algorithm can obtain specific data on the actual effect of each segmentation method, which enables the algorithm to make decisions based on actual performance, thereby improving the accuracy and efficiency of the search. Step 122324, through crossover and mutation operations, the algorithm can combine the advantages of different segmentation methods to generate new and potentially better segmentation methods. This mechanism helps the algorithm to continuously evolve and improve during the search process, gradually approaching the optimal solution. Step 122325, after the search and optimization of the genetic algorithm, the segmentation method finally obtained is the best performance under given conditions. Using this segmentation method to segment the message can achieve better performance during SHA-256 processing, such as higher processing speed, lower resource consumption, etc.
[0156] The calculation formula of the evaluation function is:
[0157] ;
[0158] in, It is The processing time of each block; It is Memory consumption per block; It is CPU consumption per block; It is The size of the block; , , , and is the weight coefficient; It is The size of the block; is the index, indicating the data blocks or elements to traverse when calculating sums and variances; It is the index used in the inner summation and is used when calculating the average value of the block size. data blocks; is the total number of blocks or data points.
[0159] In a preferred embodiment of the present invention, the above steps 122324, determining the corresponding individuals according to the processing performance, randomly selecting two individuals, exchanging the positions of some segmentation points of the individuals at a certain crossover rate, and generating new individuals, may include:
[0160] Step 1223241, determine the number of split points in the multi-point crossover, the split points are used to exchange some genes of the parent individuals, the specific implementation process: set a range of split points, this range can be set according to the length of the individual (chromosome). For example, if the length of the individual is L, the minimum value of the number of split points can be set to 1 (i.e. single-point crossover), and the maximum value can be set to L-1 (to ensure that at least one gene is not exchanged). However, in practical applications, a smaller maximum value is usually selected to avoid generating too many meaningless offspring;
[0161] Step 1223242, generate a random number corresponding to the number of selected split points, and for the two selected parent individuals, determine the split point positions to be exchanged according to the generated random number, specifically including: in each iteration, for each pair of parent individuals that need to be crossed, randomly select an integer within the above range as the number of split points. This random number can be obtained through various random number generation algorithms, such as linear congruential generator (LCG); after determining the number of split points, it is necessary to randomly select a corresponding number of positions in the gene sequence of the parent individual as split points. These split points divide the gene sequence of the parent individual into multiple segments. For example, if the length of the parent individual is 10 and the number of randomly selected split points is 2, the split point positions that may be selected are the 3rd and 7th positions, thereby dividing the individual into three segments;
[0162] Step 1223243, at each selected split point position, the corresponding split point information of the two parent individuals is exchanged, and two new individuals are obtained after multi-point crossover, which specifically includes: after the split point position is determined, the crossover operation can be performed. Specifically, the segments between the corresponding split points of the two parent individuals are exchanged to generate two new offspring individuals. Taking the above example, the segments between the 3rd and 7th positions of the two parent individuals will be exchanged. When performing the crossover operation, it is necessary to pay attention to handling some boundary conditions. For example, if the split point happens to be selected at the start or end position of the individual, the actual exchanged segment may be empty or contain the remainder of the entire individual. In addition, it is also necessary to ensure that the crossover operation does not destroy the integrity of the individual (that is, the length of the newly generated individual should be the same as that of the parent individual). Finally, the newly generated offspring individuals are added to the population, and may replace the individuals with poor performance in the population (if the population size remains unchanged), so that as the iteration proceeds, the individuals in the population will gradually adapt to the problem to be solved.
[0163] In the embodiment of the present invention, by randomly selecting the split points for exchange, multi-point crossover helps to introduce new gene combinations into the population, thereby increasing the diversity of the population. Since the new individuals inherit the excellent genes of the parent individuals, multi-point crossover helps to accelerate the convergence process of the algorithm to the optimal solution. The randomness of multi-point crossover helps the algorithm to jump out of the local optimal solution and continue to search for the global optimal solution. By adjusting the number of split points, the complexity and scope of the crossover operation can be controlled, making the algorithm more flexible and adaptable to different problem scenarios.
[0164] In a preferred embodiment of the present invention, the above step 13, determining the mode selection matrix according to the number of pipeline stages, may include:
[0165] Step 131, according to the number of pipeline stages, create a matrix, the rows of the matrix represent pipeline stages, and the columns represent SHA-256 processing units. Initially, each element in the matrix is set to a state value indicating unallocated or idle, specifically including:
[0166] Determine the number of rows of the matrix according to the number of pipeline stages. Assuming that the pipeline has N stages, the matrix will have N rows. Each row represents a specific stage in the pipeline; determine the number of columns of the matrix according to the number of SHA-256 processing units. Assuming that there are M processing units, the matrix will have M columns, each representing a specific processing unit. Create a matrix with N rows and M columns. In most programming languages, this can be implemented using a two-dimensional array or a list of lists; initialize each element in the matrix. Since all processing units are not assigned to any pipeline stage at the beginning, all elements are set to a specific state value, indicating "unassigned" or "idle". This state value can be a Boolean value False, a number 0, a specific string (such as "idle"), or any other tag that can clearly indicate the idle state. For each row of the matrix, it can be given a label or identifier to indicate the pipeline stage corresponding to the row. This label can be a simple number (from 1 to N) or a more descriptive string (such as "Stage 1", "Stage 2", etc.). Similarly, each column of the matrix can also be given a label or identifier to indicate which processing unit the column corresponds to. This can be the number of the processing unit (such as "Unit 1", "Unit 2", etc.) or any other unique identifier.
[0167] Step 132, according to the number of pipeline stages and the number of processing units, determines the pipeline stage of processing corresponding to each processing unit, specifically including:
[0168] Determine the total number of pipeline stages, which determines the number of tasks that need to be processed; determine the total number of available SHA-256 processing units, which will limit the ability to process in parallel. If there are performance differences between processing units, this also needs to be considered when allocating tasks. If the number of pipeline stages and the number of processing units are roughly equal, you can try to evenly allocate each pipeline stage to a processing unit; if some processing units have higher performance, more complex pipeline stages can be allocated to these high-performance units to increase the overall processing speed; according to the processing time and complexity of each pipeline stage, keep the workload of each processing unit balanced to avoid overloading some units while others are idle. When designing the allocation strategy, you should consider possible future expansion or changes so that the allocation can be easily adjusted when necessary. In the order of the pipeline, assign each stage to the next available processing unit in turn; to increase parallelism, you can try to stagger the allocation of pipeline stages to processing units. For example, processing unit 1 processes stage 1, stage 4, etc., processing unit 2 processes stage 2, stage 5, and so on. This approach can reduce waiting time and improve processing efficiency. Use heuristic algorithms to find the optimal allocation solution based on actual test data or performance models;
[0169] Step 133, for each processing unit, in its corresponding column, the row corresponding to the allocated pipeline stage is set to a state value indicating allocated or busy, so as to obtain a mode selection matrix, specifically including:
[0170] According to the distribution result of the pipeline stages corresponding to each processing unit, this is a list or map, where the key is the identifier of the processing unit and the value is the list or range of pipeline stages that the processing unit is responsible for. For each processing unit, traversing the list of pipeline stages it is responsible for involves two layers of loops: the outer loop traverses the processing units, and the inner loop traverses the pipeline stages that each processing unit is responsible for.
[0171] During the traversal, for each pipeline stage that each processing unit is responsible for, find the corresponding row and column in the mode selection matrix and update the state value of the position to "allocated" or "busy". The state value here can be any mark that can clearly indicate that the position has been occupied by a processing unit, such as the Boolean value True, the number 1, or a specific string (such as "allocated").
[0172] In an embodiment of the present invention, by allocating a specific pipeline stage to each processing unit, it can be ensured that all processing units are working efficiently and not idle. This avoids waste of resources and improves the utilization of overall hardware resources. The mode selection matrix allows the parallel processing strategy to be flexibly configured according to the number of pipeline stages and the number of processing units. Through reasonable allocation, the ability of parallel processing can be maximized, thereby shortening the processing time and improving the system throughput. The mode selection matrix is not static, but can be adjusted according to the actual workload and performance requirements. This means that the system can dynamically reallocate processing units according to real-time conditions to adapt to different computing requirements, thereby maintaining optimal performance. Through the mode selection matrix, the scheduling logic becomes clearer and simpler. The tasks and states of each processing unit are clear at a glance, which helps to simplify the implementation of the scheduling algorithm and reduce errors and complexity in the scheduling process. Through reasonable task allocation and state management, the system can run more stably. Conflicts and competitions between processing units are minimized, thereby reducing the risk of system crashes or performance degradation. When a problem occurs in the system, the mode selection matrix can help quickly locate the problem. By viewing the state values in the matrix, it is possible to quickly determine which processing unit or which pipeline stage has a problem, thereby speeding up troubleshooting and repair. The design of the mode selection matrix allows the system to support multiple operation modes. By adjusting the configuration of the matrix, you can easily switch between different operation modes to meet different application scenarios and requirements.
[0173] In another preferred embodiment of the present invention, the above step 14 may include:
[0174] Step 141, parse the mode selection matrix to obtain the pipeline stage to which each SHA-256 processing unit is assigned. The rows of the mode selection matrix represent different stages of the pipeline, and the columns represent different processing units. Each element in the matrix indicates whether a specific processing unit is responsible for processing a specific pipeline stage. According to the parsing results of the mode selection matrix, a task queue is prepared for each processing unit. The task queue contains the data or task identifier of the pipeline stage that the processing unit needs to process. If a processing unit is responsible for multiple pipeline stages, the tasks of these stages will be arranged in the task queue in sequence.
[0175] Before scheduling begins, the state of each processing unit is initialized. This includes setting the working mode, memory address, registers, etc. of the processing unit to ensure that they can receive and process data correctly. Once the task queue is ready and the processing unit state has been initialized, the processing unit can be started. The system sends a start signal to each processing unit and may also provide the first task of the task queue as input data at the same time. When the processing unit starts executing the tasks of the pipeline stage, the system needs to monitor their execution status. This can be achieved by polling the status register of the processing unit, interrupt signal, or using other synchronization mechanisms. The system needs to ensure that the various stages in the pipeline are executed in the correct order and that data transfer between processing units is synchronized.
[0176] If the mode selection matrix allows or the system supports dynamic scheduling, the system can dynamically adjust the task allocation during execution based on the load and performance of the processing units. For example, if a processing unit is found to be overloaded and another processing unit is relatively idle, the system can transfer some tasks from the overloaded processing unit to the idle processing unit. When the processing units complete all the tasks in their task queues, they send a completion signal to the system. The system needs to collect these processing results and integrate them into the final calculation result. This may involve merging partial hash values generated by different processing units into a complete hash value, or performing other post-processing operations. During the processing process, the system needs to ensure that all processing units execute synchronously in the predetermined order. If an error or failure occurs in a processing unit, the system needs to take appropriate error handling measures, such as restarting the processing unit, reallocating tasks to other processing units, or executing a fault recovery procedure.
[0177] In another preferred embodiment of the present invention, the above step 15 may include:
[0178] The specific implementation process of step 15 involves adding a control signal to each adjusted SHA-256 processing unit so that their operation can be started and stopped accurately. The following is a detailed description of the implementation process of this step:
[0179] First, design a standard control signal interface for the SHA-256 processing unit. This interface should be able to receive at least two types of control signals: start signal and stop signal. The interface can be in the form of pins, interrupt lines, message queues, etc., depending on the hardware design of the processing unit and the communication protocol of the system.
[0180] In the control logic inside the processing unit or closely related to it, a response mechanism for control signals needs to be implemented. This usually includes the following parts:
[0181] The startup logic, when receiving the startup signal, checks whether the processing unit is ready to perform operations (for example, whether there is enough input data, whether the internal state has been initialized correctly, etc.). If the conditions are met, the operation flow of the processing unit is started.
[0182] Stop logic: When a stop signal is received, the processing unit needs to be able to safely interrupt the current operation process and save necessary status information (such as current operation progress, intermediate results, etc.). The stop logic also needs to ensure that the processing unit can smoothly transition to the idle state so that it is ready to receive a new start signal at any time.
[0183] Integrate the control signal interface into the overall control flow of the system. This usually involves the following steps:
[0184] Hardware connection, ensure that the control signal interface of the processing unit is correctly connected to the control signal source of the system. This may involve PCB wiring, jumper connection or through backplane bus.
[0185] Software configuration, configure the control signal sending logic in the system control software. This includes defining when to send the start signal (such as when enough data is received) and when to send the stop signal (such as when system resources are insufficient or the user requests to stop the operation, etc.).
[0186] Synchronization and testing ensures that the sending and receiving of control signals are synchronized between the processing unit and the system. Perform adequate testing to verify the reliability and accuracy of the control signals.
[0187] In order to prevent false triggering or misuse of control signals, some safety mechanisms may need to be implemented. For example:
[0188] Verification is performed before sending control signals to ensure the validity and legality of the signals.
[0189] If the processing unit does not respond for a long time after receiving the start signal, the system may need to send a stop signal and take appropriate error handling measures.
[0190] Restrict control signals to only entities with appropriate permissions.
[0191] During the implementation process, setting up a monitoring mechanism to track the sending and receiving of control signals can help with debugging and troubleshooting. For example, the time of each sending and receiving of control signals, the state changes of the processing unit, and other information can be recorded in the system log.
[0192] In another preferred embodiment of the present invention, the above step 16 may include:
[0193] Step 16 aims to improve the task allocation scheme between SHA-256 processing units by optimizing the algorithm to improve the efficiency and performance of the overall system. Here, the particle swarm optimization (PSO) algorithm will be used to explain the implementation process of this step in detail.
[0194] In PSO, each particle represents a potential task allocation scheme. The position of the particle indicates the specific details of the task allocation, that is, which pipeline stage tasks each processing unit is assigned to perform. A certain number of particles are randomly generated, and an initial value is assigned to the position of each particle. These initial values can be completely random or set according to some heuristic rules to speed up the convergence process. Each particle is assigned an initial velocity, which indicates the direction and step size of the particle's movement in the search space. The initial velocity can also be random or based on some strategy.
[0195] In another preferred embodiment of the present invention, in the above step 17, the specific calculation formula of the fitness function is:
[0196] ;
[0197] in, is the standard deviation of processing time, reflecting the degree of imbalance in workload among processing units; It is the time of the unit with the longest processing time among all processing units; is the average processing time Target completion time The absolute value of the difference between , and is the weight coefficient.
[0198] The fitness function takes into account the standard deviation of the processing time , Maximum processing time and the deviation of the average processing time from the target time This comprehensive evaluation helps find task allocations that perform well on multiple performance metrics.
[0199] By minimizing the standard deviation of processing time , the fitness function encourages the task allocation scheme to achieve better load balancing. Load balancing is one of the key factors to improve the overall performance and stability of the system, because it can reduce the performance bottleneck caused by overloading some processing units. By minimizing the maximum processing time , the fitness function helps to reduce the overall completion time. In a parallel processing system, the overall completion time is usually limited by the processing unit with the longest processing time. Minimize the deviation of the average processing time from the target time It helps to find a task allocation solution that is close to but not exceeding the target completion time. Weight coefficient in fitness function , and It provides flexibility, allowing the importance of different performance indicators to be adjusted according to actual needs. This adjustability makes the fitness function applicable to a variety of optimization scenarios and goals. When using heuristic search algorithms such as particle swarm optimization (PSO), the fitness function is used as an indicator to evaluate the performance of each task allocation scheme, directly guiding the search direction of the optimization process. By iteratively updating the position and velocity of the particles to maximize the fitness value (or minimize its inverse), the algorithm can gradually approach the optimal task allocation scheme.
[0200] Taking all the above into consideration, the optimized task allocation scheme can significantly improve the overall efficiency and stability of the SHA-256 processing unit. By reducing processing time, achieving better load balancing and meeting real-time requirements, the system can better cope with the needs of high-load and complex computing tasks.
[0201] In another preferred embodiment of the present invention, the above step 18 may include:
[0202] According to the current position of the particle (i.e. the current task allocation plan), use the fitness function to calculate its fitness value; record the individual best fitness (pBest) and position of each particle, and update it if the current fitness is better than the previous record. Find the global best fitness (gBest) among the individual best fitness of all particles and record the corresponding position.
[0203] For each particle in the swarm:
[0204] Update speed. Update the speed of the particle according to the following formula:
[0205] ;
[0206] in, It is a particle In the The speed of the dimension, It is a particle In the The location of the dimension, is the inertia weight, and is the learning factor, and is a random number (between 0, 1), It is a particle In the The individual best position of dimension, is the value of the global best position in the dth dimension.
[0207] Update the position and calculate the new position based on the updated speed and current position:
[0208] ;
[0209] Repeat the steps until the preset maximum number of iterations is reached. When the iteration is completed, the global best position is the optimized task allocation scheme. Apply the scheme to the SHA-256 processing unit and perform subsequent processing as needed (such as starting the processing unit for calculation).
[0210] like Figure 2 As shown, an embodiment of the present invention further provides a system 20 for controlling and optimizing parallel operation of multiple HASH algorithms in an ASIC chip, comprising:
[0211] A creation module 21 is used to create an integration module, where the integration module is used to accommodate multiple SHA-256 processing units;
[0212] An analysis module 22, used for analyzing the pipeline of the SHA-256 processing unit to determine the number of pipeline stages;
[0213] A determination module 23, used to determine a mode selection matrix according to the number of pipeline stages;
[0214] A scheduling module 24, for scheduling different SHA-256 processing units and pipeline stages according to the mode selection matrix to obtain adjusted SHA-256 processing units;
[0215] An adding module 25 is used to add a control signal to each adjusted SHA-256 processing unit, and the control signal is used to control the start and stop of the SHA-256 processing unit operation;
[0216] An optimization module 26, used for optimizing the adjusted task allocation scheme between the SHA-256 processing units, using each particle to represent a specific task allocation scheme;
[0217] A definition module 27, used to define a fitness function for evaluating the performance of each task allocation scheme;
[0218] The calculation module 28 is used to calculate the fitness of each particle according to the fitness function, update the speed and position of the particle, and iterate until a preset number of iterations is reached to obtain a final task allocation solution.
[0219] It should be noted that the system is a system corresponding to the above method, and all implementation methods in the above method embodiment are applicable to this embodiment and can achieve the same technical effect.
[0220] The embodiment of the present invention further provides a computing device, comprising: a processor, a memory storing a computer program, wherein when the computer program is executed by the processor, the method described above is executed. All implementations in the above method embodiment are applicable to this embodiment and can achieve the same technical effect.
[0221] The embodiment of the present invention also provides a computer-readable storage medium storing instructions, which, when executed on a computer, enable the computer to execute the method described above. All implementations in the above method embodiment are applicable to this embodiment and can achieve the same technical effect.
Claims
1. A method for optimizing parallel operation control of multiple HASH algorithms in an ASIC chip, characterized in that: The method comprises the following steps: Creating an integration module, the integration module is used to accommodate multiple SHA-256 processing units; Analyze the pipeline of the SHA-256 processing unit to determine the number of pipeline stages; According to the number of pipeline stages, a mode selection matrix is determined; According to the mode selection matrix, different SHA-256 processing units and pipeline stages are scheduled to obtain an adjusted SHA-256 processing unit; A control signal is added to each adjusted SHA-256 processing unit, where the control signal is used to control the start and stop of the SHA-256 processing unit operation; The task allocation scheme between the optimized and adjusted SHA-256 processing units is represented by one task allocation scheme per particle; Define a fitness function for evaluating the performance of each task allocation scheme; According to the fitness function, the fitness of each particle is calculated, the speed and position of the particle are updated, and the operation is iterated until the preset number of iterations is reached to obtain the final task allocation plan; According to the number of pipeline stages, determine the mode selection matrix, including: According to the number of pipeline stages, create a matrix, where the rows of the matrix represent pipeline stages and the columns represent SHA-256 processing units. Initially, each element in the matrix is set to a state value indicating unassigned or free; According to the number of pipeline stages and the number of processing units, the pipeline stage of processing corresponding to each processing unit is determined; For each processing unit, in its corresponding column, set the row corresponding to the assigned pipeline stage to a state value indicating assigned or busy to obtain a mode selection matrix; According to the mode selection matrix, different SHA-256 processing units and pipeline stages are scheduled to obtain an adjusted SHA-256 processing unit, including: If the mode selection matrix allows dynamic scheduling, the task allocation is dynamically adjusted during execution according to the load and performance of the processing unit.
2. The method for optimizing parallel operation control of multiple HASH algorithms in an ASIC chip according to claim 1, characterized in that: Analyze the pipeline of the SHA-256 processing unit to determine the number of pipeline stages, including: Get the circuit diagram of the SHA-256 processing unit; Based on the circuit diagram, identify the different pipeline stages in the SHA-256 processing unit; According to different pipeline stages in the SHA-256 processing unit, the number of pipeline stages is determined, wherein the number of pipeline stages corresponds to different stages identified in the hardware implementation, each stage is regarded as one stage of the pipeline, and the number of pipeline stages is determined by analyzing the input data and output data of each stage, as well as the dependency between the input data and the output data.
3. The method for optimizing parallel operation control of multiple HASH algorithms in an ASIC chip according to claim 2, characterized in that: Based on the circuit diagram, identify the different pipeline stages in the SHA-256 processing unit, including: Based on the circuit diagram, identify logic units and modules, including registers, arithmetic logic units, and control units; According to the circuit diagram and data path, the SHA-256 processing unit is divided into different functional modules, including a message filling module, an expansion module, and a compression function module; Refine each functional module and identify the various stages that make up the pipeline; Mark the boundaries of each pipeline stage on the circuit diagram.
4. The method for optimizing parallel operation control of multiple HASH algorithms in an ASIC chip according to claim 3, characterized in that: Each functional module is refined to identify the various stages that make up the pipeline, including: According to the length of the original message in the message filling module, the message is filled according to the filling rule of SHA-256 so that its length meets the conditions, and a 64-bit integer is appended to the end of the filled message to indicate the length of the original message, so as to obtain the processed message, which specifically includes: Get the original message to be processed and record its length in bits; SHA-256 requires that the message length is a multiple of 512 bits, calculate the number of bits needed to be padded so that the total length of the original message plus the padded message reaches the closest multiple of 512 bits of the original message length and is not less than the original message length. The padded length is composed of a 1 and several 0s, with 1 at the beginning of the padded message, followed by a sufficient number of 0s to reach the required length; first add a '1' bit to the end of the original message; add a sufficient number of '0' bits so that the total length from the beginning of the original message to the end of the padded message is 64 bits less than the nearest multiple of 512 bits; convert the recorded original message length into a 64-bit unsigned integer, which represents the big endian order, that is, the most significant bit is in front; append this 64-bit integer to the end of the padded message, and the length of the entire message is a multiple of 512 bits; Split the processed message into fixed-size blocks, each of which is 512 bits in size; Expand each 512-bit message block into 64 32-bit words; Initialize a loop counter to control the number of iterations of the compression function; According to the value of the loop counter, the corresponding W word is selected for processing. A series of temporary variables are calculated using the selected W word and the current working variable. According to the temporary variables and the logic function, the value of the working variable is updated, the loop counter is incremented, and the next iteration is prepared. After all message blocks are processed, the value of the working variable is used as the final hash value, which includes: According to the current loop counter value t, starting from 0 and ending at 63, the 64 32-bit words W0 to W1 obtained from the expansion 63 Select the corresponding W t ; If this is the first message block to be processed, eight working variables a, b, c, d, e, f, g, h need to be set from the predefined initial hash value, that is, the initial value of SHA-256; If it is not the first message block, the working variable is the hash value after the previous message block is processed; Use the selected W t and the current working variable, calculate a series of temporary variables through the compression function defined in the SHA-256 algorithm, the temporary variables include T1 and T2, and the temporary variables are calculated based on the logic function and the current working variable; Use the calculated temporary variables T1 and T2, as well as the current working variables, to update the values of the working variables according to the rules defined in the SHA-256 algorithm, and increment the loop counter value t by 1 to prepare for the next iteration; if the loop counter value t is less than 64, return to continue processing the next W t ; If t is equal to 64, it means that all words of the current message block have been processed; if there are more message blocks to be processed, the value of the current working variable is used as the initial hash value for the next message block, and the process returns to start processing the next message block; when all message blocks are processed, the value of the current working variable is the final hash value; this hash value is output as the result of the SHA-256 algorithm; in the SHA-256 hash algorithm, each iteration step involves calculating temporary variables T1 and T2, and using the temporary variables and the current working variables to update the values of the working variables. The following is the specific implementation process: Set the initial working variables a, b, c, d, e, f, g, h to the initial hash value of SHA-256 or the hash value after the previous message block is processed. For each message block, perform 64 rounds of iterations, and perform the following operations in each round of iteration: Select W t , according to the loop counter t, from 0 to 63, select the corresponding 32-bit word W t , if it is the first message block, W t is obtained directly from the original message after padding and grouping; otherwise, W t It is obtained through the message expansion algorithm; Using Logical Functions , and the epicycle constant To calculate T1: ; in, is a series of bit rotations and XOR operations, is the selection function, according to The value of select output, using the logical function and To calculate T2: ; in, is another series of bit rotations and XOR operations, is a majority function, according to Return the value of the bit with the highest number of occurrences and update the working variable: ; Increment the loop counter: ; Prepare for the next iteration or complete processing: If t<64, return to continue iterating. If t=64, it means that all words of the current message block have been processed. If there are more message blocks to be processed, the value of the current working variable is used as the initial hash value for the next message block, and return to start processing the next message block. When all message blocks are processed, the value of the current working variable a, b, c, d, e, f, g, h is the final hash value.
5. The method for optimizing parallel operation control of multiple HASH algorithms in an ASIC chip according to claim 4, characterized in that: Split the processed message into fixed-size blocks, each of which is 512 bits in size, consisting of: Generate an initial population. Each individual in the initial population represents a message segmentation method, including the positions of a series of segmentation points. The size of each block at the segmentation point is ≤ 512 bits. Define an evaluation function to evaluate the performance of each split method in the SHA-256 processing process; For each individual, the message is split according to the splitting method it represents, the split message blocks are processed with SHA-256, and the processing performance is evaluated using the evaluation function; According to the processing performance, the corresponding individuals are determined, two individuals are randomly selected, and the positions of some of the individual's segmentation points are exchanged by single-point crossover to generate new individuals; the positions of the individual's segmentation points are randomly changed according to the preset mutation rate to generate a new generation of population, and iterative evolution is performed until the preset evolutionary generation is reached to generate the final segmentation method; The processed message is split into fixed-size blocks according to the final split method.
6. The method for optimizing parallel operation control of multiple HASH algorithms in an ASIC chip according to claim 5, characterized in that: According to the processing performance, the corresponding individuals are determined, two individuals are randomly selected, and the positions of some segmentation points of the individuals are exchanged by single-point crossover to generate new individuals, including: Determine the number of split points in multi-point crossover, which are used to exchange some genes of parent individuals; Generate a random number corresponding to the number of selected split points, and for the two selected parent individuals, determine the split point positions to be exchanged according to the generated random number; At each selected split point, the corresponding split point information of the two parent individuals is exchanged, and after multi-point crossover, two new individuals are obtained.
7. A parallel operation control optimization system for multiple HASH algorithms in an ASIC chip, characterized in that: Applied to the method according to any one of claims 1 to 6, comprising: A creation module is used to create an integration module, and the integration module is used to accommodate multiple SHA-256 processing units; An analysis module, for analyzing a pipeline of a SHA-256 processing unit to determine the number of pipeline stages; A determination module, used for determining a mode selection matrix according to the number of pipeline stages; A scheduling module, for scheduling different SHA-256 processing units and pipeline stages according to a mode selection matrix to obtain an adjusted SHA-256 processing unit; An adding module is used to add a control signal to each adjusted SHA-256 processing unit, and the control signal is used to control the start and stop of the SHA-256 processing unit operation; An optimization module, used to optimize the task allocation scheme between the adjusted SHA-256 processing units, using each particle to represent a task allocation scheme; A definition module is used to define a fitness function for evaluating the performance of each task allocation scheme; The calculation module is used to calculate the fitness of each particle according to the fitness function, update the speed and position of the particle, and iterate until the preset number of iterations is reached to obtain the final task allocation plan.
8. A computing device, characterized in that include: one or more processors; A storage device for storing one or more programs, when the one or more programs are executed by the one or more processors, the one or more processors implement the method as claimed in any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a program, which, when executed by a processor, implements the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Cryptographic system for post quantum cryptographic operation
CN117651949A
Prediction execution-based band state programmable data plane structure and chip
CN118259887A