Vector operation device, method, product and storage medium
By designing a vector computing device containing multiple interleaving units and computing units, dynamically adjusting the data arrangement, the problem of traditional vector processors lacking flexibility when processing computing tasks required for special arrangement is solved, efficient and flexible computing processing is achieved, and overall performance and computing efficiency are improved.
Patent Information
- Application Number
- CN202411944639.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-27
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-12-27
AI Technical Summary
Traditional vector processors lack flexibility in handling computing tasks required for special permutation and need to readjust data arrangement through additional software layers or hardware resources, resulting in reduced computing efficiency and resource utilization.
A vector computing device is designed, including an input computing component, a data computing component and an output component. Through the coordinated work of the input interleaving unit, a complex multiplication and addition unit, an operation interleaving unit, a complex addition unit and an output interleaving unit, the data arrangement is dynamically adjusted to conform to the operation logic of each unit, thereby achieving seamless data fluency.
By dynamically adjusting data arrangement, inefficiency and complex operations caused by fixed data format limitations are avoided, computing processes are simplified, overall performance is optimized, computing efficiency is improved, and dependence on additional software or hardware resources is reduced.
Smart Images

Figure CN119357542B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of vector operation technology, and in particular to a vector operation device, method, product and storage medium. Background Art
[0002] Vector processors are widely used in the field of high-performance computing, especially in tasks that require a large amount of data-level parallel processing, such as artificial intelligence, image processing, wireless communications, etc. Traditional vector processors use arithmetic logic units to perform basic operations between vectors or vectors and scalars, such as addition, subtraction, multiplication, and multiplication conjugate. These operations usually have a unified operand structure and rules, which are suitable for implementing operations through efficient hardware, thereby significantly improving processing speed. In addition, vector processors also support logical operations (such as AND, OR, NOT, shift, etc.) and comparison operations, and have good operational regularity. Traditional vector processors have significant advantages when performing such conventional operations, and can make full use of parallel computing resources to accelerate large-scale data processing.
[0003] However, many efficient applications not only rely on conventional operations, but also need to process certain operations with special operand arrangement requirements, such as small matrix multiplication, determinant operations, etc.; traditional vector processors lack flexible data arrangement and adjustment capabilities and can usually only support fixed input arrangement structures, which are not adaptable enough. When operations with special operand arrangement requirements occur, traditional vector processors must readjust the data arrangement through additional software layers or hardware resources, which significantly increases the complexity and time overhead of the operation, resulting in reduced resource utilization and limited overall performance. Summary of the invention
[0004] The present invention provides a vector operation device, method, product and storage medium to solve the defect that the prior art cannot flexibly support operation tasks with special data arrangement requirements, and requires additional software layers or hardware resources to readjust the data arrangement to complete the operation, which greatly reduces the operation efficiency.
[0005] The present invention provides a vector operation device, the device comprising an input operation component, a data operation component and an output component connected in sequence;
[0006] The input operation component comprises an input interleaving unit and a complex multiplication-addition unit connected in sequence, the input interleaving unit is used to perform interleaving processing on input data to obtain first interleaved data whose arrangement structure conforms to the operation logic of the complex multiplication-addition unit; the complex multiplication-addition unit is used to perform multiplication-addition operation on the first interleaved data sent by the input interleaving unit to obtain an initial operation result;
[0007] The data operation component comprises an operation interleaving unit and a complex number addition unit connected in sequence, the operation interleaving unit is used to perform an interleaving process on the initial operation result sent by the complex number multiplication and addition unit to obtain second interleaved data whose arrangement structure conforms to the operation logic of the complex number addition unit; the complex number addition unit is used to perform an addition operation on the second interleaved data sent by the operation interleaving unit to obtain an intermediate operation result;
[0008] The output component comprises an output interleaving unit, which is used to perform interleaving processing on the intermediate operation results sent by the complex addition unit to obtain vector operation results whose arrangement structure meets the requirements of the output interface.
[0009] According to a vector operation device provided by the present invention, the input interleaving unit includes a first configurator and a plurality of first gates;
[0010] The first configurator is connected to each of the first gates respectively, and the first configurator is used to configure configuration information of each of the first gates;
[0011] Any one of the first selectors is used to perform interleaving processing on multiple input data based on the configuration information configured by the first configurator to obtain first interleaved data.
[0012] According to a vector operation device provided by the present invention, the complex multiplication and addition unit includes multiple multipliers, multiple second selectors and multiple first adders; the first interleaved data includes first interleaved data requiring multiplication operation and first interleaved data not requiring multiplication operation.
[0013] Each of the first gates is correspondingly connected to one of the multipliers and one of the second gates; each of the first gates outputs the first interleaved data requiring multiplication operation to one of the multipliers, and outputs the first interleaved data not requiring multiplication operation to one of the second gates.
[0014] Each of the multipliers is used for performing a multiplication operation on the first interleaved data requiring multiplication operation to obtain a multiplication operation result, and outputting the multiplication operation result to the second selector.
[0015] Each of the second gates is used for interleaving the multiplication result sent by the multiplier and the first interleaved data that does not require multiplication and is sent by the first gate to obtain third interleaved data.
[0016] Each of the first adders is used for performing an addition operation on the third interleaved data sent by the second selector to obtain an initial operation result.
[0017] According to a vector operation device provided by the present invention, the operation interleaving unit includes a second configurator and a plurality of third gates;
[0018] The second configurator is connected to each of the third gates respectively, and the second configurator is used to configure the configuration information of each of the third gates;
[0019] Any of the third selectors is used to perform interleaving processing on the initial operation result sent by the complex multiplication and addition unit based on the configuration information configured by the second configurator to obtain second interleaved data.
[0020] According to a vector operation device provided by the present invention, the complex addition unit includes a plurality of second adders, the third gate is connected to the second adders, and the third gate sends the second interleaved data to the second adders;
[0021] Each of the second adders is used for performing an addition operation on the second interleaved data to obtain an intermediate operation result, and sending the intermediate operation result to the output interleaving unit.
[0022] According to a vector operation device provided by the present invention, the output interleaving unit includes a third configurator and a plurality of fourth gates;
[0023] The third configurator is connected to each of the fourth gates respectively, and the third configurator is used to configure the configuration information of each of the fourth gates;
[0024] Any of the fourth selectors is used to interleave the intermediate operation results sent by the second adder based on the configuration information configured by the third configurator to obtain a vector operation result.
[0025] The present invention also provides a vector operation method, which is applied to the above-mentioned vector operation device, and the method comprises:
[0026] Performing interleaving processing on the input data to obtain first interleaved data; performing multiplication and addition operations on the first interleaved data to obtain an initial operation result;
[0027] Performing interleaving processing on the initial operation result to obtain second interleaved data; performing addition operation on the second interleaved data to obtain an intermediate operation result;
[0028] Interleaving processing is performed on the intermediate operation results to obtain vector operation results.
[0029] According to a vector operation method provided by the present invention, the first interleaved data includes first interleaved data requiring multiplication operation and first interleaved data not requiring multiplication operation, and performing multiplication and addition operation on the first interleaved data to obtain an initial operation result includes:
[0030] Performing a multiplication operation on the first interleaved data requiring a multiplication operation to obtain a multiplication operation result;
[0031] Interleaving the multiplication result and the first interleaved data that does not require multiplication to obtain third interleaved data;
[0032] An addition operation is performed on the third interleaved data to obtain an initial operation result.
[0033] The present invention also provides a computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements any of the above-mentioned vector operation methods.
[0034] The present invention also provides a non-transitory computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the vector operation method described above is implemented.
[0035] The vector operation device provided by the present invention comprises an input operation component, a data operation component and an output component connected in sequence; the input operation component comprises an input interleaving unit and a complex multiplication and addition unit connected in sequence, the input interleaving unit is used to perform interleaving processing on the input data to obtain first interleaved data whose arrangement structure conforms to the operation logic of the complex multiplication and addition unit; the complex multiplication and addition unit is used to perform multiplication and addition operations on the first interleaved data sent by the input interleaving unit to obtain an initial operation result; the data operation component comprises an operation interleaving unit and a complex addition unit connected in sequence, the operation interleaving unit is used to perform interleaving processing on the initial operation result sent by the complex multiplication and addition unit to obtain second interleaved data whose arrangement structure conforms to the operation logic of the complex addition unit; the complex addition unit is used to perform addition operation on the second interleaved data sent by the operation interleaving unit to obtain an intermediate operation result; the output component comprises an output interleaving unit, the output interleaving unit is used to perform interleaving processing on the intermediate operation result sent by the complex addition unit to obtain a vector operation result whose arrangement structure conforms to the requirements of the output interface. The present invention solves the problem of lack of flexibility of traditional vector processors when processing computing tasks with special arrangement requirements. Specifically, the input interleaving unit interleaves the input data to generate first interleaved data, so that the arrangement structure of the first interleaved data can adapt to the computing logic of the complex multiplication and addition unit, thereby ensuring that the data smoothly enters the next step of processing; further, the computing interleaving unit rearranges the initial computing results generated by the complex multiplication and addition unit to generate second interleaved data to meet the computing logic requirements of the complex addition unit; further, the output interleaving unit interleaves the intermediate computing results, and the generated vector computing results meet the requirements of the output interface, thereby ensuring the uniformity and compatibility of the data format; through the above mechanism, the present invention realizes dynamic adjustment of data arrangement, so that the input and output data between each unit are seamlessly connected, avoiding inefficiency and complex operations caused by fixed data format restrictions, and effectively reducing the dependence on additional software or hardware resources for data format conversion, thereby simplifying the computing process, optimizing overall performance, and improving computing efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0037] Figure 1 It is a structural schematic diagram of the vector operation device provided by the present invention.
[0038] Figure 2It is a structural schematic diagram of an input interleaving unit of a vector operation device provided by the present invention.
[0039] Figure 3 This is one of the data flow diagrams of the vector operation device provided by the present invention.
[0040] Figure 4 This is the second data flow diagram of the vector operation device provided by the present invention.
[0041] Figure 5 This is one of the schematic diagrams of the operation data flow of the vector operation device provided by the present invention.
[0042] Figure 6 This is the second schematic diagram of the operation data flow of the vector operation device provided by the present invention.
[0043] Figure 7 It is a schematic diagram of the structure of a computing system using a vector operation device provided by the present invention.
[0044] Figure 8 It is a flow chart of the vector operation method provided by the present invention. DETAILED DESCRIPTION
[0045] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0046] With the widespread application of vector processors in high-performance computing fields such as artificial intelligence, image processing, and wireless communications, parallel processing technology for large-scale data has made significant progress. Traditional vector processors rely on arithmetic logic units to perform efficient basic operations, especially addition, multiplication, and multiplication conjugation between vectors and scalars or between vectors. They can greatly improve the operation speed through a unified operand structure and efficient hardware design. However, with the diversification of application requirements, processors not only need to perform routine operations, but also need to adapt to special tasks (such as small matrix multiplication, determinant calculation, etc.) Special requirements for data arrangement.
[0047] The applicant of this patent has found through a large number of practical application studies that traditional vector processors perform well when processing fixed input data arrangement formats, but when the input data does not conform to the established logical arrangement of the processor, the data format must be readjusted with the help of external software layers or hardware resources. This adjustment process brings significant complexity, increases time overhead, and occupies additional system resources, seriously affecting computing efficiency. In some special tasks, the cost of data format adjustment may even exceed the actual computing resource consumption, becoming a system bottleneck.
[0048] In view of the above problems, the present invention proposes the following embodiments.
[0049] Figure 1 Schematic diagram of the structure of the vector operation device provided by the present invention. Figure 1 As shown, the device includes an input operation component, a data operation component and an output component connected in sequence; it is applied to data operations in artificial intelligence, image processing, wireless communication and other aspects.
[0050] The input operation component includes an input interleaving unit and a complex multiplication-addition unit connected in sequence, the input interleaving unit is used to perform interleaving processing on the input data to obtain first interleaved data whose arrangement structure conforms to the operation logic of the complex multiplication-addition unit; the complex multiplication-addition unit is used to perform multiplication-addition operations on the first interleaved data sent by the input interleaving unit to obtain an initial operation result.
[0051] The data operation component includes an operation interleaving unit and a complex addition unit connected in sequence, the operation interleaving unit is used to interleave the initial operation result sent by the complex multiplication and addition unit to obtain second interleaved data whose arrangement structure conforms to the operation logic of the complex addition unit; the complex addition unit is used to perform addition operation on the second interleaved data sent by the operation interleaving unit to obtain an intermediate operation result.
[0052] The output component comprises an output interleaving unit, which is used to perform interleaving processing on the intermediate operation results sent by the complex addition unit to obtain vector operation results whose arrangement structure meets the requirements of the output interface.
[0053] It should be noted that the vector operation device can be provided with multiple layers of data operation components, each layer of which includes an operation interleaving unit and a complex number addition unit, which are progressively connected to each other to form a modular, recursive operation structure. The multi-layer data operation components can process complex vector calculation tasks layer by layer and support recursive or phased algorithm execution. For example, in complex problems such as matrix multiplication or polynomial solving, the multi-layer structure allows the calculation process to be gradually refined to ensure the accuracy of the calculation results and the integrity of the algorithm logic. An efficient data interaction design is adopted between each layer of components to avoid congestion or synchronization delay problems that may occur in the traditional single-layer structure due to excessive data volume. The multi-layer structure allows data to flow in an orderly manner between different components, improves the fluency of data processing, and ensures the stable operation of the system.
[0054] In addition, the multi-layer design can not only flexibly adapt to complex computing requirements and ensure the integrity and correctness of the algorithm, but also optimize parallel computing efficiency by dynamically allocating computing resources, reduce redundancy and resource waste, and overall improve the device's performance in computing efficiency, adaptability and scalability.
[0055] Here, the input interleaving unit performs interleaving processing on the input data, including but not limited to expanding scalar data into vector data, vector interleaving, etc. This processing can make the arrangement structure of the input data meet the calculation requirements of the subsequent complex multiplication and addition unit, that is, dynamically adjust the data arrangement format, avoiding the complexity of the traditional operator requiring additional resources for adjustment when the data format does not match. The complex multiplication and addition unit receives the first interleaved data and performs a composite operation of complex multiplication and addition to generate an initial operation result. Complex multiplication and addition is a basic but important operation logic, which is widely used in signal processing and matrix operations.
[0056] Here, the operation interleaving unit generates second interleaved data by interleaving the received initial operation result again to adapt the operation logic of the complex addition unit, so as to ensure that the arrangement of the second interleaved data can be smoothly passed to the next step of processing and avoid unnecessary complex data adjustment steps.
[0057] Here, the output interleaving unit generates a vector operation result that meets the requirements of the output interface by receiving the intermediate operation results of the complex addition unit and performing the final interleaving processing. Its function is to ensure that the format and order of the output data can be directly connected to the external system interface, thereby improving the compatibility and consistency of the overall operation.
[0058] The vector operation device provided by the embodiment of the present invention solves the problem that traditional vector processors lack flexibility when processing operation tasks with special arrangement requirements. Specifically, the input interleaving unit interleaves the input data to generate first interleaved data, so that the arrangement structure of the first interleaved data can adapt to the operation logic of the complex multiplication and addition unit, thereby ensuring that the data smoothly enters the next step of processing; further, the operation interleaving unit rearranges the initial operation result generated by the complex multiplication and addition unit to generate second interleaved data to meet the operation logic requirements of the complex addition unit; further, the output interleaving unit interleaves the intermediate operation result, and the generated vector operation result meets the output interface requirements, thereby ensuring the uniformity and compatibility of the data format; through the above mechanism, the present invention realizes dynamic adjustment of data arrangement, so that the input and output data between each unit are seamlessly connected, avoiding inefficiency and complex operations caused by fixed data format restrictions, and effectively reducing the dependence on additional software or hardware resources for data format conversion, thereby simplifying the operation process, optimizing overall performance, and improving operation efficiency.
[0059] like Figure 2 As shown, Figure 2 A schematic diagram of the structure of the input interleaving unit is shown.
[0060] Based on any of the above embodiments, in the device, the input interleaving unit includes a first configurator and a plurality of first gates;
[0061] The first configurator is connected to each of the first gates respectively, and the first configurator is used to configure configuration information of each of the first gates;
[0062] Any one of the first selectors is used to perform interleaving processing on multiple input data based on the configuration information configured by the first configurator to obtain first interleaved data.
[0063] Here, the first configurator is a configuration module responsible for dynamically controlling and setting the configuration parameters of multiple first selectors, that is, responsible for providing configuration information to multiple first selectors. Specifically, the first configurator will decide how to interleave the input data according to the system requirements and the characteristics of the current computing task. It controls the behavior of each first selector to ensure that the input data is interleaved in the expected order and structure; specifically, the existence of the first configurator ensures that the interleaving process of the input data can be flexibly adjusted according to the different requirements of the task, thereby supporting various types of data arrangements and computing logic.
[0064] It should be noted that the first selector is a hardware unit that performs interleaving operations. Each first selector will perform interleaving processing on the input data based on the configuration information transmitted by the first configurator. Interleaving processing refers to rearranging multiple input data in a specific way to form a data format that meets the requirements of subsequent data operation components. Through interleaving processing, data can be more efficiently adapted to subsequent computing units, avoiding efficiency losses caused by mismatched data arrangements in traditional systems. The arrangement of data usually depends on the current computing task. For example, different tasks such as vector operations and matrix operations have different requirements for data arrangement.
[0065] The vector operation device provided by the embodiment of the present invention can realize dynamic interleaving processing of input data and meet the flexible requirements of subsequent data operation components for data arrangement by introducing the collaborative work of a first configurator and a plurality of first selectors; the first configurator provides dynamic configuration information to guide the first selector to alternately or reorganize the data according to specific task requirements, and generate first interleaved data that conforms to the logic of the complex multiplication and addition unit. This process effectively solves the problem of fixed input data arrangement of traditional vector processors when facing diversified tasks, and avoids the complex operation of additional software layers or hardware resources for data format adjustment. By directly optimizing data arrangement at the hardware layer, the adaptability of the operation device is improved, and processing delays and resource overhead are significantly reduced, thereby improving overall operation efficiency and system performance while supporting complex operation tasks.
[0066] like Figure 3 As shown, Figure 3 One of the diagrams showing the data flow of a vector operation device.
[0067] Based on any of the above embodiments, in the device, the complex multiplication and addition unit includes a plurality of multipliers, a plurality of second selectors and a plurality of first adders; the first interleaved data includes first interleaved data requiring multiplication operation and first interleaved data not requiring multiplication operation;
[0068] Each of the first gates is connected to one of the multipliers and one of the second gates; each of the first gates outputs the first interleaved data requiring multiplication operation to one of the multipliers, and outputs the first interleaved data not requiring multiplication operation to one of the second gates;
[0069] Each of the multipliers is used for performing a multiplication operation on the first interleaved data requiring a multiplication operation to obtain a multiplication operation result, and outputting the multiplication operation result to the second gate;
[0070] Each of the second gates is used for interleaving the multiplication result sent by the multiplier and the first interleaved data sent by the first gate without multiplication operation to obtain third interleaved data;
[0071] Each of the first adders is used for performing an addition operation on the third interleaved data sent by the second selector to obtain an initial operation result.
[0072] Here, the first interleaved data includes the first interleaved data that needs multiplication operation and the first interleaved data that does not need multiplication operation. The first interleaved data that needs multiplication operation refers to the data that needs to be processed by a multiplier, such as the product calculation part involved in a specific operation task, such as the real part or imaginary part in the complex product calculation. The first interleaved data that does not need multiplication operation refers to the data that does not need to be processed by a multiplier and can directly enter the subsequent interleaving or addition steps, such as constants or directly transmitted data. The third interleaved data refers to the integrated data stream output by the second selector, which ensures the consistency of the data structure and provides input for the subsequent addition operation.
[0073] Here, the multiplier refers to a hardware module that performs complex multiplication on input data, and converts the part of the first interleaved data that needs multiplication into a multiplication result output. The first adder is a hardware module that performs addition operation on the third interleaved data, generates an initial operation result as the output of the entire input operation component. The second selector refers to a hardware module that integrates and interleaves the multiplication result and the first interleaved data that does not need multiplication operation, and the output data structure meets the logic requirements of the subsequent first adder.
[0074] In one embodiment, the first gate divides the first interleaved data according to the data type, and the first interleaved data that requires multiplication operation is sent to the multiplier for calculation, and the first interleaved data that does not require multiplication operation is directly sent to the second gate; further, the multiplier performs corresponding calculations on the first interleaved data that requires multiplication operation, generates a multiplication result and passes it to the second gate; further, the second gate receives the multiplication result from the multiplier and the directly transmitted first interleaved data that does not require multiplication operation, interleaves these two parts of data, generates third interleaved data with the same format, and ensures that the data meets the input requirements of the subsequent adder; finally, the first adder receives the third interleaved data, completes the addition operation, generates an initial operation result, and provides output data that meets the logic requirements for the subsequent computing unit.
[0075] Exemplarily, a multiplier can implement complex number multiplication and real number multiplication; complex number multiplication is accomplished by decomposing a complex number into a real part and an imaginary part, using a combination of multiple multiplication operations and addition and subtraction. Real number multiplication, as a special case of complex number multiplication, only needs to operate the real part, and the imaginary part does not participate in the operation. The multiplier can support both operations. The adder can implement complex number addition and real number addition. Complex number addition directly adds the real part and the imaginary part respectively. Real number addition, as a special case of complex number addition, only adds the real part, and the imaginary part does not participate in the operation. Through simple data selection and configuration, seamless switching between the two operation modes can be achieved.
[0076] Specifically, by integrating the calculation functions of complex numbers and real numbers in the complex multiplication and addition unit, the versatility and computing efficiency of the system are significantly improved. On the one hand, the multiplier and the first adder can support the multiplication and addition operations of complex numbers and real numbers respectively, realizing the unified processing of multiple computing requirements. On the other hand, the design reduces the redundancy of hardware resources, saves costs through hardware reuse, simplifies the device structure, and at the same time reduces the data transmission delay during the operation process, thereby improving the overall computing performance. In addition, the system has good scalability and compatibility, can adapt to a wider range of application scenarios, and provides efficient and flexible solutions for complex operations such as complex signal processing and vector calculations. Therefore, the vector operation device provided by the present invention has significant comprehensive advantages in function, performance and resource optimization.
[0077] Exemplarily, the number of the first selector, the multiplier, the second selector and the first adder is allowed to be flexibly increased or decreased according to specific computing requirements. In computing tasks, the amount and complexity of input data are usually not fixed. For example, when the data stream is large or multiple complex operations need to be processed in parallel, the computing load can be shared by adding multiple modules of the same type to achieve higher throughput. For simple operations, only fewer units need to be configured to meet the requirements, saving hardware resources. Through the configurable design principle, the number of the first selector, the multiplier, the second selector and the first adder can be flexibly adjusted according to the parallelism requirements of the actual task, thereby taking into account performance, resource utilization and hardware cost.
[0078] For example, if the input is complex numbers. If the number of complex numbers to be output is N, then a total of complex multiply-add units, total multipliers and adders, a total of Layer data operation components, the first layer requires complex addition units, the last layer requires complex addition units, and a total of adders are required: An adder.
[0079] Each multiplier independently processes a part of the first interleaved data that needs multiplication, allowing multiple multiplication operations to be performed simultaneously, making full use of the hardware parallel computing capabilities and shortening the operation time; the first interleaved data that does not need multiplication directly enters the second selector, further reducing the data transmission delay; the second selector interleaves the multiplication results from the multiplier and the first interleaved data that does not need to be calculated, ensuring that the structure of the output third interleaved data matches the needs of the subsequent calculation unit, unifying the data format, and the parallel calculation of the first adder further improves the operation efficiency. It can be seen that the device can flexibly respond to a variety of computing tasks, while significantly improving the computing efficiency and hardware resource utilization, reducing system complexity and optimizing overall performance.
[0080] Based on any of the above embodiments, in the device, the operation interleaving unit includes a second configurator and a plurality of third gates;
[0081] The second configurator is connected to each of the third gates respectively, and the second configurator is used to configure the configuration information of each of the third gates;
[0082] Any of the third selectors is used to perform interleaving processing on the initial operation result sent by the complex multiplication and addition unit based on the configuration information configured by the second configurator to obtain second interleaved data.
[0083] Here, the second configurator is responsible for providing specific configuration information to each third selector, that is, determining the interleaving processing rules, including how to rearrange the data, how to interleave, and the specific format of the output data. Its flexibility allows adaptation to the needs of different computing tasks. The third selector is used to perform a specific interleaving operation, and rearranges the initial computing result into a format that conforms to the subsequent computing logic according to the configuration information of the second configurator.
[0084] In one embodiment, the third selector receives configuration information sent by the second configurator, dynamically selects, crosses, reorganizes or rearranges the initial operation results sent by the complex multiplication and addition units according to the configuration information, and outputs the interleaved data as the second interleaved data.
[0085] It should be noted that the design purpose of the operation interleaving unit is to adjust the initial operation result sent by the complex multiplication and addition unit so that it meets the input requirements of the subsequent complex addition unit.
[0086] Exemplarily, the arrangement of the initial operation results is usually fixed, but the subsequent data operation components may require a specific input arrangement structure. The operation interleaving unit can adjust the data format in real time through the collaboration of the second configurator and the third selector, so as to achieve seamless connection with the subsequent units and avoid data incompatibility problems. In traditional solutions, format adjustment often needs to rely on additional software or hardware resources, which not only increases the complexity, but also brings additional time delays and energy consumption. The hardware implementation scheme of this step integrates the interleaving process into the vector operation device, eliminating the additional time delay.
[0087] In the vector operation device provided by the embodiment of the present invention, the operation interleaving unit uses the second configurator to generate dynamic configuration information, and processes the initial operation results in parallel through multiple third selectors to complete the data interleaving adjustment; it not only adapts to the subsequent operation logic requirements and avoids the calculation bottleneck caused by incompatible data arrangement, but also effectively reduces the additional software processing and hardware resource consumption. The parallel processing capability of multiple third selectors significantly improves the data flow efficiency and ensures the optimization of the overall operation performance.
[0088] like Figure 4 As shown, Figure 4 The second diagram shows the data flow of the vector operation device.
[0089] Based on any of the above embodiments, in the device, the complex addition unit includes a plurality of second adders, the third gate is connected to the second adders, and the third gate sends the second interleaved data to the second adders;
[0090] Each of the second adders is used for performing an addition operation on the second interleaved data to obtain an intermediate operation result, and sending the intermediate operation result to the output interleaving unit.
[0091] In one embodiment, a plurality of second adders receive second interleaved data sent by a third selector, wherein the second interleaved data includes two parts of data, one requiring addition calculation and the other not requiring addition calculation; the second adder performs an addition operation on the second interleaved data requiring addition calculation to obtain a first addition operation result; the second adder performs an addition of 0 operation on the second interleaved data not requiring addition calculation to obtain a second addition operation result; an intermediate operation result is determined based on the first addition operation result and the second addition operation result; further, the second adder sends the intermediate operation result to an output interleaving unit.
[0092] Exemplarily, the second interleaved data entering the second adder is a part that needs to be added, such as a real part and an imaginary part of a complex number.
[0093] The vector operation device provided by the embodiment of the present invention realizes efficient parallel processing of complex addition operations by designing multiple second adders in the complex addition unit and flexibly distributing the second interleaved data to the second adders through the third selector. After the second adder performs addition operation on the input data, the intermediate operation result is sent to the output interleaving unit, which lays the foundation for subsequent interleaving processing and result output. In addition, the structure of the complex addition unit ensures the flexibility and reliability of the operation process; the multi-adder architecture enhances the scalability of the device and can adapt to operation requirements of different scales and complexities.
[0094] Based on any of the above embodiments, in the device, the output interleaving unit includes a third configurator and a plurality of fourth gates;
[0095] The third configurator is connected to each of the fourth gates respectively, and the third configurator is used to configure the configuration information of each of the fourth gates;
[0096] Any of the fourth selectors is used to interleave the intermediate operation results sent by the second adder based on the configuration information configured by the third configurator to obtain a vector operation result.
[0097] Here, the third configurator provides configuration information for the fourth selector to ensure that each fourth selector can select and interleave data in an expected manner when processing data; through configuration, the third selector can decide how to interleave intermediate operation results to meet the final output format requirements.
[0098] Here, the fourth selector interleaves the intermediate operation result of the second adder according to the configuration information provided by the third configurator. The purpose of the interleaving operation is to reorganize the data in a specific order or rule to obtain a structure that meets the output requirements. During the entire operation process, the role of the fourth selector is to ensure that the final vector operation result meets the output interface requirements.
[0099] The vector operation device provided in the embodiment of the present invention realizes efficient and flexible interleaving processing through the effective cooperation of the third configurator and the fourth selector, so that the vector operation device can accurately and quickly generate vector operation results that meet the output requirements. This process optimizes the calculation process, improves system performance, and can adapt to complex data processing requirements.
[0100] For example, Figure 5 As shown, a matrix multiplication is calculated using the vector operation device provided by the present invention; the first operand is , the second operand is , the calculation result is:
[0101] ;
[0102] in, Figure 5 The multiple multiplier-adders in represent multiple multipliers and multiple first adders.
[0103] For example, Figure 6 As shown, the vector operation device provided by the present invention is used to calculate the operation of two sets of complex vectors; the first operand is {rs1[0], rs1[1], rs1[2], rs1[3], rs1[4], rs1[5], rs1[6],rs1[7]}; the second operand is {rs2[0], rs2[1], rs2[2], rs2[3]}; the operation result operand is {rd[0],rd[1]};
[0104] rd[0]=rs1[0] rs2[0]+rs1[1] rs2[1]+rs1[2] rs2[2]+rs1[3] rs2[3];
[0105] rd[1]=rs1[4] rs2[0]+rs1[5] rs2[1]+rs1[6] rs2[2]+rs1[7] rs2[3];
[0106] in, Figure 6 The multiple multiplier-adders in represent multiple multipliers and multiple first adders. Figure 6 The method of using two-layer data operation components is demonstrated, each layer of data operation components includes an operation interleaving unit and a complex addition unit; the number of adders in the complex addition unit can be configured according to needs.
[0107] For example, Figure 7 As shown, the vector operation device provided in this embodiment can also be used in a computing system; the computing system includes a program controller, a decoder, a register file, a program memory, a storage management unit, a data memory and the vector operation device provided in this embodiment.
[0108] Specifically, the program controller is responsible for managing the instruction count and loading instructions from the program memory one by one. The loaded instructions are parsed into control signals by the decoder and processed differently according to the instruction type: for storage instructions, the control signal guides the storage management unit to complete the data interaction between the data storage and the register file; for the operation unit configuration instruction, the configuration signal is sent to the vector operation device to dynamically configure its interleaving unit to meet specific operation requirements; for the operation instruction, the vector operation device reads data from the register file according to the control information, completes the calculation and writes the result back to the register file. This process not only realizes flexible and efficient operation processing through clear instruction allocation and module collaboration, but also ensures the accuracy of data interaction and system scalability.
[0109] like Figure 8 As shown, Figure 8 A flow chart of a vector operation method is shown, which is applied to the above-mentioned vector operation device, and the method includes:
[0110] Step 810, interleave the input data to obtain first interleaved data; perform multiplication and addition operations on the first interleaved data to obtain an initial operation result.
[0111] Step 820, performing interleaving processing on the initial operation result to obtain second interleaved data; performing addition operation on the second interleaved data to obtain an intermediate operation result.
[0112] Step 830, interleaving the intermediate operation results to obtain vector operation results.
[0113] In the vector operation method provided by the embodiment of the present invention, the input data is first interleaved to obtain the first interleaved data. Subsequently, the first interleaved data is multiplied and added to obtain the initial operation result. The interleaving process ensures the effective arrangement of the data, thereby improving the accuracy and efficiency of the multiplication and addition operation; further, the initial operation result is interleaved again to obtain the second interleaved data, and then the second interleaved data is added to obtain the intermediate operation result. This process ensures that the data is arranged in the correct order and optimizes the addition operation; further, the intermediate operation result is interleaved again to finally generate a vector operation result that meets the requirements of the output interface. The data order is adjusted by interleaving to ensure that the final result is both accurate and compatible with the external system. Through multiple interleaving processes and precise operation steps of the input data, the vector operation can be completed more efficiently and accurately. Each interleaving not only ensures the correctness of the operation logic, but also optimizes the data flow between each step, avoids redundant steps in the calculation, and ensures the efficiency and stability of the operation result.
[0114] Based on the vector operation method provided in the above embodiment, the first interleaved data includes first interleaved data requiring multiplication operation and first interleaved data not requiring multiplication operation, and performing multiplication and addition operation on the first interleaved data to obtain an initial operation result includes:
[0115] Performing a multiplication operation on the first interleaved data requiring a multiplication operation to obtain a multiplication operation result;
[0116] Interleaving the multiplication result and the first interleaved data that does not require multiplication to obtain third interleaved data;
[0117] An addition operation is performed on the third interleaved data to obtain an initial operation result.
[0118] In the vector operation method provided by the embodiment of the present invention, the first interleaved data is divided into a part that needs multiplication operation and a part that does not need multiplication operation. First, the first interleaved data that needs multiplication operation is multiplied to obtain a multiplication result; then, the multiplication result is interleaved with the first interleaved data that does not need multiplication operation to obtain third interleaved data. This step ensures the correct arrangement of data and the accuracy of subsequent addition operations through interleaving operations; finally, the third interleaved data is added to obtain an initial operation result. This method avoids unnecessary multiplication operations and saves computing resources. At the same time, the interleaving process can optimize the calculation process and improve the computing efficiency while ensuring the correct data order.
[0119] On the other hand, the present invention also provides a computer program product, which includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the vector operation method provided by the above methods, which includes: interleaving the input data to obtain first interleaved data; performing multiplication and addition operations on the first interleaved data to obtain an initial operation result; interleaving the initial operation result to obtain second interleaved data; performing addition operations on the second interleaved data to obtain an intermediate operation result; and interleaving the intermediate operation result to obtain a vector operation result.
[0120] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is implemented to execute the vector operation method provided by the above-mentioned methods. The method includes: interleaving the input data to obtain first interleaved data; performing multiplication and addition operations on the first interleaved data to obtain an initial operation result; interleaving the initial operation result to obtain second interleaved data; performing addition operations on the second interleaved data to obtain an intermediate operation result; interleaving the intermediate operation result to obtain a vector operation result.
[0121] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A vector operation device, characterized in that: The device comprises an input operation component, a data operation component and an output component connected in sequence; The input operation component comprises an input interleaving unit and a complex multiplication-addition unit connected in sequence, the input interleaving unit is used to perform interleaving processing on input data to obtain first interleaved data whose arrangement structure conforms to the operation logic of the complex multiplication-addition unit; the complex multiplication-addition unit is used to perform multiplication-addition operation on the first interleaved data sent by the input interleaving unit to obtain an initial operation result; The data operation component comprises an operation interleaving unit and a complex number addition unit connected in sequence, the operation interleaving unit is used to perform an interleaving process on the initial operation result sent by the complex number multiplication and addition unit to obtain second interleaved data whose arrangement structure conforms to the operation logic of the complex number addition unit; the complex number addition unit is used to perform an addition operation on the second interleaved data sent by the operation interleaving unit to obtain an intermediate operation result; The output component comprises an output interleaving unit, and the output interleaving unit is used to perform interleaving processing on the intermediate operation results sent by the complex addition unit to obtain a vector operation result whose arrangement structure meets the requirements of the output interface; The input interleaving unit comprises a first configurator and a plurality of first gates; the first configurator is connected to each of the first gates respectively, and the first configurator is used to configure configuration information of each of the first gates; any of the first gates is used to perform interleaving processing on a plurality of input data based on the configuration information configured by the first configurator to obtain first interleaved data; The complex multiplication and addition unit includes multiple multipliers, multiple second gates and multiple first adders; the first interleaved data includes first interleaved data that needs multiplication operation and first interleaved data that does not need multiplication operation; each first gate is correspondingly connected to one of the multipliers and one of the second gates; each of the first gates outputs the first interleaved data that needs multiplication operation to one of the multipliers, and outputs the first interleaved data that does not need multiplication operation to one of the second gates; each of the multipliers is used to multiply the first interleaved data that needs multiplication operation to obtain a multiplication result, and output the multiplication result to the second gate; each of the second gates is used to interleave the multiplication result sent by the multiplier and the first interleaved data that does not need multiplication operation sent by the first gate to obtain third interleaved data; each of the first adders is used to add the third interleaved data sent by the second gate to obtain an initial operation result.
2. The vector operation device according to claim 1, characterized in that: The operation interleaving unit includes a second configurator and a plurality of third gates; The second configurator is connected to each of the third gates respectively, and the second configurator is used to configure the configuration information of each of the third gates; Any of the third selectors is used to perform interleaving processing on the initial operation result sent by the complex multiplication and addition unit based on the configuration information configured by the second configurator to obtain second interleaved data.
3. The vector operation device according to claim 2, characterized in that: The complex adding unit comprises a plurality of second adders, the third gate is connected to the second adders, and the third gate sends the second interleaved data to the second adders; Each of the second adders is used for performing an addition operation on the second interleaved data to obtain an intermediate operation result, and sending the intermediate operation result to the output interleaving unit.
4. The vector operation device according to claim 3, characterized in that: The output interleaving unit includes a third configurator and a plurality of fourth gates; The third configurator is connected to each of the fourth gates respectively, and the third configurator is used to configure the configuration information of each of the fourth gates; Any of the fourth selectors is used to interleave the intermediate operation results sent by the second adder based on the configuration information configured by the third configurator to obtain a vector operation result.
5. A vector operation method, characterized in that: Applied to the vector operation device according to any one of claims 1 to 4, the method comprises: Performing interleaving processing on the input data to obtain first interleaved data; performing multiplication and addition operations on the first interleaved data to obtain an initial operation result; Performing interleaving processing on the initial operation result to obtain second interleaved data; performing addition operation on the second interleaved data to obtain an intermediate operation result; Interleaving processing is performed on the intermediate operation results to obtain vector operation results.
6. The vector operation method according to claim 5, characterized in that: The first interleaved data includes first interleaved data requiring multiplication operation and first interleaved data not requiring multiplication operation, and performing multiplication and addition operation on the first interleaved data to obtain an initial operation result includes: Performing a multiplication operation on the first interleaved data requiring a multiplication operation to obtain a multiplication operation result; Interleaving the multiplication result and the first interleaved data that does not require multiplication to obtain third interleaved data; An addition operation is performed on the third interleaved data to obtain an initial operation result.
7. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the vector operation method according to any one of claims 5 to 6 is implemented.
8. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the vector operation method according to any one of claims 5 to 6 is implemented.
Citation Information
Patent Citations
Mixed-radix FFT (Fast Fourier Transform) processor
CN106372034A