Data processing system

CN120821501BActive Publication Date: 2026-09-25SHANGHAI QI ZHI INSTITUTE
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510923063.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-04
Publication Date
2026-09-25
Estimated Expiration
2045-07-04

AI Technical Summary

Technical Problem

[0004]通过Goldilocks域等特殊素数模数进行模乘运算,未能充分利用现代处理器架构的并行计算能力,尤其是高级向量扩展指令集(如AVX512)所提供的宽位寄存器和并行处理能力,内存访问较高,进而导致内存资源的浪费

Benefits of technology

[0010]本公开的上述各个实施例中具有如下有益效果:通过本公开的一些实施例的数据处理系统,避免了进行模乘运算时内存资源的浪费。具体来说,造成内存资源的浪费的原因在于:通过Goldilocks域等特殊素数模数进行模乘运算,未能充分利用现代处理器架构的并行计算能力,尤其是高级向量扩展指令集(如AVX512)所提供的宽位寄存器和并行处理能力,内存访问较高,进而导致内存资源的浪费。基于此,本公开的一些实施例的数据处理系统,首先,获取第一模乘数据和第二模乘数据。由此,可以确定进行模乘运算的数据。其次,初始化寄存器组,以及基于上述寄存器组,分别对上述第一模乘数据和上述第二模乘数据进行重排处理,以生成第一高位数据、第一低位数据、第二高位数据和第二低位数据。由此,可以将数据按照高位和低位进行重排。然后,基于上述第一高位数据、上述第一低位数据、上述第二高位数据和上述第二低位数据,生成模乘中间数据组。由此,可以通过并行算法准确计算子寄存器中四种部分的乘积,为后续组合成完整的64位乘法结果奠定基础,也通过并行算法,充分利用了现代处理器架构的并行计算能力,减少了内存访问,从而避免了内存资源的浪费。之后,基于上述模乘中间数据组,生成至少一个进位标志,以及将上述至少一个进位标志存储至向量掩码寄存器中。由此,可以确定模乘过程中是否发生进位。最后,基于向量掩码寄存器存储的至少一个进位标志,对上述模乘中间数据组中的各个模乘中间数据进行进位处理,以生成全乘法结果;将上述全乘法结果存储至寄存器组包括的寄存器中。由此,通过并行的方式,完成寄存器中的模乘运算,避免了进行模乘运算时内存资源的浪费。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120821501B_ABST
    Figure CN120821501B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure disclose a data processing system. A specific implementation of the system includes: a data processing chip configured to acquire first and second modulo multiplication data stored by an edge terminal included in the data processing system; initialize a register group, and rearrange the first and second modulo multiplication data based on the register group respectively; generate a modulo multiplication intermediate data group; generate at least one carry flag based on the modulo multiplication intermediate data group, and store the at least one carry flag in a vector mask register included in the register group; perform carry processing on each modulo multiplication intermediate data in the modulo multiplication intermediate data group based on the at least one carry flag stored in the vector mask register to generate a full multiplication result; and store the full multiplication result in a register included in the register group. The implementation avoids waste of memory resources when performing modulo multiplication operations.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments disclosed herein relate to the field of computer technology, and more specifically to data processing systems. Background Technology

[0002] Modular multiplication is a core operation in modern cryptography (including post-quantum cryptography), and its efficiency directly affects the overall performance of various algorithms. Computation over finite fields, especially operations based on special prime modulo operations, forms the basis of many advanced applications. Currently, the common approach to modular multiplication is to perform it using special prime modulo operations such as Goldilocks fields.

[0003] However, when performing modular multiplication using the above method, the following technical problems often arise:

[0004] Modular multiplication using special prime moduloes such as Goldilocks fields fails to fully utilize the parallel computing capabilities of modern processor architectures, especially the wide-bit registers and parallel processing capabilities provided by advanced vector extension instruction sets (such as AVX512). This results in high memory access rates and a waste of memory resources.

[0005] The information disclosed in this background section is only intended to enhance the understanding of the background of the inventive concept, and therefore may contain information that does not constitute prior art known to those skilled in the art. Summary of the Invention

[0006] The summary portion of this disclosure is intended to provide a brief overview of the concepts, which will be described in detail in the detailed description portion. This summary portion is not intended to identify key or essential features of the claimed technical solutions, nor is it intended to limit the scope of the claimed technical solutions.

[0007] Some embodiments of this disclosure propose a data processing system to solve one or more of the technical problems mentioned in the background section above.

[0008] In a first aspect, some embodiments of this disclosure provide a data processing system, comprising: a data processing chip, an edge terminal, and a register set, wherein the data processing chip is configured to acquire first modular multiplication data and second modular multiplication data stored in the edge terminal included in the data processing system, wherein the first modular multiplication data and the second modular multiplication data are stored in different registers respectively; the data processing chip is configured to initialize the register set, and based on the register set, to perform rearrangement processing on the first modular multiplication data and the second modular multiplication data respectively to generate first high-order data, first low-order data, second high-order data, and second low-order data; the data processing... The chip is configured to generate a modular multiplication intermediate data set based on the first high-order data, the first low-order data, the second high-order data, and the second low-order data; the data processing chip is configured to generate at least one carry flag based on the modular multiplication intermediate data set, and to store the at least one carry flag in a vector mask register included in the register set; the data processing chip is configured to perform carry processing on each modular multiplication intermediate data in the modular multiplication intermediate data set based on the at least one carry flag stored in the vector mask register to generate a full multiplication result; the data processing chip is configured to store the full multiplication result in a register included in the register set.

[0009] Secondly, some embodiments of this disclosure provide a data processing method, the method comprising: acquiring first modular multiplication data and second modular multiplication data stored in an edge terminal of a data processing system, wherein the first modular multiplication data and the second modular multiplication data are stored in different registers; initializing a register group, and rearranging the first modular multiplication data and the second modular multiplication data based on the register group to generate first high-order data, first low-order data, second high-order data and second low-order data; generating a modular multiplication intermediate data group based on the first high-order data, the first low-order data, the second high-order data and the second low-order data; generating at least one carry flag based on the modular multiplication intermediate data group, and storing the at least one carry flag in a vector mask register; performing carry processing on each modular multiplication intermediate data in the modular multiplication intermediate data group based on the at least one carry flag stored in the vector mask register to generate a full multiplication result; and storing the full multiplication result in a register included in the register group.

[0010] The various embodiments of this disclosure have the following beneficial effects: the data processing system of some embodiments of this disclosure avoids the waste of memory resources when performing modular multiplication. Specifically, the reason for the waste of memory resources is that modular multiplication using special prime modulo operations such as Goldilocks fields fails to fully utilize the parallel computing capabilities of modern processor architectures, especially the wide-bit registers and parallel processing capabilities provided by advanced vector extension instruction sets (such as AVX512), resulting in high memory access and thus wasting memory resources. Based on this, the data processing system of some embodiments of this disclosure first acquires first modular multiplication data and second modular multiplication data. This determines the data to be used for modular multiplication. Second, it initializes a register set and, based on the register set, rearranges the first modular multiplication data and the second modular multiplication data to generate first high-order data, first low-order data, second high-order data, and second low-order data. This rearranges the data according to the high and low orders. Then, based on the first high-order data, the first low-order data, the second high-order data, and the second low-order data, an intermediate modular multiplication data set is generated. Therefore, the product of the four parts in the sub-register can be accurately calculated using a parallel algorithm, laying the foundation for the subsequent combination into a complete 64-bit multiplication result. This parallel algorithm also fully utilizes the parallel computing capabilities of modern processor architectures, reducing memory accesses and thus avoiding wasted memory resources. Next, based on the aforementioned modular multiplication intermediate data set, at least one carry flag is generated and stored in a vector mask register. This determines whether a carry occurs during the modular multiplication process. Finally, based on the at least one carry flag stored in the vector mask register, carry processing is performed on each of the intermediate modular multiplication data in the aforementioned intermediate data set to generate the full multiplication result; this full multiplication result is then stored in the registers included in the register group. Thus, the modular multiplication operation in the registers is completed in parallel, avoiding wasted memory resources during the modular multiplication operation. Attached Figure Description

[0011] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and elements are not necessarily drawn to scale.

[0012] Figure 1 This is a system architecture diagram based on some embodiments of the data processing system disclosed herein;

[0013] Figure 2 This is a flowchart of some embodiments of the data processing method according to the present disclosure. Detailed Implementation

[0014] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0015] It should also be noted that, for ease of description, only the parts relevant to the invention are shown in the accompanying drawings. Unless otherwise specified, the embodiments and features described in this disclosure can be combined with each other.

[0016] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.

[0017] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0018] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.

[0019] This disclosure will now be described in detail with reference to the accompanying drawings and embodiments.

[0020] Figure 1 A schematic diagram of the structure of some embodiments of the data processing system according to this disclosure is shown.

[0021] like Figure 1 As shown, the data processing system of this disclosure includes: a data processing chip 101, an edge terminal 102, and a register group 103. The data processing chip can be a chip used for Montgomery modular multiplication. Here, the data processing chip 101 can be a processor chip equipped with an AVX-512 instruction set architecture. As an example, the data processing chip 101 can be a Xeon 8490H processor chip. The edge terminal 102 can be a user terminal connected to the data processing chip 101 via a wired or wireless connection. The registers in the register group 103 can be memories used for storing binary data. The registers in the register group 103 can be shift registers.

[0022] The following is for reference. Figure 2The diagram illustrates a flow 200 of some embodiments of a data processing system according to the present disclosure. The data processing system includes the following steps:

[0023] Step 201: The data processing chip acquires the first and second modular multiplication data stored in the edge terminal of the data processing system.

[0024] In some embodiments, the data processing chip acquires first and second modular multiplication data stored in the edge terminal of the data processing system. The first and second modular multiplication data are stored in different registers. Both the first and second modular multiplication data can be 64-bit data. The registers can be AVX-512 registers, enabling the multiplication of eight pairs of 64-bit integers. The data processing chip can also receive the first and second modular multiplication data sent by the edge terminal.

[0025] Step 202: The data processing chip initializes the register group, and based on the register group, rearranges the first modular multiplication data and the second modular multiplication data to generate the first high-order data, the first low-order data, the second high-order data, and the second low-order data.

[0026] In some embodiments, the data processing chip can initialize a register group and, based on the register group, rearrange the first modular multiplication data and the second modular multiplication data to generate first high-order data, first low-order data, second high-order data, and second low-order data.

[0027] In practice, the data processing chip can be further configured to initialize the register set through the following steps, and based on the register set, rearrange the first modular multiplication data and the second modular multiplication data to generate the first high-order data, the first low-order data, the second high-order data, and the second low-order data:

[0028] The first step is to create a preset number of sub-registers as a register set. This preset number can be created using the `vpshufd` instruction. The preset number can be a pre-defined number of sub-registers to be generated. For example, the preset number could be 4.

[0029] The second step is to initialize the parameters of the above register group to generate the initialized sub-registers, thus obtaining the initialized register group.

[0030] The third step is to rearrange the first modular multiplication data to generate the first high-order data and the first low-order data. Here, the 64-bit first modular multiplication data can be rearranged into the first high-order data (high 32 bits) and the first low-order data (low 32 bits).

[0031] The fourth step is to rearrange the second modular multiplication data to generate the second high-order data and the second low-order data. Here, the 64-bit second modular multiplication data can be rearranged into the second high-order data (high 32 bits) and the second low-order data (low 32 bits).

[0032] The fifth step is to store the first high-order data, the first low-order data, the second high-order data, and the second low-order data into their respective sub-registers.

[0033] Step 203: The data processing chip generates a modular multiplication intermediate data set based on the first high-order data, the first low-order data, the second high-order data, and the second low-order data.

[0034] In some embodiments, the data processing chip can generate a modular multiplication intermediate data set based on the first high-order data, the first low-order data, the second high-order data, and the second low-order data.

[0035] In practice, the data processing chip can be further configured to generate a modular multiplication intermediate data set based on the first high-order data, the first low-order data, the second high-order data, and the second low-order data through the following steps:

[0036] The first step is to determine the first intermediate data and the second intermediate data by multiplying the first high-order data, the second high-order data, and the second low-order data.

[0037] The second step is to determine the third and fourth intermediate data by multiplying the first low-order data with the second high-order data and the second low-order data respectively.

[0038] The third step is to combine the first intermediate data, the second intermediate data, the third intermediate data, and the fourth intermediate data into a modular multiplication intermediate data group.

[0039] Step 204: Based on the modular multiplication intermediate data set, generate at least one carry flag and store at least one carry flag in the vector mask register included in the register set.

[0040] In some embodiments, the data processing chip can generate at least one carry flag based on the aforementioned modular multiplication intermediate data set, and store the at least one carry flag in a vector mask register included in the register set. The vector mask register is a mask register that stores vector data and is used to control whether a certain element in the register is masked, thereby preventing its use.

[0041] In the process of adopting technical solutions to address the aforementioned technical problems, the following technical issues often arise: when performing full multiplication operations in Montgomery modular multiplication using traditional algorithms, it is necessary to perform operations on the higher-digit data step by step, which consumes a considerable amount of time for the full multiplication operation. Considering the above technical problems and the current state of available technology, the following solution can be adopted.

[0042] In some alternative implementations of certain embodiments, the data processing chip may be further configured to generate at least one carry flag based on the aforementioned modular multiplication intermediate data set, and to store the at least one carry flag in a vector mask register via the following steps:

[0043] The first step is to determine the sum of the second and third intermediate data included in the above modular multiplication intermediate data group as the first modular multiplication cross value.

[0044] The second step is to determine the data size relationship between the first modular cross value and the third intermediate data.

[0045] Third, in response to the first modular multiplication cross value being less than the third intermediate data, a first carry flag is generated. This first carry flag indicates that a carry occurred when the first modular multiplication cross value was generated.

[0046] The fourth step is to decompose the first modular cross value to generate the first decomposed value and the second decomposed value.

[0047] The fifth step involves shifting the second decomposition value to generate a shifted decomposition value. This shifting process can involve shifting the second decomposition value 32 bits to the left.

[0048] The sixth step is to determine the second modular cross value by summing the above fourth intermediate data with the above moved decomposition value.

[0049] Step 7: Determine the data size relationship between the second modular cross value and the fourth intermediate data.

[0050] Step 8: In response to the second modular multiplication cross value being less than the fourth intermediate data, a second carry flag is generated. This second carry flag indicates that a carry occurred when the second modular multiplication cross value was generated.

[0051] The ninth step is to store the first carry flag and the second carry flag into different vector mask registers.

[0052] The first to ninth steps and related content described above, as an inventive point of this disclosure, solve the technical problem that "when performing full multiplication in Montgomery modular multiplication using traditional algorithms, it is necessary to perform operations on the high-digit data step by step, which consumes a long time for the full multiplication operation." The factors that cause the long time required for full multiplication operations are often as follows: when performing full multiplication in Montgomery modular multiplication using traditional algorithms, it is necessary to perform operations on the high-digit data step by step, which consumes a long time for the full multiplication operation. If these factors are resolved, the time required for full multiplication operations can be reduced. To achieve this effect, firstly, the sum of the second and third intermediate data included in the above modular multiplication intermediate data group is determined as the first modular multiplication cross value; the data size relationship between the first modular multiplication cross value and the third or second intermediate data is determined. Thus, whether a carry has occurred can be determined by comparing the size relationship. Secondly, in response to the first modular multiplication cross value being less than the third or second intermediate data, a first carry flag is generated; the first modular multiplication cross value is decomposed to generate a first decomposed value and a second decomposed value. Therefore, the cross value can be decomposed into data with fewer bits by using high and low bits. Third, the second decomposed value is shifted to generate a shifted decomposed value. This allows for calculation by shifting the numerical value. Fourth, the sum of the fourth intermediate data and the shifted decomposed value is determined as the low-order full multiplication result; the data size relationship between the low-order full multiplication result and the fourth intermediate data or the shifted decomposed value is determined. This allows for determining whether a carry has occurred by comparing their size relationships. Fifth, in response to the low-order full multiplication result being less than the fourth intermediate data or the shifted decomposed value, a second carry flag is generated. This allows for generating a corresponding carry flag when a carry occurs. Sixth, the first carry flag and the second carry flag are stored in different vector mask registers. This allows for full multiplication by cross-multiplication, reducing the computational load and time required for full multiplication.

[0053] Step 205: The data processing chip performs carry processing on each modular multiplication intermediate data in the modular multiplication intermediate data group based on at least one carry flag stored in the vector mask register to generate the full multiplication result.

[0054] In some embodiments, the data processing chip can perform carry processing on each of the modular multiplication intermediate data in the above-mentioned modular multiplication intermediate data group based on at least one carry flag stored in the vector mask register to generate a full multiplication result.

[0055] In practice, the data processing chip can be further configured to perform carry processing on each intermediate modular multiplication data in the intermediate modular multiplication data group based on at least one carry flag stored in the vector mask register through the following steps to generate a full multiplication result:

[0056] The first step is to determine the sum of the first intermediate data and the first decomposed value as the initial value of the high-order full multiplication result.

[0057] The second step involves carrying over the initial value of the high-order full multiplication result based on at least one carry flag included in the vector mask register, to generate a carried-over initial value for the high-order full multiplication result, which serves as the full multiplication result. In practice, in response to the vector mask register including a first carry flag, a first preset value can be added to the corresponding position of the initial value of the high-order full multiplication result. In response to the vector mask register including a second carry flag, a second preset value can be added to the corresponding position of the initial value of the high-order full multiplication result. The first preset value can be 1. The second preset value can be 2. 32 .

[0058] Step 206: The data processing chip stores the full multiplication result into the registers included in the register group.

[0059] In some embodiments, the execution entity may store the full multiplication result into a register included in the register group.

[0060] In the process of adopting technical solutions to address the aforementioned technical problems, the following technical issues often arise: when using traditional algorithms to perform Montgomery modular multiplication, multiple full multiplication operations, multiple vector multiplications, and vector addition and subtraction operations are typically required, which consumes a considerable amount of time. Considering the above technical problems and the current state of available technology, the following solution can be adopted.

[0061] Optionally, after step 206, the data processing chip can be further configured to perform the following steps:

[0062] The first step is for the data processing chip to split the above full multiplication result into high-order full multiplication result and low-order full multiplication result.

[0063] In some embodiments, the data processing chip can split the full multiplication result to generate a high-order full multiplication result and a low-order full multiplication result. The full multiplication result can be 128 bits of data; after splitting, both the high-order and low-order full multiplication results are 64 bits of data.

[0064] The second step is that the data processing chip performs a first shift operation on the above low-order full multiplication result to generate a shifted low-order result as the first shift result, and the sum of the above low-order full multiplication result and the above first shift result is determined as the first modulo reduction intermediate result.

[0065] In some embodiments, the data processing chip can perform a first shift operation on the low-order full multiplication result to generate a shifted low-order result, which is then used as the first shift result. The sum of the low-order full multiplication result and the first shift result is determined as the intermediate result of the first modulo reduction. The first shift operation can be achieved by left-shifting the low-order full multiplication result by 32 bits.

[0066] The third step involves the data processing chip generating a third carry flag corresponding to the intermediate result of the first mode reduction, and storing the third carry flag in the mask register.

[0067] In some embodiments, the data processing chip can generate a third carry flag corresponding to the intermediate result of the first modulus reduction, and store the third carry flag in a mask register.

[0068] Fourth, the data processing chip performs a second displacement process on the intermediate result of the first mode reduction to generate a low-order displacement result after displacement, which is used as the second displacement result.

[0069] In some embodiments, the data processing chip can perform a second shifting process on the intermediate result of the first modulus reduction to generate a lower-order shifted result, which is then used as the second shifted result. This second shifting process can be a right shift of 32 bits.

[0070] Fifth, the data processing chip generates the second modulus reduction intermediate result based on the second displacement result and the mask register.

[0071] In some embodiments, the data processing chip can generate a second modulus reduction intermediate result based on the second displacement result and the mask register. In practice, the difference between the first modulus reduction intermediate result, the second displacement result, and the mask register can be determined as the second modulus reduction intermediate result.

[0072] The sixth step involves the data processing chip determining the difference between the above high-order full multiplication result and the above second modular reduction intermediate result as the intermediate difference, and generating a borrow flag corresponding to the above intermediate difference.

[0073] In some embodiments, the data processing chip can determine the difference between the high-order full multiplication result and the intermediate result of the second modular reduction as an intermediate difference, and generate a borrow flag corresponding to the intermediate difference.

[0074] Step 7: The data processing chip determines the Montgomery modular multiplication result based on the borrow flag and the intermediate difference mentioned above.

[0075] In some embodiments, the data processing chip can determine the Montgomery modular multiplication result based on the borrow flag and the intermediate difference. In practice, in response to the borrow flag indicating the presence of a borrow, the difference between the intermediate difference and the preset constant is determined as the Montgomery modular multiplication result. In response to the borrow flag indicating the absence of a borrow, the intermediate difference is directly determined as the Montgomery modular multiplication result. The preset constant can be 0xffffffff.

[0076] The aforementioned steps one through seven and related content, as an inventive point of this disclosure, solve the technical problem that "when performing Montgomery modular multiplication using traditional algorithms, multiple full multiplication operations, multiple vector multiplications, and vector addition and subtraction operations are usually required, which takes a long time." The factors that cause Montgomery modular multiplication to take a long time are often as follows: when performing Montgomery modular multiplication using traditional algorithms, multiple full multiplication operations, multiple vector multiplications, and vector addition and subtraction operations are usually required, which takes a long time. If these factors are solved, the time required for Montgomery modular multiplication can be reduced. To achieve this effect, firstly, the full multiplication result is split to generate a high-order full multiplication result and a low-order full multiplication result. Thus, the first full multiplication result can be split according to the high and low orders. Second, the low-order full multiplication result is subjected to a first shift operation to generate a shifted low-order result, which is used as the first shift result. The sum of the low-order full multiplication result and the first shift result is determined as the first modular reduction intermediate result. A third carry flag corresponding to the first modular reduction intermediate result is generated, and the third carry flag is stored in the mask register. Thus, the low-order result can be shifted to generate the first modular reduction intermediate result, and whether a carry occurs is recorded. Third, the first modular reduction intermediate result is subjected to a second shift operation to generate a shifted low-order shift result, which is used as the second shift result. Based on the second shift result and the mask register, a second modular reduction intermediate result is generated. Thus, a second modular reduction intermediate result for modular reduction operation can be generated. Fourth, the difference between the high-order full multiplication result and the second modular reduction intermediate result is determined as the intermediate difference, and a borrow flag corresponding to the intermediate difference is generated. Based on the borrow flag and the intermediate difference, the Montgomery modular multiplication result is determined. Thus, Montgomery modular multiplication can be performed, and because it only requires one full multiplication operation, four vector subtractions, and two vector shifts, the time required for Montgomery modular multiplication is reduced.

[0077] The various embodiments of this disclosure have the following beneficial effects: the data processing system of some embodiments of this disclosure avoids the waste of memory resources when performing modular multiplication. Specifically, the reason for the waste of memory resources is that modular multiplication using special prime modulo operations such as Goldilocks fields fails to fully utilize the parallel computing capabilities of modern processor architectures, especially the wide-bit registers and parallel processing capabilities provided by advanced vector extension instruction sets (such as AVX512), resulting in high memory access and thus wasting memory resources. Based on this, the data processing system of some embodiments of this disclosure first acquires first modular multiplication data and second modular multiplication data. This determines the data to be used for modular multiplication. Second, it initializes a register set and, based on the register set, rearranges the first modular multiplication data and the second modular multiplication data to generate first high-order data, first low-order data, second high-order data, and second low-order data. This rearranges the data according to the high and low orders. Then, based on the first high-order data, the first low-order data, the second high-order data, and the second low-order data, an intermediate modular multiplication data set is generated. Therefore, the product of the four parts in the sub-register can be accurately calculated using a parallel algorithm, laying the foundation for the subsequent combination into a complete 64-bit multiplication result. This parallel algorithm also fully utilizes the parallel computing capabilities of modern processor architectures, reducing memory accesses and thus avoiding wasted memory resources. Next, based on the aforementioned modular multiplication intermediate data set, at least one carry flag is generated and stored in a vector mask register. This determines whether a carry occurs during the modular multiplication process. Finally, based on the at least one carry flag stored in the vector mask register, carry processing is performed on each of the intermediate modular multiplication data in the aforementioned intermediate data set to generate the full multiplication result; this full multiplication result is then stored in the registers included in the register group. Thus, the modular multiplication operation in the registers is completed in parallel, avoiding wasted memory resources during the modular multiplication operation.

[0078] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.

[0079] Computer program code for performing operations of some embodiments of this disclosure can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0080] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0081] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.

[0082] The above description is merely a selection of preferred embodiments of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of the invention involved in the embodiments of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described inventive concept. For example, technical solutions formed by substituting the above-described features with (but not limited to) technical features with similar functions disclosed in the embodiments of this disclosure.

Claims

1. A data processing system, the data processing system comprising: The system includes a data processing chip, an edge terminal, and a register set. The registers are AVX-512 registers. The data processing chip is a processor chip equipped with an AVX-512 instruction set architecture. The data processing chip is configured to acquire first modular multiplication data and second modular multiplication data stored in the edge terminal of the data processing system, wherein the first modular multiplication data and the second modular multiplication data are stored in different registers respectively; The data processing chip is configured to initialize a register group and, based on the register group, rearrange the first modular multiplication data and the second modular multiplication data to generate first high-order data, first low-order data, second high-order data, and second low-order data. The data processing chip is configured to generate a modular multiplication intermediate data set based on the first high-order data, the first low-order data, the second high-order data, and the second low-order data; The data processing chip is configured to generate at least one carry flag based on the modular multiplication intermediate data set, and to store the at least one carry flag in the vector mask register included in the register set; The data processing chip is configured to perform carry processing on each modular multiplication intermediate data in the modular multiplication intermediate data group based on at least one carry flag stored in the vector mask register to generate a full multiplication result. The data processing chip is configured to store the full multiplication result into the registers included in the register group; Specifically, based on the modular multiplication intermediate data set, at least one carry flag is generated, and the at least one carry flag is stored in a vector mask register through the following steps: The sum of the second and third intermediate data included in the modular multiplication intermediate data group is determined as the first modular multiplication cross value; Determine the data size relationship between the first modular cross value and the third intermediate data; In response to the first modular multiplication cross value being less than the third intermediate data, a first carry flag is generated; The first modular cross value is decomposed to generate a first decomposed value and a second decomposed value; The second decomposed value is shifted to generate a shifted decomposed value; The sum of the fourth intermediate data and the moved decomposed value is determined as the second modular cross value; Determine the data size relationship between the second modular cross value and the fourth intermediate data; In response to the second modular cross value being less than the fourth intermediate data, a second carry flag is generated; The first carry flag and the second carry flag are stored in different vector mask registers respectively; The data processing chip is further configured to perform the following steps: The data processing chip splits the full multiplication result to generate a high-order full multiplication result and a low-order full multiplication result; The data processing chip performs a first shift operation on the low-order full multiplication result to generate a shifted low-order result, which is used as the first shift result. The sum of the low-order full multiplication result and the first shift result is determined as the first modulo reduction intermediate result. The data processing chip generates a third carry flag corresponding to the intermediate result of the first modulus reduction, and stores the third carry flag in the mask register; The data processing chip performs a second displacement process on the intermediate result of the first modulus reduction to generate a low-order displacement result after displacement, which is used as the second displacement result. The data processing chip generates a second modulus reduction intermediate result based on the second displacement result and the mask register; The data processing chip determines the difference between the high-order full multiplication result and the intermediate result of the second modular reduction as the intermediate difference, and generates a borrow flag corresponding to the intermediate difference; The data processing chip determines the Montgomery modular multiplication result based on the borrow flag and the intermediate difference. The data processing chip is further configured as follows: Create a preset number of sub-registers as a register group using the vpshufd instruction; The register group is initialized with parameters to generate initialized registers, which serve as the initialization register group. The first modular multiplication data is rearranged to generate the first high-order data and the first low-order data; The second modular multiplication data is rearranged to generate the second high-order data and the second low-order data; The first high-order data, the first low-order data, the second high-order data, and the second low-order data are stored in their respective sub-registers.

2. The data processing system according to claim 1, wherein, The data processing chip is further configured to: The products of the first high-order data and the second high-order data and the second low-order data are respectively determined as the first intermediate data and the second intermediate data; The products of the first low-order data and the second high-order data and the second low-order data are respectively determined as the third intermediate data and the fourth intermediate data; The first intermediate data, the second intermediate data, the third intermediate data, and the fourth intermediate data are combined into a modular multiplication intermediate data group.

3. The data processing system according to claim 2, wherein, The data processing chip is further configured to: The sum of the first intermediate data and the first decomposed value is determined as the initial value of the high-order full multiplication result; Based on at least one carry flag included in the vector mask register, the initial value of the high-order full multiplication result is subjected to carry processing to generate the initial value of the high-order full multiplication result after carry, which is used as the full multiplication result.

4. The data processing system according to claim 3, wherein, The data processing chip is further configured to: In response to the vector mask register including a first carry flag, a first preset value is added to the corresponding position of the initial value of the high-order full multiplication result; In response to the vector mask register including a second carry flag, a second preset value is added to the corresponding position of the initial value of the high-order full multiplication result.

5. A data processing method, comprising: The data processing system acquires first and second modular multiplication data stored in the edge terminal, wherein the first and second modular multiplication data are stored in different registers, which are AVX-512 registers. Initialize the register group, and based on the register group, rearrange the first modular multiplication data and the second modular multiplication data respectively to generate the first high-order data, the first low-order data, the second high-order data and the second low-order data; Based on the first high-order data, the first low-order data, the second high-order data, and the second low-order data, a modular multiplication intermediate data group is generated; Based on the modular multiplication intermediate data set, at least one carry flag is generated, and the at least one carry flag is stored in a vector mask register; Based on at least one carry flag stored in the vector mask register, carry processing is performed on each of the modular multiplication intermediate data in the modular multiplication intermediate data group to generate a full multiplication result. The result of the full multiplication is stored in the registers included in the register group; Specifically, based on the modular multiplication intermediate data set, at least one carry flag is generated, and the at least one carry flag is stored in a vector mask register through the following steps: The sum of the second and third intermediate data included in the modular multiplication intermediate data group is determined as the first modular multiplication cross value; Determine the data size relationship between the first modular cross value and the third intermediate data; In response to the first modular multiplication cross value being less than the third intermediate data, a first carry flag is generated; The first modular cross value is decomposed to generate a first decomposed value and a second decomposed value; The second decomposed value is shifted to generate a shifted decomposed value; The sum of the fourth intermediate data and the moved decomposed value is determined as the second modular cross value; Determine the data size relationship between the second modular cross value and the fourth intermediate data; In response to the second modular cross value being less than the fourth intermediate data, a second carry flag is generated; The first carry flag and the second carry flag are stored in different vector mask registers respectively; The data processing chip splits the full multiplication result to generate a high-order full multiplication result and a low-order full multiplication result. The data processing chip is a processor chip equipped with the AVX-512 instruction set architecture. The data processing chip performs a first shift operation on the low-order full multiplication result to generate a shifted low-order result, which is used as the first shift result. The sum of the low-order full multiplication result and the first shift result is determined as the first modulo reduction intermediate result. The data processing chip generates a third carry flag corresponding to the intermediate result of the first modulus reduction, and stores the third carry flag in the mask register; The data processing chip performs a second displacement process on the intermediate result of the first modulus reduction to generate a low-order displacement result after displacement, which is used as the second displacement result. The data processing chip generates a second modulus reduction intermediate result based on the second displacement result and the mask register; The data processing chip determines the difference between the high-order full multiplication result and the intermediate result of the second modular reduction as the intermediate difference, and generates a borrow flag corresponding to the intermediate difference; The data processing chip determines the Montgomery modular multiplication result based on the borrow flag and the intermediate difference. The data processing chip creates a preset number of sub-registers as a register group, and creates the preset number of sub-registers using the vpshufd instruction; initializes the parameters of the register group to generate initialized registers, which are then used as an initialized register group; rearranges the first modular multiplication data to generate first high-order data and first low-order data; rearranges the second modular multiplication data to generate second high-order data and second low-order data; and stores the first high-order data, the first low-order data, the second high-order data, and the second low-order data into the respective sub-registers.

Citation Information

Patent Citations

  • High-performance Montgomery modular multiplication method based on NLP representation

    CN115016765A

  • Processing method and device for modulo multiplication

    CN118426737A