Data processing system

By rearranging modular multiplication data and processing the carry flag, the problem of memory resource waste in the prior art is solved, and efficient modular multiplication operation is achieved.

CN120821501APending Publication Date: 2025-10-21SHANGHAI QI ZHI INSTITUTE
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510923063.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-04
Publication Date
2025-10-21

AI Technical Summary

Technical Problem

Existing technologies fail to fully utilize the parallel computing capabilities of modern processor architectures when performing modular multiplication operations, especially the wide-bit registers and parallel processing capabilities provided by the Advanced Vector Extensions instruction set (such as AVX512), resulting in a waste of memory resources.

Method used

The modular multiplication data is rearranged through the data processing system to generate high-order and low-order data, the carry flag is stored in the vector mask register, the modular multiplication intermediate data is calculated in parallel, the full multiplication result is generated, and the memory access is reduced.

Benefits of technology

Effectively utilize the parallel computing capabilities of the processor, reduce memory resource waste, and improve modular multiplication efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120821501A_ABST
    Figure CN120821501A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a data processing system. According to a specific embodiment of the system, a data processing chip is configured to obtain first modular multiplication data and second modular multiplication data stored by an edge terminal included in the data processing system; initializing a register group, and performing rearrangement processing on the first modular multiplication data and the second modular multiplication data based on the register group; generating a modular multiplication intermediate data set; generating at least one carry flag based on the modular multiplication intermediate data set, and storing the at least one carry flag in a vector mask register included in the register set; carrying out carry processing on each modular multiplication intermediate data in the modular multiplication intermediate data group on the basis of at least one carry flag stored in the vector mask register so as to generate a full multiplication result; and storing the full multiplication result into the registers included in the register group. According to the embodiment, waste of memory resources during modular multiplication operation is avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present disclosure relate to the field of computer technology, and in particular to a data processing system. Background Art

[0002] Modular multiplication is a core operation in modern cryptography (including post-quantum cryptography), and its efficiency directly impacts the overall performance of various algorithms. Computations over finite fields, particularly those based on special prime moduli, form the foundation of numerous advanced applications. Currently, modular multiplication is commonly performed using special prime moduli, such as Goldilocks fields.

[0003] However, when using the above method to perform modular multiplication, the following technical problems often occur:

[0004] Modular multiplication operations using special prime moduli such as the Goldilocks field fail to fully utilize the parallel computing capabilities of modern processor architectures, especially the wide registers and parallel processing capabilities provided by the Advanced Vector Extensions instruction set (such as AVX512). Memory access is high, which leads to a waste of memory resources.

[0005] The above information disclosed in this Background section is only for enhancement of understanding of the background of the inventive concept and therefore it may contain information that does not form the prior art that is already known in this country to a person of ordinary skill in the art. Summary of the Invention

[0006] The content of this disclosure is used to briefly introduce concepts that will be described in detail in the detailed description section below. The content of this disclosure is not intended to identify key features or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.

[0007] Some embodiments of the present disclosure provide a data processing system to solve one or more of the technical problems mentioned in the above background technology section.

[0008] In a first aspect, some embodiments of the present disclosure provide a data processing system, comprising: a data processing chip, an edge terminal, and a register group, wherein the data processing chip is configured to obtain first modular multiplication data and second modular multiplication data stored in the edge terminal included in the data processing system, wherein the first modular multiplication data and the second modular multiplication data are respectively stored in different registers; the data processing chip is configured to initialize the register group, and based on the register group, rearrange the first modular multiplication data and the second modular multiplication data to generate first high-order data, first low-order data, second high-order data, and second low-order data; the data processing chip The chip is configured to generate a modular multiplication intermediate data group based on the above-mentioned first high-order data, the above-mentioned first low-order data, the above-mentioned second high-order data and the above-mentioned second low-order data; the above-mentioned data processing chip is configured to generate at least one carry flag based on the above-mentioned modular multiplication intermediate data group, and store the above-mentioned at least one carry flag in the vector mask register included in the above-mentioned register group; the above-mentioned data processing chip is configured to perform carry processing on each modular multiplication intermediate data in the above-mentioned modular multiplication intermediate data group based on the at least one carry flag stored in the vector mask register to generate a full multiplication result; the above-mentioned data processing chip is configured to store the above-mentioned full multiplication result in the register included in the register group.

[0009] In a second aspect, some embodiments of the present disclosure provide a data processing method, the method comprising: obtaining first modular multiplication data and second modular multiplication data stored in an edge terminal included in a data processing system, wherein the first modular multiplication data and the second modular multiplication data are respectively stored in different registers; initializing a register group, and based on the register group, rearranging the first modular multiplication data and the second modular multiplication data respectively to generate first high-order data, first low-order data, second high-order data and second low-order data; generating a modular multiplication intermediate data group based on the first high-order data, the first low-order data, the second high-order data and the second low-order data; generating at least one carry flag based on the modular multiplication intermediate data group, and storing the at least one carry flag in a vector mask register; performing carry processing on each modular multiplication intermediate data in the modular multiplication intermediate data group based on the at least one carry flag stored in the vector mask register to generate a full multiplication result; and storing the full multiplication result in the register included in the register group.

[0010] The above-described various embodiments of the present disclosure have the following beneficial effects: Through the data processing system of some embodiments of the present disclosure, the waste of memory resources when performing modular multiplication operations is avoided. Specifically, the reason for the waste of memory resources is that modular multiplication operations using special prime moduli such as the Goldilocks field fail to fully utilize the parallel computing capabilities of modern processor architectures, especially the wide registers and parallel processing capabilities provided by the Advanced Vector Extensions instruction set (such as AVX512), resulting in high memory access requirements and thus leading to waste of memory resources. Based on this, the data processing system of some embodiments of the present disclosure first obtains first modular multiplication data and second modular multiplication data. Thus, the data for modular multiplication operations can be determined. Next, a register group is initialized, and based on the register group, the first modular multiplication data and the second modular multiplication data are respectively rearranged to generate first high-order data, first low-order data, second high-order data, and second low-order data. Thus, the data can be rearranged according to high-order and low-order data. Then, a modular multiplication intermediate data group is generated based on the first high-order data, the first low-order data, the second high-order data, and the second low-order data. Thus, the products of the four parts in the subregisters can be accurately calculated through a parallel algorithm, laying the foundation for the subsequent combination into a complete 64-bit multiplication result. The parallel algorithm also fully utilizes the parallel computing capabilities of modern processor architectures, reduces memory access, and thus avoids the waste of memory resources. Afterwards, based on the above-mentioned modular multiplication intermediate data group, at least one carry flag is generated, and the at least one carry flag is stored in the vector mask register. Thus, it can be determined whether a carry occurs during the modular multiplication process. Finally, based on the at least one carry flag stored in the vector mask register, carry processing is performed on each modular multiplication intermediate data in the above-mentioned modular multiplication intermediate data group to generate a full multiplication result; the full multiplication result is stored in the register included in the register group. Thus, the modular multiplication operation in the register is completed in a parallel manner, avoiding the waste of memory resources when performing the modular multiplication operation. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The above and other features, advantages, and aspects of the various embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and that components and elements are not necessarily drawn to scale.

[0012] Figure 1 is a system structure diagram of some embodiments of the data processing system according to the present disclosure;

[0013] Figure 2 is a flow chart of some embodiments of the data processing method according to the present disclosure. DETAILED DESCRIPTION

[0014] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments described herein. On the contrary, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.

[0015] It should also be noted that, for ease of description, only the parts related to the invention are shown in the drawings. In the absence of conflict, the embodiments and features in the embodiments of the present disclosure may be combined with each other.

[0016] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.

[0017] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, they should be understood as "one or more".

[0018] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.

[0019] The present disclosure will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments.

[0020] Figure 1 A schematic structural diagram of some embodiments of a data processing system according to the present disclosure is shown.

[0021] like Figure 1 As shown, the data processing system of the present disclosure includes: a data processing chip 101, an edge terminal 102, and a register group 103. The data processing chip can be a chip for performing Montgomery modular multiplication. Here, the data processing chip 101 can be a processor chip equipped with the AVX-512 instruction set architecture. As an example, the data processing chip 101 can be a Xeon 8490H processor chip. The edge terminal 102 can be a user terminal connected to the data processing chip 101 via a wired or wireless connection. The registers in the register group 103 can be a memory for storing binary data. The registers in the register group 103 can be shift registers.

[0022] Reference below Figure 2, shows a process 200 of some embodiments of a data processing system according to the present disclosure. The data processing system includes the following steps:

[0023] In step 201 , a data processing chip obtains first modular multiplication data and second modular multiplication data stored in an edge terminal included in a data processing system.

[0024] In some embodiments, a data processing chip obtains first modular multiplication data and second modular multiplication data stored by an edge terminal included in a data processing system. The first modular multiplication data and the second modular multiplication data are stored in different registers. The first modular multiplication data and the second modular multiplication data can be 64-bit data. The registers can be AVX-512 registers, thereby enabling multiplication operations on eight pairs of 64-bit integers. The data processing chip can also receive the first modular multiplication data and the second modular multiplication data sent by the edge terminal.

[0025] In step 202 , the data processing chip initializes the register group, and based on the register group, rearranges the first modular multiplication data and the second modular multiplication data to generate first high-order data, first low-order data, second high-order data, and second low-order data.

[0026] In some embodiments, the data processing chip can initialize a register group, and based on the register group, rearrange the first modular multiplication data and the second modular multiplication data to generate first high-order data, first low-order data, second high-order data and second low-order data.

[0027] In practice, the data processing chip may be further configured to initialize the register group through the following steps, and based on the register group, rearrange the first modular multiplication data and the second modular multiplication data to generate first high-order data, first low-order data, second high-order data, and second low-order data:

[0028] The first step is to create a preset number of subregisters as a register group. Here, the preset number of subregisters can be created using the vpshufd instruction. The preset number can be a pre-set number of generated subregisters. As an example, the preset number can be 4.

[0029] The second step is to initialize the parameters of the register group to generate initialized sub-registers and obtain an initialized register group.

[0030] The third step is to rearrange the first modular multiplication data to generate first high-order data and first low-order data. Here, the 64-bit first modular multiplication data can be rearranged into first high-order data of the upper 32-bit data and first low-order data of the lower 32-bit data.

[0031] The fourth step is to rearrange the second modular multiplication data to generate second high-order data and second low-order data. Here, the 64-bit second modular multiplication data can be rearranged into the second high-order data of the upper 32-bit data and the second low-order data of the lower 32-bit data.

[0032] In a fifth step, the first high-order data, the first low-order data, the second high-order data, and the second low-order data are stored in respective sub-registers.

[0033] Step 203 : The data processing chip generates a modular multiplication intermediate data group based on the first high-order data, the first low-order data, the second high-order data, and the second low-order data.

[0034] In some embodiments, the data processing chip may generate a modular multiplication intermediate data group based on the first high-order data, the first low-order data, the second high-order data, and the second low-order data.

[0035] In practice, the data processing chip may be further configured to generate a modular multiplication intermediate data group based on the first high-order data, the first low-order data, the second high-order data, and the second low-order data through the following steps:

[0036] In the first step, the products of the first high-order data, the second high-order data, and the second low-order data are respectively determined as first intermediate data and second intermediate data.

[0037] In the second step, the products of the first low-order data, the second high-order data, and the second low-order data are respectively determined as third intermediate data and fourth intermediate data.

[0038] The third step is to combine the first intermediate data, the second intermediate data, the third intermediate data and the fourth intermediate data into a modular multiplication intermediate data group.

[0039] Step 204: Generate at least one carry flag based on the modular multiplication intermediate data group, and store the at least one carry flag in a vector mask register included in the register group.

[0040] In some embodiments, the data processing chip may generate at least one carry flag based on the modular multiplication intermediate data group, and store the at least one carry flag in a vector mask register included in the register group. The vector mask register is a mask register that stores vector data and is used to control whether an element in the register is masked to avoid use.

[0041] While employing technical solutions to address the aforementioned technical issues, the following technical problem often arises: When performing full multiplication in Montgomery modular multiplication using traditional algorithms, operations on higher-order data must be performed incrementally, resulting in a long time-consuming full multiplication. Considering these technical issues and considering the current state of technology, the following solution was chosen.

[0042] In some optional implementations of some embodiments, the data processing chip may be further configured to generate at least one carry flag based on the modular multiplication intermediate data group, and store the at least one carry flag in a vector mask register through the following steps:

[0043] In the first step, the sum of the second intermediate data and the third intermediate data included in the modular multiplication intermediate data group is determined as a first modular multiplication intersection value.

[0044] The second step is to determine the data size relationship between the first modular multiplication cross value and the third intermediate data.

[0045] In the third step, in response to the first modular multiplication cross value being less than the third intermediate data, a first carry flag is generated, wherein the first carry flag can indicate that a carry occurs when generating the first modular multiplication cross value.

[0046] The fourth step is to decompose the first modular multiplication cross value to generate a first decomposition value and a second decomposition value.

[0047] Step 5: Performing a bit shifting process on the second decomposed value to generate a shifted decomposed value. The bit shifting process may be shifting the second decomposed value to the left by 32 bits.

[0048] In the sixth step, the sum of the fourth intermediate data and the shifted decomposition value is determined as the second modular multiplication intersection value.

[0049] The seventh step is to determine the data size relationship between the second modular multiplication cross value and the fourth intermediate data.

[0050] In step 8, in response to the second modular multiplication cross value being less than the fourth intermediate data, a second carry flag is generated, wherein the second carry flag may indicate that a carry occurs when generating the second modular multiplication cross value.

[0051] In the ninth step, the first carry flag and the second carry flag are stored in different vector mask registers respectively.

[0052] The above-mentioned steps 1-9 and their related contents, as an inventive feature of an embodiment of the present disclosure, address the technical problem that "when performing full multiplication operations in Montgomery modular multiplication using conventional algorithms, operations on higher-order data must be performed step by step, resulting in a long time required for the full multiplication operation." Factors that often lead to a long time required for full multiplication operations are as follows: When performing full multiplication operations in Montgomery modular multiplication using conventional algorithms, operations on higher-order data must be performed step by step, resulting in a long time required for the full multiplication operation. If these factors are addressed, the time required for the full multiplication operation can be reduced. To achieve this, first, the sum of the second intermediate data and the third intermediate data included in the modular multiplication intermediate data group is determined as a first modular multiplication cross value; and the data size relationship between the first modular multiplication cross value and the third intermediate data or the second intermediate data is determined. Thus, the size relationship can be compared to determine whether a carry occurred. Second, in response to the first modular multiplication cross value being less than the third intermediate data or the second intermediate data, a first carry flag is generated; and the first modular multiplication cross value is decomposed to generate a first decomposed value and a second decomposed value. Thus, the cross value can be decomposed into data with fewer bits by high and low bits. Third, the second decomposed value is subjected to a bit shift process to generate a shifted decomposed value. Thus, calculations can be performed by shifting the values. Fourth, the sum of the fourth intermediate data and the shifted decomposed value is determined as the low-bit full multiplication result; and the data size relationship between the low-bit full multiplication result and the fourth intermediate data or the shifted decomposed value is determined. Thus, whether a carry occurs can be determined by comparing the size relationship. Fifth, in response to the low-bit full multiplication result being less than the fourth intermediate data or the shifted decomposed value, a second carry flag is generated. Thus, a corresponding carry flag can be generated when a carry occurs. Sixth, the first carry flag and the second carry flag are stored in different vector mask registers. Thus, when performing full multiplication of data, full multiplication can be performed through cross multiplication, thereby reducing the computational complexity of the full multiplication and the time required to perform the full multiplication operation.

[0053] Step 205: The data processing chip performs carry processing on each modular multiplication intermediate data in the modular multiplication intermediate data group based on at least one carry flag stored in the vector mask register to generate a full multiplication result.

[0054] In some embodiments, the data processing chip may perform carry processing on each modular multiplication intermediate data in the modular multiplication intermediate data group based on at least one carry flag stored in the vector mask register to generate a full multiplication result.

[0055] In practice, the data processing chip may be further configured to perform carry processing on each modular multiplication intermediate data in the modular multiplication intermediate data group based on at least one carry flag stored in the vector mask register through the following steps to generate a full multiplication result:

[0056] In the first step, the sum of the first intermediate data and the first decomposed value is determined as the initial value of the high-order full multiplication result.

[0057] In the second step, based on at least one carry flag included in the vector mask register, the high-order full multiplication result initial value is subjected to carry processing to generate the high-order full multiplication result initial value after the carry as the full multiplication result. In practice, in response to the vector mask register including the first carry flag, the first preset value can be added to the corresponding position of the high-order full multiplication result initial value. In response to the vector mask register including the second carry flag, the second preset value can be added to the corresponding position of the high-order full multiplication result initial value. The first preset value can be 1. The second preset value can be 2. 32 .

[0058] In step 206 , the data processing chip stores the full multiplication result in a register included in the register group.

[0059] In some embodiments, the execution entity may store the full multiplication result in a register included in the register file.

[0060] While adopting technical solutions to address the aforementioned issues, the following technical problem often arises: Using traditional algorithms to perform Montgomery modular multiplications typically requires multiple full multiplications, multiple vector multiplications, and vector addition and subtraction operations, resulting in a lengthy Montgomery modular multiplication process. Considering these technical issues and considering the current state of technology, the following solution was chosen.

[0061] Optionally, after step 206, the data processing chip may be further configured to perform the following steps:

[0062] In the first step, the data processing chip splits the full multiplication result to generate a high-bit full multiplication result and a low-bit full multiplication result.

[0063] In some embodiments, the data processing chip may split the full multiplication result to generate a high-order full multiplication result and a low-order full multiplication result. The full multiplication result may be 128-bit data. After the splitting process, the high-order full multiplication result and the low-order full multiplication result are both 64-bit data.

[0064] In the second step, the data processing chip performs a first shift processing on the above low-order full multiplication result to generate a shifted low-order result as the first shift result, and determines the sum of the above low-order full multiplication result and the above first shift result as the first modular reduction intermediate result.

[0065] In some embodiments, the data processing chip may perform a first shift on the low-order full multiplication result to generate a shifted low-order result as the first shifted result. The sum of the low-order full multiplication result and the first shifted result may be determined as the first modular reduction intermediate result. The first shift may include left-shifting the low-order full multiplication result by 32 bits.

[0066] In the third step, the data processing chip generates a third carry flag corresponding to the intermediate result of the first modular reduction, and stores the third carry flag in the mask register.

[0067] In some embodiments, the data processing chip may generate a third carry flag corresponding to the first modular reduction intermediate result, and store the third carry flag in a mask register.

[0068] In the fourth step, the data processing chip performs a second shift processing on the intermediate result of the first modular reduction to generate a low-bit shift result after shifting as the second shift result.

[0069] In some embodiments, the data processing chip may perform a second shift on the first modular reduction intermediate result to generate a shifted low-bit result as the second shift result, wherein the second shift may be a right shift of 32 bits.

[0070] In the fifth step, the data processing chip generates a second modular reduction intermediate result based on the second shift result and the mask register.

[0071] In some embodiments, the data processing chip may generate a second modulo reduction intermediate result based on the second shift result and the mask register. In practice, the second modulo reduction intermediate result may be determined as the difference between the first modulo reduction intermediate result, the second shift result, and the mask register.

[0072] In the sixth step, the data processing chip determines the difference between the high-order full multiplication result and the second modular reduction intermediate result as an intermediate difference, and generates a borrow flag corresponding to the intermediate difference.

[0073] In some embodiments, the data processing chip may determine a difference between the high-order full multiplication result and the second modular reduction intermediate result as an intermediate difference, and generate a borrow flag corresponding to the intermediate difference.

[0074] In the seventh step, the data processing chip determines the Montgomery modular multiplication result based on the borrow flag and the intermediate difference.

[0075] In some embodiments, the data processing chip may determine a Montgomery modular multiplication result based on the borrow flag and the intermediate difference. In practice, in response to the borrow flag indicating a borrow, the difference between the intermediate difference and the preset constant is determined as the Montgomery modular multiplication result. In response to the borrow flag indicating no borrow, the intermediate difference is directly determined as the Montgomery modular multiplication result. The preset constant may be 0xffffffff.

[0076] The above-mentioned first to seventh steps and their related contents serve as an inventive point of an embodiment of the present disclosure, and solve the technical problem that "when using a traditional algorithm to perform a Montgomery modular multiplication operation, it is usually necessary to perform multiple full multiplication operations, multiple vector multiplications, and vector addition and subtraction operations, and it takes a long time to perform the Montgomery modular multiplication operation." The factors that lead to a long time to perform a Montgomery modular multiplication operation are often as follows: When using a traditional algorithm to perform a Montgomery modular multiplication operation, it is usually necessary to perform multiple full multiplication operations, multiple vector multiplications, and vector addition and subtraction operations, and it takes a long time to perform a Montgomery modular multiplication operation. If the above-mentioned factors are solved, the effect of reducing the time for performing the Montgomery modular multiplication operation can be achieved. In order to achieve this effect, first, the above-mentioned full multiplication result is split to generate a high-order full multiplication result and a low-order full multiplication result. Thus, the first full multiplication result can be split according to the high order and the low order. Second, the lower-order full multiplication result is subjected to a first shift process to generate a shifted lower-order result as the first shifted result; the sum of the lower-order full multiplication result and the first shifted result is determined as the first modular reduction intermediate result; a third carry flag corresponding to the first modular reduction intermediate result is generated, and the third carry flag is stored in a mask register. Thus, the lower-order result can be shifted to generate the first modular reduction intermediate result, and whether a carry occurs is recorded. Third, the first modular reduction intermediate result is subjected to a second shift process to generate a shifted lower-order shifted result as the second shifted result; a second modular reduction intermediate result is generated based on the second shifted result and the mask register. Thus, a second modular reduction intermediate result for performing a modular reduction operation can be generated. Fourth, the difference between the upper-order full multiplication result and the second modular reduction intermediate result is determined as an intermediate difference; a borrow flag corresponding to the intermediate difference is generated; and a Montgomery modular multiplication result is determined based on the borrow flag and the intermediate difference. Thus, the Montgomery modular multiplication operation can be completed. Also, because the Montgomery modular multiplication operation is completed through only one full multiplication operation, four vector subtractions and two vector shifts, the time for performing the Montgomery modular multiplication operation is reduced.

[0077] The above-described various embodiments of the present disclosure have the following beneficial effects: Through the data processing system of some embodiments of the present disclosure, the waste of memory resources when performing modular multiplication operations is avoided. Specifically, the reason for the waste of memory resources is that modular multiplication operations using special prime moduli such as the Goldilocks field fail to fully utilize the parallel computing capabilities of modern processor architectures, especially the wide registers and parallel processing capabilities provided by the Advanced Vector Extensions instruction set (such as AVX512), resulting in high memory access requirements and thus leading to waste of memory resources. Based on this, the data processing system of some embodiments of the present disclosure first obtains first modular multiplication data and second modular multiplication data. Thus, the data for modular multiplication operations can be determined. Next, a register group is initialized, and based on the register group, the first modular multiplication data and the second modular multiplication data are respectively rearranged to generate first high-order data, first low-order data, second high-order data, and second low-order data. Thus, the data can be rearranged according to high-order and low-order data. Then, a modular multiplication intermediate data group is generated based on the first high-order data, the first low-order data, the second high-order data, and the second low-order data. Thus, the products of the four parts in the subregisters can be accurately calculated through a parallel algorithm, laying the foundation for the subsequent combination into a complete 64-bit multiplication result. The parallel algorithm also fully utilizes the parallel computing capabilities of modern processor architectures, reduces memory access, and thus avoids the waste of memory resources. Afterwards, based on the above-mentioned modular multiplication intermediate data group, at least one carry flag is generated, and the at least one carry flag is stored in the vector mask register. Thus, it can be determined whether a carry occurs during the modular multiplication process. Finally, based on the at least one carry flag stored in the vector mask register, carry processing is performed on each modular multiplication intermediate data in the above-mentioned modular multiplication intermediate data group to generate a full multiplication result; the full multiplication result is stored in the register included in the register group. Thus, the modular multiplication operation in the register is completed in a parallel manner, avoiding the waste of memory resources when performing the modular multiplication operation.

[0078] In some embodiments, the client and server can communicate using any currently known or future developed network protocol, such as HTTP (HyperText Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or future developed network.

[0079] Computer program code for performing the operations of some embodiments of the present disclosure may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0080] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0081] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.

[0082] The above description is only an illustration of some preferred embodiments of the present disclosure and the technical principles used. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by the specific combination of the above-mentioned technical features, but should also cover other technical solutions formed by any combination of the above-mentioned technical features or their equivalent features without departing from the above-mentioned inventive concept. For example, the above-mentioned features are replaced with (but not limited to) technical features with similar functions disclosed in the embodiments of the present disclosure.

Claims

1. A data processing system, comprising: Data processing chip, edge terminal and register group, among which, The data processing chip is configured to obtain first modular multiplication data and second modular multiplication data stored in an edge terminal included in the data processing system, wherein the first modular multiplication data and the second modular multiplication data are respectively stored in different registers; The data processing chip is configured to initialize a register group, and based on the register group, rearrange the first modular multiplication data and the second modular multiplication data to generate first high-order data, first low-order data, second high-order data, and second low-order data; The data processing chip is configured to generate a modular multiplication intermediate data group based on the first high-order data, the first low-order data, the second high-order data, and the second low-order data; The data processing chip is configured to generate at least one carry flag based on the modular multiplication intermediate data group, and store the at least one carry flag in a vector mask register included in the register group; The data processing chip is configured to perform carry processing on each modular multiplication intermediate data in the modular multiplication intermediate data group based on at least one carry flag stored in the vector mask register to generate a full multiplication result; The data processing chip is configured to store the full multiplication result in a register included in the register group.

2. The data processing system according to claim 1, wherein: The data processing chip is further configured to: respectively determining products of the first upper-order data, the second upper-order data, and the second lower-order data as first intermediate data and second intermediate data; respectively determining products of the first low-order data, the second high-order data, and the second low-order data as third intermediate data and fourth intermediate data; The first intermediate data, the second intermediate data, the third intermediate data, and the fourth intermediate data are combined into a modular multiplication intermediate data group.

3. The data processing system according to claim 2, wherein: The data processing chip is further configured to: Determine the sum of the first intermediate data and the first decomposed value as the initial value of the high-bit full multiplication result; Based on at least one carry flag included in the vector mask register, carry processing is performed on the high-order full multiplication result initial value to generate a high-order full multiplication result initial value after carry as the full multiplication result.

4. The data processing system according to claim 3, wherein: The data processing chip is further configured to: In response to the vector mask register including a first carry flag, adding a first preset value to a corresponding position of the high-order full multiplication result initial value; In response to the vector mask register including the second carry flag, a second preset value is added to a corresponding position of the high-order full multiplication result initial value.

5. The data processing system according to claim 1, wherein: The data processing chip is further configured to: Initializing parameters of the register group to generate initialized registers as the initialization register group; Rearranging the first modular multiplication data to generate first high-order data and first low-order data; Rearranging the second modular multiplication data to generate second high-order data and second low-order data; The first high-order data, the first low-order data, the second high-order data, and the second low-order data are stored in respective sub-registers.

6. A data processing method comprising: Acquire first modular multiplication data and second modular multiplication data stored in an edge terminal included in a data processing system, wherein the first modular multiplication data and the second modular multiplication data are respectively stored in different registers; Initializing a register group, and based on the register group, rearranging the first modular multiplication data and the second modular multiplication data to generate first high-order data, first low-order data, second high-order data, and second low-order data; Generate a modular multiplication intermediate data group based on the first high-order data, the first low-order data, the second high-order data, and the second low-order data; generating at least one carry flag based on the modular multiplication intermediate data group, and storing the at least one carry flag in a vector mask register; Based on at least one carry flag stored in the vector mask register, perform carry processing on each modular multiplication intermediate data in the modular multiplication intermediate data group to generate a full multiplication result; The full multiplication result is stored in a register included in the register group.

Citation Information

Patent Citations

  • Encryption and decryption method and device, electronic equipment and computer readable storage medium

    CN112711395A

  • High-performance Montgomery modular multiplication method based on NLP representation

    CN115016765A

  • Hardware accelerator implementation method for Montgomery modular multiplication and hardware accelerator

    CN118312138A

  • Processing method and device for modulo multiplication

    CN118426737A

  • Computationally efficient modular multiplication method and apparatus

    US20010010077A1