A C++ fusion programming method based on heterogeneous many-core architecture
By providing C++ fusion programming methods on heterogeneous multi-core processors, using the athreadcxx class and slavecxx.h header files, the symbol multi-definition, memory access and one-way call difficulties of C++ programming in heterogeneous multi-core processors are solved, and the high-performance development of C++ language on domestic heterogeneous multi-core chips is realized.
Patent Information
- Application Number
- CN202110325186.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-03-26
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2041-03-26
AI Technical Summary
In C++ programming, heterogeneous multi-core processors have multiple symbol definitions, problems with global variable access between slave cores, and difficulties in one-way calling between master cores and slave cores.
It provides a C++ fusion programming method based on heterogeneous multi-core architecture. Through the athreadcxx class and slavecxx.h header file, symbol address management and parameter transfer between the master and slave cores are realized. The thread_local keyword and compiler options -mhost, -mslave, -mhybrid are used to generate mixed executable codes of different instruction sets.
It solves the multi-definition of symbols, memory access problems and one-way call difficulties of C++ programming in heterogeneous multi-core processors, and realizes the effective use and high-performance development of C++ language on domestic heterogeneous multi-core chips.
Smart Images

Figure CN114217770B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a C++ fusion programming method based on a heterogeneous many-core architecture, and belongs to the technical field of high performance computing. Background Art
[0002] Domestic heterogeneous many-core processors are suitable for the field of high-performance computing, and are increasingly suitable for the field of artificial intelligence. There is a large demand for C++ programming in these two major fields to speed up the programming pace and simplify the programming complexity, especially generic programming and the large number of basic containers and algorithms provided by the C++ library.
[0003] The heterogeneous many-core architecture includes two cores, the master core and the slave core. The instruction set of each core is inconsistent. Therefore, in C++ programming, it is necessary to generate the target code of the two instruction sets and mix and link the target codes of different instruction sets together. Since the generated global symbols are stored in the core group shared space, both the master and slave cores can access them, and multiple definitions of symbols may occur; at the same time, since each slave core will share a piece of code at the same time, and the global variables in the code are stored in the core group shared space, each slave core can access and modify the variables, and there are memory access problems between slave cores such as read-after-write and write-after-read.
[0004] Currently, heterogeneous many-core processors generally only support C language and FORTRAN language, which makes programming and development difficult. The many-core heterogeneous architecture and multi-level storage structure make it impossible to directly apply the general C++ language and C++ compiler to domestic heterogeneous many-core chips, limiting the development capabilities of domestic heterogeneous chip programming users, and making it impossible to use the advanced features of domestic heterogeneous many-core chips to achieve high-performance program development. Summary of the invention
[0005] The purpose of the present invention is to provide a C++ fusion programming method based on a heterogeneous many-core architecture to solve the problems of multiple definitions of mixed link symbols of instruction sets of different architectures, the problem of global variable memory access between slave cores, and the difficulty of one-way calling between master cores and between master cores and slave cores, thus filling the gap in C++ programming in the domestic heterogeneous many-core chip ecosystem.
[0006] To achieve the above object, the technical solution adopted by the present invention is: to provide a C++ fusion programming method based on a heterogeneous multi-core architecture, comprising the following steps:
[0007] S1. The main core provides an object of the athreadcxx class in the form of a header file "athreadcxx.h" and stores the object in the core group shared space so that the main cores with different symbol addresses do not affect each other;
[0008] S2. The object of the athreadcxx class initializes the slave core resources through the constructor and recycles the slave core resources through the destructor;
[0009] The object of the athreadcxx class provides a member variable cgid, which is used to save the core group number of the current core group;
[0010] The object of athreadcxx class provides a member structure variable core.info, which is used to save the symbolic address of the master-slave core transfer parameters;
[0011] The object of the athreadcxx class provides a member function spawn, which is used to call the slave kernel function, specifically:
[0012] S21, add the slave_ prefix to the slave core function name and pass it to the slave core as the first pointer parameter of the member function spawn;
[0013] S22, pack the parameters to be transferred into a structure, and pass the structure pointer to the slave core as the second parameter of the member function spawn;
[0014] S3. The compiler compiles the main core program programmed with the object of the athreadcxx class through the option -mhost. In the process of processing the symbol address, the C++ compiler renames the function name according to the general rules. After the renaming is completed, the slave_ prefix is identified, and the information of the renamed function name is extracted to generate a symbolic address containing the slave_ prefix without affecting the original function information, so as to remove the influence of the slave_ prefix on the renaming;
[0015] S4. The slave core provides thread-private global variables PEN, COL, and ROW in the form of a header file "slavecxx.h" to save the number and row and column information of the current slave core;
[0016] The slave core provides a global function getArg in the form of a header file "slavecxx.h". The return value of this function is the second parameter pointer passed from the master core to the slave core in S22. By deconstructing the return value, the parameter that the master core wants to pass to the slave core is obtained;
[0017] The slave core uses the thread_local keyword to declare the slave core's private global variables, declaring that the variables are stored in the slave core's private space, while ordinary global variables are stored in the core group shared space;
[0018] S5. The compiler compiles the slave program including the header file "slavecxx.h" through the option -mslave. When the thread_local keyword is recognized, the symbolic address of the global variable is added with section information. When linking, the symbolic address of the variable including the section information is addressed as the address format of the slave private space. In the process of processing the symbolic address, the C++ compiler renames the function names of all slave symbols according to the general rules. After the renaming is completed, the slave_ prefix is added to distinguish the symbolic address of the master.
[0019] S6. The compiler links the master program symbol address, the symbol address containing the slave_ prefix in the master program, and all addresses containing slave_ in the slave program through the option -mhybrid to generate hybrid executable codes of different instruction sets, so that the master core can call the slave core only through the prefix slave_.
[0020] Due to the application of the above technical solution, the present invention has the following advantages compared with the prior art:
[0021] The present invention adopts a language extension method to retain the language characteristics of the C++ language itself to the greatest extent, reducing the difficulty of program development and transplantation, and at the same time providing an extended language to effectively use the multi-core features supported by the hardware; under the domestic heterogeneous multi-core architecture with multi-level storage, the storage structure of different spaces is fully utilized, and the architecture codes of two different instruction sets are effectively integrated together, which can ensure the correctness of the program while making good use of the computing resources of the master core and the slave core. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Attached Figure 1 It is a schematic diagram of a C++ fusion programming method based on a heterogeneous many-core architecture of the present invention. DETAILED DESCRIPTION
[0023] Embodiment: The present invention provides a C++ fusion programming method based on a heterogeneous multi-core architecture, which specifically includes the following steps:
[0024] S1. The main core provides an object of the athreadcxx class in the form of a header file "athreadcxx.h" and stores the object in the core group shared space so that the main cores with different symbol addresses do not affect each other;
[0025] S2. The object of the athreadcxx class initializes the slave core resources through the constructor and recycles the slave core resources through the destructor;
[0026] The object of the athreadcxx class provides a member variable cgid, which is used to save the core group number of the current core group;
[0027] The object of athreadcxx class provides a member structure variable core.info, which is used to save the symbolic address of the master-slave core transfer parameters;
[0028] The object of the athreadcxx class provides a member function spawn, which is used to call the slave kernel function, specifically:
[0029] S21. Add the slave_ prefix to the slave function name and pass it to the slave core as the first pointer parameter of the member function spawn. The prefix is used to distinguish the master core function from the slave core function.
[0030] S22, pack the parameters to be transferred into a structure, and pass the structure pointer to the slave core as the second parameter of the member function spawn;
[0031] S3. The compiler compiles the main core program programmed with the object of the athreadcxx class through the option -mhost. In the process of processing the symbol address, the C++ compiler renames the function name according to the general rules. After the renaming is completed, the slave_ prefix is identified, and the information of the renamed function name is extracted to generate a symbolic address containing the slave_ prefix without affecting the original function information, so as to remove the influence of the slave_ prefix on the renaming;
[0032] S4. The slave core provides thread-private global variables PEN, COL, and ROW in the form of a header file "slavecxx.h" to save the number and row and column information of the current slave core;
[0033] The slave core provides a global function getArg in the form of a header file "slavecxx.h". The return value of this function is the second parameter pointer passed from the master core to the slave core in S22. By deconstructing the return value, the parameter that the master core wants to pass to the slave core is obtained;
[0034] The slave core uses the thread_local keyword to declare the slave core's private global variables, declaring that the variables are stored in the slave core's private space, while ordinary global variables are stored in the core group shared space;
[0035] S5. The compiler compiles the slave program including the header file "slavecxx.h" through the option -mslave. When the thread_local keyword is recognized, the symbolic address of the global variable is added with section information. When linking, the symbolic address of the variable including the section information is addressed as the address format of the slave private space. In the process of processing the symbolic address, the C++ compiler renames the function names of all slave symbols according to the general rules. After the renaming is completed, the slave_ prefix is added to distinguish the symbolic address of the master.
[0036] S6. The compiler links the master program symbol address, the symbol address containing the slave_ prefix in the master program, and all addresses containing slave_ in the slave program through the option -mhybrid to generate hybrid executable codes of different instruction sets, so that the master core can call the slave core only through the prefix slave_.
[0037] The further explanation of the above embodiment is as follows:
[0038] The present invention realizes the combination of software and hardware through language extension support, realizes C++ many-core fusion programming on domestic heterogeneous many-core architecture chips, fills the gap in the ecological chain, and provides a new programming method for the application development of domestic heterogeneous many-core architecture.
[0039] By extending the C++ language and programming model, and making the programming language correspond to the underlying chip design, a C++ fusion programming method is designed that can utilize both C++ features and heterogeneous multi-core chips. This is of great significance in real-world high-performance and artificial intelligence applications.
[0040] The present invention extends the C++ language features and programming mode, effectively maps the language features to the chip architecture, so that the programming language corresponds to the underlying chip design, and realizes the implementation of the extended programming mode by applying the features of the architecture;
[0041] A C++ fusion programming method is designed that can utilize both C++ features and heterogeneous many-core chips. This allows C++ features to be used for programming on domestic heterogeneous many-core chips, and ensures that the architectural features of domestic heterogeneous many-core chips are fully utilized, giving the program certain performance advantages. This is of great significance in real-world high-performance and artificial intelligence applications.
[0042] The specific process is as follows:
[0043] 1. The main core provides the object of the athreadcxx class in the form of the header file "athreadcxx.h". The object is stored in the shared space of the core group. The main cores with different symbolic addresses do not affect each other.
[0044] 2. The athreadcxx class initializes the slave core resources through the constructor and recycles the slave core resources through the destructor.
[0045] 3. athreadcxx provides a member variable cgid to save the current core group related information, such as the core group number.
[0046] 4. athreadcxx provides a member structure variable core.info to store the symbolic address of the master-slave core transfer parameters.
[0047] 5. athreadcxx provides a member function spawn to call the slave core function. The slave core function name is prefixed with slave_ and passed to the slave core as the first pointer parameter of the member function spawn. The prefix method is used to distinguish the master core function from the slave core function. The parameters to be passed are packaged into a structure, and the structure pointer is passed to the slave core as the second parameter of the member function spawn.
[0048] 6. The compiler compiles the main core program programmed with athreadcxx objects through the option -mhost. During the symbol address processing, the C++ compiler needs to rename the function name according to general rules. After the renaming is completed, the slave_ prefix is identified, and the information of the renamed function name is extracted to remove the influence of the slave_ prefix on the renaming, and a symbol address containing the slave_ prefix without affecting the original function information is generated.
[0049] 7. The slave core provides thread-private global variables PEN, COL, and ROW in the form of a header file "slavecxx.h" to save the number and row and column information of the current slave core.
[0050] 8. The slave core provides a global function getArg in the form of a header file "slavecxx.h", and the return value of this function is the second parameter pointer passed by the master core in step 5. By deconstructing the structure pointer, specific information is obtained.
[0051] 9. The slave core uses the thread_local keyword to declare a global variable, declaring that the variable is stored in the slave core's private space, while ordinary global variables are stored in the core group shared space.
[0052] 10. The compiler compiles the slave program containing the header file "slavecxx.h" through the option -mslave. When the thread_local keyword is recognized, the symbolic address of the global variable is added with section information. When linking, the symbolic address of the variable containing the section information is addressed as the address format of the slave private space. In the process of processing the symbolic address, the C++ compiler renames the function name of all slave symbols according to the general rules. After the renaming is completed, the slave_ prefix is added to distinguish the symbolic address of the master.
[0053] 11. The compiler uses the option -mhybrid to link the symbolic address of the master program, the symbolic address containing the slave_ prefix in the master program, and all the addresses containing slave_ in the slave program, ensuring that the master can only call the slave through the prefix slave_. Finally, a hybrid executable code with different instruction sets is generated.
[0054] From the above steps, it can be found that steps 1-5 determine the programming method of the main core in the form of header files, step 6 uses the compiler renaming method to solve the problem of inconsistent symbolic addresses of the main core calling the slave core, steps 7-8 provide information management and parameter transfer mechanisms for the slave core, step 9 provides the implementation of multi-level storage keywords, step 10 solves the problem of multiple definitions of the symbolic addresses of the master and slave cores, and realizes the basis of mixed linking, and step 10 merges the codes of different instruction sets into one executable code through linking.
[0055] When adopting the above-mentioned C++ fusion programming method based on heterogeneous many-core architecture, it adopts the language extension method to retain the language characteristics of the C++ language itself to the greatest extent, reduce the difficulty of program development and porting, and at the same time provide an extended language to effectively use the many-core features supported by the hardware; under the domestic heterogeneous many-core architecture with multi-level storage, it fully utilizes the storage structure of different spaces and effectively integrates the architecture codes of two different instruction sets, which can ensure the correctness of the program while making good use of the computing resources of the master core and the slave core.
[0056] In order to facilitate a better understanding of the present invention, the terms used in this article are briefly explained below:
[0057] Heterogeneous many-core architecture: In chip design, each architecture will determine an instruction set. If the same chip has two or more architectures and two or more instruction sets, this design is called a heterogeneous architecture. If a chip design contains hundreds of cores, this design is called a many-core architecture.
[0058] Domestic heterogeneous many-core chips: a high-performance heterogeneous central processing unit that integrates a small number of general-purpose main cores that perform management, communication and computing functions and a large number of streamlined slave cores that perform computing functions on a complete chip; the general-purpose main core runs a general-purpose operating system, mainly performs the management and control functions of the entire chip, and also performs certain computing functions and the communication functions between the chip and the outside world; the slave core plays the role of accelerating computing; the master core and the slave core have different architectures and use different instruction sets.
[0059] Core group: A heterogeneous many-core chip contains several core groups. Each core group contains a master core and several slave cores. The core groups share the core group shared space.
[0060] Multi-level storage space: divided into three levels of storage space: full-chip space, core group global space, and slave core private space; the full-chip space is divided into several core group global spaces, and each core group space has exclusive access rights to the corresponding core group space; the core group global space is shared by a master core and several slave cores in the core group; each slave core has a copy of the slave core private space, and each slave core has exclusive access rights.
[0061] Symbolic address: refers to the address identifier generated by the compiler for global variables, function names, and other types in the program. When linked into an executable code, the symbolic address must be uniquely determined.
[0062] The above embodiments are only for illustrating the technical concept and features of the present invention, and their purpose is to enable people familiar with the technology to understand the content of the present invention and implement it accordingly, and they cannot be used to limit the protection scope of the present invention. Any equivalent changes or modifications made according to the spirit of the present invention should be included in the protection scope of the present invention.
Claims
1. A C++ fusion programming method based on heterogeneous multi-core architecture, characterized in that: The following steps are involved: S1. The main core provides an object of the athreadcxx class in the form of a header file "athreadcxx.h" and stores the object in the core group shared space so that the main cores with different symbol addresses do not affect each other; S2. The object of the athreadcxx class initializes the slave core resources through the constructor and recycles the slave core resources through the destructor; The object of the athreadcxx class provides a member variable cgid, which is used to save the core group number of the current core group; The object of athreadcxx class provides a member structure variable core.info, which is used to save the symbolic address of the master-slave core transfer parameters; The object of the athreadcxx class provides a member function spawn, which is used to call the slave kernel function, specifically: S21, add the slave_ prefix to the slave core function name and pass it to the slave core as the first pointer parameter of the member function spawn; S22, pack the parameters to be transferred into a structure, and pass the structure pointer to the slave core as the second parameter of the member function spawn; S3. The compiler compiles the main core program programmed with the object of the athreadcxx class through the option -mhost. In the process of processing the symbol address, the C++ compiler renames the function name according to the general rules. After the renaming is completed, the slave_ prefix is identified, and the information of the renamed function name is extracted to generate a symbolic address containing the slave_ prefix without affecting the original function information, so as to remove the influence of the slave_ prefix on the renaming; S4. The slave core provides thread-private global variables PEN, COL, and ROW in the form of a header file "slavecxx.h" to save the number and row and column information of the current slave core; The slave core provides a global function getArg in the form of a header file "slavecxx.h". The return value of this function is the second parameter pointer passed from the master core to the slave core in S22. By deconstructing the return value, the parameter that the master core wants to pass to the slave core is obtained; The slave core uses the thread_local keyword to declare the slave core's private global variables, declaring that the variables are stored in the slave core's private space, while ordinary global variables are stored in the core group shared space; S5. The compiler compiles the slave program including the header file "slavecxx.h" through the option -mslave. When the thread_local keyword is recognized, the symbolic address of the global variable is added with section information. When linking, the symbolic address of the variable including the section information is addressed as the address format of the slave private space. In the process of processing the symbolic address, the C++ compiler renames the function names of all slave symbols according to the general rules. After the renaming is completed, the slave_ prefix is added to distinguish the symbolic address of the master. S6. The compiler links the master program symbol address, the symbol address containing the slave_ prefix in the master program, and all addresses containing slave_ in the slave program through the option -mhybrid to generate hybrid executable codes of different instruction sets, so that the master core can call the slave core only through the prefix slave_.
Citation Information
Patent Citations
Compiling and generation method for heterogeneous code fusion
CN105426226A
Algorithm parallel processing method and system based on heterogeneous many-core processor
CN112306678A