Compiler back-end development assisting method and device based on artificial intelligence

By applying AI programming models and multilingual data sets in compiler back-end development, the problems of low development efficiency and high cost caused by different processor architectures are solved, and high-precision code completion, generation and repair functions are achieved, improving the level and efficiency of development automation.

CN120215908APending Publication Date: 2025-06-27INST OF COMPUTING TECH CHINESE ACAD OF SCI
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510235261.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

During the process of compiler backend development, the diversity of different processor architectures is facing the need to customize the compiler backend for each architecture, and the lack of effective automation tool support leads to low development efficiency and high cost.

Method used

Design and implement an AI programming model dedicated to compiler backend development. By building multilingual compiler backend data sets, developing dedicated large-scale language models, and introducing a searcher based on In-Context Learning, providing high-precision support for code completion, generation and repair.

Benefits of technology

It improves the automation level of compiler backend development, reduces manual coding workload, improves development efficiency, and reduces developers' learning costs, and realizes end-to-end support for compiler backend development.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120215908A_ABST
    Figure CN120215908A_ABST
Patent Text Reader

Abstract

According to the artificial intelligence-based compiler back-end development assisting method and device, fine tuning training is performed on an existing LLM, a large language model specially aiming at compiler back-end development optimization is obtained, and high-precision support can be provided for tasks such as code completion, generation and repair. A retrieval method is designed, an example most similar to user input can be automatically retrieved, and therefore the performance of LLM is improved in the environment that computing resources are limited. According to the method, the automation level of the rear-end development of the compiler is improved, the manual coding workload is reduced, the development efficiency is improved, and the learning cost of developers is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of processor development, compiler backend code generation, completion, and code repair, and particularly relates to a compiler backend development assistance method, device, electronic device, computer-readable storage medium, and computer program product based on artificial intelligence. Background Art

[0002] Compiler technology is the link connecting new processors and domain applications, and is usually laid out in the same period as the design of new processors. The quality of compilation technology largely determines whether the personalized design of the processor and the advantages of the architecture can be fully utilized. However, the R & D cycle of compilers is relatively long, and the R & D process is of high difficulty, which is the "slow and difficult" problem in compiler R & D.

[0003] Modern compilers usually adopt a three-stage architecture, dividing the compiler into a front end, a middle end, and a back end, which provides great flexibility for compiler transplantation. The front end is responsible for lexical analysis, syntax analysis, and semantic analysis, converting the source code into an intermediate representation that the compiler can understand. The middle end receives the intermediate representation generated by the front end and performs optimizations to improve code performance and efficiency. The back end then converts the optimized intermediate representation into machine code according to the hardware characteristics of the target platform, depending on a specific hardware architecture. Among them, when a new target structure Target needs to be supported, only a back end needs to be developed, and the front end and optimizer are reused, greatly improving the compiler transplantation efficiency. Existing compilers, such as GCC (GNU Compiler Collection), are programming language compilers developed by the open-source software organization GNU. GCC was originally a compiler specifically written for the GNU operating system and has now been adopted as the standard compiler by most Unix-like operating systems (such as Linux, BSD, MacOS X, etc.), and can even be used on Microsoft Windows. GCC supports multiple computer architecture chips, such as x86, ARM, MIPS, etc., and has been ported to many other hardware platforms; LLVM, LLVM is a framework system for building compilers, written in C++. LLVM provides compiler-related support and can be used as the backend for compilers in multiple languages. It can be regarded as a collection of modular and reusable compiler and tool technologies. The target-independent code generation mechanism has been utilized to support different microprocessors.

[0004] The working process of a compiler is to first convert the code written by a programmer into its internal representation, and then convert the internal representation into code that can run on a computer (machine code). Machine codes for different computer architectures (which can be understood as different types of computers) are different. The target-independent code generation mechanism means that the compiler provides a set of reusable components for converting the compiler's internal representation into machine code of a specified type. It is not limited to a single architecture (only being able to generate machine code for a single architecture), so it is called target-independent. The "target" refers to the type of architecture, which includes various different hardware such as chip architectures, graphics processor architectures, and digital signal processor architectures. In particular, the development of the compiler backend needs to implement two basic components: (1) the abstract target description file and (2) the implementation of the abstract function interface. These components integrate target-specific features with the compiler infrastructure to facilitate the accurate generation of machine code. The abstract target description file is used to capture the detailed target-specific features of the target. To effectively write the target description file, developers must thoroughly understand the properties of the target architecture (mainly the instruction set architecture ISA), hardware features, and the syntax rules of the programming language of the target description file, such as the machine description (.md) file in GCC and the machine target file (.td) file in LLVM. Through the abstract target description file, developers can encapsulate target-specific details in a structured format, providing a basis for the implementation of the compiler backend. Once the target description file is completed, developers will implement the abstract function interface according to the target-specific features recorded in the abstract target description file. These interfaces need to map target-specific features to the compiler infrastructure for target-specific implementations, providing specific functions for each component in the backend. Abstract function interfaces are usually divided into two categories: 1) inherited functions, which include function interfaces for the compiler infrastructure to perform specific tasks during the backend process. Programmers must inherit these interfaces and provide implementations for each target. For example, the "getRelocType" function in LLVM maps relocation variants and immediate values in the target's ISA. 2) custom functions, which are specifically designed for certain specific targets. For example, the "isImm24bit" function in the LLVM ARM backend is used to check whether the encoding length of the immediate value is 24 bits. The "isImm24bit" function is specific to ARM and does not exist in other targets.

[0005] Compared with redevelopment of the compiler backend for each architecture, the target-independent code generation mechanism greatly reduces development work. However, with the increasing diversification of processor types, programmers need to customize specific compiler backends for each processor. This is still a labor-intensive and time-consuming process.

[0006] These backends target different processor types, including general-purpose processors for CISC instruction sets (such as X86, which is a complex instruction set computer (CISC) instruction set architecture initially developed by Intel based on the Intel 8086 microprocessor and its 8088 variant, and most current desktops and laptops are based on the X86 architecture) and RISC instruction sets (such as MIPS, which is a reduced instruction set (RISC) processor architecture that emerged in 1981, developed and licensed by MIPS Technologies and widely used in many electronic products, network devices, personal entertainment devices, and commercial devices), as well as GPU types (such as AMDGPU, where the AMDGPU backend provides ISA code generation for AMD gpus, starting from the R600 series up to the current GCN series) and digital information processor types (such as Hexagon, which is the brand name of Qualcomm's digital signal processors, with each version corresponding to an instruction set architecture). And the task of compiler backend development is undoubtedly a heavy burden on R & D personnel. We take LLVM as an example to illustrate the difficulty and workload of backend development.

[0007] The LLVM 17.0.1 version contains backend codes for 25 different architectures. Among them, the AMDGPU backend alone contains 219 C / C++ files, approximately 118,500 lines of code. For the target description files of LLVM 17.0.1, the X86 architecture alone contains 53 target description files (.td files), approximately 57,900 lines of TableGen code. Moreover, due to the high flexibility of the C / C++ language and the code languages of the target description files, the C / C++ files and target description files exhibit extremely diverse personalized styles among different target platforms.

[0008] Through research and analysis of C / C++ files and target description files for different target platforms in GCC and the LLVM backend, the following problems were found: ① Although mainstream large language models (LLMs) perform excellently in general code generation, they have obvious deficiencies in the compiler backend field. Currently, there is a lack of dedicated datasets for multi-language (C / C++, TableGen, MD) and multi-task (code completion, repair, etc.) in the compiler backend, resulting in insufficient accuracy in backend code generation (for example, experiments show that the EM accuracy of the baseline model is generally lower than 1%, where EM refers to Exact Match). ② Taking LLVM 17.0.1 as an example, in the interface implementations for different targets, 57.6% are general functions inherited from the compilation framework, and the rest are customized functions. Although the structures of such codes are similar, they involve a large number of target-specific values (such as instruction encodings, immediate value ranges), and the coding styles of developers vary significantly, exacerbating the costs of code maintenance and migration. ③ Fork-Flow (the traditional compiler backend development process) relies on manual reuse of similar backends and line-by-line modification. However, this process requires developers to deeply understand target characteristics and the compilation framework. For example, developing the RISC-V backend often requires manually adjusting thousands of lines of code, which not only reduces development efficiency but also brings a large amount of redundant work.

[0009] There are already related technologies, namely the automatic generation method for the compiler backend and the automatic construction method for compiler backend code. The former uses a deep learning model to capture the semantic patterns of the code and realizes end-to-end code generation. However, the model training only depends on the code features of a single language (such as C++), and does not perform multi-language optimization for target description files (such as TableGen, MD). The latter mainly focuses on the automatic generation of target description files (such as the.td files of LLVM), and uses static analysis and rule matching to complete code synthesis, but only supports the single task of code generation and does not cover the development needs of multi-language or multi-scenarios.

[0010] In summary, the current compiler backend development mainly faces the following problems: (1) The diversity of different processor architectures leads to the need to customize the compiler backend for each architecture, lacking effective automated tool support. (2) Writing compiler backend code usually involves a large amount of C / C++ code and target description files (such as TableGen, MD files), and the development process is time-consuming and error-prone, increasing the difficulty of maintenance and migration. (3) The existing large language models (LLMs) perform poorly in compiler backend code generation, completion, and program repair tasks, lacking targeted backend development datasets. Summary of the Invention

[0011] The technical problem to be solved by the present invention is to design and implement an AI programming model dedicated to compiler backend development, improve the development efficiency of the compiler backend and reduce the development cost. The AI programming model can give developers better optimization suggestions or code designs based on the learned existing optimization solutions, improve the performance of the compiler and thus enhance the performance of the chip.

[0012] To solve the above problems, the present invention proposes a compiler backend development method based on a multi-language dataset. The dataset covers multiple programming languages such as C / C++, TableGen, and MD, contains rich processor backend instances, and provides comprehensive data support for common scenarios in backend development (code generation, completion, repair). On this basis, we fine-tune the existing LLM to develop a large language model specifically optimized for compiler backend development, which can provide high-precision support for tasks such as code completion, generation, and repair. In addition, we design a retriever based on In-Context Learning, which can automatically retrieve the examples most similar to the user input and construct few-shot prompts, thereby improving the performance of the LLM in an environment with limited computing resources. The present invention improves the automation level of compiler backend development, reduces the manual coding workload, improves the development efficiency, and reduces the learning cost of developers.

[0013] Aiming at the deficiencies of the prior art, as Figure 7 shown, the present invention proposes an AI-based compiler backend development assistance method, which includes:

[0014] A data collection step of obtaining a compiler backend code library including multiple chip architectures, extracting repair data from its historical commit records, using the code before repair in the repair data as an error segment, and the code after repair as a repair segment; and the functions in the compiler backend code have corresponding description information;

[0015] A data collection step of obtaining a compiler backend code library including multiple architectures, for the program repair function, extracting repair data from its historical commit records, using the code before repair in the repair data as an error segment, and the code after repair as a repair segment; for the code generation function, extracting the functions in the compiler backend code that have corresponding description information;

[0016] A code extraction step of dividing the code in the compiler backend code library according to the code syntax to obtain structured code;

[0017] Code recognition step: Identify the boundary positions of error fragments and repair fragments in the structured code, mark the error boundaries and repair boundaries in the structured code, and replace the numerical constants, string literals, and enumeration prefix values in the structured code with preset standardized placeholders to obtain a normalized code;

[0018] Sample construction step: According to the chip architecture category, construct the normalized code into task samples under each chip architecture. The task samples include: completion task samples, next sentence prediction task samples, code generation task samples, and program repair task samples;

[0019] Model fine-tuning step: Input the task prompts of the task samples into a large language model, and perform fine-tuning training on the large language model according to the task results and the labels of the task samples to obtain a programming model that can complete completion tasks, next sentence prediction tasks, code generation tasks, and program repair tasks;

[0020] Model deployment step: The programming model is used to execute the compiler backend programming tasks of a specified architecture, and integrate the executed programming code into the compiler backend code of the specified architecture; and during the execution process, the programming model provides code completion, next sentence suggestions, and repair suggestions according to the current programming code during the execution process. The programming model generates a corresponding function code template according to the input code description to assist in writing code.

[0021] The described compiler backend development assistance method based on artificial intelligence, wherein the sample construction step includes:

[0022] Intercept a part of the continuous statements in the normalized code as the input, and the remaining part as the label to form the completion task sample;

[0023] Use the previous continuous statements in the normalized code as the input and the subsequent continuous statements as the label to construct the next sentence prediction task sample;

[0024] Use the description information as the input and the corresponding normalized code of the description information as the label to construct the code generation task sample;

[0025] Use the code within the error boundary in the normalized code as the input and the code within the repair boundary corresponding to the error boundary in the normalized code as the label to construct the program repair task sample.

[0026] The described compiler backend development assistance method based on artificial intelligence, wherein the model fine-tuning step includes:

[0027] Select the current sample from the completion task sample, the next sentence prediction task sample, the code generation task sample, and the program repair task sample, and retrieve it through the context retriever. Select the data with the highest similarity to the current sample in the structured code and construct it with the current sample as the sample context template to provide input data for the fine-tuning training of the large language model.

[0028] The described method for assisting in the development of the compiler backend based on artificial intelligence, wherein the model deployment step includes:

[0029] Integrate the programming model into a cross-platform code editor using the model encapsulation framework and extension tools to provide users with code completion, next sentence suggestion, program repair, and code generation functions;

[0030] When the programming model implements code completion, according to the currently written code during the execution process, predict and output the subsequent code snippet of the currently written code;

[0031] When the programming model provides the next sentence suggestion, according to the context of the currently written code during the execution process, predict and output the next sentence of the currently written code;

[0032] When the programming model provides repair suggestions, according to the currently programmed code during the execution process, predict and output the error part of the currently edited code and the modification suggestions;

[0033] When the programming model provides function code template generation, according to the natural language code description input during the execution process, predict and generate function code that conforms to the natural language code description.

[0034] As Figure 8 shown, the present invention also proposes an apparatus for assisting in the development of the compiler backend based on artificial intelligence, which includes:

[0035] The data acquisition module obtains the compiler backend code library including multiple chip architectures, extracts repair data from its historical commit records, uses the code before repair in the repair data as the error fragment, and the code after repair as the repair fragment; and the functions in the compiler backend code have corresponding description information;

[0036] The data acquisition module obtains the compiler backend code library including multiple architectures. For the program repair function, extract repair data from its historical commit records, use the code before repair in the repair data as the error fragment, and the code after repair as the repair fragment; for the code generation function, extract the corresponding description information of the functions in the compiler backend code;

[0037] The code extraction module divides the code in the compiler backend code library according to the code syntax to obtain structured code;

[0038] The code recognition module identifies the boundary positions of error fragments and repair fragments in the structured code, marks the error boundaries and repair boundaries in the structured code, and replaces the numerical constants, string literals, and enumeration prefix values in the structured code with preset standardized placeholders to obtain a normalized code;

[0039] The sample construction module constructs the normalized code into task samples under each chip architecture according to the chip architecture category. The task samples include: completion task samples, next sentence prediction task samples, code generation task samples, and program repair task samples;

[0040] The model fine-tuning module inputs the task prompts of the task samples into a large language model, and fine-tunes and trains the large language model according to the task results and the labels of the task samples to obtain a programming model that can complete completion tasks, next sentence prediction tasks, code generation tasks, and program repair tasks;

[0041] The model deployment module. The programming model is used to execute the compiler backend programming tasks of the specified architecture and integrate the executed programming code into the compiler backend code of the specified architecture; and during the execution process, the programming model provides code completion, next sentence suggestions, and repair suggestions according to the current programming code during the execution process. The programming model generates a corresponding function code template according to the input code description to assist in writing code.

[0042] The described compiler backend development assistance device based on artificial intelligence, wherein the sample construction module includes:

[0043] Part of the content of consecutive statements in the normalized code is intercepted as input, and the remaining part is used as a label to form the completion task sample;

[0044] Using the previous consecutive statements in the normalized code as input and the subsequent consecutive statements as labels, the next sentence prediction task sample is constructed;

[0045] Using the description information as input and the normalized code corresponding to the description information as a label, the code generation task sample is constructed;

[0046] Taking the code within the error boundary in the normalized code as input and the code within the repair boundary corresponding to the error boundary in the normalized code as a label, the program repair task sample is constructed;

[0047] The model fine-tuning module includes:

[0048] Select the current sample from the sample for the completion task, the sample for the next sentence prediction task, the sample for the code generation task, and the sample for the program repair task, and retrieve it through the context retriever. Select the data with the highest similarity to the current sample in the structured code and construct it with the current sample as the sample context template to provide input data for the fine-tuning training of the large language model.

[0049] The described artificial intelligence-based compiler backend development assistance device, wherein the model deployment module includes:

[0050] Integrate the programming model into a cross-platform code editor using the model encapsulation framework and extension tool integration to provide users with code completion, next sentence suggestions, program repair, and code generation functions;

[0051] When the programming model implements code completion, according to the currently written code during the execution process, predict and output the subsequent code snippet of the currently written code;

[0052] When the programming model provides next sentence suggestions, according to the context of the currently written code during the execution process, predict and output the next sentence of the currently written code;

[0053] When the programming model provides repair suggestions, according to the currently programmed code during the execution process, predict and output the error part of the currently edited code and the modification suggestions;

[0054] When the programming model provides function code template generation, according to the natural language code description input during the execution process, predict and generate function code that conforms to the natural language code description.

[0055] The present invention also proposes an electronic device, which includes the described artificial intelligence-based compiler backend development assistance device. The electronic device is either connected to an information display device, and the information display device is used to display code completion, next sentence suggestions, repair suggestions, or code descriptions with the display parameters, attributes set by the user, or through an artificial intelligence model.

[0056] The present invention also proposes a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps of the described artificial intelligence-based compiler backend development assistance method.

[0057] The present invention also proposes a computer program product, including a computer program. When the computer program is executed by a processor, it implements the steps of the described artificial intelligence-based compiler backend development assistance method.

[0058] As can be seen from the above solution, the advantages of the present invention are as follows: By constructing the above-mentioned multi-language compiler backend dataset, developing a dedicated large language model, and introducing a retriever based on In-Context Learning and its prompt construction method, a complete AI programming assistant is provided for compiler backend development. The multi-language, multi-target, and multi-task characteristics of the above-mentioned multi-language compiler backend dataset provide rich data resources for compiler backend development and support the generalization ability of the model on different processor types and programming languages. The dedicated large language model can provide a specially optimized language model for compiler backend development tasks by fine-tuning on the above dataset, improving the performance of code completion, generation, and repair. The retriever based on In-Context Learning and its prompt construction method provide an effective solution for users with limited resources through context learning technology, further improving the performance of the large language model in compiler backend development tasks. Generally speaking, the present invention can improve the efficiency and accuracy of compiler backend development, reduce development costs, and provide a multi-functional AI programming assistance tool for compiler backend developers. Description of the Drawings

[0059] Figure 1 Four task diagrams in the multi-language multi-task compiler backend dataset;

[0060] Figure 2 Prompt templates for four tasks and a diagram of the synthesized prompt template after retrieval by the retriever;

[0061] Figure 3 Diagram of program repair data annotation;

[0062] Figure 4 Workflow diagram of the AI programming assistant;

[0063] Figure 5 Workflow diagram of the retriever based on context learning;

[0064] Figure 6 Diagram of the deployment of the invention in the IDE and its working effect;

[0065] Figure 7 Flowchart of the method of the present invention;

[0066] Figure 8 Module diagram of the device of the present invention;

[0067] Figure 9 Schematic diagram of the structure of the first electronic device of the present invention;

[0068] Figure 10 Schematic diagram of the application environment structure of the first electronic device of the present invention;

[0069] Figure 11 This is a schematic diagram of the structure of the second electronic device of the present invention.

[0070] Reference numerals:

[0071] A - The first electronic device;

[0072] B - An artificial intelligence-based compiler backend development assistance device;

[0073] C - A data acquisition device;

[0074] D - An information display device;

[0075] 1000 - The second electronic device;

[0076] Ⅰ - A calculation unit;

[0077] Ⅱ - ROM;

[0078] Ⅲ - RAM;

[0079] Ⅳ - A bus;

[0080] Ⅴ - An interface;

[0081] Ⅵ - An input unit;

[0082] Ⅶ - An output unit;

[0083] Ⅷ - A storage medium;

[0084] Ⅸ - A communication unit. Detailed implementation manners

[0085] It should be noted that in this application, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device.

[0086] Without further limitation, an element defined by the statement "comprising one..." does not exclude the existence of another identical element in the process, method, article or device comprising the said element.

[0087] The processor described in the present invention is the control center of an electronic device, which can be a single processor or a collective term for multiple processing elements. For example, it can be one or more central processing units (CPUs), or an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present invention. For example: one or more digital signal processors (DSPs), or one or more field programmable gate arrays (FPGAs).

[0088] Optionally, the processor can execute various functions of the electronic device by running or executing software programs stored in the memory and invoking data stored in the memory.

[0089] In a specific implementation, as an embodiment, the processor can include one or more CPUs. Each of these processors can be a single-core processor (single-CPU) or a multi-core processor (multi-CPU). Here, the processor can refer to one or more devices, circuits, and / or processing cores for processing data (such as computer program instructions). The electronic device can include: servers, desktop computers, laptop computers, smartphones, tablets, embedded computers, etc., where the embedded computer includes vehicles and robots, etc.

[0090] The memory is used to store the software program for implementing the solution of the present invention and is controlled by the processor for execution. The specific implementation manner can refer to the above method embodiments and will not be elaborated here.

[0091] It should be noted that the structure of the electronic device shown in the drawings of the present invention does not limit it. The actual knowledge structure recognition device may include more or fewer components than shown in the drawings, or combine certain components, or have different component arrangements.

[0092] The above embodiments can be implemented in whole or in part by software, hardware (such as circuits), firmware, or any combination thereof. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wired (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that contains one or more collections of available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, or magnetic tape), an optical medium (such as a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state drive.

[0093] It should also be understood that the term "and / or" in this document is merely a description of the association relationship between associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. Here, A and B can be singular or plural. In addition, the character " / " in this document generally represents an "or" relationship between the associated objects before and after, but it may also represent an "and / or" relationship, which can be specifically understood with reference to the context before and after.

[0094] In the present invention, "at least one" means one or more, and "a plurality" means two or more. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b, or c can represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, and c can be single or multiple.

[0095] It should also be understood that in various embodiments of the present invention, the magnitudes of the sequence numbers of the above processes do not mean the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0096] In several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the couplings or direct couplings or communication connections shown or discussed with each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.

[0097] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0098] In addition, in each embodiment of the present invention, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.

[0099] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.

[0100] A core problem faced in the field of compiler backend development is the differences in different processor architectures, which forces developers to customize the compiler backend for each architecture and lack effective automated tools to support this process.

[0101] In the past few years, the machine learning community has achieved great success in AI code assistance for artificial intelligence (AI), from AI painting, AI chatbots to AI programming. As one of the representative works, GitHub and OpenAI collaborated to develop GitHub Copilot, and by fine-tuning GPT-3 on 54 million code repositories from GitHub, totaling over 159GB of code, it demonstrated the ability to assist in writing high-level computer programs. It can automatically complete code blocks, repetitive code, and read the user's problem description and automatically generate runnable code. However, these general code generation tools perform poorly in the field of compiler backend development because they cannot adapt to the programming characteristics of the compiler backend. The purpose of this invention is to create an AI programming model specifically designed for compiler backend development, improve development efficiency, and simplify the complex process of manually writing code. It uses an AI model to learn the existing C / C++ code of the compiler backend and the target description file code (TableGen, MD), assist in developing the compiler backend code for a new target platform, and ultimately achieve end-to-end support for compiler backend development.

[0102] To address this challenge, we created a dataset containing multiple programming languages, rich processor backend instances, and specific backend tasks, providing comprehensive data resources for training large language models (LLMs) suitable for backend development.

[0103] Another technical difficulty lies in how to make the LLM understand and generate code specific to the backend. To solve this problem, we proposed an LLM specifically for backend development. By fine-tuning the existing model on the above dataset, we developed two different parameter versions of the LLM, which can provide high-precision code completion, generation, and repair functions for compiler backend development tasks. Through sequence-to-sequence prediction training in the dataset, our LLM learned the structure and patterns of backend code, thus providing effective support for programmers in actual development.

[0104] In addition, to improve the performance of the LLM in an environment with limited computing resources, we also designed a retriever based on In-Context Learning. It automatically retrieves the most similar examples from the dataset based on context learning, constructs few-shot prompts, and further improves the accuracy of the LLM in backend development tasks.

[0105] To achieve the above technical effects, the present invention proposes the following key technical points:

[0106] Key Point 1. A Multilingual Dataset for Compiler Backend Development

[0107] The present invention proposes a multilingual dataset for compiler backend development. By integrating 183 backend codes of two major compilers, GCC and LLVM, covering three programming languages, C / C++, Machine Description (MD), and TableGen (TD), and designing structured data for four tasks in backend development (statement completion, next statement suggestion, code generation, program repair). Through steps such as code extraction, target feature value replacement, and program repair data annotation, the dataset realizes the standardized representation and semantic abstraction of compiler backend codes, providing end-to-end, high-quality, and multi-scenario domain-specific data support for model training.

[0108] Key Point 2. AI Programming Assistant for Compiler Backend Development

[0109] The present invention proposes a dedicated large language model based on a multilingual compiler backend, aiming to improve the automation degree and efficiency of compiler backend development. By fine-tuning on the above dataset, the model can provide a specially optimized language model for compiler backend development tasks. The model has two parameter versions, 1.5B (Billion, indicating the number of model parameters) and 7B, to adapt to different computing resource requirements. Through the LoRA (Low-Rank Adaptation) fine-tuning technique, while maintaining the model performance, the model reduces the computing resource requirements. The model can provide significant performance improvements in various compiler backend tasks, including statement-level code completion, next statement suggestion, code generation, and program repair. By training on the above dataset, the model can learn the specific patterns and structures of compiler backend codes, thus providing more accurate code completion and generation suggestions in actual development. In addition, the model also provides an easy-to-use tool for developers in the form of an integrated development environment extension (such as Visual Studio Code, a cross-platform source code editor released by Microsoft), further improving the efficiency of compiler backend development.

[0110] Key Point 3. Context-Based Retriever

[0111] The present invention introduces a retriever based on In-Context Learning and its prompt construction method, aiming to improve the performance of large language models in compiler backend development tasks by constructing a small number of example prompts. The retriever retrieves code snippets or functions similar to the user input from the above dataset, generates prompts containing examples, and thus improves the model's performance using context learning technology without parameter updates. The working process of the retriever includes extracting the task and target category from the prompt provided by the user, extracting data from the above dataset according to the task and target category, and performing data retrieval according to the code snippet or function description input by the user. In this way, the retriever can provide relevant context information for the model, thereby improving the accuracy of code completion, generation, and repair. In addition, the design of the retriever takes into account the limitations of computing resources, enabling it to operate effectively in resource-constrained environments, providing a practical solution for users without local computing resources.

[0112] To make the above features and effects of the present invention more clearly and understandably described, specific embodiments are hereinafter given and detailed descriptions are made in conjunction with the accompanying drawings of the specification. This specification discloses one or more embodiments containing the features of the present invention. The disclosed embodiments are only for illustrative purposes. The protection scope of the present invention is not limited to the disclosed embodiments, and the present invention is defined by the appended claims.

[0113] The working process of the present invention is divided into three stages: the dataset construction stage, the model training stage, and the deployment stage, plus an optional prompt construction and context enhancement stage. The dataset construction stage includes multi-source data collection and task design, the model training and deployment stage covers dedicated model fine-tuning and tool integration, and the compiler backend development assistance stage combines retrieval enhancement and context learning technologies.

[0114] In the dataset construction stage, first, multi-language backend codes are collected from mainstream compiler frameworks (LLVM, GCC), including C / C++, Machine Description (MD), and TableGen (TD) codes. And cross-platform task samples are constructed through structured processing. Taking the RISC-V backends of LLVM and GCC as an example, target description files and function implementations are extracted through code parsing tools, target-specific values are identified, and intermediate representations are constructed, such as numerical constants (such as instruction encodings), string literals (such as register names), and enumeration prefix values. Four types of tasks (statement completion, next suggestion, code generation, program repair) are designed for the development scenario, and repair data is constructed by analyzing error-fix pairs in the Git (a version control system that can manage different versions of code in a code repository) commit history. Specifically, keywords filtering (such as 'fix', 'bug') is used to identify backend-related modifications, and through <bugs> / <buge>and <fixs> / <fixe>Mark errors and repair boundaries, and finally form a dedicated dataset for the multi-language multi-task compiler backend covering 183 backends.

[0115] In the model training and deployment stages, based on the above-mentioned multi-language multi-task dataset, the pre-trained open-source large language models Qwen2.5-Coder-1.5B and Qwen2.5-Coder-7B are used as the basis. By designing dedicated task prompt templates, the model is fine-tuned end-to-end so that it can accurately understand the syntax structure of the compiler backend code and the mapping relationship of target specific values, thus achieving a high accuracy rate in tasks such as statement completion, next suggestion, code generation, and program repair; in this process, the parameter update is optimized through the LoRA technology, and multiple rounds of iterative adjustment are carried out using the training set and the validation set to ensure that the model can achieve good generalization ability and stable performance in various tasks. In the model deployment stage, the specially fine-tuned model is integrated locally, and real-time code completion, next suggestion, code generation, and program repair functions are realized through the mainstream IDE (Integrated Development Environment) extension plugin;

[0116] In the prompt construction and context enhancement stage, a retriever based on the In-Context Learning technology is introduced to automatically retrieve relevant examples according to the user's current input, construct Few-Shot Prompts, and provide sufficient context information for the model, thereby further improving the accuracy and adaptability of the generated results, and continuously optimizing the system performance through real-time interaction and user feedback mechanism. Finally, efficient and intelligent auxiliary support is provided for the compiler backend development.

[0117] The working process of the AI programming assistant is as Figure 4 shown. After the developer inputs different contents, the AI programming assistant will automatically generate the code for the corresponding task. The working process of the retriever based on In-Context Learning is as Figure 5 shown

[0118] Step 1: Multi-source Data Collection

[0119] Through a crawler script, open-source repositories are collected on code hosting platforms such as GitHub with the keyword "GCC / LLVM + backend", and incomplete code libraries are filtered out. Finally, 81 GCC backends and 102 LLVM backend codes are selected. At the same time, the official GCC source code (version range from 3.0 to 14.0) and the official LLVM source code (version range from 2.0.1 to 19.0.1) are downloaded synchronously to build a complete original code library.

[0120] In terms of program repair data collection, a crawler script is also used to extract repair data from the historical commit records of the code repository. Among them, the code before modification in each commit is regarded as the error fragment, while the code after modification is regarded as the repair fragment.

[0121] Function description information is mainly collected through two channels: on the one hand, it is directly extracted from the comments in the source code file; on the other hand, for LLVM, the information is extracted using the documents provided on its official Doxygen website.

[0122] Step 2: Structured Code Extraction

[0123] First, we preprocess the source code of each architecture Target, delete duplicate files and comments, and retain the function-level code units. Then, we use Tree-sitter (a tool for building syntax analysis generators) to parse the backend code, such as C / C++ function code, and divide the code into separate statements according to line terminators (such as ";", "{", "}"). For target description files (such as the MD file of GCC and the TableGen file of LLVM), they are split according to their respective domain syntax rules. For example, statements starting with the pattern "(define_*" are extracted from the.md file, and for the.td file, ";" is used as the statement terminator for separation.

[0124] Step 3: Target-specific Value Normalization

[0125] Identify three types of target-specific values in the code: numeric constants (such as instruction encodings), string literals (such as register names), and enumeration prefix values (such as RISCVMCExpr::VK_RISCV_LO). Replace them with standardized placeholders through regular rules: numeric → <NUM_LIT>, string → <STR_LIT>, enumeration → <ISA_LIT>. For example, replace the code statement "OS.write("\x20",1);" with "OS.write(<STR_LIT>,<NUM_LIT>);", and replace the code statement "case RISCVMCExpr::VK_RISCV_LO:" with "case <ISA_LIT>. Also, as Figure 3 shown, corresponding to the error boundary annotation of program repair data in Step 1 ( <bugs> ... <buge>) and repair the boundary( <fixs> ... <fixe>).

[0126] Step 4: Multi-task Sample Construction

[0127] Generate task data for four types of development scenarios:

[0128] 1) Code completion: intercept 5 consecutive statements in the function, use the first n-1 statements + the first 50%-90% of the nth statement as input, and the rest as output ( Figure 1 (a)).

[0129] 2) Next sentence prediction: the first n-1 sentences are used as input, and the nth complete sentence is output ( Figure 1 (b)).

[0130] 3) Code generation: Pairing a function description with a code template that removes the target value ( Figure 1 (c)).

[0131] 4) Program fixes: Convert the error-fix pairs in the GitHub commit history into annotation format ( Figure 1 (d)).

[0132] Step 5: Dataset Partition & Validation

[0133] Divide the dataset by target type: RISC-V (CPU), ARC (MPU), and NVPIX (GPU) as test sets (CPU is the central processing unit, which is responsible for performing general computing tasks in the computer; MPU is a microprocessor unit, which can be understood as a CPU optimized for embedded applications; GPU is a graphics processor, used for image rendering and parallel computing) and the rest are divided into training / validation sets at a ratio of 9:1. LLVM's RISCY backend is excluded when dividing the training set (because it is a RISC-V customized variant with code overlap), but it is retained as an iterative extension verification case

[0134] Step 6: Specialized Model Fine-tuning

[0135] Based on the open source models Qwen2.5-Coder-1.5B and Qwen2.5-Coder-7B, the LoRA low-rank adapter is used to fine-tune the model for multiple tasks. During the fine-tuning process, the input consists of a prompt template consisting of a system instruction prompt (for example, "You are a RISC-V backend expert and proficient in C / C++ code"), a task description, and a code context (for example, Figure 2 As shown, the Input code is the input code, the Input Description is the description of the input function, the Ground Truth is the output code, and the Example is the retrieved result. The underlined part is the input information that needs to be distinguished, such as the programming language (C / C++ code or target description language), the type of target architecture, the type of compiler, the input code, the function description, or the error code. The retriever will retrieve a series of results. Due to space limitations, Example (2-n) is omitted, and the combined results after retrieval of other tasks are the same as Figure 2 (e) similar). The key parameters set during model fine-tuning are: rank r = 32, scaling factor α = 16, learning rate 5×10 -4 , maximum input length of 1024 tokens, and batch size of 4. (Among them, the rank r is the rank of the LoRA adapter, which determines the number of parameters of the adapter; the scaling factor α is the learning rate of the LoRA adapter; the learning rate is the learning rate of model fine-tuning, which determines the amplitude of model parameter updates; the batch size is the number of samples for each iteration of training) After training for 50 rounds on NVIDIA A100 GPUs, specialized models of 1.5B and 7B versions suitable for different computing resource scenarios are generated respectively.

[0136] Among them, the Qwen2.5-Coder-1.5B and Qwen2.5-Coder-7B models are two versions of the code large model, where 1.5b and 7b refer to the number of model parameters. The present invention can adopt multiple models to facilitate users to select models according to their own computing resources.

[0137] Step 7: Development Tool Integration (IDE Tool Integration)

[0138] Using tools such as a model packaging framework (Ollama, used to package and deploy the fine-tuned and trained model) and an extension plugin (Continue, a tool for developing with local models), the fine-tuned model is encapsulated and integrated into the Visual Studio Code extension of the cross-platform code editor (see Figure 6 ). This cross-platform code editor can be installed on a local server or on the developer's personal computer. During the code editing (compiler backend development) process, this extension can provide real-time code completion ( Figure 6 ②) and next sentence suggestions ( Figure 6 ③) to help developers quickly write code; when the user marks an error code segment, the system will automatically generate interactive repair suggestions ( Figure 6 ④), effectively reducing the debugging difficulty; after the user inputs a natural language description, the extension can also automatically generate the corresponding function template ( Figure 6 ⑤). After development is completed, integrate the code after testing into the compiler framework or toolchain to ensure seamless connection with the front-end and middle-end of the compiler, facilitating users to use it on a specific Target (architecture).

[0139] Step 8: Retrieval-Augmented Inference (CB-Retriever Enhanced Inference)

[0140] Build a retriever based on the dataset partitioned in Step 5 to automatically screen context examples in the compiler backend code data in the dataset that match the task type and target category input by the user. The specific process is as follows: First, calculate the similarity with the samples in the dataset based on the code snippet or natural language description provided by the user, and select the three closest examples among the samples with the same task type and target category (this value can be adjusted according to computing resources); subsequently, append these retrieved examples to the original prompt to construct a few-shot context (see the workflow in Figure 5 , and the generated prompt template is as Figure 2 (e) shown), thereby improving the accuracy of the code generated by the model.

[0141] The following is a system embodiment corresponding to the above method embodiment. This embodiment can be implemented in cooperation with the above embodiment. The relevant technical details mentioned in the above embodiment are still valid in this embodiment. To avoid repetition, they will not be elaborated here. Correspondingly, the relevant technical details mentioned in this embodiment can also be applied in the above embodiment.

[0142] As Figure 8 shown, the present invention also proposes an artificial intelligence-based compiler backend development assistance device, which includes:

[0143] A data acquisition module that obtains a compiler backend code library including multiple chip architectures, extracts repair data from its historical commit records, uses the code before repair in the repair data as error fragments, and the code after repair as repair fragments; and the functions in the compiler backend code have corresponding description information;

[0144] A data acquisition module that obtains a compiler backend code library including multiple architectures, for the program repair function, extracts repair data from its historical commit records, uses the code before repair in the repair data as error fragments, and the code after repair as repair fragments; for the code generation function, extracts the functions in the compiler backend code that have corresponding description information;

[0145] A code extraction module that divides the code in the compiler backend code library according to the code syntax to obtain structured code;

[0146] A code recognition module identifies the boundary positions of error segments and repair segments in the structured code to mark the error boundaries and repair boundaries in the structured code, and replaces the numerical constants, string literals and enumeration prefix values ​​in the structured code with preset standardized placeholders to obtain standardized code;

[0147] A sample construction module, which constructs the normalized code into task samples under each chip architecture according to the chip architecture category. The task samples include: completion task samples, next sentence prediction task samples, code generation task samples, and program repair task samples;

[0148] The model fine-tuning module inputs the task prompt of the task sample into the large language model, and fine-tunes the large language model according to the task result and the label of the task sample to obtain a programming model that can complete the completion task, the next sentence prediction task, the code generation task, and the program repair task;

[0149] Model deployment module, the programming model is used to execute compiler back-end programming tasks of the specified architecture, and integrate the programming code obtained by execution into the compiler back-end code of the specified architecture; and during the execution process, the programming model provides code completion, next sentence suggestions and repair suggestions based on the current programming code in the execution process. The programming model generates the corresponding function code template based on the input code description to assist in code writing.

[0150] The artificial intelligence-based compiler backend development auxiliary device, wherein the sample construction module includes:

[0151] Part of the continuous sentences in the normalized code is intercepted as input, and the remaining part is the label, which constitutes the completion task sample;

[0152] Taking the previous continuous sentence in the normalized code as input and the subsequent continuous sentence as label, construct the next sentence prediction task sample;

[0153] Taking the description information as input and the normalized code corresponding to the description information as a label, constructing the code generation task sample;

[0154] The code in the normalized code that is within the error boundary is used as input, and the code in the repair boundary corresponding to the error boundary in the normalized code is used as a label to construct a sample of the program repair task;

[0155] The model fine-tuning module includes:

[0156] Select the current sample from the completion task sample, the next sentence prediction task sample, the code generation task sample, and the program repair task sample, retrieve it through the context retriever, and select the data with the highest similarity to the current sample in the structured code to construct a sample context template with the current sample, providing input data for the fine-tuning training of the large language model.

[0157] The described compiler backend development assistance device based on artificial intelligence, wherein the model deployment module includes:

[0158] Integrate the programming model into a cross-platform code editor using the model encapsulation framework and extension tool integration, providing users with code completion, next sentence suggestions, program repair, and code generation functions;

[0159] When the programming model implements code completion, according to the currently written code during the execution process, predict and output the subsequent code snippet of the currently written code;

[0160] When the programming model provides next sentence suggestions, according to the context of the currently written code during the execution process, predict and output the next line of code of the currently written code;

[0161] When the programming model provides repair suggestions, according to the currently programmed code during the execution process, predict and output the error part of the currently edited code and the modification suggestions;

[0162] When the programming model provides function code template generation, according to the natural language code description input during the execution process, predict and generate function code that conforms to the natural language code description.

[0163] As Figure 9 shown, in another embodiment of the present invention, a first electronic device A is also proposed, including the described compiler backend development assistance device based on artificial intelligence.

[0164] As Figure 10 shown, the first electronic device A can also be connected to the data acquisition device C and the information display device D through a wired or wireless information transmission scheme. The data acquisition device C is used to collect and obtain the video to be identified and classified, such as the sorting video described in the embodiments of the present invention. The information display device D is used to display the code completion, next sentence suggestions, repair suggestions, or code descriptions analyzed by the present invention.

[0165] Among them, the information display device D can process the data output by the first electronic device A based on the information display mechanism to improve the readability of the data output by the first electronic device A. The information display mechanism can be preset manually. For example, the data output by the first electronic device A is visually displayed, and it can display according to the display parameters and / or attributes set by the user. The display parameters can be, for example, the display data range, and the display attributes can be, for example, the display font, color, whether to scroll and play, etc. Present the key information specified by the user to the user, enabling the user to understand this information more timely without having to access the secondary page or scroll the page, saving the user's operations. Or the information display mechanism can be an artificial intelligence AI display model, which can learn the user's key focus information based on the user's previous usage habits, such as viewing duration, click times, editing times, etc., and then automatically present rich and necessary key information to the user.

[0166] The present invention also provides a computer program product, the computer program product includes a computer program, the computer program can be stored on a readable storage medium, and when the computer program is executed by a processor, the computer can execute the artificial intelligence-based compiler backend development assistance method provided by each of the above methods.

[0167] In another embodiment, the present invention further provides a storage medium VIII for storing a computer program for executing the method for assisting in the development of the compiler backend based on artificial intelligence. It should be understood that the storage medium in the embodiments of the present invention may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of random access memory (RAM) are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchlink dynamic random access memory (SLDRAM), and direct rambus random access memory (DR RAM).

[0168] Figure 11 FIG. shows a schematic block diagram of a second electronic device 1000 that can be used to implement the embodiments of the present invention. The second electronic device 1000 is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The second electronic device 1000 may also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed herein. The second electronic device 1000 may be the same as or different from the first electronic device A.

[0169] The second electronic device 1000 includes a computing unit I, which can perform various appropriate actions and processes according to computer programs stored in a read-only memory II (ROM) or computer programs loaded from a storage medium VIII into a random access memory (RAM) III. In the RAM III, various programs and data required for the operation of the device 1000 can also be stored. The computing unit I, the ROM II, and the RAM III are connected to each other through a bus IV. An input / output (I / O) interface V is also connected to the bus IV.

[0170] Multiple components in the second electronic device 1000 are connected to the I / O interface V, including: an input unit VI, such as a keyboard, a mouse, etc.; an output unit VII, such as various types of displays, speakers, etc.; a storage medium VIII, such as a magnetic disk, an optical disc, etc.; and a communication unit IX, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit IX allows the second electronic device 1000 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0171] The computing unit I can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit I include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit I executes the various methods and processes described above, such as method steps S1 - S6. For example, in some embodiments, the method can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage medium VIII. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 1000 via the ROM II and / or the communication unit IX. When the computer program is loaded into the RAM III and executed by the computing unit I, one or more steps of the method described above can be executed. Alternatively, in other embodiments, the computing unit I can be configured to execute the method in any other appropriate way (e.g., by means of firmware).

[0172] Although the embodiments of the present invention have been disclosed as above, they are not limited only to the applications listed in the specification and the embodiments. It can be fully applied to various fields suitable for the present invention. For those familiar with the art, additional modifications can be easily made. Therefore, without departing from the general concept defined by the claims and the equivalent scope, the present invention is not limited to the specific details and the illustrated and described examples here.< / fixe> < / fixs> < / buge> < / bugs> < / fixe> < / fixs> < / buge> < / bugs>

Claims

1. A compiler backend development auxiliary method based on artificial intelligence, characterized in that: include: In the data collection step, a compiler backend code base including multiple architectures is obtained. For the program repair function, repair data is extracted from its historical submission records. The code before the repair in the repair data is used as the error fragment, and the code after the repair is used as the repair fragment. For the code generation function, the corresponding description information of the function in the compiler backend code is extracted; The code extraction step divides the code in the backend code base of the compiler according to the code syntax to obtain structured code; A code recognition step, identifying the boundary positions of the error segments and the repair segments in the structured code, marking the error boundaries and the repair boundaries in the structured code, and replacing the numerical constants, string literals and enumeration prefix values ​​in the structured code with preset standardized placeholders to obtain a standardized code; The sample construction step constructs the normalized code into task samples under each chip architecture according to the chip architecture category. The task samples include: completion task samples, next sentence prediction task samples, code generation task samples, and program repair task samples; The model fine-tuning step is to input the task prompt of the task sample into the large language model, and fine-tune the large language model according to the task result and the label of the task sample to obtain a programming model that can complete the completion task, the next sentence prediction task, the code generation task, and the program repair task; Model deployment step, the programming model is used to execute the compiler back-end programming task of the specified architecture, and integrate the programming code obtained by execution into the compiler back-end code of the specified architecture; and during the execution process, the programming model provides code completion, next sentence suggestions and repair suggestions based on the current programming code in the execution process. The programming model generates the corresponding function code template based on the input code description to assist in code writing.

2. The artificial intelligence-based compiler backend development auxiliary method according to claim 1, characterized in that: The sample build steps include: Part of the continuous sentences in the normalized code is intercepted as input, and the remaining part is the label, which constitutes the completion task sample; Taking the previous continuous sentence in the normalized code as input and the subsequent continuous sentence as label, construct the next sentence prediction task sample; Taking the description information as input and the normalized code corresponding to the description information as a label, constructing the code generation task sample; The code in the normalized code that is within the error boundary is used as input, and the code in the repair boundary corresponding to the error boundary in the normalized code is used as a label to construct a sample of the program repair task.

3. The compiler backend development auxiliary method based on artificial intelligence as claimed in claim 2, characterized in that: The model fine-tuning steps include: The current sample is selected from the completion task sample, the next sentence prediction task sample, the code generation task sample and the program repair task sample, and is retrieved through the context retriever. The data with the highest similarity to the current sample in the structured code is selected and constructed as a sample context template with the current sample to provide input data for fine-tuning training of the large language model.

4. The artificial intelligence-based compiler backend development auxiliary method according to claim 1, 2 or 3, characterized in that: The model deployment steps include: Use the model encapsulation framework and extension tools to integrate the programming model into the cross-platform code editor to provide users with code completion, next sentence suggestions, program repair, and code generation functions; When the programming model implements code completion, it predicts and outputs subsequent code snippets of the currently written code based on the currently written code during execution; When the programming model provides next sentence suggestions, it predicts and outputs the next sentence of the currently written code based on the context of the currently written code during execution; When the programming model provides repair suggestions, it predicts and outputs the erroneous parts of the current editing code and modification suggestions based on the current programming code during execution; This programming model provides function code template generation, which predicts and generates function code that conforms to the natural language code description based on the natural language code description input during the execution process.

5. A compiler backend development auxiliary device based on artificial intelligence, characterized in that: include: A data acquisition module obtains a compiler backend code library including multiple chip architectures, extracts repair data from its historical submission records, wherein the code before repair in the repair data is used as an error fragment, and the code after repair is used as a repair fragment; and the functions in the compiler backend code have corresponding description information; The data acquisition module obtains the compiler backend code base including multiple architectures, extracts the repair data from its historical submission records for the program repair function, and uses the code before the repair as the error fragment and the code after the repair as the repair fragment in the repair data; extracts the corresponding description information of the function in the compiler backend code for the code generation function; The code extraction module divides the code in the backend code base of the compiler according to the code syntax to obtain structured code; A code recognition module identifies the boundary positions of error segments and repair segments in the structured code to mark the error boundaries and repair boundaries in the structured code, and replaces the numerical constants, string literals and enumeration prefix values ​​in the structured code with preset standardized placeholders to obtain standardized code; A sample construction module, which constructs the normalized code into task samples under each chip architecture according to the chip architecture category. The task samples include: completion task samples, next sentence prediction task samples, code generation task samples, and program repair task samples; The model fine-tuning module inputs the task prompt of the task sample into the large language model, and fine-tunes the large language model according to the task result and the label of the task sample to obtain a programming model that can complete the completion task, the next sentence prediction task, the code generation task, and the program repair task; Model deployment module, the programming model is used to execute compiler back-end programming tasks of the specified architecture, and integrate the programming code obtained by execution into the compiler back-end code of the specified architecture; and during the execution process, the programming model provides code completion, next sentence suggestions and repair suggestions based on the current programming code in the execution process. The programming model generates the corresponding function code template based on the input code description to assist in code writing.

6. The artificial intelligence-based compiler backend development auxiliary device according to claim 1, characterized in that: This sample building block includes: Part of the continuous sentences in the normalized code is intercepted as input, and the remaining part is the label, which constitutes the completion task sample; Taking the previous continuous sentence in the normalized code as input and the subsequent continuous sentence as label, construct the next sentence prediction task sample; Taking the description information as input and the normalized code corresponding to the description information as a label, constructing the code generation task sample; The code in the normalized code that is within the error boundary is used as input, and the code in the repair boundary corresponding to the error boundary in the normalized code is used as a label to construct a sample of the program repair task; The model fine-tuning module includes: The current sample is selected from the completion task sample, the next sentence prediction task sample, the code generation task sample and the program repair task sample, and is retrieved through the context retriever. The data with the highest similarity to the current sample in the structured code is selected and constructed as a sample context template with the current sample to provide input data for fine-tuning training of the large language model.

7. The artificial intelligence-based compiler backend development auxiliary device according to claim 5 or 6, characterized in that: The model deployment module includes: Use the model encapsulation framework and extension tools to integrate the programming model into the cross-platform code editor to provide users with code completion, next sentence suggestions, program repair, and code generation functions; When the programming model implements code completion, it predicts and outputs subsequent code snippets of the currently written code based on the currently written code during execution; When the programming model provides next sentence suggestions, it predicts and outputs the next sentence of the currently written code based on the context of the currently written code during execution; When the programming model provides repair suggestions, it predicts and outputs the erroneous parts of the current editing code and modification suggestions based on the current programming code during execution; This programming model provides function code template generation, which predicts and generates function code that conforms to the natural language code description based on the natural language code description input during the execution process.

8. An electronic device, characterized in that: It includes an artificial intelligence-based compiler back-end development assistant device as described in claims 5-7, and the electronic device is connected to an information display device, which is used to display code completion, next sentence suggestions, repair suggestions or code descriptions based on display parameters and attributes set by the user or through an artificial intelligence model.

9. A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the artificial intelligence-based compiler back-end development auxiliary method as described in any one of claims 1 to 4.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the compiler back-end development auxiliary method based on artificial intelligence described in any one of claims 1-4 are implemented.

Citation Information

Cited By

  • End-to-end intelligent system development method and system based on AI large model

    CN120950041A

  • Automatic code generation system and method

    CN121387267A