A compilation method, device, and medium for multi-language mixed code
Through progressive scanning and semantic correlation map construction, the problem of inefficient compilation of multilingual hybrid code is solved, efficient and accurate compilation and optimization are achieved, and the quality and performance of multilingual hybrid code is improved.
Patent Information
- Application Number
- CN202510099630.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-01-22
AI Technical Summary
Existing compiler technology is difficult to effectively handle multilingual mixed code, resulting in low compilation efficiency, lots of redundant code, insufficient optimization strategies, affecting development efficiency and quality.
Through progressive scanning, multilingual code snippets are identified, language identification is determined, semantic association maps are constructed, deep semantic analysis and transformation are performed, intermediate forms are generated, and the final code format is optimized according to the target platform environment.
Improves compilation accuracy and efficiency, reduces redundant steps and resource consumption, discovers deep logic errors, and improves the quality and performance of multilingual hybrid code.
Smart Images

Figure CN119536743B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular, to a compilation method, device, and medium for multi-language mixed code. Background Art
[0002] In today's digital age, the complexity of software systems is increasing at an unprecedented rate. To cope with the increasingly complex business requirements and technical challenges, multi-language mixed programming has become the norm in the field of software development. However, most of the existing compiler technologies are still limited to the design and development of single programming languages and are unable to handle multi-language mixed code effectively.
[0003] During the compilation process, traditional compilers only perform static code analysis and transformation based on the language syntax rules and semantic specifications they support, lacking a deep understanding of the complex semantic relationships when multiple languages interact. This results in the compiler's difficulty in accurately capturing the deep semantic associations between different language code segments when dealing with multi-language mixed code, such as cross-language function calls, maintaining the consistency of shared data, and the collaborative work between different language control flows and data structures.
[0004] This lack of semantic understanding brings a series of serious problems. First, in terms of compilation efficiency, since the compiler cannot optimize multi-language mixed code as a whole, it often performs isolated and inefficient compilation processing on different language code segments, resulting in a large amount of redundant intermediate code and unnecessary runtime overhead, making the compilation process time-consuming and seriously affecting development efficiency. Second, in terms of optimization strategies, existing compilers also have difficulty formulating effective optimization measures for the characteristics of multi-language mixed code. Summary of the Invention
[0005] To solve the above problems, this application proposes a compilation method for multi-language mixed code, including: scanning a multi-language mixed code file line by line to identify code segments corresponding to multiple languages and determining the language identifiers corresponding to the code segments; determining the corresponding parsing rules according to the language identifiers and performing semantic parsing on the code segments according to the parsing rules to determine the semantic association graph between the multiple code segments; converting the multiple code segments according to the semantic association graph to determine an intermediate form; determining the operating environment of the target platform and converting the intermediate form according to the operating environment to convert it into the code format corresponding to the target platform.
[0006] In one example, a multi - language mixed - code file is scanned line by line, specifically including: determining the feature fields corresponding to the multi - language, scanning the code file line by line according to the feature fields to determine whether the code file contains the feature fields; if the code file contains the feature fields, determining the code snippets according to the feature fields, determining the corresponding language identifiers according to the feature fields, and marking the code snippets according to the language identifiers.
[0007] In one example, determining the corresponding parsing rules according to the language identifiers and performing semantic parsing on the code snippets according to the parsing rules, specifically including: determining the corresponding grammar according to the language identifiers to determine the corresponding parsing rules according to the grammar; constructing a model according to the parsing rules to determine a semantic parsing model, and the functions of the semantic parsing model include variable declaration and usage analysis, function call parsing, data type derivation, control - flow analysis, interface and dependency relationship analysis between code segments.
[0008] In one example, determining the semantic association graph between multiple code snippets, specifically including: performing semantic parsing on multiple code snippets through the semantic parsing model to determine the function call relationships between multiple code snippets, and determining semantic associations according to the function call relationships; determining the data sharing mechanism and interface specifications between multiple code snippets, and determining the semantic association graph according to the semantic associations, the data sharing mechanism and the interface specifications to display the interaction paths and dependency networks between multiple code snippets through the semantic association graph.
[0009] In one example, converting multiple code snippets according to the semantic association graph to determine an intermediate form, specifically including: determining the syntactic structure information of the code snippets according to the semantic association graph and determining the semantic relationships between multiple code snippets; performing mapping conversion according to the syntactic structure and the semantic relationships to determine the intermediate form.
[0010] In one example, the method further includes: determining the usage frequency and declaration cycle of variables in the code snippets, and determining the registers corresponding to the code snippets according to the usage frequency and the declaration cycle.
[0011] In one example, the method further includes: determining the object reference mechanism corresponding to the multi - language and laying out the code snippets according to the object reference mechanism.
[0012] In one example, after determining the language identifier corresponding to the code snippet, the method further includes: pre - processing the code snippet, and the pre - processing process includes removing comments and processing whitespace characters.
[0013] On the other hand, the present application also provides a compilation device for multi-language hybrid code, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the compilation device for multi-language hybrid code is enabled to perform: scanning a multi-language hybrid code file line by line to identify code segments corresponding to multiple languages and determining language identifiers corresponding to the code segments; determining corresponding parsing rules according to the language identifiers and semantically parsing the code segments according to the parsing rules to determine a semantic association graph between the multiple code segments; converting the multiple code segments according to the semantic association graph to determine an intermediate form; determining the operating environment of a target platform and converting the intermediate form according to the operating environment to convert it into a code format corresponding to the target platform.
[0014] On the other hand, the present application also provides a non-volatile computer storage medium storing computer-executable instructions, which are configured to: scan a multi-language hybrid code file line by line to identify code segments corresponding to multiple languages and determine language identifiers corresponding to the code segments; determine corresponding parsing rules according to the language identifiers and semantically parse the code segments according to the parsing rules to determine a semantic association graph between the multiple code segments; convert the multiple code segments according to the semantic association graph to determine an intermediate form; determine the operating environment of a target platform and convert the intermediate form according to the operating environment to convert it into a code format corresponding to the target platform.
[0015] This application identifies multi-language code snippets through line-by-line scanning and accurately determines their language identifiers, laying a solid foundation for subsequent processing. This process effectively avoids parsing errors caused by language confusion and improves the accuracy and efficiency of compilation. This application selects corresponding parsing rules based on the language identifier for semantic parsing, which can deeply analyze the semantic relationships between code snippets and construct a detailed semantic association graph. This not only helps to understand the complex interactions of the code but also provides strong support for subsequent optimization and transformation. Converting multi-language code snippets into a unified intermediate form effectively integrates the syntax and semantic information of different languages, creating favorable conditions for the optimization of cross-language code and cross-platform deployment. This application fully considers the operating environment of the target platform to ensure that the intermediate form can be smoothly converted into the code format corresponding to the target platform, enhancing the flexibility and applicability of the compilation method and enabling it to be widely applied to various platforms and environments. This application has many advantages such as high accuracy, strong flexibility, and good optimization effect, providing strong guarantee for cross-language programming and cross-platform deployment. Through semantic perception and a unified compilation process, this application can co-process multi-language code during compilation, reducing unnecessary compilation steps and resource consumption, thus significantly improving the compilation efficiency. Based on semantic analysis, these errors can be accurately detected during the compilation stage, and through cross-language semantic association analysis, some deep-seated logical errors can also be discovered, effectively improving the quality of multi-language mixed code. Through intermediate representation optimization techniques such as cross-language function inlining, shared data access optimization, and instruction scheduling optimization, the performance potential of multi-language mixed code is fully exploited, improving the execution speed and resource utilization rate of the program. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The drawings described herein are used to provide a further understanding of the present application and form a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings:
[0017] Figure 1 is a schematic flow chart of a compilation method for multi-language mixed code in an embodiment of the present application;
[0018] Figure 2 is a schematic diagram of a compilation device for multi-language mixed code in an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0019] To make the objectives, technical solutions, and advantages of the present application clearer, the technical solutions of the present application will be clearly and completely described below in conjunction with the specific embodiments of the present application and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0020] The technical solutions provided by each embodiment of the present application will be described in detail below with reference to the accompanying drawings.
[0021] The deficiencies of existing compiler technologies in the scenario of mixed programming of multiple languages have become the key bottleneck restricting the improvement of software system development efficiency, quality, and performance. There is an urgent need for an innovative method and system that can effectively process mixed-language code and perform precise compilation based on semantics to meet the increasingly complex modern software development needs.
[0022] As Figure 1 shown, to solve the above problems, a compilation method for mixed-language code provided by an embodiment of the present application includes:
[0023] S101. Scan the mixed-language code file line by line to identify code segments corresponding to multiple languages and determine the language identifiers corresponding to the code segments.
[0024] The core of mixed-language code lies in the mutual cooperation of different programming languages at the semantic level and the runtime environment level. From a semantic perspective, the key lies in identifying and matching functionally and logically equivalent expression forms in different languages to ensure that they can seamlessly integrate into the overall logical flow of the program and cooperate. For example, function calls in Python and function calls in C++ both aim to execute specific code blocks to complete tasks at the semantic level, and through interface definitions and conversion mechanisms, these semantics can be effectively transmitted between different languages. Considering from the dimension of the runtime environment, the essence of mixed-language code is reflected in how the runtime systems of different languages share and coordinate resources provided by the operating system, such as memory, processor time, etc. Although the code segments of different languages have their own unique layouts and management mechanisms in memory, through means such as shared libraries and memory-mapped files, the smooth transmission of data and control flow between different language codes is achieved, and then the various functions of the software system are jointly realized.
[0025] First, preprocess the code and mark the language identification. When processing a mixed-language code file, the system will scan the input content line by line to accurately identify and distinguish code segments in different languages, and then attach corresponding language identification tags to each identified code segment. During this process, preliminary code preprocessing operations will also be performed, including removing comments, simplifying whitespace characters, etc., aiming to simplify subsequent analysis steps while ensuring that the structural information and semantic clues of the code are retained.
[0026] In one embodiment, during the code preprocessing and language identification stage, for a mixed file containing C++ and Python code, it is first scanned line by line. When a line starting with the ".cpp" extension or containing C++ specific syntax keywords, such as "class", "template", etc., is detected, the system determines that the segment is a C++ code snippet and attaches a "C++" identification tag. When encountering a line that conforms to Python syntax features, such as a format that relies on indentation, containing keywords "def", "import", etc., it is determined to be a Python code snippet and attached with a "Python" identification tag. In this process, the comments in the code are removed, and consecutive blank characters are merged or uniformly formatted, but structural symbols such as brackets and semicolons in the code are carefully retained to ensure that the integrity of the logical structure of the code is not affected.
[0027] S102: Determine a corresponding parsing rule according to the language identifier, and perform semantic parsing on the code snippet according to the parsing rule to determine a semantic association graph between the plurality of code snippets.
[0028] In order to cope with a variety of common programming languages, we have specially designed their own semantic analysis units. Each unit builds a semantic analysis model based on the corresponding language's grammatical rules, semantic specifications, and standard library function definitions and other professional knowledge. These models have the ability to perform in-depth semantic analysis of the code snippets of the language they belong to, covering the analysis of variable declarations and uses, parsing of function calls, derivation of data types, combing of control flows, and detailed analysis of interfaces and dependencies between code segments.
[0029] After completing the semantic analysis of the code snippets in each language, a semantic association graph between code snippets in different languages is constructed based on the function call relationships, data sharing mechanisms, and language-specific interaction protocols in the code, such as the cross-language interface standards defined in some multi-language frameworks. This graph can clearly reveal the semantic interaction paths and dependency networks between the various parts of the multi-language mixed code, providing comprehensive and detailed semantic information support for subsequent compilation and optimization steps.
[0030] In one embodiment, during the process of constructing the semantic analysis unit, based on the C++ standard grammar rules, a semantic model capable of deeply parsing the complex grammar structure of C++ is built. This model includes classes, templates, overloaded functions, etc. Specifically, when parsing variable declarations in C++ code, the model accurately derives the type and lifecycle of variables based on details such as type modifiers and scope qualifiers of the variables. When performing function call parsing, it properly handles complex situations such as function overloading and default parameters. For Python, according to its dynamic type system and unique indentation grammar rules, a semantic analysis model is also constructed. In the data type derivation process, the model dynamically and accurately determines the type of a variable based on its assignment operation. When dealing with function calls, it flexibly copes with the diverse function parameter passing methods in Python, such as positional arguments, keyword arguments, and variable arguments, etc.
[0031] In one embodiment, during the cross - language semantic association and integration phase, when a function in C++ is called by Python code, it is necessary to analyze the function call statement in C++ and the corresponding function call definition in Python to accurately establish the semantic association between the two. For example, if there is a call like lib.add_numbers(5, 3) in Python code, a direct semantic connection needs to be established with the add_numbers function in C++; similarly, the call to lib.multiply_numbers(2,4) needs to be associated with the multiply_numbers function in C++. In addition, for the data shared between C++ and Python, such as global_variable, it is necessary to deeply analyze its access patterns and sharing mechanisms in the two languages and integrate this information into the semantic association graph to ensure the accuracy and efficiency of cross - language data interaction.
[0032] S103. Convert the multiple code segments according to the semantic association graph to determine the intermediate form.
[0033] Based on the constructed semantic association graph, the multi - language mixed code is converted into a unified intermediate representation form. This intermediate representation not only comprehensively retains the syntax structure information of the code but also deeply incorporates cross - language semantic relationships. After successfully generating the intermediate representation, a series of optimization operations are carried out according to the semantic information therein. For example, for functions frequently called across languages, in - line optimization is implemented to improve efficiency; for access to shared data, optimization is carried out to reduce unnecessary data transmission overhead; at the same time, according to the execution flow and semantic logic of the code, fine - grained instruction scheduling optimization is carried out in order to further improve the overall performance.
[0034] In one embodiment, during the intermediate representation generation and optimization phase, a unified intermediate representation form is designed. This form uses a graph-based data structure to accurately map the syntax structures and semantic relationships of C++ and Python codes. In this graph, nodes represent core elements such as variables, functions, and code blocks, while edges depict data dependencies, control flow associations, and cross-language function call relationships, etc. In the optimization process, Python functions that are frequently called in C++ are analyzed and identified, especially those with relatively simple functions. For these functions, an inlining strategy is adopted to directly incorporate them into the corresponding intermediate representation part of the C++ code, thereby effectively reducing the additional overhead of function calls. At the same time, in-depth optimization is carried out for access operations on shared data. When multiple redundant data transmissions are detected, the data access path is optimized by introducing an intermediate caching mechanism or directly using memory mapping technology to reduce the number of data transmissions, thereby improving the overall performance of the system.
[0035] S104. Determine the operating environment of the target platform to convert the intermediate form according to the operating environment into the code format corresponding to the target platform.
[0036] According to the instruction set architecture of the target platform and the specific requirements of the runtime environment, the optimized intermediate representation is accurately converted into target machine code or a code format directly executable by the target platform. In this conversion process, the unique operating characteristics of multi-language hybrid code on the target platform are considered, and based on this, meticulous work such as register allocation, memory layout adjustment, and adaptation of the exception handling mechanism is carried out to ensure that the generated target code can run efficiently and stably on the target platform.
[0037] In one embodiment, if the target platform is an operating system based on the x86 architecture, the optimized intermediate representation is accurately converted into x86 machine code. When performing register allocation, the usage frequency and life cycle of variables are considered, so as to scientifically and reasonably allocate general-purpose registers and special-purpose registers. For the adjustment of the memory layout, the storage requirements of C++ and Python codes in memory are comprehensively analyzed. For example, the layout method of objects in C++ and the reference mechanism of objects in Python, and based on this, an optimized layout is carried out in order to improve the efficiency of memory access. At the same time, strictly following the requirements of the exception handling mechanism of the target platform, corresponding processing code segments are added for possible cross-language exception situations. This measure aims to ensure a high degree of stability and reliability of the program during operation.
[0038] As Figure 2 shown, an embodiment of the present application also provides a compilation device for multi-language hybrid code, including:
[0039] At least one processor; and,
[0040] A memory communicatively connected to at least one processor; wherein,
[0041] The memory stores instructions executable by at least one processor, and the instructions are executed by at least one processor to enable a compilation device for multi-language hybrid code to execute:
[0042] Scan a multi-language hybrid code file line by line to identify code segments corresponding to multiple languages and determine the language identifiers corresponding to the code segments;
[0043] Determine the corresponding parsing rules according to the language identifiers, and semantically parse the code segments according to the parsing rules to determine the semantic association graph between the multiple code segments;
[0044] Convert the multiple code segments according to the semantic association graph to determine an intermediate form;
[0045] Determine the operating environment of the target platform, and convert the intermediate form according to the operating environment to convert it into the code format corresponding to the target platform.
[0046] An embodiment of the present application further provides a non-volatile computer storage medium storing computer-executable instructions, and the computer-executable instructions are set to:
[0047] Scan a multi-language hybrid code file line by line to identify code segments corresponding to multiple languages and determine the language identifiers corresponding to the code segments;
[0048] Determine the corresponding parsing rules according to the language identifiers, and semantically parse the code segments according to the parsing rules to determine the semantic association graph between the multiple code segments;
[0049] Convert the multiple code segments according to the semantic association graph to determine an intermediate form;
[0050] Determine the operating environment of the target platform, and convert the intermediate form according to the operating environment to convert it into the code format corresponding to the target platform.
[0051] In the 1990s, it was obvious to distinguish whether an improvement to a technology was an improvement in hardware (e.g., improvement to the circuit structure of diodes, transistors, switches, etc.) or an improvement in software (improvement to the method flow). However, with the development of technology, many improvements to method flows today can be regarded as direct improvements to the hardware circuit structure. Almost all designers obtain the corresponding hardware circuit structure by programming the improved method flow into the hardware circuit. Therefore, it cannot be said that an improvement to a method flow cannot be implemented with a hardware entity module. For example, a programmable logic device (PLD) (such as a field programmable gate array (FPGA)) is such an integrated circuit whose logical function is determined by the user programming the device. Designers can program themselves to "integrate" a digital system on a single PLD, without having to ask a chip manufacturer to design and produce a dedicated integrated circuit chip. Moreover, nowadays, instead of manually fabricating integrated circuit chips, this programming is mostly implemented using "logic compiler" software, which is similar to the software compiler used in program development and writing. The original code before compilation also has to be written in a specific programming language, which is called a hardware description language (HDL). There is not only one type of HDL, but many types, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used ones currently are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also be aware that by simply performing a little logical programming on the method flow with the above-mentioned several hardware description languages and programming it into the integrated circuit, it is easy to obtain the hardware circuit that implements the logical method flow.
[0052] The controller can be implemented in any suitable manner. For example, the controller can take the form of, for example, a microprocessor or a processor and a computer-readable medium storing computer-readable program code (such as software or firmware) executable by the (micro)processor, logic gates, switches, an application specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller. Examples of the controller include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art also know that in addition to implementing the controller in the form of pure computer-readable program code, it is entirely possible to implement the same function by logically programming the method steps so that the controller is in the form of logic gates, switches, application specific integrated circuits, programmable logic controllers, embedded microcontrollers, etc. Therefore, such a controller can be considered a hardware component, and the devices included therein for implementing various functions can also be regarded as the structures within the hardware component. Or even, the devices for implementing various functions can be regarded as either software modules for implementing the method or structures within the hardware component.
[0053] The systems, devices, modules, or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer can be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or any combination of these devices.
[0054] For the convenience of description, when describing the above devices, they are described separately as various units according to their functions. Of course, when implementing this specification, the functions of each unit can be implemented in one or more software and / or hardware.
[0055] Each embodiment in this application is described in a progressive manner. The same or similar parts among the embodiments can be referred to each other, and the key point of each embodiment is to illustrate the differences from other embodiments. In particular, for the device and medium embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiments.
[0056] The devices, media, and methods provided by the embodiments of the present application correspond one by one. Therefore, the devices and media also have beneficial technical effects similar to those of their corresponding methods. Since the beneficial technical effects of the methods have been described in detail above, the beneficial technical effects of the devices and media will not be elaborated here.
[0057] Those skilled in the art should understand that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) that contain computer-usable program code.
[0058] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be realized by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to generate a machine, so that the instructions executed by the processors of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.
[0059] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means implements the functions specified in Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.
[0060] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.
[0061] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.
[0062] Memory may include non-permanent storage in a computer-readable medium, in the form of random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.
[0063] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.
[0064] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0065] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included within the scope of the claims of the present application.
Claims
1. A compilation method for multi-language mixed code, characterized in that, Including: Scanning a multi - language mixed code file line by line to identify code segments corresponding to multiple languages and determining the language identifiers corresponding to the code segments; Determining corresponding parsing rules according to the language identifiers and semantically parsing the code segments according to the parsing rules to determine a semantic association graph among the multiple code segments; Converting the multiple code segments according to the semantic association graph to determine an intermediate form; Determining the operating environment of the target platform to convert the intermediate form according to the operating environment into the code format corresponding to the target platform; Determining corresponding parsing rules according to the language identifiers and semantically parsing the code segments according to the parsing rules, specifically including: Determining the corresponding grammar according to the language identifier to determine the corresponding parsing rules according to the grammar; Constructing a model according to the parsing rules to determine a semantic parsing model, and the functions of the semantic parsing model include variable declaration and usage analysis, function call parsing, data type derivation, control flow analysis, interface and dependency relationship analysis between code segments; Determining a semantic association graph among the multiple code segments, specifically including: Semantically parsing the multiple code segments through the semantic parsing model to determine the function call relationships among the multiple code segments, and determining semantic associations according to the function call relationships; Determining the data sharing mechanism and interface specifications among the multiple code segments, and determining the semantic association graph according to the semantic associations, the data sharing mechanism and the interface specifications to display the interaction paths and dependency networks among the multiple code segments through the semantic association graph; Determining a semantic model based on the C++ standard grammar, and the semantic model includes classes, templates and overloaded functions; when parsing variable declarations in C++ code, the semantic model derives the type and lifecycle of variables based on the type modifiers and scope qualifiers of the variables; when performing function call parsing, it processes function overloading and default parameters; For Python, a semantic analysis model is constructed according to the dynamic type system and indentation syntax rules; in the data type derivation section, the semantic analysis model determines the type based on the variable assignment operation; when processing function calls, it determines the function parameter passing method in Python; In the intermediate representation generation and optimization stage, determining an intermediate representation form, and the intermediate representation form adopts a graph - based data structure to map the syntax structures and semantic relationships of C++ and Python codes; the nodes of the graph represent variables, functions, and code blocks; the edges of the graph represent data dependencies, control flow associations, and cross - language function call relationships; in the optimization process, analyze and identify Python functions called in C++, and adopt an inlining strategy for the functions to integrate the functions into the intermediate representation part corresponding to the C++ code.
2. The method according to claim 1, wherein Scanning a multi - language mixed code file line by line, specifically including: Determine the feature fields corresponding to the multiple languages, and scan the code file line by line according to the feature fields to determine whether the code file contains the feature fields; If the code file contains the feature fields, determine the code snippets according to the feature fields, determine the corresponding language identifiers according to the feature fields, and mark the code snippets according to the language identifiers.
3. The method according to claim 1, characterized in that, Convert multiple code snippets according to the semantic association graph to determine an intermediate form, specifically including: Determine the syntactic structure information of the code snippets according to the semantic association graph, and determine the semantic relationships between the multiple code snippets; Perform mapping conversion according to the syntactic structure and the semantic relationships to determine the intermediate form.
4. The method according to claim 1, wherein The method further includes: Determine the usage frequency and declaration period of variables in the code snippet, and determine the register corresponding to the code snippet according to the usage frequency and the declaration period.
5. The method according to claim 1, characterized in that, The method further includes: Determine the object reference mechanism corresponding to the multiple languages, and layout the code snippets according to the object reference mechanism.
6. The method according to claim 1, characterized in that, After determining the language identifier corresponding to the code snippet, the method further includes: Preprocess the code snippet, and the preprocessing process includes removing comments and processing whitespace characters.
7. A compilation device for multi-language mixed code, characterized in that, Including: At least one processor; And, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that a compilation device for multi-language mixed code can execute: Scan the multi-language mixed code file line by line to identify code snippets corresponding to multiple languages, and determine the language identifiers corresponding to the code snippets; Determine the corresponding parsing rules according to the language identifiers, and perform semantic parsing on the code snippets according to the parsing rules to determine the semantic association graph between the multiple code snippets; Convert multiple code snippets according to the semantic association graph to determine an intermediate form; Determine the operating environment of the target platform, and convert the intermediate form according to the operating environment to convert it into the code format corresponding to the target platform.
8. A non-volatile computer storage medium stores computer-executable instructions, characterized in that, The computer-executable instructions are set to: Scan the multi-language mixed code file line by line to identify code snippets corresponding to multiple languages, and determine the language identifiers corresponding to the code snippets; Determine the corresponding parsing rules according to the language identifiers, and perform semantic parsing on the code snippets according to the parsing rules to determine the semantic association graph between the multiple code snippets; Convert multiple code snippets according to the semantic association graph to determine an intermediate form; Determine the operating environment of the target platform, and convert the intermediate form according to the operating environment to convert it into the code format corresponding to the target platform.
Citation Information
Patent Citations
Cross-chip platform compiling tool chain method
CN119322619A