A method and system for integrating multimodal code structure into large models

By obtaining the code attribute diagram and performing multimodal coding and alignment, the problem of large models insufficient understanding of code structure is solved, and the performance of code analysis and related tasks is significantly improved.

CN119576389BActive Publication Date: 2025-05-16武汉金银湖实验室
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510136490.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-07
Publication Date
2025-05-16
Estimated Expiration
2045-02-07

AI Technical Summary

Technical Problem

When the existing large models directly use the data flow and control flow information of the code structure, there are false positives problems, indicating that their understanding of the code structure still needs to be improved.

Method used

By obtaining the code attribute diagram corresponding to the code text, input the graph encoder to obtain the graph structure representation, and input a large language model in combination with the code text to perform multimodal encoding and alignment to improve the model's understanding of the code structure.

Benefits of technology

By deeply modeling the syntax information and structured features of the code, the model can not only capture the superficial semantics of the code, but also understand its internal logic and structural relationships, thereby improving its performance in tasks such as code analysis, vulnerability detection, and automatic repair.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119576389B_ABST
    Figure CN119576389B_ABST
Patent Text Reader

Abstract

The present application belongs to the field of computer technology, and specifically discloses a method and system for integrating a multimodal code structure into a large model, the method comprising: obtaining a code attribute graph corresponding to a code text; inputting the code attribute graph into a graph encoder, and obtaining a graph structure representation output by the graph encoder; inputting the code text and the graph structure representation into a large language model, and obtaining a code processing result output by the large language model, wherein the large language model is used to process the code according to the code processing task based on the code text and the graph structure representation; wherein the graph structure representation and the serialized language representation are aligned in the representation space, and the serialized language representation is obtained by serializing the code text by the text encoder in the large language model. By introducing a multimodal encoder, the present application can fuse and process different modal information of the code, such as syntax, semantics, control flow, data flow, etc., and can significantly improve the large model's ability to understand the code structure.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of computer technology, and more specifically, relates to a method and system for integrating a multimodal code structure into a large model. Background Art

[0002] With the widespread use of large-scale language models (abbreviated as large language models or large models), they have also shown promising results in the field of code. Currently, the most common method of using large models is to directly query the large language model by constructing prompts. Code with structural information is placed in the input text prompts. When the large model processes the task input, the code is converted into serialized text. Despite this, it has shown considerable results in some downstream software engineering tasks (such as vulnerability detection, automatic program repair, and vulnerability localization). However, when directly using large models to query information such as data flow and control flow of the code structure, some false positives still occur, indicating that there is still room for improvement in the large model's ability to understand code structure. How to effectively improve the large model's understanding of code structure is a technical problem that needs to be solved in this field. Summary of the Invention

[0003] In response to the shortcomings of the existing technology, the purpose of this application is to effectively improve the ability of large models to understand code structure.

[0004] To achieve the above objectives, in a first aspect, the present application provides a method for integrating a code structure into a large model based on multimodality, the method comprising: obtaining a code attribute graph corresponding to a code text;

[0005] Input the code attribute graph to the graph encoder and obtain the graph structure representation output by the graph encoder;

[0006] Input code text and graph structure representation into the large language model, and obtain the code processing results output by the large language model. The large language model is used to process the code according to the code processing task based on the code text and graph structure representation;

[0007] Among them, the graph structure representation and the serialized language representation are aligned in the representation space. The serialized language representation is obtained by serializing the code text by the text encoder in the large language model.

[0008] In a possible implementation, obtaining the code attribute graph corresponding to the code text includes:

[0009] Use code parsing tools to parse the code text and generate multiple code graph structure information corresponding to the code text;

[0010] Based on multiple pieces of code graph structure information corresponding to the code text, a code attribute graph corresponding to the code text is generated.

[0011] In a possible implementation, the multiple pieces of code graph structure information include: an abstract syntax tree, a control flow graph, and a data flow graph.

[0012] In a possible implementation, the code attribute graph corresponding to the code text is generated based on the multiple pieces of code graph structure information corresponding to the code text, including:

[0013] Based on the abstract syntax tree, control flow graph and data flow graph corresponding to the code text, fusion processing is performed to generate a code attribute graph corresponding to the code text.

[0014] In one possible implementation, the model structure of the graph encoder adopts a graph neural network.

[0015] Exemplarily, the graph neural network can be a graph convolutional network (GCN), a graph attention network (GAT), a graph autoencoder (GAE) or a graph transformer, etc.

[0016] In one possible implementation, the graph encoder is obtained by the following steps:

[0017] Obtain code samples by querying open source code repositories;

[0018] Obtain the code attribute graph corresponding to the code sample through the code parsing tool;

[0019] Based on the code attribute graph and target label corresponding to the code sample, the graph neural network is trained to obtain the trained graph neural network, and the trained graph neural network is used as the graph encoder.

[0020] Optionally, the above-mentioned process of obtaining a graph encoder can also be referred to as graph encoder pre-training. Graph encoder pre-training can include tasks such as code structure understanding, code structure completion, node prediction, and edge prediction. Code structure understanding: Output the structural representation of the code and generate a code summary. Code structure completion: Randomly mask the information in the structure and let the encoder complete it. Node prediction: Includes the prediction of the parent / child nodes of a specified node. For example, given a graph and a specified node, predict the parent / child nodes of the specified node in the graph. Edge prediction: Given a graph and two nodes in the graph, predict whether there is a directly connected edge between the two nodes, and other tasks.

[0021] In one possible implementation, the method further includes: adjusting model parameters of the graph encoder and the model parameters of the large language model through two-stage training to align the graph structure representation with the serialized language representation in the representation space;

[0022] Among them, the first stage training is to fix the parameters of the large language model and adjust the model parameters of the graph encoder. The second stage training is to fix the parameters of the graph encoder and adjust the model parameters of the large language model.

[0023] In a second aspect, the present application provides a multimodal code structure integrated into a large model system, including:

[0024] A graph structure information acquisition module is used to obtain the code attribute graph corresponding to the code text;

[0025] A graph structure information processing module is used to input a code attribute graph into a graph encoder and obtain a graph structure representation output by the graph encoder;

[0026] The code processing task execution module is used to input code text and graph structure representation into the large language model and obtain the code processing results output by the large language model. The large language model is used to process the code according to the code processing task based on the code text and graph structure representation;

[0027] Among them, the graph structure representation and the serialized language representation are aligned in the representation space. The serialized language representation is obtained by serializing the code text by the text encoder in the large language model.

[0028] In a third aspect, the present application provides an electronic device comprising: at least one memory for storing programs; and at least one processor for executing the programs stored in the memory. When the program stored in the memory is executed, the processor is used to execute the method described in the first aspect or any possible implementation of the first aspect.

[0029] In a fourth aspect, the present application provides a computer-readable storage medium, which stores a computer program. When the computer program runs on a processor, the processor executes the method described in the first aspect or any possible implementation of the first aspect.

[0030] In general, the above technical solutions conceived by this application have the following beneficial effects compared with the existing technologies:

[0031] The textual representation of the code is uniformly encoded with its graph structure (such as an abstract syntax tree, control flow graph, data dependency graph, etc.), and the grammatical information and structural features of the code are deeply modeled with the help of the characteristics of the multimodal encoder. Through this process, the model can not only capture the surface semantics of the code, but also understand its internal logic and structural relationships, thereby improving its performance in tasks such as code analysis, vulnerability detection, and automatic repair. Ultimately, this application can improve the large model's ability to understand code structure while enhancing its cross-modal processing capabilities, providing strong support for intelligent analysis in complex code scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] Figure 1 This is a flowchart of a method for integrating a multimodal code structure into a large model provided by an embodiment of the present application;

[0033] Figure 2 This is a schematic diagram of extracting a code property graph using a code analysis tool provided in an embodiment of the present application;

[0034] Figure 3 This is a schematic diagram of encoding using a graph attention neural network provided in an embodiment of the present application;

[0035] Figure 4 is a schematic diagram of modal alignment of text and graphic structures provided in an embodiment of the present application;

[0036] Figure 5 This is a schematic diagram of the structure of a multimodal code structure integrated into a large model system according to an embodiment of the present application;

[0037] Figure 6 It is a structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0038] In order to facilitate a clearer understanding of the various embodiments of the present application, some relevant background knowledge is first introduced as follows.

[0039] If the structural information of the code can be encoded at the stage of input to the large model, the large model's ability to understand the input code can be enhanced, thereby improving the performance of downstream tasks that require understanding of the code. Multimodal large models are an emerging deep learning method that can fuse data from other modalities when the pre-training data is only text and serialized data, so as to achieve the purpose of using the understanding of the large model to complete more complex and diverse downstream tasks. Based on this, this application proposes a method and system for integrating multimodal code structure into a large model through aligned fusion training.

[0040] The upstream data used for pre-training large-scale language models is entirely serialized. Enabling models trained solely on serialized data to acquire and understand the structural information of code graphs is a key challenge in multimodal fusion. This requires not only extracting the logical and syntactic structure of the code from a linear sequence but also requiring the model to capture higher-dimensional graph structure information, such as function calls, control flow, and data flow. Code graph structures (such as abstract syntax trees, control flow graphs, and dependency graphs) are typically used to represent relationships between program elements. However, serialized data only displays the linear order of these elements and cannot directly reflect complex dependencies. Therefore, combining graph structure representations with traditional serialized language representations in a multimodal environment, fully integrating the characteristics of these different data forms and thereby improving the model's understanding and reasoning capabilities, remains a pressing challenge in the current state of the art. This requires not only effective representation of multimodal information at the input level but also the ability for the model to collaboratively process information from different modalities to achieve a comprehensive understanding and accurate inference of code semantics. This application uses a multimodal fusion alignment method to encode the code graph structure and integrate it into the large model, and aligns different modal information so that the module used for encoding and the large model module can collaborate in reasoning to complete the task.

[0041] In order to make the purpose, technical solutions and advantages of this application more clearly understood, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0042] In order to make the purpose, technical solutions and advantages of this application more clearly understood, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0043] In the embodiments of this application, words such as "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described as "exemplary" or "for example" in the embodiments of this application should not be interpreted as being preferred or advantageous over other embodiments or designs. Rather, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.

[0044] In the description of the embodiments of the present application, unless otherwise specified, "multiple" means two or more, for example, multiple processing units means two or more processing units, etc.; multiple elements means two or more elements, etc.

[0045] The embodiments of the present application are described below in conjunction with the drawings in the embodiments of the present application.

[0046] Figure 1This is a flow chart of a method for integrating a multimodal code structure into a large model provided by an embodiment of the present application, such as Figure 1 As shown, the method includes the following steps S101, S102 and S103.

[0047] Step S101, obtaining a code attribute graph corresponding to the code text;

[0048] Step S102: input the code attribute graph to the graph encoder, and obtain the graph structure representation output by the graph encoder;

[0049] Step S103: inputting the code text and graph structure representation into the large language model, and obtaining the code processing result output by the large language model. The large language model is used to process the code according to the code processing task based on the code text and graph structure representation;

[0050] Among them, the graph structure representation and the serialized language representation are aligned in the representation space. The serialized language representation is obtained by serializing the code text by the text encoder in the large language model.

[0051] Here, an example is given of a code processing task, which may be a code analysis, vulnerability detection, automatic repair, or other tasks. When the code processing task is a code analysis task, the large language model performs corresponding analysis on the code based on the code text and graph structure representation, and the code processing result output by the large model is specifically the code analysis result. When the code processing task is a vulnerability detection task, the large language model performs vulnerability detection on the code based on the code text and graph structure representation, and the code processing result output by the large model is specifically the vulnerability detection result. When the code processing task is an automatic repair task, the large language model repairs the code based on the code text and graph structure representation, and the code processing result output by the large model is specifically the repaired code.

[0052] By introducing multimodal encoders (text encoders and graph encoders), different modal information of the code (such as syntax, semantics, control flow, data flow, etc.) can be fused and processed, which can significantly improve the large model's ability to understand the code structure.

[0053] Specifically, this application uniformly encodes the textual representation of the code and its graph structure (such as abstract syntax trees, control flow graphs, data dependency graphs, etc.), and uses the characteristics of multimodal encoders to deeply model the grammatical information and structural features of the code. Through this process, the model can not only capture the surface semantics of the code, but also understand its internal logic and structural relationships, thereby improving its performance in tasks such as code analysis, vulnerability detection, and automatic repair. Ultimately, this application can enhance the large model's ability to understand code structure while enhancing its cross-modal processing capabilities, providing strong support for intelligent analysis in complex code scenarios.

[0054] The following is an illustrative example of the multimodal code structure integration method provided by this application into a large model. In this method, first, a static code parsing tool processes the code extraction to obtain a structural representation of the graph (such as an abstract syntax tree, control flow graph, dependency graph, etc.). Then, an independent graph encoder is designed to process the structural information of the graph and encode it. Finally, an alignment module is designed to align the encoding results obtained by the graph encoder into the representation space of the large model. This method mainly includes a data sorting stage, an encoding stage, and an alignment stage.

[0055] Figure 2 This is a schematic diagram of extracting a code attribute graph using a code analysis tool provided in an embodiment of the present application, such as Figure 2 As shown in the figure, the data collation stage includes the following steps: data collection and code graph structure extraction.

[0056] Regarding the data collection phase, specifically, a large amount of single-function programming codes are collected through open source code library websites (such as Github, etc.), including programming language, function name, function parameters and other information.

[0057] For code graph structure extraction, we use static code parsing tools (such as Joern) to parse the collected function code to generate a Code Property Graph (CPG) containing the code graph structure information. The CPG unifies multiple code representations, such as the Abstract Syntax Tree (AST), Control Flow Graph (CFG), and Data Flow Graph (DFG), into a single graph structure, enabling analysis tools to gain a deeper understanding and process the code.

[0058] The steps of code graph structure extraction include the following sub-steps: (1) parsing the code into an abstract syntax tree (AST), which helps capture the basic structure and syntax elements of the code; (2) further constructing a control flow graph (CFG) to represent the execution order of statements in the program, and a data flow graph (DFG) to track the flow of variables and data in the program; (3) fusing these graph structures into a unified code property graph (CPG), which can effectively represent the semantics and logical structure of the program.

[0059] This article explains the Code Property Graph (CPG). CPG is a comprehensive graph structure used to represent program code. It combines the features of an abstract syntax tree (AST), a control flow graph (CFG), and a data flow graph (DFG), and can provide rich information for code analysis, static analysis, and security auditing.

[0060] This section explains the Abstract Syntax Tree (AST). An AST is a tree structure used to represent the grammatical structure of source code. Each node represents a construct in the source code (such as a variable, operator, or statement), while the tree's hierarchical structure reflects the grammatical relationships between these constructs.

[0061] This section explains a control flow graph (CFG). A CFG is a directed graph used to represent the paths of control flow in a program. Nodes in the graph represent basic blocks (i.e., continuous sections of code) in the program, while edges represent control flow transitions (such as conditional branches and loops).

[0062] This section explains a data flow graph (DFG), a graphical representation used to describe the flow and dependencies of data in a program. Nodes in a DFG represent computational operations or variables, while edges represent data dependencies.

[0063] Here, we provide an example of the multiple code graph structure information corresponding to the above-mentioned generated code text. (1) Parse the source code to generate AST: Use a code parsing tool to parse the source code into an abstract syntax tree (AST). The nodes of the AST represent the syntax elements in the code, such as variables, functions, control structures, etc. (2) Construct a control flow graph (CFG): Extract control flow information from the AST and construct a control flow graph (CFG). Basic blocks (i.e., a continuous section of code) are identified as nodes of the CFG, while control flow transfers (such as conditional branches, loops, etc.) are represented as edges. This step usually involves analyzing the control structure of the program to determine the execution path between different code blocks. (3) Construct a data flow graph (DFG): Analyze the AST and CFG, extract data dependencies, and construct a data flow graph (DFG).

[0064] Here, we will illustrate the fusion processing of the abstract syntax tree, control flow graph, and data flow graph corresponding to the code text to generate the code attribute graph corresponding to the code text. (1) Node fusion: The nodes of CPG can simultaneously represent the syntax elements in AST, the basic blocks in CFG, and the data operations in DFG. For example, a function node can contain its control flow information and data dependency information at the same time. (2) Edge definition: The edges in CPG can represent multiple relationships, such as control flow edges (from CFG), data dependency edges (from DFG), and syntax structure edges (from AST). This multiple edge definition enables CPG to capture the complex interactions of the code. (3) Attribute annotation: Add attributes to the nodes and edges in CPG to facilitate subsequent analysis and query. For example, nodes can contain attributes such as type information, scope, and definition location, while edges can contain information such as conditional expressions and data flow direction.

[0065] Figure 3This is a schematic diagram of encoding using a graph attention neural network provided in an embodiment of the present application, such as Figure 3 As shown, the encoding phase mainly encodes structural information through the designed graph encoder (graph attention neural network), allowing graph information that cannot be input into the large model to be input in the form of vectors. The encoding phase includes the following steps S301, S302, and S303.

[0066] Step S301: Graph preprocessing and feature extraction. This involves preprocessing and feature extraction of the graph. Taking the Graph Attention Network (GAT) as an example, the GAT receives a graph structure with node features and edges. Therefore, it is necessary to generate a feature vector for each node and use the edge type as the graph's connectivity information.

[0067] Graph preprocessing and feature extraction include the following sub-steps. (1) Node feature extraction: Feature representations can be generated for each node based on the attributes of code elements (such as variable names, data types, operator types, etc.), usually through word embedding. (2) Edge information representation: Edges in the graph represent different types of relationships, such as syntactic connections, data dependencies, or control flow. These edges need to be incorporated into the model through specific graph representations.

[0068] Step S302: Construct a Graph Attention Network (GAT) model. GAT is a neural network that automatically learns the relationships between nodes in a graph structure. It achieves information propagation by calculating the attention weights of neighboring nodes for each node on the graph. GAT consists of multiple graph attention layers. Each layer aggregates the features of neighboring nodes through an attention mechanism and generates a new node representation.

[0069] The construction of a graph attention network consists of the following sub-steps. (1) Initialize node features and edges. (2) In each layer of the network, calculate the attention weights between the target node and its neighboring nodes. These weights are based on the similarity of node features. (3) Use these weights to perform a weighted summation of the features of neighboring nodes and update the representation of each node. (4) Multi-layer GAT can capture deeper structural information.

[0070] Step S303: The constructed GAT model is used for training. The training data typically includes CPGs and their corresponding target labels (e.g., function classifications, vulnerability detection labels, etc.). During training, the model gradually optimizes the attention mechanism and node feature representation to minimize the loss function.

[0071] The alignment stage mainly aligns the output of the graph attention neural network independent of the large model to the representation space of the large model through training, so that the two models can work better together.

[0072] This is achieved through a two-step training process (also known as two-stage training). First, the parameters of the large model are fixed, and only the parameters of the graph encoder are trained. During training, methods such as contrastive learning and transfer learning are used to optimize the parameters of the graph encoder, so that the output of the graph encoder and the output of the text encoder are as close as possible under a specified metric. The goal is to align the output of the graph encoder with the output of the text encoder in the large model, making the output of the graph encoder more consistent with the representational understanding of the large model. Then, the parameters of the graph encoder are fixed, and the parameters of the large model continue to be trained. The goal is to explore the potential capabilities of the large model by combining the input of the graph structure.

[0073] Figure 4 This is a schematic diagram of modal alignment of text and graph structure provided by the embodiment of the present application, such as Figure 4 As shown, the output of the graph encoder remains aligned with the output of the text encoder in the large model.

[0074] This application introduces multimodal fusion technology to inject the structured information of the code as input into the large model, allowing the large model, which originally only accepted plain text sequence data, to expand its perception range and pay attention to more complex modal information. This technology breakthrough combines the multi-dimensional representation of the code (such as grammatical structure, control flow, data flow, dependencies, etc.) with the language processing capabilities of the large model, so that the model not only relies on traditional linear serialization input, but can also capture the deep structure and logical associations in the code. In this way, the model's code understanding ability is significantly improved, and it can more accurately identify logical patterns, potential problems and functional implementation details in the code, thereby demonstrating a stronger level of intelligence and processing capabilities in complex tasks such as code auditing, vulnerability detection, and automatic program repair. At the same time, this method enhances the large model's ability to integrate cross-modal information, giving it a wider range of adaptability and application scenarios, further promoting the advancement of intelligent programming technology.

[0075] The following describes the multimodal-based code structure integration into a large model system provided by this application. The multimodal-based code structure integration into a large model system described below and the multimodal-based code structure integration into a large model method described above can be referenced to each other.

[0076] Figure 5 This is a schematic diagram of the structure of the multimodal code structure integrated into the large model system provided by the embodiment of the present application, such as Figure 5 As shown, the system includes: a graph structure information acquisition module 10, a graph structure information processing module 20 and a code processing task execution module 30. Among them:

[0077] A graph structure information acquisition module 10 is used to obtain a code attribute graph corresponding to the code text;

[0078] A graph structure information processing module 20 is used to input a code attribute graph into a graph encoder and obtain a graph structure representation output by the graph encoder;

[0079] A code processing task execution module 30 is configured to input code text and graph structure representation into a large language model and obtain code processing results output by the large language model. The large language model is configured to process the code according to the code processing task based on the code text and graph structure representation.

[0080] Among them, the graph structure representation and the serialized language representation are aligned in the representation space. The serialized language representation is obtained by serializing the code text by the text encoder in the large language model.

[0081] It is understandable that the detailed functional implementation of each of the above units / modules can be found in the introduction of the aforementioned method embodiment, and will not be repeated here.

[0082] It should be understood that the above-mentioned system is used to execute the method in the above-mentioned embodiment. The implementation principle and technical effect of the corresponding program module in the system are similar to those described in the above-mentioned method. The working process of the system can refer to the corresponding process in the above-mentioned method and will not be repeated here.

[0083] Based on the method in the above embodiment, an embodiment of the present application provides an electronic device, Figure 6 is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application, such as Figure 6 As shown, the electronic device may include: a processor (Processor) 810, a communication interface (Communications Interface) 820, a memory (Memory) 830 and a communication bus 840, wherein the processor 810, the communication interface 820, and the memory 830 communicate with each other via the communication bus 840. The processor 810 can call the logic instructions in the memory 830 to execute the method in the above embodiment.

[0084] In addition, the logic instructions in the aforementioned memory 830 can be implemented in the form of a software functional unit and, when sold or used as an independent product, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the portion that contributes to the prior art, or the portion of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application.

[0085] Based on the method in the above embodiment, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program runs on a processor, the processor executes the method in the above embodiment.

[0086] Based on the method in the above embodiment, an embodiment of the present application provides a computer program product. When the computer program product runs on a processor, the processor executes the method in the above embodiment.

[0087] It is understood that the processor in the embodiments of the present application may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field programmable gate arrays (FPGA), other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. The general-purpose processor may be a microprocessor or any conventional processor.

[0088] The method steps in the embodiments of the present application can be implemented by hardware or by a processor executing software instructions. The software instructions can be composed of corresponding software modules, which can be stored in random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, mobile hard disks, CD-ROMs, or any other form of storage medium known in the art. An exemplary storage medium is coupled to the processor so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be an integral part of the processor. The processor and storage medium can be located in an ASIC.

[0089] The above embodiments can be implemented in whole or in part using software, hardware, firmware, or any combination thereof. When implemented using software, they can be implemented in whole or in part in the form of a computer program product. The computer program product comprises one or more computer instructions. When loaded and executed on a computer, the computer program instructions fully or partially produce the processes or functions described in the embodiments of this application. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted via the computer-readable storage medium. The computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium accessible by a computer or a data storage device such as a server or data center that integrates one or more available media. The available medium can be magnetic media (e.g., floppy disk, hard disk, tape), optical media (e.g., DVD), or semiconductor media (e.g., solid-state drive (SSD)).

[0090] It will be understood that the various numerical numbers involved in the embodiments of the present application are merely distinctions for the convenience of description and are not intended to limit the scope of the embodiments of the present application.

[0091] It is easy for those skilled in the art to understand that the above is only a preferred embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present application should be included in the scope of protection of the present application.

Claims

1. A method for integrating multimodal code structure into a large model, characterized in that: include: Get the code attribute graph corresponding to the code text; Input the code attribute graph to the graph encoder and obtain the graph structure representation output by the graph encoder; Input code text and graph structure representation to the large language model, and obtain the code processing result output by the large language model. The large language model is used to process the code according to the code processing task based on the code text and graph structure representation; The graph structure representation and the serialized language representation are aligned in the representation space. The serialized language representation is obtained by serializing the code text by the text encoder in the large language model. It also includes: adjusting the model parameters of the graph encoder and the model parameters of the large language model through two-stage training, the first stage training is to fix the large language model parameters and adjust the model parameters of the graph encoder, and the second stage training is to fix the graph encoder parameters and adjust the model parameters of the large language model.

2. The method for integrating multimodal code structure into a large model according to claim 1, characterized in that: The step of obtaining a code attribute graph corresponding to the code text includes: The code text is parsed through the code parsing tool to generate multiple code graph structure information corresponding to the code text; Based on multiple pieces of code graph structure information corresponding to the code text, a code attribute graph corresponding to the code text is generated.

3. The method for integrating multimodal code structure into a large model according to claim 2, characterized in that: The multiple pieces of code graph structure information include: an abstract syntax tree, a control flow graph, and a data flow graph.

4. The method for integrating multimodal code structure into a large model according to claim 3, characterized in that: The generating of the code attribute graph corresponding to the code text based on the multiple code graph structure information corresponding to the code text includes: Based on the abstract syntax tree, control flow graph and data flow graph corresponding to the code text, fusion processing is performed to generate a code attribute graph corresponding to the code text.

5. The method for integrating multimodal code structure into a large model according to any one of claims 1 to 4, characterized in that: The model structure of the graph encoder adopts a graph neural network.

6. The method for integrating multimodal code structure into a large model according to claim 5, characterized in that: The graph encoder is obtained by the following steps: Get code samples by querying open source code repositories; Obtain the code attribute graph corresponding to the code sample through the code parsing tool; Based on the code attribute graph and target label corresponding to the code sample, the graph neural network is trained to obtain the trained graph neural network, and the trained graph neural network is used as a graph encoder.

7. A multi-modal code structure integrated into a large model system, characterized in that: include: A graph structure information acquisition module, used to obtain a code attribute graph corresponding to a code text; A graph structure information processing module, used to input a code attribute graph into a graph encoder and obtain a graph structure representation output by the graph encoder; A code processing task execution module is used to input code text and graph structure representation into the large language model and obtain the code processing result output by the large language model. The large language model is used to process the code according to the code processing task based on the code text and graph structure representation; The graph structure representation and the serialized language representation are aligned in the representation space. The serialized language representation is obtained by serializing the code text by the text encoder in the large language model. It also includes: adjusting the model parameters of the graph encoder and the model parameters of the large language model through two-stage training, the first stage training is to fix the large language model parameters and adjust the model parameters of the graph encoder, and the second stage training is to fix the graph encoder parameters and adjust the model parameters of the large language model.

8. An electronic device, characterized in that: include: at least one memory for storing a computer program; At least one processor is used to execute the program stored in the memory. When the program stored in the memory is executed, the processor is used to execute the method according to any one of claims 1 to 6.

9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program runs on a processor, the processor is caused to execute the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Software vulnerability detection method based on multi-modal feature adaptive fusion

    CN119004491A

  • Intelligent dynamic prompting method and system for online code writing

    CN119149001A