Code data management method based on artificial intelligence and related device

By converting project code into a code graph and performing graph search, the problem of difficult location in code data governance is solved, achieving efficient and accurate code data governance.

CN122018956APending Publication Date: 2026-05-12BEIJING VOLCANO ENGINE TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING VOLCANO ENGINE TECH CO LTD
Filing Date
2026-04-13
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

As software development projects grow in scale, the amount of code increases, and the logic becomes more complex, existing technologies struggle to quickly locate specific modules that require code data governance, leading to low development efficiency, code redundancy, and difficulties in troubleshooting.

Method used

The project code is converted into a code graph representation, and nodes strongly related to code requests are obtained from the graph through graph search. Data governance is then performed in conjunction with subgraphs, providing a set of general and efficient code data governance logic.

Benefits of technology

By using graph search, you can quickly focus on the parts of the project code that are related to code processing requests, thus improving the efficiency and accuracy of code data governance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122018956A_ABST
    Figure CN122018956A_ABST
Patent Text Reader

Abstract

One or more situations of the invention provide a code data governance method based on artificial intelligence and a related device. The code data governance method comprises the following steps: obtaining a first code processing request; searching the first code graph based on the first code processing request to obtain a first node set; and processing the first code processing request according to first sub-graphs corresponding to the plurality of first nodes in the first code graph to obtain a data governance result corresponding to the first code processing request. Complex project codes are converted into code graphs for expression, graph search is carried out in the code graphs to obtain nodes strongly related to code requests, data management is carried out on the code requests in combination with sub-graphs formed by the strongly related nodes, and a set of universal and efficient code data management logic is provided for different types of code processing requests. Parts of codes related to the code processing request in the project codes are quickly focused, the code data treatment efficiency is improved, and the code data treatment accuracy is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] One or more scenarios described herein relate to an AI-based code data governance method, an AI-based code data governance device, an electronic device, a computer-readable storage medium, and a computer program product. Background Technology

[0002] As software development projects grow in scale, the amount of code increases, and the code logic becomes more complex, the process of code data governance for developers becomes increasingly complicated. Therefore, improving the efficiency of code data governance is particularly important. Summary of the Invention

[0003] This summary section is provided to briefly introduce the concepts, which will be described in detail in the detailed description section below. This summary section is not intended to identify key or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.

[0004] This paper provides at least one scenario of an AI-based code data governance method, comprising: obtaining a first code processing request, wherein the first code processing request is used to perform data governance on the code of a first project; searching a first code graph based on the first code processing request to obtain a first node set, wherein nodes in the first code graph represent code entities in the code of the first project, edges in the first code graph represent relationships between the code entities, the first node set includes multiple first nodes, and the degree of association between each of the multiple first nodes and the first code processing request satisfies a first predetermined condition; and processing the first code processing request according to a first subgraph corresponding to the multiple first nodes in the first code graph to obtain a data governance result corresponding to the first code processing request.

[0005] This paper provides at least one scenario of an AI-based code data governance device, comprising: an acquisition module configured to acquire a first code processing request, wherein the first code processing request is used for data governance of first project code; a search module configured to search a first code graph based on the first code processing request to obtain a first node set, wherein nodes in the first code graph represent code entities in the first project code, edges in the first code graph represent relationships between the code entities, the first node set includes multiple first nodes, and the degree of association between each of the multiple first nodes and the first code processing request satisfies a first preset condition; and a processing module configured to process the first code processing request according to a first subgraph corresponding to the multiple first nodes in the first code graph to obtain a data governance result corresponding to the first code processing request.

[0006] At least one scenario of this document provides an electronic device, including: at least one processor; and at least one memory, including one or more computer program instructions; wherein the one or more computer program instructions are executed by the processor to perform the AI-based code data governance method provided in at least one scenario of this document.

[0007] At least one aspect of this paper provides a computer-readable storage medium for non-transitory storage of computer-readable instructions, wherein the AI-based code data governance method provided by at least one aspect of this paper is implemented when the computer-readable instructions are executed by a processor.

[0008] At least one aspect of this document provides a computer program product, including a computer program that, when executed by a processor, implements the AI-based code data governance method provided by at least one aspect of this document.

[0009] In one of the AI-based code data governance methods presented in this paper, complex project code is converted into a code graph representation. For code processing requests, a graph search is performed within the code graph to identify nodes strongly related to the request. Then, a subgraph formed by these strongly related nodes is used to perform data governance on the code request. This provides a universal and efficient code data governance logic for different types of code processing requests. By using graph search, the code relevant to the processing request is quickly focused, improving code data governance efficiency. Furthermore, the contextual information from the subgraph enhances the accuracy of code data governance. Attached Figure Description

[0010] The above and other features, advantages, and aspects of the various scenarios herein will become more apparent when taken in conjunction with the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale.

[0011] Figure 1 This illustration shows an application scenario diagram of an AI-based code data governance method provided in at least one of the cases presented in this paper;

[0012] Figure 2 This paper schematically illustrates a system architecture diagram of a code data governance system provided in at least one scenario.

[0013] Figure 3 The illustration schematically depicts a flowchart of constructing a first code graph, provided for at least one scenario described herein;

[0014] Figure 4 The illustration schematically depicts a flowchart of constructing a first code graph, provided for at least one scenario described herein;

[0015] Figure 5 The diagram illustrates a flowchart of an AI-based code data governance method provided in at least one scenario of this paper.

[0016] Figure 6 The diagram illustrates a flowchart of a graph search provided in at least one scenario of this paper;

[0017] Figure 7 The illustration shows a schematic diagram of a data governance outcome provided in at least one scenario of this paper;

[0018] Figure 8 The schematic diagram illustrates the structure of an artificial intelligence-based code data governance device provided in at least one scenario of this paper; and

[0019] Figure 9 A schematic diagram of the structure of an electronic device suitable for implementing at least one of the scenarios described herein is shown. Detailed Implementation

[0020] One or more scenarios described herein will now be described in more detail with reference to the accompanying drawings. While some scenarios are shown in the drawings, it should be understood that this document can be implemented in various forms and should not be construed as limited to the scenarios set forth herein; rather, these scenarios are provided to provide a more thorough and complete understanding of this document. It should be understood that the accompanying drawings and scenarios are for illustrative purposes only and are not intended to limit the scope of this document.

[0021] It should be understood that the steps described in the method embodiments herein may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this document is not limited in this respect.

[0022] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one situation" means "at least one situation"; the term "another situation" means "at least one additional situation"; the term "some situations" means "at least some situations". Definitions of other terms will be given in the following description.

[0023] It should be noted that the concepts of "first" and "second" mentioned in this article are only used to distinguish different devices, modules or units, and are not used to limit the order of the functions performed by these devices, modules or units or their interdependencies.

[0024] It should be noted that the terms "one" and "more" used in this document are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0025] The names of the messages or information exchanged between the various devices in the embodiments herein are for illustrative purposes only and are not intended to limit the scope of these messages or information.

[0026] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition, use, storage or deletion of the data) shall comply with the requirements of relevant laws, regulations and related provisions.

[0027] It is understood that before using the technical solutions disclosed in each scenario in this article, relevant users should be informed of the type, scope of use, and usage scenarios of the information involved in this article and their authorization should be obtained through appropriate means in accordance with relevant laws and regulations. Relevant users may include any type of rights holder, such as individuals, enterprises, or groups.

[0028] For example, in response to receiving an active request from a user, a prompt message is sent to the relevant user to clearly inform the user that the requested operation will require obtaining and using the user's information, thereby enabling the relevant user to choose whether to provide information to the software or hardware such as electronic devices, applications, servers, or storage media that perform the operation of any of the technical solutions described herein.

[0029] As an optional but non-restrictive implementation, in response to a user's active request, a prompt message can be sent to the user, such as a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide information to the electronic device.

[0030] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation method described in this article. Other methods that comply with relevant laws and regulations may also be applied to the implementation method described in this article.

[0031] A software development project can be understood as a project that delivers a software product. During the development process, developers can write project code, which can be stored in a software repository. As the scale of a software development project increases, the amount of project code gradually increases, and the code repository becomes increasingly complex.

[0032] In software development projects, code and data governance is often required. For example, when software development requirements change, the project code needs to be modified; when the project code's performance is poor, it needs to be optimized; and when the project code malfunctions, it needs to be troubleshooted.

[0033] However, due to the large volume and complex logic of project code, developers struggle to quickly locate specific modules or functions requiring code data governance within the vast amount of code, leading to difficulties in code discovery. Consequently, the difficulty in quickly locating code within a large amount of project code results in developers frequently writing functionally similar code, leading to low code reuse rates, low development efficiency, and code redundancy. Furthermore, when code defects occur, troubleshooting requires developers to spend a significant amount of time reading the project code, and new developers joining the software development project require a lengthy learning period to understand the code logic.

[0034] To address the aforementioned issues, the industry typically employs the following methods for code data governance: For example, keyword searches are used to identify the parts of the project code involved in code processing requests. This can be achieved by using the `grep` command to match keywords in the code processing requests with the project code. However, keyword searches lack an understanding of the semantics of the project code, resulting in redundant and inaccurate search results. Another example is the use of static code analysis tools. While static code analysis tools can understand code structure, they are typically used for code inspection and are less effective at proactively and intelligently responding to code processing requests. Yet another example is the use of semantic search tools to identify the parts of the project code involved in code processing requests. This can be achieved by using the semantic search functionality provided by integrated development environment (IDE) plugins. However, semantic search tools have relatively limited search strategies, making efficient and accurate exploration difficult.

[0035] To at least partially solve the above-mentioned technical problems, this paper provides an AI-based code data governance method in at least one scenario. The method includes: obtaining a first code processing request, which is used to perform data governance on the code of a first project; searching a first code graph based on the first code processing request to obtain a first node set, wherein the nodes in the first code graph represent code entities in the code of the first project, the edges in the first code graph represent the relationships between code entities, the first node set includes multiple first nodes, the degree of association between each of the multiple first nodes and the first code processing request satisfies a first set condition, and processing the first code processing request according to the first subgraph corresponding to the multiple first nodes in the first code graph to obtain the data governance result corresponding to the first code processing request.

[0036] Based on the AI-based code data governance method provided in at least one of the embodiments described herein, at least one of the embodiments described herein also provides an AI-based code data governance device, electronic device, computer-readable storage medium, and computer program product.

[0037] In one of the AI-based code data governance methods presented in this paper, complex project code is converted into a code graph representation. For code processing requests, a graph search is performed within the code graph to identify nodes strongly related to the request. Then, a subgraph formed by these strongly related nodes is used to perform data governance on the code request. This provides a universal and efficient code data governance logic for different types of code processing requests. By using graph search, the code relevant to the processing request is quickly focused, improving code data governance efficiency. Furthermore, the contextual information from the subgraph enhances the accuracy of code data governance.

[0038] The following detailed description, with reference to the accompanying drawings, illustrates one or more scenarios and some examples thereof.

[0039] Figure 1 The illustration shows an application scenario diagram of an AI-based code data governance method provided in at least one of the cases presented in this paper.

[0040] like Figure 1 As shown, the application scenario provided in this case may include user 101, terminal device 102, server 103, and database 104. Terminal device 102 can be various electronic devices capable of providing interactive pages, such as smart wearable devices, smart appliances, smart cars, mobile phones, tablets, laptops, or desktop computers.

[0041] In one or more scenarios described herein, a client may be installed on the terminal device 102. This client may be a client of a code data governance system. The server 103 may be a server that provides support for the operation of the client installed on the terminal device 102. That is, the server 103 may be a server of a code data governance system. For example, the server 103 may be a server for a local area network or a wide area network, or it may be a cloud server, etc. The one or more scenarios described herein are not limited in this respect.

[0042] User 101 can be a user of a client installed on terminal device 102. For example, user 101 can be a developer using a code data governance system.

[0043] The server 103 can communicate with the terminal device 102, for example, by providing the client installed on the terminal device 102 with relevant data (such as project code) required by the terminal device to run the client; or, for example, the server 103 can also receive relevant data returned by the terminal device 102 during the running of the client (such as operation data of the project code triggered by user 101 on the client).

[0044] Database 104 can communicate with server 103. For example, database 104 can be a storage component with persistent storage and data management capabilities, such as a relational database, a non-relational database, or a distributed database. Database 104 can store data related to the code data governance system. For example, database 104 can include a code repository to store project code; it can also store code graphs. Server 103 can perform read and write operations on the data stored in database 104.

[0045] For example, the AI-based code data governance methods provided in one or more scenarios of this paper can be implemented in software, hardware, firmware, or any combination thereof.

[0046] For example, the AI-based code data governance methods provided in one or more scenarios of this paper are applicable to terminal devices that can load and execute these methods. This paper does not impose any limitations on these scenarios. For instance, the terminal device may include a central processing unit (CPU), graphics processing unit (GPU), digital signal processor (DSP), neural network processing unit (NPU), or other processing units with data processing and / or instruction execution capabilities, as well as storage units. The server or terminal device may also have an operating system and various types of application programming interfaces (APIs) installed, implementing the AI-based code data governance methods provided in one or more scenarios of this paper by running code or instructions.

[0047] Figure 2 The diagram illustrates a system architecture diagram of a code data governance system provided in at least one scenario of this paper.

[0048] like Figure 2 As shown, the code data governance system can be logically divided into an application layer 201, a service layer 202, a search engine layer 203, a data layer 204, and a code layer 205.

[0049] Application layer 201 can be understood as the logic layer that interacts directly with the user. Application layer 201 can provide a code data governance interface. For example, users can send code processing requests (such as code writing, code optimization, problem investigation, code retrieval, etc.) in the code data governance interface provided by application layer 201. After the code processing request is completed, the data governance results can also be presented in the code data governance interface provided by application layer 201.

[0050] Service layer 202 can be understood as a logic layer that encapsulates code data governance business capabilities. For example, service layer 202 can parse the received code processing request, communicate with search engine layer 203, receive the node set returned by search engine layer 203, process the code processing request according to the subgraph corresponding to the node set, and generate data governance results.

[0051] The search engine layer 203 can be understood as a logical layer used to execute the graph search process. For example, the search engine layer 203 can find a set of nodes that are strongly related to the code processing request from the code graph based on the code processing request, and send the set of nodes to the service layer 202.

[0052] Data layer 204 can be understood as a logical layer that stores data related to the code graph. For example, data layer 204 can store the code graph and its multidimensional index. During the graph search process in search engine layer 203, data layer 204 can provide the code graph and its multidimensional index to search engine layer 203.

[0053] The code layer 205 can be understood as a logical layer that stores project code. For example, the code layer 205 can store project code from software repositories corresponding to different software development projects. During the process of building the code graph in the data layer 204, the code layer 205 can provide project code to the data layer 204.

[0054] The following will combine Figures 3 to 7 This paper provides a detailed description of an AI-based code data governance method for at least one scenario.

[0055] In one or more scenarios described in this paper, a code graph is constructed for the project code of a software development project, mapping the project code into a queryable and searchable code space.

[0056] Taking the first project code as an example, the first project code can be understood as the project code of a software development project. For example, the first project code may include code representing business logic, etc. The code graph corresponding to the first project code can be called the first code graph. The first code graph includes multiple nodes and multiple edges. The nodes in the first code graph represent code entities in the first project code, and the edges in the first code graph represent the relationships between code entities.

[0057] A code entity can be understood as the smallest unit in the code of a first project that has independent meaning, can be identified and manipulated. A code entity can have a clear identifier, independent semantics or functions, and can be manipulated or reused. For example, a code entity can be one or more of the following: variable, constant, function, method, class, object, module, package, interface and abstract class.

[0058] The relationships between code entities can be understood as associations between different code entities. For example, two code entities with an association are connected by an edge, while two code entities without an association are not connected by an edge. Furthermore, based on the specific relationship type, the edges in the first code graph can be directed edges. For instance, when code entity A calls code entity B, the directed edge between the node corresponding to code entity A and the node corresponding to code entity B can be a path from the node corresponding to code entity A to the node corresponding to code entity B.

[0059] In some cases, the first code graph can also represent richer information about the first project's code. For example, nodes in the first code graph can be associated with and store attribute information of code entities. The attribute information of a code entity can include one or more of the following: the identifier of the code entity, the code content corresponding to the code entity, and the structural information of the code entity. The identifier of the code entity can be the name of the code entity (such as a function name, class name, etc.), the code content corresponding to the code entity can include at least one of the source code text and comment text, and the structural information of the code entity can include at least one of the following: the file information to which the code entity belongs, the location information of the code entity, and the type of the code entity.

[0060] For example, the edges in the first code graph can be associated with the relationship types between code entities. The relationship types between code entities can be used to describe how different code entities establish connections. For example, the relationship between code entities can be one or more of the following: call relationship, inheritance relationship, containment relationship, and parameter passing relationship.

[0061] In this way, the complex and multidimensional code logic in the first project code is represented by the first code diagram, and the large number of code entities and the messy relationships between them are clearly presented in the first code diagram, providing a good foundation for efficient code data governance.

[0062] Figure 3 The illustration shows a flowchart of constructing a first code graph for at least one scenario described herein.

[0063] like Figure 3 As shown, the process of constructing the first code graph of the first project code may include the following steps:

[0064] Step S301: Obtain the code for the first project.

[0065] This article does not restrict the method of obtaining the first project code in one or more scenarios. In some scenarios, obtaining the first project code includes: obtaining the first project code from a local source, for example, when the code files of the software development project are stored locally, obtaining the code files of the software development project from the local source to obtain the first project code.

[0066] In other cases, obtaining the first project code includes: scanning the code repository to obtain multiple code files, and selecting the code file whose file type belongs to source code file as the first project code based on the file types of the multiple code files.

[0067] A code repository can be understood as a centralized data storage unit used to store and manage code files in a software development project. For example, a code repository can correspond to a software development project; that is, the code files generated in a software development project are stored in a code repository. Therefore, by scanning a code repository, multiple code files of a software development project can be obtained.

[0068] Considering that code files stored in a code repository can have different file types, such as source code files or configuration files, configuration files usually do not involve specific business logic and are only used to describe external information such as system running parameters, behavior rules, and environment configuration information. Code data governance is also usually unrelated to configuration files. Therefore, by parsing multiple code files to obtain the file types of multiple code files, we can filter out the code files that belong to source code files and use the code files that belong to source code files as the first project code.

[0069] In this way, the code of the first project is strongly associated with the specific business logic, reducing the negative impact of configuration files that are less related to the business logic on the construction of the first code graph, and thus on the subsequent use of the first code graph for code data governance.

[0070] Step S302: Extract multiple code entities from the first project code based on the abstract syntax tree of the first project code.

[0071] Step S303: Construct a first code graph based on multiple code entities of the first project code.

[0072] An abstract syntax tree (AST) can be understood as a structured abstract representation of the code of a first project. An AST can present the syntactic structure of the code of a first project in a tree-like data structure. By generating an AST of the code of a first project, multiple code entities of the first project can be extracted from the AST. For example, each node in the AST of the first project can be used as multiple code entities of the first project. Then, based on the AST of the first project, relationships between multiple code entities can be established. For example, edges in the AST of the first project can be used as relationships between multiple code entities to construct a first code graph.

[0073] In this way, by generating an abstract syntax tree, the linear first-item code is transformed into structured data, reducing the difficulty of analyzing the first-item code. Using the abstract syntax tree of the first-item code, the first code graph can be constructed quickly and efficiently.

[0074] Figure 4 The illustration shows a flowchart of constructing a first code graph for at least one scenario described herein.

[0075] like Figure 4 As shown, by scanning the code repository, multiple code files are obtained. These multiple code files are parsed to determine their file types. Then, it is determined whether the file type is a source code file. For the first project code that belongs to the source code file, an abstract syntax tree is generated, multiple code entities are extracted, and a first code graph is constructed.

[0076] Continue as Figure 4 As shown, in one or more scenarios in this paper, since the business logic of the first project code is represented by the first code graph, in order to perform efficient searching and accurate positioning in the first code graph when using the first code graph for code data governance in the future, a multi-dimensional index of the first code graph can be constructed, that is, it supports searching in the first code graph using indexes of different dimensions.

[0077] In some cases, the multidimensional index of the first code graph may include one or more of text indexes, structural indexes, and semantic indexes.

[0078] For example, a text index can be understood as a text-dimensional index, that is, using the text information of code entities as an index to search in the first code graph. A text index can include mappings between attribute information of multiple code entities and node identifiers of multiple nodes, such as mappings between identifiers of multiple code entities and node identifiers of multiple nodes, mappings between keywords in the code content of multiple code entities and node identifiers of multiple nodes, and mappings between comments in the code content of multiple code entities and node identifiers of multiple nodes, etc.

[0079] For example, a structural index can be understood as an index of the structural dimension, that is, using the relationships between code entities as an index to search in the first code graph. A structural index can include a mapping between the relationship types between code entities and the edge identifiers of multiple edges, such as a mapping between call relationships between code entities and the edge identifiers of multiple edges, or a mapping between inheritance relationships between code entities and the edge identifiers of multiple edges, etc.

[0080] For example, semantic indexing can be understood as an index based on semantic dimensions, that is, using the semantic information of code entities as an index to search in the first code graph. Semantic indexing can include the mapping relationship between multiple spatial vectors and the node identifiers of multiple nodes. Multiple spatial vectors can be obtained by converting the attribute information of multiple code entities into vectors, for example, by using a code embedding model (i.e., a vector model) to convert the attribute information (such as code content) of multiple code entities into multiple spatial vectors, and then constructing the mapping relationship between multiple spatial vectors and the node identifiers of multiple nodes to form a semantic index.

[0081] Thus, by constructing a multi-dimensional index, it is possible to search for information in the first code graph using multiple dimensions such as text, structure, and semantics. Compared with simple keyword search, a more comprehensive search can be achieved in the first code graph, thereby improving the accuracy of code data governance using the first code graph.

[0082] The process of using the first code graph for code data governance is explained below.

[0083] Figure 5 The illustration shows a flowchart of an AI-based code data governance method provided in at least one scenario of this paper.

[0084] like Figure 5As shown, the AI-based code data governance method in this scenario includes steps S501 to S503. In some cases, the executing entity of this AI-based code data governance method can be an electronic device with a client deployed, an electronic device with a server deployed, or any electronic device that connects the client and the server. This document does not limit the scope of one or more scenarios. The steps included in this AI-based code data governance method are described below:

[0085] Step S501: Obtain the first code to process the request.

[0086] In one or more scenarios described herein, a first code processing request can be used for data governance of the first project code. A first code processing request can be understood as a user's (e.g., a developer's) data governance requirement for the first project code. A first code processing request can be described in natural language. That is, when a user has a data governance requirement for the first project code, the data governance requirement is described in natural language to form a first code processing request.

[0087] In some cases, the first code processing request may correspond to a task type; that is, the first code processing request can be used to perform data governance on the code of the first project for a certain task type. For example, the task type corresponding to the first code processing request includes at least one of the following: code writing task, code optimization task, and problem investigation task.

[0088] Code writing tasks can be understood as tasks used to generate new code. For example, when the first code processing request is "Help me write a linker function," the task type of the first code processing request can be a code writing task. Code optimization tasks can be understood as tasks used to modify existing code. For example, when the first code processing request is "Help me find code with performance bottlenecks and optimize it," the task type of the first code processing request can be a code optimization task. Troubleshooting tasks can be understood as tasks used to locate parts of the code based on code problems. For example, when the first code processing request is "Why did the Ray task fail?", the task type of the first code processing request can be a troubleshooting task.

[0089] Thus, it supports data governance for code processing requests of different task types. That is, the AI-based code data governance methods provided in one or more scenarios in this paper can be flexibly applied to a variety of different code data governance scenarios and have strong versatility.

[0090] Step S502: Based on the first code processing request, search the first code graph to obtain the first node set.

[0091] In one or more scenarios described herein, the set of first nodes may include multiple first nodes, and the degree of association between each of the multiple first nodes and the first code processing request satisfies a first set condition.

[0092] The first setting condition can be understood as a condition used to describe the strong correlation between a node and a first code processing request. For example, the degree of correlation between a first node and a first code processing request can be presented in the form of a score. The first setting condition can be that the degree of correlation between a node and a first code processing request is greater than a first setting threshold.

[0093] In other words, by searching in the first code graph, multiple first nodes that are strongly related to the first code processing request are found in the first code graph, thereby locating the part of the code involved in the first code processing request in the first project code.

[0094] In some cases, the degree of association between each of the multiple first nodes and the first code processing request can be measured by at least one of the following: the degree of text matching between the first node and the first code processing request; the degree of association between the first node and other nodes in the first code graph; the distance between the first node and other first nodes in the first code graph; the frequency of the first node in historical data governance results; and the degree of association between external data associated with the first node and the first code processing request.

[0095] The text matching degree between the first node and the first code processing request can be understood as the text relevance between them. For example, calculating the text similarity between the first node (such as the code content of the code entity represented by the first node) and the first code processing request yields the text matching degree. The higher the text matching degree, the stronger the relevance between the first node and the first code processing request. Therefore, the text matching degree between the first node and the first code processing request can measure the degree of association between them.

[0096] The degree of association between the first node and other nodes in the first code graph can be used to represent the graph centrality of the first node. For example, the degree of association between the first node and other nodes in the first code graph can be obtained based on the number of edges of the first node. The more edges the first node has, the higher the degree of association between the first node and other nodes in the first code graph. A higher degree of association between the first node and other nodes in the first code graph indicates that the first node is more important in the first code graph. Therefore, the degree of association between the first node and other nodes in the first code graph can measure the degree of association between the first node and the first code processing request.

[0097] The distance between a first node and other first nodes in the first code graph can be used to represent the structural proximity of that first node to the other first nodes. For example, the distance between a first node and other first nodes in the first code graph can be obtained based on the number of edges between the first node and other first nodes. The fewer the edges between a first node and other first nodes, the closer the first node is to the other first nodes in the first code graph. For example, when node A is connected to node B, the distance between node A and node B is 1; when node A is connected to node B and node B is connected to node C, the distance between node A and node C is 2. The closer the distance between a first node and other first nodes in the first code graph, the more relevant that first node is to the other first nodes. Therefore, the distance between a first node and other first nodes in the first code graph can measure the degree of association between the first node and the first code processing request.

[0098] Historical data governance results can be understood as the data governance results corresponding to historical code processing requests. These historical code processing requests correspond to the same task type as the first code processing request. The frequency of the first node appearing in the historical data governance results can be understood as the frequency with which the first node is operated on or processed in other code processing requests of the same task type. For example, by obtaining historical code processing requests within a historical time period and obtaining historical data governance results, the number of times the first node appears in the historical data governance results is determined. Then, based on the ratio between the number of times the first node appears in the historical data governance results and the total number of historical data governance results, the frequency of the first node in the historical data governance results is obtained. The higher the frequency of the first node in the historical data governance results, the more relevant the first node is to that task type. Therefore, the frequency of the first node in the historical data governance results can measure the degree of correlation between the first node and the first code processing request.

[0099] External data associated with the first node can be understood as external data related to the runtime status of the code entity represented by the first node. The degree of correlation between the external data associated with the first node and the first code processing request can be understood as the degree of influence of the external data associated with the first node on the external data type involved in the first code processing request. For example, when the external data type involved in the first code processing request is runtime, the external data associated with the first node can be the runtime of the code content of the code entity represented by the first node. The longer the runtime of the code content, the higher the degree of correlation between the external data associated with the first node and the first code processing request. The higher the degree of correlation between the external data associated with the first node and the first code processing request, the greater the influence of the first node on the first code processing request. Therefore, the degree of correlation between the external data associated with the first node and the first code processing request can measure the degree of correlation between the first node and the first code processing request.

[0100] In this way, by measuring the degree of correlation between nodes and first code processing requests from different dimensions and multiple aspects, the correlation between nodes and first code processing requests can be comprehensively evaluated, and multiple first nodes that are strongly related to first code processing requests can be accurately identified from the multiple nodes included in the first code graph.

[0101] In some possible implementations, the degree of association between the external data associated with the first node and the first code processing request can be obtained through the following steps: identifying the semantic information of the first code processing request, obtaining the external data type that matches the semantic information of the first code processing request, calling the tool corresponding to the external data type, obtaining the external data associated with the first node, analyzing the external data based on the semantic information of the first code processing request, and obtaining the degree of association between the external data associated with the first node and the first code processing request.

[0102] External data types can be understood as the type of external data that needs to be acquired. By identifying the semantic information of the first code processing request, the external data types involved in the processing of the first code processing request can be determined. For example, external data types can be runtime performance data, data source data, etc.

[0103] Tools corresponding to external data types can be understood as tools used to acquire external data. For example, tools corresponding to external data types can be data acquisition plugins, application programming interfaces (APIs), etc.

[0104] By calling the tool corresponding to the external data type, the external data associated with the first node is obtained. Then, based on the semantic information of the first code processing request, the relationship between the external data and the degree of association between the external data associated with the first node and the first code processing request is determined. For example, the larger the external data, the greater the degree of association between the external data associated with the first node and the first code processing request, or the smaller the external data, the greater the degree of association between the external data associated with the first node and the first code processing request.

[0105] Thus, when selecting the first node that is strongly related to the first code processing request, we combine external data to judge the degree of influence of the node in code data governance, and select the first node in a targeted manner based on the first code processing request.

[0106] In one or more scenarios described herein, the first code graph can be searched using a heuristic search approach. For example, based on a first code processing request, at least one initial node is selected from the first code graph. Starting from the at least one initial node, the first code graph is searched to obtain at least one associated node connected to the at least one initial node. Based on the at least one initial node and the at least one associated node, a first node set is obtained, wherein the degree of association between each associated node and the first code processing request satisfies a first predefined condition.

[0107] In other words, firstly, the initial node related to the first code processing request is found in the first code graph. Then, starting from the initial node, the exploration is carried out in the first code graph. In each step of the exploration, the associated nodes strongly related to the first code processing request are found. Finally, the initial node and associated nodes are combined to determine the first node set strongly related to the first code processing request.

[0108] In this way, we first perform preliminary positioning in multiple nodes in the first code graph, and then perform precise search starting from the initial node. Without traversing all nodes in the first code graph, we can efficiently search for the set of first nodes that are strongly related to the first code processing request.

[0109] In some possible implementations, at least one initial node is selected from the first code graph based on the first code processing request, including: using at least one of text index, structure index and semantic index, to perform a search in the first code graph based on the first code processing request to obtain at least one initial node.

[0110] Since initial positioning needs to be performed across multiple nodes in the first code graph, a multidimensional index is used to retrieve at least one initial node related to the first code processing request. This initial positioning is performed from different dimensions within the first code graph, resulting in a more comprehensive retrieval of initial nodes that better match the first code processing request (i.e., the user's needs for code data governance), thus improving the effectiveness of code data governance.

[0111] The following describes the specific graph search process.

[0112] For example, starting from at least one initial node, a search is performed on the first code graph to obtain at least one associated node connected to the at least one initial node, and a first node set is obtained based on the at least one initial node and the at least one associated node, including: starting from at least one initial node, performing a multi-round search process in the first code graph, and obtaining the first node set based on the search results of the multi-round search process, wherein each round of search process is used to determine whether to add a second node to the first node set based on the degree of association between a second node and a first code processing request, wherein the second node is at least one initial node or at least one associated node.

[0113] In other words, the graph search process involves multiple rounds of searching. In each round, it is determined whether the initial node or the associated node connected to the initial node (e.g., directly connected or indirectly connected through other nodes) needs to be added to the first node set as feedback for the final graph search result.

[0114] It should be noted that although the degree of association between the associated node and the first code processing request meets the first set condition, it does not mean that all associated nodes should be added to the first node set. For example, when the number of nodes in the first node set is sufficient, associated nodes will no longer be added to the first node set.

[0115] During graph search, depth-first search or breadth-first search strategies can be used simultaneously or alternately to reduce the possibility of getting stuck in local optima or missing neighboring strongly related nodes.

[0116] For example, based on the degree of association between at least one initial node and the first code processing request, at least one initial node is sorted in descending order, and the sorted initial node is added to the first queue in sequence. The following steps are executed repeatedly until the termination condition is met: obtain the second node at the head of the first queue; in response to the degree of association between the second node and the first code processing request meeting the first set condition, add the second node to the first node set; traverse the neighbor nodes of the second node; in response to the neighbor node being the first visited node and the degree of association between the neighbor node and the first code processing request meeting the second set condition, add the neighbor node to the first position of the first queue based on the degree of association between the neighbor node and the first code processing request, so that the degree of association between the nodes from the head to the tail of the first queue and the first code processing request gradually decreases.

[0117] Figure 6 The diagram illustrates a flowchart of a graph search provided in at least one of the cases described herein.

[0118] like Figure 6 As shown, after selecting at least one initial node from the first code graph, an empty first queue is initialized. According to the degree of association between each initial node and the first code processing request, at least one initial node is added to the first queue in sequence. The initial nodes are arranged in descending order in the first queue, so that the initial node with the highest degree of association with the first code processing request is located at the head of the first queue.

[0119] Next, the search process is executed in multiple rounds. In each round: it is checked whether the first queue is empty. If the first queue is empty, it indicates that the multi-round search process has ended and there are no nodes to explore. The first node set is output. If the first queue is not empty, the second node at the head of the first queue is obtained. The second node is the node in the first queue with the highest correlation to the first code processing request. It is checked whether the correlation between the second node and the first code processing request meets the first set condition. If the first set condition is not met, it means that the second node is not strongly correlated with the first code processing request. The second node is skipped and not added to the first node set. If the first set condition is met, it means that the second node is strongly correlated with the first code processing request. The second node is added to the first node set. Then, the neighbor nodes of the second node are traversed. For the currently traversed neighbor node, it is checked whether the correlation between the neighbor node and the first code processing request meets the second set condition. The second set condition can be understood as the condition used to describe whether to add the node to the first queue. The second set condition can be more stringent than the first set condition. The first condition is more lenient. For example, when the correlation between a node and the first code processing request can be represented by a score, the first condition can be that the correlation between the node and the first code processing request is greater than a first threshold. The second condition can be that the correlation between the node and the first code processing request is greater than a second threshold. If the first threshold is greater than the second threshold, the next neighbor node is traversed. If the second condition is met, the neighbor node is added to the first queue, so that the correlation between the nodes and the first code processing request gradually decreases from the head to the tail of the first queue. That is, the node with the highest correlation with the first code processing request is located at the head of the first queue, and the node with the lowest correlation with the first code processing request is located at the tail of the first queue. At the end of this round of search, it is determined whether the search process meets the termination condition. For example, the termination condition can be that the number of first nodes in the first node set reaches a number threshold, the maximum search depth or maximum search breadth is reached, or the search time threshold is exceeded.

[0120] In this way, by using graph search algorithms, the set of first nodes that are strongly related to the first code processing request can be logically found in the first code graph, saving computing resources and improving the search efficiency of the first node set.

[0121] Step S503: Process the first code processing request according to the first subgraph corresponding to the multiple first nodes in the first code graph to obtain the data governance result corresponding to the first code processing request.

[0122] In one or more scenarios described in this paper, since the first set of nodes is obtained by searching the first code graph, the first code processing request is not processed using isolated code fragments, but rather using the first subgraphs corresponding to multiple first nodes in the first code graph. In this way, code data governance is performed by combining rich contextual information, thereby improving the accuracy of code data governance.

[0123] The first subgraph can be understood as a subgraph composed of multiple first nodes in the first code graph and the relationships between multiple first nodes. The first subgraph can not only represent code entities that are strongly related to the first code processing request, but also represent the relationships between code entities that are strongly related to the first code processing request. For example, it includes multiple code entities represented by multiple first nodes in the first node set and code paths associated with multiple code entities.

[0124] The first subgraph is used to process the first code processing request, generating the corresponding data governance result. The processing procedure for the first code processing request can be related to the task type. For example, for a code writing task, the multiple code entities represented by the first subgraph and the code paths associated with them can serve as examples to generate new code, ensuring that the generated code can run smoothly within the multiple code entities represented by the first subgraph and the associated code paths. Similarly, for a code optimization task, the multiple code entities represented by the first subgraph and the associated code paths can be understood as the code that needs optimization; modifying these entities yields the optimized code. Furthermore, for a problem investigation task, the multiple code entities represented by the first subgraph and the associated code paths can be understood as the code that caused the problem; displaying these entities helps developers quickly locate the problem.

[0125] Thus, in code data governance scenarios, by converting project code into a code graph and searching within that graph, one can quickly focus on code areas strongly related to code processing requests without traversing the entire project code, thereby improving search efficiency and effectiveness, and ultimately enhancing the efficiency and accuracy of code data governance.

[0126] The following explanation uses specific code data governance requests as an example.

[0127] Figure 7 The illustration shows a schematic diagram of a data governance outcome provided in at least one scenario of this paper.

[0128] like Figure 7As shown, the first code processing request can be "Why is the Spark task running slowly, causing data skew?", meaning the task type corresponding to the first code processing request is a problem investigation task. Based on the first code processing request, the initial node "main.py" is selected from the first code graph, and then multiple rounds of search are performed to obtain the first node set. The first nodes in the first node set include "main.py", "data_processing.py", "group_by_aggregation", "join_transformation", "groupBy.agg", "dataFrame.join", ".groupBy(user_id).agg", and "df1.join(df2,user_id)". When judging the degree of association between a node and the first code processing request, the degree of association between the node and the first code processing request is measured by the degree of association between the external data associated with the node and the first code processing request. For example, the obtained external data can include performance index data and data source partition data.

[0129] After completing the graph search and obtaining the first set of nodes, the first code processing request is analyzed by combining the first subgraphs corresponding to multiple first nodes, identifying data skew patterns, finding operations that may cause data skew, and providing specific optimization suggestions, such as using broadcast connections, adding random prefixes, and filtering outliers.

[0130] In this way, by combining the first code diagram, automatic and efficient data governance can be achieved for the first code processing request, ensuring the accuracy and comprehensiveness of the data governance results.

[0131] Based on the AI-based code data governance method provided in at least one scenario of this paper, an AI-based code data governance device is also provided in at least one scenario of this paper. The following will combine... Figure 8 A detailed description of an AI-based code data governance device is provided.

[0132] Figure 8 The schematic diagram illustrates the structure of an AI-based code data governance device provided in at least one scenario of this paper.

[0133] like Figure 8As shown, the AI-based code data governance device 800 includes an acquisition module 801, a search module 802, and a processing module 803. For example, the acquisition module 801, search module 802, and processing module 803 can be implemented using hardware (e.g., circuit) modules or software modules. The following scenarios are similar and will not be repeated. For example, the acquisition module 801, search module 802, and processing module 803 can be implemented using a central processing unit (CPU), a general-purpose graphics processing unit (GPGPU), a graphics processing unit (GPU), a tensor processor (TPU), a field-programmable gate array (FPGA), or other forms of processing units with data processing capabilities and / or instruction execution capabilities, along with corresponding computer instructions.

[0134] The acquisition module 801 is configured to acquire a first code processing request, wherein the first code processing request is used to perform data governance on the code of the first project. For example, the acquisition module 801 can be configured to execute step S501 described above, and its specific implementation principle can be referred to the relevant description of step S501, which will not be repeated here.

[0135] The search module 802 is configured to: search the first code graph based on the first code processing request to obtain a first node set, wherein the nodes in the first code graph represent code entities in the first project code, the edges in the first code graph represent the relationships between the code entities, and the first node set includes multiple first nodes, wherein the degree of association between each of the multiple first nodes and the first code processing request satisfies a first preset condition. For example, the search module 802 can be configured to execute step S502 described above; its specific implementation principle can be found in the relevant description of step S502, and will not be repeated here.

[0136] The processing module 803 is configured to process the first code processing request according to the first subgraph corresponding to the plurality of first nodes in the first code graph, and obtain the data governance result corresponding to the first code processing request. For example, the processing module 803 can be configured to execute step S503 described above. The specific implementation principle can be referred to the relevant description of step S503, which will not be repeated here.

[0137] In at least one of the embodiments described herein, the degree of association between each of the plurality of first nodes and the first code processing request is measured by at least one of the following: the degree of text matching between the first node and the first code processing request; the degree of association between the first node and other nodes in the first code graph; the distance between the first node and other first nodes in the first code graph; the frequency of the first node appearing in historical data governance results, wherein the historical data governance results are the data governance results corresponding to historical code processing requests, and the historical code processing requests correspond to the same task type as the first code processing request; and the degree of association between external data associated with the first node and the first code processing request.

[0138] In at least one scenario described herein, the degree of association between the external data associated with the first node and the first code processing request is obtained through the following steps: identifying the semantic information of the first code processing request; obtaining an external data type that matches the semantic information of the first code processing request; calling a tool corresponding to the external data type to obtain the external data associated with the first node; and analyzing the external data based on the semantic information of the first code processing request to obtain the degree of association between the external data associated with the first node and the first code processing request.

[0139] In at least one scenario described herein, the search module 802 is further configured to: filter at least one initial node from the first code graph based on the first code processing request; search the first code graph starting from the at least one initial node to obtain at least one associated node connected to the at least one initial node; and obtain the first node set based on the at least one initial node and the at least one associated node, wherein the degree of association between each associated node in the at least one and the first code processing request satisfies the first set condition.

[0140] In at least one embodiment of this paper, the search module 802 is further configured to: use at least one of a text index, a structure index, and a semantic index to perform a retrieval in the first code graph based on the first code processing request to obtain the at least one initial node; wherein the text index includes a mapping relationship between attribute information of multiple code entities and node identifiers of multiple nodes, the structure index includes a mapping relationship between the relationship type of code entities and edge identifiers of multiple edges, and the semantic index includes a mapping relationship between multiple spatial vectors and node identifiers of multiple nodes, wherein the multiple spatial vectors are obtained by converting the attribute information of the multiple code entities into vectors.

[0141] In at least one scenario described herein, the search module 802 is further configured to: perform multiple rounds of search processes in the first code graph, starting from the at least one initial node, and obtain the first node set based on the search results of the multiple rounds of search processes; wherein each round of the search process is used to determine whether to add the second node to the first node set based on the degree of association between the second node and the first code processing request, wherein the second node is the at least one initial node or the at least one associated node.

[0142] In at least one scenario described herein, the search module 802 is further configured to: sort the at least one initial node in descending order based on the degree of association between the at least one initial node and the first code processing request, and add the at least one initial node in descending order to a first queue in sequence; cyclically execute the following steps until a termination condition is met: obtain the second node at the head of the first queue; in response to the degree of association between the second node and the first code processing request satisfying the first set condition, add the second node to the first node set; traverse the neighbor nodes of the second node, and in response to the neighbor node being the first visited node and the degree of association between the neighbor node and the first code processing request satisfying the second set condition, add the neighbor node to the first position of the first queue based on the degree of association between the neighbor node and the first code processing request, so that the degree of association between the nodes from the head to the tail of the first queue and the first code processing request gradually decreases.

[0143] In at least one scenario described herein, the AI-based code data governance device 800 further includes a construction module configured to: acquire first project code; extract multiple code entities of the first project code based on an abstract syntax tree of the first project code; and construct the first code graph based on the multiple code entities of the first project code.

[0144] In at least one scenario described herein, the build module is further configured to: scan a code repository to obtain multiple code files; and, based on the file type of the multiple code files, select the code file whose file type belongs to source code files as the first project code.

[0145] In at least one of the scenarios described herein, the task type corresponding to the first code processing request includes at least one of the following: code writing task, code optimization task, and problem investigation task.

[0146] It should be noted that, for clarity and brevity, at least one scenario herein does not present all the constituent units of the AI-based code data governance device 800. To achieve the necessary functions of the AI-based code data governance device 800, those skilled in the art may provide or configure other constituent units (not shown) according to specific needs, and one or more scenarios herein do not impose any limitations on this.

[0147] The AI-based code data management device 800 provided in at least one aspect of this document and the AI-based code data management method provided in at least one aspect of this document are based on the same inventive concept and can achieve the same technical effect and the same technical purpose as the AI-based code data management method provided in at least one aspect of this document. For details, please refer to the relevant description above, which will not be repeated here.

[0148] At least one embodiment of this document also provides an electronic device including a processing device and a storage device, the storage device including one or more computer program modules; wherein the one or more computer program modules are stored in the storage device and configured to be executed by the processing device, the one or more computer program modules being used to implement the AI-based code data governance method provided in any embodiment of this document.

[0149] For example, the processing device may be a processor, such as a central processing unit (CPU), digital signal processor (DSP), image processor (GPU), general-purpose graphics processor (GPGPU), or other form of processing unit with data processing capabilities and / or instruction execution capabilities. It may be a general-purpose processor or a dedicated processor and may control other components in the electronic device to perform the desired functions.

[0150] For example, the storage device may be a memory, which may include one or more computer program products. These computer program products may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may, for example, include random access memory (RAM) and / or cache memory. The non-volatile memory may, for example, include read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and a processing device may execute these program instructions to implement the function (implemented by the processing device) in at least one of the embodiments described herein, and / or other desired functions. Various application programs and various data may also be stored on the computer-readable storage medium, which is not limited in the embodiments described herein.

[0151] The following is for reference. Figure 9The diagram illustrates a structural schematic of an electronic device (e.g., a terminal device or a server) 900 suitable for implementing at least one of the embodiments described herein. The terminal device in at least one embodiment may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital radio receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 9 The electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of at least one of the scenarios described herein.

[0152] like Figure 9 As shown, electronic device 900 may include a processing device (e.g., a central processing unit, a graphics processor, etc.) 901, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 902 or a program loaded from storage device 908 into random access memory (RAM) 903. RAM 903 also stores various programs and data required for the operation of electronic device 900. Processing device 901, ROM 902, and RAM 903 are interconnected via bus 904. Input / output (I / O) interface 905 is also connected to bus 904.

[0153] Typically, the following devices can be connected to I / O interface 905: input devices 906 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 907 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 908 including, for example, magnetic tapes, hard disks, etc.; and communication devices 909. Communication device 909 allows electronic device 900 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 9 An electronic device 900 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.

[0154] In particular, according to one or more embodiments herein, the process described in the above-referenced flowchart can be implemented as a computer software program. For example, one or more embodiments herein include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via communication device 909, or installed from storage device 908, or installed from ROM 902. When the computer program is executed by processing device 901, it performs the functions defined in the method of at least one embodiment herein.

[0155] The electronic device 900 provided in at least one aspect of this article and the code data governance method based on artificial intelligence provided in at least one aspect of this article are based on the same inventive concept and can achieve the same technical effect and the same technical purpose as the code data governance method based on artificial intelligence provided in at least one aspect of this article. For details, please refer to the relevant description above, which will not be repeated here.

[0156] It should be noted that the computer-readable medium described above can be a computer-readable signal medium, a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this document, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this document, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0157] The computer-readable storage medium provided in at least one aspect of this article and the AI-based code data governance method provided in at least one aspect of this article are based on the same inventive concept and can achieve the same technical effect and the same technical purpose as the AI-based code data governance method provided in at least one aspect of this article. For details, please refer to the relevant descriptions above, which will not be repeated here.

[0158] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.

[0159] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.

[0160] The aforementioned computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to perform the aforementioned AI-based code data governance method.

[0161] Computer program code for performing the operations described herein may be written in one or more programming languages ​​or a combination thereof, including but not limited to object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code may execute entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer may be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or may be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0162] One or more scenarios described herein also provide a computer program product comprising one or more computer instructions. When these computer instructions are loaded and executed on a computing device, all or part of the flow or function described in any of these scenarios is generated.

[0163] The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, or data center to another website, computer, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means.

[0164] When the computer program product is executed by a computer, the computer executes any of the aforementioned methods of the AI-based code data governance method. The computer program product can be a software installation package; when any of the aforementioned AI-based code data governance methods needs to be used, the computer program product can be downloaded and executed on the computer.

[0165] The computer program product provided in at least one of the embodiments described herein and the AI-based code data governance method provided in at least one of the embodiments described herein are based on the same inventive concept and can achieve the same technical effect and the same technical purpose as the AI-based code data governance method provided in at least one of the embodiments described herein. For details, please refer to the relevant descriptions above, which will not be repeated here.

[0166] The flowcharts and block diagrams in the accompanying figures illustrate the architecture, functionality, and operation of possible implementations of the systems, methods, and computer program products according to the various scenarios described herein. In this respect, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the figures. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0167] The units or modules described in at least one of the scenarios herein can be implemented in software or hardware. The names of the units or modules do not, in some cases, constitute a limitation on the unit or module itself.

[0168] The functions described above in this document can be performed at least in part by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), and so on.

[0169] In the context of this document, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0170] Based on one or more scenarios described in this paper, Example 1 provides an AI-based code data governance approach, including:

[0171] Obtain a first code processing request, wherein the first code processing request is used to perform data governance on the code of the first project;

[0172] Based on the first code processing request, a first code graph is searched to obtain a first node set, wherein the nodes in the first code graph represent code entities in the code of the first project, the edges in the first code graph represent the relationship between the code entities, the first node set includes multiple first nodes, and the degree of association between each of the multiple first nodes and the first code processing request satisfies a first set condition.

[0173] The first code processing request is processed according to the first subgraph corresponding to the multiple first nodes in the first code graph to obtain the data governance result corresponding to the first code processing request.

[0174] Based on one or more scenarios in this paper, Example 2 provides that the degree of association between each of the multiple first nodes in Example 1 and the first code processing request is measured by at least one of the following:

[0175] The degree of text matching between the first node and the first code processing request;

[0176] The degree of association between the first node and other nodes in the first code graph;

[0177] The distance between the first node and other first nodes in the first code graph;

[0178] The frequency of the first node appearing in historical data governance results, wherein the historical data governance results are the data governance results corresponding to historical code processing requests, and the historical code processing requests correspond to the same task type as the first code processing request; and

[0179] The degree of association between the external data associated with the first node and the first code processing request.

[0180] Based on one or more scenarios in this paper, Example 3 provides that the degree of association between the external data associated with the first node in Example 2 and the first code processing request is obtained through the following steps:

[0181] Identify the semantic information of the request processed by the first code;

[0182] Obtain the external data type that matches the semantic information of the request processed by the first code;

[0183] Call the tool corresponding to the external data type to obtain the external data associated with the first node;

[0184] Based on the semantic information of the first code processing request, the external data is analyzed to obtain the degree of correlation between the external data associated with the first node and the first code processing request.

[0185] Based on one or more scenarios described in this paper, Example 4 provides the method described in Example 1, which involves searching the first code graph based on the first code processing request to obtain a first set of nodes, including:

[0186] Based on the first code processing request, at least one initial node is selected from the first code graph;

[0187] Starting from the at least one initial node, the first code graph is searched to obtain at least one associated node connected to the at least one initial node, and the first node set is obtained based on the at least one initial node and the at least one associated node, wherein the degree of association between each associated node in the at least one and the first code processing request satisfies the first set condition.

[0188] Based on one or more scenarios described herein, Example 5 provides the method from Example 4 for processing requests using the first code, filtering at least one initial node from the first code graph, including:

[0189] Using at least one of text index, structure index and semantic index, based on the first code processing request, a search is performed in the first code graph to obtain the at least one initial node;

[0190] The text index includes a mapping relationship between the attribute information of multiple code entities and the node identifiers of multiple nodes; the structure index includes a mapping relationship between the relationship type between code entities and the edge identifiers of multiple edges; and the semantic index includes a mapping relationship between multiple spatial vectors and the node identifiers of multiple nodes. The multiple spatial vectors are obtained by converting the attribute information of the multiple code entities into vectors.

[0191] Based on one or more scenarios in this document, Example Six provides the method described in Example Four, which involves searching the first code graph starting from the at least one initial node to obtain at least one associated node connected to the at least one initial node, and obtaining the first node set based on the at least one initial node and the at least one associated node, including:

[0192] Starting from the at least one initial node, a multi-round search process is performed in the first code graph, and the first node set is obtained based on the search results of the multi-round search process;

[0193] In each round of the search process, the second node is used to determine whether to add the second node to the first node set based on the degree of association between the second node and the first code processing request. The second node is either the at least one initial node or the at least one associated node.

[0194] Based on one or more scenarios described herein, Example 7 provides the method described in Example 6, which involves performing a multi-round search process in the first code graph starting from the at least one initial node, and obtaining the first set of nodes based on the search results of the multi-round search process, including:

[0195] Based on the degree of association between the at least one initial node and the first code processing request, the at least one initial node is sorted in descending order, and the at least one initial node in descending order is added to the first queue in sequence;

[0196] The following steps are executed repeatedly until a termination condition is met: The second node at the head of the first queue is obtained; in response to the association degree between the second node and the first code processing request satisfying the first set condition, the second node is added to the first node set; the neighboring nodes of the second node are traversed, and in response to the neighboring node being the first visited node and the association degree between the neighboring node and the first code processing request satisfying the second set condition, the neighboring node is added to the first position of the first queue based on the association degree between the neighboring node and the first code processing request, so that the association degree between the nodes from the head to the tail of the first queue and the first code processing request gradually decreases.

[0197] Based on one or more scenarios in this article, Example 8 provides the first code graph from Example 1 constructed through the following steps:

[0198] Get the code for the first project;

[0199] Based on the abstract syntax tree of the first project code, extract multiple code entities from the first project code;

[0200] The first code graph is constructed based on multiple code entities of the first project code.

[0201] Based on one or more scenarios in this article, Example 9 provides the code for obtaining the first item from Example 8, including:

[0202] Scan the code repository to obtain multiple code files;

[0203] Based on the file types of the multiple code files, the code file whose file type belongs to source code file is selected as the first project code.

[0204] Based on one or more scenarios in this article, Example 10 provides that the task type corresponding to the first code processing request in any of the examples from Example 1 to Example 9 includes at least one of the following: code writing task, code optimization task, and problem investigation task.

[0205] Based on one or more scenarios described herein, Example 11 provides an AI-based code data governance apparatus, comprising:

[0206] The acquisition module is configured to: acquire a first code processing request, wherein the first code processing request is used to perform data governance on the code of the first project;

[0207] The search module is configured to: search the first code graph based on the first code processing request to obtain a first node set, wherein the nodes in the first code graph represent code entities in the code of the first project, the edges in the first code graph represent the relationship between the code entities, the first node set includes multiple first nodes, and the degree of association between each of the multiple first nodes and the first code processing request satisfies a first set condition.

[0208] The processing module is configured to process the first code processing request according to the first subgraph corresponding to the plurality of first nodes in the first code graph, and obtain the data governance result corresponding to the first code processing request.

[0209] According to one or more of the scenarios described herein, Example Twelve provides an electronic device comprising:

[0210] At least one processor; and

[0211] At least one memory, including one or more computer program instructions;

[0212] The one or more computer program instructions are executed by the processor at runtime, which provides an AI-based code data governance method for at least one scenario described herein.

[0213] According to one or more scenarios described herein, Example Thirteen provides a computer-readable storage medium that non-transitory stores computer-readable instructions, wherein the AI-based code data governance method provided by at least one scenario described herein is implemented when the computer-readable instructions are executed by a processor.

[0214] According to one or more scenarios described herein, Example Fourteen provides a computer program product including a computer program that, when executed by a processor, implements the AI-based code data governance method provided in at least one scenario described herein.

[0215] The above description is merely a preferred embodiment and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of disclosure herein is not limited to technical solutions formed by specific combinations of the above-described technical features, but also includes other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-disclosed concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed herein that have similar functions.

[0216] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain contexts, multitasking and parallel processing may be advantageous. Similarly, while some specific implementation details are included in the above discussion, these should not be interpreted as limiting the scope of this paper. Certain features described in the context of a single case can also be implemented in combination within that single case. Conversely, various features described in the context of a single case can also be implemented individually or in any suitable sub-combination in multiple cases.

[0217] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.

Claims

1. An AI-based code data governance method, comprising: Obtain a first code processing request, wherein the first code processing request is used to perform data governance on the code of the first project; Based on the first code processing request, a first code graph is searched to obtain a first node set, wherein the nodes in the first code graph represent code entities in the code of the first project, the edges in the first code graph represent the relationship between the code entities, the first node set includes multiple first nodes, and the degree of association between each of the multiple first nodes and the first code processing request satisfies a first set condition. The first code processing request is processed according to the first subgraph corresponding to the multiple first nodes in the first code graph to obtain the data governance result corresponding to the first code processing request.

2. The method according to claim 1, wherein, The degree of association between each of the plurality of first nodes and the first code processing request is measured by at least one of the following: The degree of text matching between the first node and the first code processing request; The degree of association between the first node and other nodes in the first code graph; The distance between the first node and other first nodes in the first code graph; The frequency of the first node appearing in the historical data governance results, wherein the historical data governance results are the data governance results corresponding to the historical code processing requests, and the historical code processing requests correspond to the same task type as the first code processing request; as well as The degree of association between the external data associated with the first node and the first code processing request.

3. The method according to claim 2, wherein, The degree of association between the external data associated with the first node and the first code processing request is obtained through the following steps: Identify the semantic information of the request processed by the first code; Obtain the external data type that matches the semantic information of the request processed by the first code; Call the tool corresponding to the external data type to obtain the external data associated with the first node; Based on the semantic information of the first code processing request, the external data is analyzed to obtain the degree of correlation between the external data associated with the first node and the first code processing request.

4. The method according to claim 1, wherein, The step of searching the first code graph based on the first code processing request to obtain a first set of nodes includes: Based on the first code processing request, at least one initial node is selected from the first code graph; Starting from the at least one initial node, the first code graph is searched to obtain at least one associated node connected to the at least one initial node, and the first node set is obtained based on the at least one initial node and the at least one associated node, wherein the degree of association between each associated node in the at least one and the first code processing request satisfies the first set condition.

5. The method according to claim 4, wherein, The step of selecting at least one initial node from the first code graph based on the first code processing request includes: Using at least one of text index, structure index and semantic index, based on the first code processing request, a search is performed in the first code graph to obtain the at least one initial node; The text index includes a mapping relationship between the attribute information of multiple code entities and the node identifiers of multiple nodes; the structure index includes a mapping relationship between the relationship type between code entities and the edge identifiers of multiple edges; and the semantic index includes a mapping relationship between multiple spatial vectors and the node identifiers of multiple nodes. The multiple spatial vectors are obtained by converting the attribute information of the multiple code entities into vectors.

6. The method according to claim 4, wherein, Starting from the at least one initial node, the first code graph is searched to obtain at least one associated node connected to the at least one initial node. Based on the at least one initial node and the at least one associated node, the first node set is obtained, including: Starting from the at least one initial node, perform multiple rounds of search in the first code graph, and obtain the first node set based on the search results of the multiple rounds of search. In each round of the search process, the second node is used to determine whether to add the second node to the first node set based on the degree of association between the second node and the first code processing request. The second node is either the at least one initial node or the at least one associated node.

7. The method according to claim 6, wherein, Starting from the at least one initial node, a multi-round search process is performed in the first code graph, and the first node set is obtained based on the search results of the multi-round search process, including: Based on the degree of association between the at least one initial node and the first code processing request, the at least one initial node is sorted in descending order, and the at least one initial node in descending order is added to the first queue in sequence; Repeat the following steps until the termination condition is met: Obtain the second node located at the head of the first queue; In response to the fact that the degree of association between the second node and the first code processing request meets the first set condition, the second node is added to the first node set; Traverse the neighbor nodes of the second node. In response to the neighbor node being the first visited node and the degree of association between the neighbor node and the first code processing request satisfying the second set condition, add the neighbor node to the first position of the first queue based on the degree of association between the neighbor node and the first code processing request, so that the degree of association between the nodes from the head to the tail of the first queue and the first code processing request gradually decreases.

8. The method according to claim 1, wherein, The first code graph is constructed through the following steps: Get the code for the first project; Based on the abstract syntax tree of the first project code, extract multiple code entities from the first project code; The first code graph is constructed based on multiple code entities of the first project code.

9. The method according to claim 8, wherein, The process of obtaining the first project code includes: Scan the code repository to obtain multiple code files; Based on the file types of the multiple code files, the code file whose file type belongs to source code file is selected as the first project code.

10. The method according to any one of claims 1 to 9, wherein, The task type corresponding to the first code processing request includes at least one of the following: code writing task, code optimization task, and problem investigation task.

11. An artificial intelligence-based code data governance device, comprising: The acquisition module is configured to: acquire a first code processing request, wherein the first code processing request is used to perform data governance on the code of the first project; The search module is configured to: search the first code graph based on the first code processing request to obtain a first node set, wherein the nodes in the first code graph represent code entities in the code of the first project, the edges in the first code graph represent the relationship between the code entities, the first node set includes multiple first nodes, and the degree of association between each of the multiple first nodes and the first code processing request satisfies a first set condition. The processing module is configured to process the first code processing request according to the first subgraph corresponding to the plurality of first nodes in the first code graph, and obtain the data governance result corresponding to the first code processing request.

12. An electronic device, comprising: At least one processor; as well as At least one memory, including one or more computer program instructions; The one or more computer program instructions are executed by the processor to perform the method according to any one of claims 1 to 10.

13. A computer-readable storage medium for non-transitory storage of computer-readable instructions, wherein, The method of any one of claims 1 to 10 is implemented when the computer-readable instructions are executed by a processor.

14. A computer program product comprising a computer program that, when executed by a processor, implements the method of any one of claims 1 to 10.