Code text processing method, code supplementing method, and computing device
By parsing the pending code text and looking for valid reference code metadata from the target database, the problem of code generation models relying on text similarity to find reference code in the prior art is solved, and high-accurate code text processing is achieved.
Patent Information
- Application Number
- PCT/CN2024/122293
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-28
- Filing Date
- 2024-09-29
- Publication Date
- 2025-06-05
AI Technical Summary
When processing code text, existing code generation models rely on text similarity to find reference codes, resulting in the inability to guarantee the validity of reference codes, causing model illusions and reducing the accuracy of code generation.
By parsing the pending code text, obtaining the first code metadata of each code element, determining the target query information, and finding the corresponding second code metadata from the pre-built target database, as effective reference code, and guiding the code generation model to generate the target code text.
Improves the accuracy of code text processing, avoids model illusions, and ensures high accuracy of generated target code text.
Smart Images

Figure CN2024122293_05062025_PF_FP_ABST
Abstract
Description
Code text processing method, code supplement method and computing device
[0001] This disclosure claims priority to the Chinese patent application filed with the China Patent Office on November 28, 2023, with application number 2023116137944 and application name “Code text processing method, code supplementation method and computing device”, the entire content of which is incorporated by reference in this disclosure. Technical Field
[0002] The embodiments of the present disclosure relate to the field of deep learning technology, and in particular to a code text processing method, a code supplementation method, and a computing device. Background Art
[0003] With the development of deep learning technology, the code generation model obtained through targeted training is used to generate code to assist developers in writing code text, improve development efficiency and reduce development costs.
[0004] Currently, code generation models rely on reference code text for code text processing. Based on the reference code, the code generation model needs to generate target code text corresponding to the code text to be processed, completing code generation, code supplementation, or code rewriting.
[0005] However, reference code is retrieved through text similarity, which doesn't guarantee its validity. For example, if the code to be processed includes the variable "ABC," and the same variable exists in another project's project file, while the two have high textual similarity, their definitions, logic, and many other aspects differ. Introducing such invalid reference code prevents the code generation model from accurately generating code, resulting in insufficient accuracy in the generated target code and processing.
[0006] Summary of the Invention
[0007] In view of this, embodiments of the present disclosure provide a code text processing method. One or more embodiments of the present disclosure also include a code supplementation method, a code text processing apparatus, a code supplementation apparatus, a computing device, a computer-readable storage medium, and a computer program to address technical deficiencies in the prior art.
[0008] According to a first aspect of an embodiment of the present disclosure, a code text processing method is provided, comprising:
[0009] Get the code text to be processed of the target project;
[0010] Parsing the code text to be processed to obtain first code metadata of each code element, wherein each code element has corresponding element information;
[0011] Determine target query information based on element information of each code element;
[0012] Searching for second code metadata corresponding to the target query information in a target database, wherein the target database records a correspondence between reference query information and reference code metadata, and the reference query information is constructed based on element information of each code element in a project file of the target project;
[0013] Based on the first code metadata and the second code metadata, a target code text is generated using a code generation model, wherein the code generation model is obtained by training a text processing model based on the sample code text.
[0014] A second aspect of an embodiment of the present disclosure provides a code supplement method, which is applied to a cloud-side device and includes:
[0015] The code text to be supplemented of the target item input by the receiving device;
[0016] Parsing the code text to be supplemented to obtain first code metadata of each code element, wherein each code element has corresponding element information;
[0017] Determining target query information according to element information of each code element;
[0018] searching, from a target database, for second code metadata corresponding to the target query information, wherein the target database records a correspondence between reference query information and reference code metadata, the reference query information being constructed based on element information of each code element in a project file of the target project;
[0019] generating supplementary code text using a code generation model based on the first code metadata and the second code metadata, wherein the code generation model is trained on a text processing model based on sample code text;
[0020] The supplementary code text is sent to the end-side device, so that the end-side device supplements the code text to be supplemented by using the supplementary code text.
[0021] According to a third aspect of an embodiment of the present disclosure, a computing device is provided, including: a memory and a processor;
[0022] The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the method provided in any one of the above aspects are implemented.
[0023] According to a fourth aspect of the embodiments of the present disclosure, a computer-readable storage medium is provided, which stores computer-executable instructions. When the computer-executable instructions are executed by a processor, the steps of the method provided in any of the above aspects are implemented.
[0024] According to a fifth aspect of the embodiments of the present disclosure, a computer program is provided, wherein when the computer program is executed in a computer, the computer is caused to execute the steps of the method provided in any one of the above aspects.
[0025] In one embodiment of the present disclosure, a code text to be processed of a target project is obtained; the code text to be processed is parsed to obtain first code metadata for each code element, wherein each code element has corresponding element information; target query information is determined based on the element information of each code element; second code metadata corresponding to the target query information is searched from a target database, wherein the target database records a correspondence between reference query information and reference code metadata, wherein the reference query information is constructed based on the element information of each code element in a project file of the target project; and target code text is generated using a code generation model based on the first code metadata and the second code metadata, wherein the code generation model is trained on a text processing model based on sample code text. By preparating the project file of the target project, reference query information is constructed based on the element information of each code element and stored in the target database in correspondence with the reference code metadata. During code text processing, the code text to be processed is parsed to obtain first code metadata for each code element, and then target query information is determined based on the element information of each code element. This clear code reference relationship is used to query the target database to obtain valid second code metadata, which is used as a reference code to guide the code generation model to generate a highly accurate target code text, thereby improving the accuracy of code text processing. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] FIG1 is a flowchart of a code text processing method provided by one embodiment of the present disclosure;
[0027] FIG2 is a flowchart of constructing a target database in a code text processing method provided by one embodiment of the present disclosure;
[0028] FIG3 is a schematic diagram of an abstract syntax tree in a code text processing method provided by an embodiment of the present disclosure;
[0029] FIG4 is a schematic diagram of a first reference dictionary and a second reference dictionary in a code text processing method provided by an embodiment of the present disclosure;
[0030] FIG5 is a flowchart of updating a code text sequence in a code text processing method provided by one embodiment of the present disclosure;
[0031] FIG6 is a flowchart of real-time analysis of code text in a code text processing method provided by one embodiment of the present disclosure;
[0032] FIG7 is a flowchart of code text parsing in a code text processing method provided by one embodiment of the present disclosure;
[0033] FIG8 is a front-end schematic diagram of a code text processing method provided by one embodiment of the present disclosure;
[0034] FIG9 is a flowchart of a code supplement method provided by one embodiment of the present disclosure;
[0035] FIG10 is a flowchart of a method for processing code text applied to an integrated development environment according to an embodiment of the present disclosure;
[0036] FIG11 is a schematic structural diagram of a code text processing device provided by an embodiment of the present disclosure;
[0037] FIG12 is a schematic structural diagram of a code supplement device provided by an embodiment of the present disclosure;
[0038] FIG13 is a structural block diagram of a computing device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION
[0039] The following description sets forth many specific details to facilitate a full understanding of the present disclosure. However, the present disclosure can be implemented in many other ways than those described herein, and those skilled in the art can make similar generalizations without violating the scope of the present disclosure. Therefore, the present disclosure is not limited to the specific implementations disclosed below.
[0040] The terms used in one or more embodiments of the present disclosure are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of the present disclosure. The singular forms "a", "the", and "the" used in one or more embodiments of the present disclosure and the appended claims are also intended to include plural forms, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of the present disclosure refers to and includes any or all possible combinations of one or more associated listed items.
[0041] It should be understood that although the terms first, second, etc. may be used to describe various information in one or more embodiments of the present disclosure, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of one or more embodiments of the present disclosure, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".
[0042] In addition, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in one or more embodiments of the present disclosure are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0043] In one or more embodiments of the present disclosure, a large model refers to a deep learning model with large-scale model parameters, which typically contains hundreds of millions, tens of billions, hundreds of billions, trillions, or even more than ten trillion model parameters. A large model can also be called a cornerstone model / foundation model (Foundation Model), which is pre-trained by using large-scale unlabeled corpus to produce a pre-trained model with more than 100 million parameters. This model can adapt to a wide range of downstream tasks, and the model has good generalization ability, such as a large-scale language model (LLM) and a multi-modal pre-training model.
[0044] When large models are used in practice, only a small number of samples are needed to fine-tune the pre-trained model and it can be applied to different tasks. Large models can be widely used in natural language processing (NLP), computer vision and other fields. Specifically, they can be applied to computer vision tasks such as visual question answering (VQA), image caption (IC), and image generation, as well as natural language processing tasks such as text-based sentiment classification, text summary generation, and machine translation. The main application scenarios of large models include digital assistants, intelligent robots, search, online education, office software, e-commerce, and intelligent design.
[0045] First, the terms involved in one or more embodiments of the present disclosure are explained.
[0046] Code generation: This technology uses deep learning technology to enable text processing models to learn from massive amounts of sample code text and generate targeted code text for developers.
[0047] Model hallucination: The problem of confusing reference knowledge from different sources due to limited model performance, resulting in low-accuracy code text generation, which can be compared to humans' misunderstanding of homographs.
[0048] Abstract Syntax-code Tree (AST): It is a tree representation of the grammatical structure of a code file. Each node on the tree represents a code element in the code file, representing the grammatical structure of the entire code file (text). For example: module 1 - object A - function a - variable x, where modules, objects, functions, and variables are code elements. The abstract syntax tree represents the above hierarchical structure. The abstract syntax tree does not depend on the specific code language. The abstract syntax tree records the location information of the code element corresponding to each node.
[0049] Symbol table: The symbol table stores various symbols in the source code text and their related information, such as variables, labels, functions, types, etc., and records the location of the symbols in the entire program.
[0050] Java package: Java package is a namespace mechanism that can be used to organize related classes and interfaces to avoid class name conflicts.
[0051] Java Import Type: In Java, if you want to use a class or interface from another package, you need to import it using the import statement.
[0052] Java class definition: A class is the basic building block in Java that defines the data and behavior of an object. A class definition includes the class name, member variables, constructor, and various access control modifiers.
[0053] Java inheritance class: Java supports single inheritance, that is, a class can only directly inherit the properties and behaviors of another class. This is done to improve code reusability.
[0054] Java implements interface: Interface is another way to provide functional features in Java. A class can declare that it implements an interface by using the implements keyword.
[0055] Java class attributes: Class attributes are also called fields or member variables. They represent the state of the class object and can be shared among all objects of the class.
[0056] Java class methods: Class methods are methods defined at the class level and can be called by any object of the class. They are used to define the operations or behaviors that class objects can perform.
[0057] Python modules: Python modules are files that contain related Python objects (such as functions, classes, and variables). These modules can be imported and used in programs using the import statement.
[0058] Python Import Types: In Python, you can import a specific module or import a specific object (such as a function, class, or variable) from a module. This is done using the import or from...import statements.
[0059] Python function definition: A function is a group of statements encapsulated together and can be reused as many times as needed. In Python, you can use the def keyword to define a function.
[0060] Python global variables: Global variables are variables defined outside a function and can be accessed by all functions in the entire program.
[0061] Python class: A Python class is a data type that can have attributes and methods. It is the basis of object-oriented programming.
[0062] Python class definition: A class definition includes the class name, attributes, methods, and special methods. A class definition begins with the class keyword.
[0063] Python class attributes: Class attributes are variables that belong to a class rather than a specific object. Class attributes are shared among all objects.
[0064] Python class method: A class method is a method that belongs to a class rather than a specific object. It is mainly used to handle class-level operations.
[0065] Python inheritance class: In Python, a class can inherit from an existing class and get all its attributes and methods. This can be done by using the extends keyword in the class definition.
[0066] JavaScript modules: In JavaScript, a "module" is a self-contained collection of functionality that can be exported for use in other files using the exports keyword.
[0067] JavaScript import type: You can use the import keyword to import the exported module content.
[0068] JavaScript function definition: In JavaScript, a function is a special type of variable declared by the keyword function.
[0069] JavaScript Global Variables / Constants: Global variables and constants are variables and constants defined outside any function and can be accessed throughout the program.
[0070] JavaScript class: A JavaScript class is a template for creating objects. It describes the properties and behaviors that an object should have.
[0071] JavaScript class definition: A class definition includes the class name, properties, methods, and special methods. A class definition begins with the class keyword.
[0072] JavaScript class properties: Class properties are variables that belong to a class rather than a specific object. Class properties are shared among all objects.
[0073] JavaScript class method: A class method is a method that belongs to a class rather than a specific object. It is mainly used to handle class-level operations.
[0074] JavaScript inheritance class: In JavaScript, a class can inherit from an existing class and get all its properties and methods. This can be done by using the extends keyword in the class definition.
[0075] JavaScript object: A JavaScript object is a collection of related values (usually variables and functions) whose contents can be accessed through dot notation or square bracket notation.
[0076] JavaScript object definition: The object definition is enclosed in curly braces {}, and inside is a series of key-value pairs with property names as keys and property values as values.
[0077] JavaScript object properties: Object properties are specific object states that can be accessed using dot notation or square bracket notation.
[0078] JavaScript object methods: Object methods are specified object behaviors that can be called using dot notation or square bracket notation.
[0079] An integrated development environment (IDE), also known as an integrated development platform, is a software application suite that combines programming functionality with other development tools. It typically includes a code editor, compiler / interpreter, code parser, debugger, build automation system, version control system, and deployment tools.
[0080] Code metadata (source code): is information used to characterize the organization, data domains, and relationships of data in code text, and is used to support functions such as indicating storage location, historical data, resource search, and file records.
[0081] Abstract Syntax Tree Parser (AST Parser): A special code parser whose goal is to extract code elements from the code text of a code file and construct an Abstract Syntax Tree (AST). Through the AST parser, the code file will be converted into a data structure that is easy to process and analyze, making it easier for developers to understand the code.
[0082] Deep Self-Attention Model (Transformer Model): A deep learning architecture based on the attention mechanism for processing sequential data such as natural language.
[0083] Bidirectional Encoder Representations from Transformers (BERT): A specialized Transformer model trained using a bidirectional Transformer encoder and large-scale unlabeled text data. BERT's outstanding performance has made it a standard baseline for many NLP tasks.
[0084] Large Language Model (LLM): A deep learning model trained on a large corpus for natural language processing tasks. These models typically consist of a multi-layer neural network. Their input is a text sequence, which they generate. Their output is the generated text result of a specific natural language processing task performed on that text sequence. Pre-training means that the model has been trained and learned to process large amounts of language data before a specific task. By pre-training the models, they can capture more complex linguistic and semantic rules, enabling them to excel in various natural language processing tasks and reducing the large-scale data requirements for specific tasks.
[0085] Resource files: A collection of files containing various non-executable data or code required for program execution. These can be images, audio, video, or other types of multimedia data, or text, configuration information, or other types of non-media data. Resource files facilitate loading and accessing required external data while avoiding hard-coding large amounts of data into program code, reducing program size and enhancing portability. Typically, a program packages various external resource data into one or more resource files, and loads and decompresses this data on demand during execution.
[0086] Binary file: A computer file that contains data or program instructions encoded in binary form. These files are often used to store graphics, audio, video, and other non-text data and cannot be viewed or edited with a simple text editor.
[0087] Log file: A file or collection of files that records system operation events, primarily used to track and monitor system performance. Log files can include system startup and shutdown events, error messages, warnings, and debugging information. Log files are typically in text format and can be viewed and analyzed using a simple text editor.
[0088] Path file: A file that records file or directory paths. It is often used to share file path information between computer systems or store the paths of frequently accessed files. Path files can have different formats, such as CSV, XML, or plain text, depending on their purpose and requirements.
[0089] JsonRpc protocol: a stateless and lightweight Remote Procedure Call (RPC) protocol that uses the JSON format for data exchange. It is an implementation of a service-oriented architecture (SOA) that supports cross-language and cross-platform interactions between applications. The core concept of JsonRpc is the request / response model, where the client sends a request to the server, and the server processes the request and returns a response to the client. Both requests and responses are represented in JSON format, containing fields such as method name, parameters, and id. The advantages of JsonRpc are simplicity, efficiency, and ease of implementation, and it is applicable to multiple programming languages and platforms. At the same time, JsonRpc also has good scalability, and custom methods and parameters can be easily added.
[0090] WebSocket communication connection: A communication protocol that establishes a full-duplex, persistent connection between a client and a server, enabling real-time, two-way communication. It allows the client and server to send data to each other without having to wait for a response from the other party.
[0091] The n-gram approach is a language modeling method that can be used to calculate the probability distribution of a given text sequence. The key feature of the n-gram model is that it assumes that each word is only related to the n-1 words that precede it, without considering subsequent words. Therefore, it ignores long-term correlations between texts.
[0092] At present, reference codes are obtained by searching through text similarity. This method cannot guarantee the validity of the reference codes. Introducing such invalid reference codes will cause model hallucinations in the code generation model, and will not be able to guide the code generation model to perform accurate code generation, resulting in insufficient accuracy of the generated target code text and insufficient accuracy of code text processing.
[0093] In response to the above problems, the present disclosure provides a code text processing method. The present disclosure also involves a code supplement method, a code text processing device, a code supplement device, a computing device, a computer-readable storage medium and a computer program, which are described in detail one by one in the following embodiments.
[0094] Referring to FIG1 , FIG1 shows a flowchart of a code text processing method provided by an embodiment of the present disclosure, which includes the following specific steps:
[0095] Step 102: Obtain the code text to be processed of the target project.
[0096] The disclosed embodiments are applied to applications, websites, or mini-programs that have code text processing capabilities. They can be the client or server of the application, website, or mini-program. For example, an integrated development environment calls the method through an interface. Another example is an application that directly provides code development and implements the method on the server of the application. Applications to code text processing task scenarios include, but are not limited to, code text generation task scenarios, code text supplementation task scenarios, and code text rewriting task scenarios.
[0097] The target project is a development project where the code text needs to be processed, such as an application, interface, website, plug-in, or database management tool.
[0098] The code text to be processed in the target project is the code text in the target project that needs to be processed. It can be the code text to be generated, the code text to be supplemented, or the code text to be rewritten, without limitation. The code text to be processed is written in a specific code language, including but not limited to: Java, Python, JavaScript, and TypeScript. For example, the code text to be supplemented written in JavaScript is: "var demo = {name: "John", age: 25, add".
[0099] Obtaining the target project's pending code text can be receiving the target project's pending code text uploaded by the front end, for example, a user uploads a plug-in code file on the front end, and the code file includes the code text to be rewritten. It can also be identifying the target project's pending code text input by the front end, for example, identifying the code text to be supplemented that the user is entering during the plug-in development process on the front-end interface of the integrated development environment. It can also be obtaining the target project's pending code text from a database, for example, using an index to obtain a plug-in's code file from a code database, and the code file includes the code text to be rewritten. This is not limited here.
[0100] For example, a code generation interface is deployed on a JavaScript integrated development environment. After the user starts the code generation interface on the integrated development environment, he begins to write the code text of a plug-in. In the process of writing a JavaScript object, the code generation interface recognizes that the current code text to be supplemented is: "var demo = {name: "John", age: 25, add".
[0101] Obtaining the target project's code text to be processed provides a code text basis for subsequent parsing to obtain the first code metadata.
[0102] Step 104: Parse the code text to be processed to obtain first code metadata of each code element, wherein each code element has corresponding element information.
[0103] Code elements are the building blocks of code text. Code text consists of code elements at different levels. These code elements at different levels form specific structural data that represents the grammatical structure of the code text. This specific structural data includes, but is not limited to, an abstract syntax tree and a symbol table. Code elements have different definitions and different grammatical structures depending on the code language: for the Java language, the Java class parent code element includes: Java package definition, Java import type, Java class definition, Java inherited class, Java implementation interface, Java class attribute and Java class method and other child code elements; for the Python language, the Python module parent code element includes: Python import type, Python function definition and Python global variable and other child code elements; the Python class parent code element includes: Python import type, Python class definition, Python class attribute, Python class method and Python inherited class and other child code elements; for the JavaScript language, the JavaScript module parent code element includes: JavaScript import type, JavaScript function definition and JavaScript global variable / constant and other child code elements; the JavaScript class parent code element includes: JavaScript import type, JavaScript class definition, JavaScript class definition, JavaScript class attribute, JavaScript class method and JavaScript inherited class and other child code elements; the JavaScript object parent code element includes: JavaScript import type, JavaScript object definition, JavaScript object attribute and JavaScript object method and other child code elements. For the convenience of expression, the present disclosure adopts three levels of code elements: object, function and variable for expression.
[0104] Code metadata is information used to characterize the organization, data domains, and relationships of data in the code text, and is used to support functions such as indicating storage location, historical data, resource search, and file records. The code metadata of a code element is information that describes the organization, data domains, and relationships of a code element. The first code metadata is the code metadata of each code element in the code text to be processed. For example, the code text to be supplemented is: "var demo = {name: "John", age: 25, add", which includes an object element ("demo") and a function element ("add"). The first code metadata of the object element is: "name": "demo", "signature": "var demo", "full_name": "project.path.demo", "fields", and the first code metadata of the function element is: "methods": {"add": {"method_name": "add", "signature": "add: function()"}}.
[0105] The element information of a code element refers to its attribute information, including but not limited to: the name of the code element, its location, and its type. For example, for a code element ("demo"), the name is demo, its location is [0, 1], its starting row and column numbers are [0, 8], its scope is within a method body, within a class, in a method definition, or in a class definition, and its type is a JavaScript object.
[0106] The code text to be processed is parsed to obtain the first code metadata of each code element, specifically by performing grammatical structure analysis on the code text to be processed to obtain the first code metadata of each code element. This step is specifically implemented by a code parser, for example, a Parser parser for Java and JavaScript languages, or a Pygments parser for Python. Syntax structure analysis can be implemented through specific structural data, including but not limited to: an abstract syntax tree and a symbol table. The implementation method through the abstract syntax tree can be seen in the following Figure 3.
[0107] Exemplarily, using the Parser, the grammatical structure of the supplementary code text "var demo={name:"John", age:25, add" is parsed through the structural data of the abstract syntax tree, and the first code metadata of the object element ("demo") is obtained as: "methods":{"add":{"method_name":"add","signature":"add:function()"}}, and the first code metadata of the function element ("add") is: "methods":{"add":{"method_name":"add","signature":"add:function()"}}.
[0108] The code text to be processed is parsed to obtain the first code metadata of each code element. The code text to be processed is parsed based on its own structure to obtain the first code metadata of each code element, which lays the foundation for the code elements for the subsequent determination of the target query information and provides code text support for the subsequent generation of the target code text.
[0109] Step 106: Determine target query information based on the element information of each code element.
[0110] The target query information is identifiable query information used to query code metadata. The query information is constructed based on the structural information between each code element, and the structural information between each code element is the hierarchical structure information between each code element, for example, a hierarchical structure such as module-object-function-variable. For the target query information, for example, for the function element of the code element ("add"), there are a large number of objects including this function in the code file of the entire target project. Therefore, considering the parent-child code element relationship with the object element ("demo"), the target query information is determined to be demo+add. This form of "the name of the parent type code element + the name of the child type code element" is identifiable in the code file of the entire target project. In the disclosed embodiments, since the storage path of the code element is identifiable, the storage path of the code element can be determined as query information. For example, for the Java language, the storage path of the Java package element name + the Java class element name is determined as query information. For another example, for the Python language, the storage path of the Python module path + the name of the Python module element, or the storage path of the Python module path + the name of the Python module element + the name of the Python class element is determined as query information. For another example, for the JavaScript language, the storage path of the JavaScript module path + the name of the JavaScript module element, or the storage path of the JavaScript module path + the name of the JavaScript module element + the name of the JavaScript class element, or the storage path of the JavaScript module path + the name of the JavaScript module element + the name of the JavaScript object element is determined as query information. The specific format is " / workspace / project / path$demo".
[0111] According to the element information of each code element, the target query information is determined, specifically in the following manner: according to the element information of the target code element in each code element, the target query information is determined. The target code element is a code unit in each code element of the code text to be processed that is used to search for the second code metadata, and there is an element structure between the target code elements. Further, according to the element information of the target code element in each code element, the structural information between the target code elements is determined, and according to the structural information between the target code elements, the target query information is determined. For example, the code text to be processed includes 3 code elements: object element ("demo"), function element ("add") and function element ("mul"), and the code element to be supplemented is function element ("mul"), and the target code elements are determined to be object element ("demo") and function element ("mul"), and the structural information is: object element ("demo") + function element ("mul"), and function element ("add") is not a target code element. The target query information can be determined by querying an information table pre-recorded with the query information based on the element information of the target code element. For example, a pre-recorded reference dictionary records the correspondence between element information and query information, and the reference dictionary is queried based on the element information of the target code element to obtain the corresponding target query information. The target query information can also be determined by directly determining the element information of the target code element as the target query information, which is not limited here. The reference dictionary can be described in FIG4 below.
[0112] Exemplarily, according to the element names of the target code elements (object elements ("demo") and function elements ("add")) in each code element, a reference dictionary that pre-records the correspondence between element information and query information is queried to obtain the target query information: module path + module name + demo + add.
[0113] The target query information is determined based on the element information of each code element, and a clear code reference relationship is obtained, which provides a query basis for subsequent queries to obtain the second code metadata.
[0114] Step 108: Search the target database for the second code metadata corresponding to the target query information, wherein the target database records the correspondence between the reference query information and the reference code metadata, and the reference query information is constructed based on the element information of each code element in the project file of the target project.
[0115] The target database is a pre-built database of code metadata for each code element in the project file of the target project. The target database also records the correspondence between reference query information and reference code metadata, serving as a basis for querying the code metadata. The target database can be constructed after or before the startup code text is processed. The reference query information and reference code metadata are provided as key-value pairs for code metadata queries.
[0116] The reference query information is identifiable query information recorded in the target database for searching code metadata. For details, see the content of the target query information in step 106. The reference code metadata is the code metadata of each code element in the project file of the target project. For details, see the content of the first code metadata in step 104. The second code metadata is the code metadata recorded in the target database corresponding to the target query information. For details, see the content of the first code metadata in step 104.
[0117] It should be noted that the target database is obtained by pre-parsing the project files of the target project, obtaining the reference code metadata of each code element, and then storing the reference code metadata based on the reference query information constructed according to the element information of each code element. The project files of the target project involved in the parsing can be all the project files in the target project or part of the project files in the target project. Correspondingly, the reference code metadata can be the code metadata of all the code elements or part of the code elements, and can be dynamically updated according to the actual query scenario. For example, the reference code metadata is written to a queue and dynamically updated according to the actual query scenario.
[0118] Searching the target database for the second code metadata corresponding to the target query information, specifically by searching for the second code metadata corresponding to the target query information based on a correspondence between the reference query information and the reference code metadata recorded in the target database. Furthermore, searching for the second code metadata corresponding to the target query information using the target query information as a key, using a key-value pair constructed based on the correspondence between the reference query information and the reference code metadata recorded in the target database.
[0119] Exemplarily, based on the key-value pairs of the correspondence between the reference query information and the reference code metadata recorded in the target database, the target query information "module path + module name + demo + add" is used as the key to find the second code metadata of the corresponding value: {"name":"demo","signature":"vardemo","full_name":"project.path.demo","fields":{"name":{"field_name":"name","field_value":"John","signature":"name:'John'"},"age":{"field_name":"age","field_value":"25","signature":"age:25"}},"methods":{"add":{"method_name":"add","signature":"add:function()"}}}.
[0120] The target database is searched for the second code metadata corresponding to the target query information. The target database records the correspondence between reference query information and reference code metadata. The reference query information is constructed based on the element information of each code element in the project file of the target project. Using the correspondence between the query information and the code metadata determined by the element information, this clear code reference relationship is used to query the target database and obtain valid second code metadata, providing reference code for the subsequent code generation model to generate the target code file.
[0121] Step 110: Based on the first code metadata and the second code metadata, a target code text is generated using a code generation model, wherein the code generation model is obtained by training a text processing model based on the sample code text.
[0122] The code generation model is a deep learning model capable of generating code text. It is trained on a text processing model based on sample code text. This model is a deep learning model fine-tuned for code text processing tasks. The code generation model can be a Transformer model, a BERT model, or a large language model, without limitation.
[0123] The target code text is the code text obtained by performing code text processing on the code text to be processed. It is the processing result of the code text processing. It can be the generated code text generated by the executed code text, the supplementary code text supplemented by the executed code text, or the rewritten code text rewritten by the executed code text, which is not limited here. The target code text is written in a specific code language and can be consistent with the code text to be processed, or inconsistent with the code text to be processed. For example, the code text to be supplemented written in JavaScript: "var demo = {name: "John", age: 25, add", executing the code supplementation processing, the obtained supplementary code text written in JavaScript is: "var demo = {name: "John", age: 25, add: function () {console.log (this.name);},}; demo.add", executing the code rewriting processing, the rewritten code text written in Python is: "class demo: def __init__ (self): self.name = "John" self.age = 25 def add (self): print (self.name) demo = demo () demo.add ().
[0124] Based on the first code metadata and the second code metadata, a target code text is generated using a code generation model, specifically by performing code text processing on the first code metadata based on the second code metadata using the code generation model to generate the target code text. Based on the second code metadata, code text processing is performed on the first code metadata, specifically by using the second code metadata as a reference code and performing code text processing on the first code metadata to generate the target code text.
[0125] Optionally, before step 110, the following specific steps are further included:
[0126] Preprocessing the code text to be processed to obtain context information, wherein the preprocessing is context information extraction;
[0127] Correspondingly, step 110 includes the following specific steps:
[0128] The context information, the first code metadata and the second code metadata are input into a code generation model, and based on the context information and the second code metadata, code text processing is performed on the first code metadata to generate a target code text.
[0129] The context information is the context code text in the code text to be processed, which is information such as writing habits, variable definitions, parameter values, etc. in the code text to be processed.
[0130] For example, the supplementary code text is: "var demo={name:"John",age:25,add” is preprocessed to obtain context information such as writing habits, variable definitions, parameter values, etc., and the context information, the first code metadata "methods":{"add":{"method_name":"add","signature":"add:function()"}} and the second code metadata {"name":"demo","signature":"vardemo","full_name":"project.path.demo","fields":{"name":{"field_name":"name","field_value":"John","signature":"name:'John'"},"age":{"field_name":"age","field_value":"25","signature":"age:25"}},"methods":{"add":{"method_name":"add","signature":"add:function()"}}} are input into a large language model with code text processing function. The large language model is used to perform code supplementation processing on the first code metadata based on the context information and the second code metadata to generate supplementary code text: "var demo={name:"John",age:25,add:function(){console.log(this.name);},};demo.add".
[0131] In the embodiment of the present disclosure, a code text to be processed of a target project is obtained; the code text to be processed is parsed to obtain first code metadata of each code element, wherein each code element has corresponding element information; target query information is determined based on the element information of each code element; second code metadata corresponding to the target query information is searched from a target database, wherein the target database records a correspondence between reference query information and reference code metadata, wherein the reference query information is constructed based on the element information of each code element in the project file of the target project; based on the first code metadata and the second code metadata, a target code text is generated using a code generation model, wherein the code generation model is trained on a text processing model based on sample code text. By pre-parsing the project file of the target project, reference query information is constructed based on the element information of each code element and stored in the target database corresponding to the reference code metadata. During the code text processing process, the code text to be processed is parsed to obtain first code metadata of each code element, and then target query information is determined based on the element information of each code element. This clear code reference relationship is used to query the target database to obtain valid second code metadata, which is used as a reference code to guide the code generation model to generate a highly accurate target code text, thereby improving the accuracy of code text processing.
[0132] In an optional embodiment of the present disclosure, before step 108, the following specific steps are further included:
[0133] Get the project file of the target project;
[0134] Parse the project files to obtain reference code metadata for each code element;
[0135] Construct reference query information based on element information of each code element;
[0136] Based on the correspondence between reference query information and reference code metadata, a target database is constructed.
[0137] The project files of the target project are project code files, which refer to code files related to the development project. They include but are not limited to: source code files, resource files, binary files, log files, and path files.
[0138] Parse the project file to obtain reference code metadata for each code element. This is done by parsing the project file's code text for syntax and structure to obtain reference code metadata for each code element. This step is typically performed using a code parser, such as Parser for Java and JavaScript, or Pygments for Python. Syntax parsing can be achieved using specific structured data, including but not limited to abstract syntax trees and symbol tables.
[0139] According to the element information of each code element, reference query information is constructed. Specifically, the structural information between each code element is determined according to the element information of each code element, and the reference query information is constructed based on the structural information between each code element.
[0140] Based on the correspondence between the reference query information and the reference code metadata, a target database is constructed. Specifically, a key-value pair is constructed based on the correspondence between the reference query information and the reference code metadata, and the reference code metadata is stored based on the constructed key-value pair to obtain the target database.
[0141] Optionally, after obtaining the project file of the target project, the following specific steps are further included:
[0142] Filter out invalid project files in project files.
[0143] Invalid project files are project files that cannot be parsed or used for code text processing, such as files that are too long, binary files, log files, and path files.
[0144] It should be noted that the embodiment of the present disclosure can be automatically executed after determining that code text processing is required during code development. Correspondingly, obtaining the project file of the target project is specifically: obtaining the project file of the target project in response to the code text processing request. For example, after starting the target project, the developer determines to start the code text processing service process, and establishes a websocket communication connection with the remote code text processing server based on the jsonrpc protocol. After the connection is completed, the plug-in end sends an initialized code text processing request to the local code text processing process. After receiving the initialized code text processing request, the local code text processing process begins to obtain the project file of the target project in an indexed manner.
[0145] FIG2 shows a flowchart of constructing a target database in a code text processing method provided by one embodiment of the present disclosure, as shown in FIG2 :
[0146] Obtain the project file of the target project; filter invalid project files; call the parser of the corresponding code language to parse the project file and obtain the reference code metadata of each code element; construct reference query information based on the element information of each code element, and construct the target database based on the correspondence between the reference query information and the reference code metadata.
[0147] Exemplarily, in response to a code text processing request, a project file of a certain plug-in is obtained, overly long files, binary files, log files and path files in the project file are filtered, and the project file is parsed using the Parser through the structural data of the abstract syntax tree to obtain the reference code metadata of each code element (JavaScript module, JavaScript import type, JavaScript function definition... JavaScript object method), and according to the element name of each code element, the structural information between each code element is determined: (JavaScript module: JavaScript import type, JavaScript function definition and JavaScript global variable / constant); (JavaScript class: JavaScript import type, JavaScript class definition, JavaScript class definition, JavaScriptScrip t class properties, JavaScript class methods and JavaScript inherited classes); (JavaScript objects: JavaScript import types, JavaScript object definitions, JavaScript object properties and JavaScript object methods), based on the structural information between each code element, construct reference query information: (JavaScript module path + JavaScript module element name); (JavaScript module path + JavaScript module element name + JavaScript class element name); (JavaScript module path + JavaScript module element name + JavaScript object element name), based on the correspondence between the reference query information and the reference code metadata, construct key-value pairs, based on the constructed key-value pairs, store the reference code metadata, and obtain the target database.
[0148] Obtain the project file for the target project; parse the project file to obtain reference code metadata for each code element; construct reference query information based on the element information of each code element; and construct the target database based on the correspondence between the reference query information and the reference code metadata. By parsing the project file for the target project and utilizing the correspondence between the query information and code metadata determined by the element information, this clear code reference relationship is constructed to construct the target database, laying the foundation for determining the corresponding code metadata in subsequent code text processing.
[0149] In an optional embodiment of the present disclosure, parsing a project file to obtain reference code metadata of each code element includes the following specific steps:
[0150] Perform syntax structure analysis on the code in the project file to obtain the project syntax tree of the project file;
[0151] Parse the project syntax tree to obtain reference code metadata for each code element.
[0152] The project syntax tree of a project file is an abstract syntax tree (AST) that describes the syntax structure of the source code within the project file. This tree structure describes the dependencies between code elements within the project file and their locations within the project file.
[0153] FIG3 shows a schematic diagram of an abstract syntax tree in a code text processing method provided by an embodiment of the present disclosure, as shown in FIG3 :
[0154] In the project file, the code text is "var demo = {name: "John", age: 25, add: function() {console.log(this.name);},}; demo.add." The abstract syntax tree is constructed as follows: Node 1 connects Node 2 and Node 3, Node 2 connects Node 4 and Node 5, and Node 3 connects Node 6 and Node 7. By parsing the abstract syntax tree, the following abstract syntax tree parsing results are obtained: Object element (Node 1) - demo; Function element (Node 2) - add; Function element (Node 3) - mul; Variable element (Node 4) - a; Variable element (Node 5) - b; Variable element (Node 6) - a; Variable element (Node 7) - b, and reference code metadata is obtained for each code element.
[0155] Exemplarily, the Parser is used to perform grammatical structure analysis on the code of a plug-in's project file to obtain a project syntax tree of the project file, and the project syntax tree is parsed to obtain reference code metadata of each code element (JavaScript module, JavaScript import type, JavaScript function definition...JavaScript object method).
[0156] Perform grammatical structure analysis on the code in the project file to obtain the project syntax tree of the project file; parse the project syntax tree to obtain reference code metadata for each code element. Using the abstract syntax tree, we achieve highly accurate grammatical structure analysis of the code in the project file and obtain highly accurate reference code metadata for each code element. This provides the data foundation for the subsequent correspondence between query information and code metadata determined using element information, creating a clear code reference relationship.
[0157] In an optional embodiment of the present disclosure, constructing reference query information according to element information of each code element includes the following specific steps:
[0158] Determining structural information between code elements based on element information of each code element;
[0159] Based on the structural information between each code element, a storage path of the reference code metadata of each code element is constructed as reference query information.
[0160] The structural information between the code elements is the grammatical structure information between the code elements, and is the structural information of the hierarchical grammatical structure between the code elements, which is expressed as a hierarchical structure of parent type code element + child type code element. Different grammatical structure information exists for different code languages. For example, for the Java language, the parent code element "Java class" includes child code elements such as Java package definition, Java import type, Java class definition, Java inherited class, Java implementation interface, Java class attribute, and Java class method; for the Python language, the parent code element "Python module" includes child code elements such as Python import type, Python function definition, and Python global variable; the parent code element "Python class" includes child code elements such as Python import type, Python class definition, Python class attribute, Python class method, and Python inherited class; for the JavaScript language, the parent code element "JavaScript module" includes child code elements such as JavaScript import type, JavaScript function definition, and JavaScript global variable / constant; the parent code element "JavaScript class" includes child code elements such as JavaScript import type, JavaScript class definition, JavaScript class definition, JavaScript class attribute, JavaScript class method, and JavaScript inherited class; the parent code element "JavaScript object" includes child code elements such as JavaScript import type, JavaScript object definition, JavaScript object attribute, and JavaScript object method.
[0161] The storage path of the reference code metadata is the storage path of the parent type code element when the project file is imported. It is obtained when the project file of the target project is obtained. For example, for the module code element, there is a corresponding module path. On this basis, the structural information between the code elements is spliced to obtain the reference query information. For example, the module path + module name + object name + function name + variable name is determined as the reference query information. For the query information of two variable elements, even if the module name, object name, function name and variable name are the same, due to the different storage paths, the constructed reference query information is also different and identifiable.
[0162] Based on the element information of each code element, the structural information between the code elements is determined. Specifically, based on the element information of each code element, each sub-type code element is assigned to a corresponding parent-type code element to obtain the structural information between the code elements. Furthermore, based on the element information and code language of each code element, each sub-type code element is assigned to a corresponding parent-type code element to obtain the structural information between the code elements.
[0163] Exemplarily, according to the element name and code language (JavaScript) of each code element, each sub-type code element is attributed to the corresponding parent-type code element, and the structural information between each code element is obtained: (JavaScript module: JavaScript import type, JavaScript function definition and JavaScript global variable / constant); (JavaScript class: JavaScript import type, JavaScript class definition, JavaScript class definition, JavaScript class property, JavaScript class method and JavaScript inheritance class); (JavaScript object: JavaScript import type, JavaScript object definition, JavaScript object property and JavaScript object method), the storage path of the parent-type code element is spliced, the structural information between each code element is obtained, and the reference query information is obtained: (JavaScript module path + name of JavaScript module element); (JavaScript module path + name of JavaScript module element + name of JavaScript class element); (JavaScript module path + name of JavaScript module element + name of JavaScript object element).
[0164] Based on the element information of each code element, the structural information between each code element is determined. Based on this structural information, the storage path of the reference code metadata of each code element is constructed as reference query information. Based on the structural information and storage path between each code element, identifiable reference query information is constructed, and a highly accurate correspondence between query information and code metadata is obtained. This clear code reference relationship lays the foundation for building the target database.
[0165] In an optional embodiment of the present disclosure, before step 104, the following specific steps are further included:
[0166] Obtain target position information of the code element to be processed in the code text to be processed;
[0167] Correspondingly, step 104 includes the following specific steps:
[0168] Perform grammatical structure analysis on the code in the code text to be processed to obtain a basic grammar tree of the code text to be processed;
[0169] Parse the basic syntax tree to obtain the first code metadata of each code element;
[0170] Correspondingly, step 106 includes the following specific steps:
[0171] Based on the target position information, each code element recorded in the basic syntax tree is traversed to determine the target code element;
[0172] Target query information is determined based on element information of the target code element.
[0173] The code element to be processed is the code element that requires code text processing. The target location information of the code element to be processed is the location information of the code element to be processed within the code text to be processed, including but not limited to: line number, column number, and range. The location information of the code element to be processed can be selected location information, such as the current cursor position, or the location information of the currently written code element, without limitation.
[0174] The base syntax tree of the code text to be processed is the abstract syntax tree of the code text to be processed. It is used to describe the specific structural data of the grammatical structure of the code text to be processed. The tree structure describes the dependencies between the various code elements in the code text to be processed and their location information in the code text to be processed.
[0175] The target code element is a code unit used to search for the second code metadata in each code element of the code text to be processed, and is obtained by traversing the basic syntax tree using the target position information. For example, taking the abstract syntax tree in Figure 3 as an example, the target position information is the starting row and column number [0, 1]; the ending row and column number [0, 5]; the range: within the method body, the code element corresponding to the target position information is determined to be node 5 from the position information of the code elements recorded in the basic syntax tree. Node 5 is connected to node 2, and node 2 is connected to node 1. The code elements corresponding to nodes 1, 2, and 5 are determined to be the target code elements, that is, the object element ("demo"), the function element ("add"), and the variable element ("b") are the target code elements.
[0176] Based on the target position information, each code element recorded in the basic syntax tree is traversed to determine the target code element. The specific method is: based on the target position information, the position information of each code element recorded in the basic syntax tree is traversed to determine the target code element.
[0177] Based on the element information of the target code element, the target query information is determined. Specifically, based on the element information of the target code element, an information table pre-recorded with query information is searched to obtain the target query information.
[0178] Exemplarily, the target position information of the code element ("add") to be supplemented in the code text to be supplemented: "var demo = {name: "John", age: 25, add" is obtained: the starting row and column number [0, 1]; the ending row and column number [0, 2]; the scope: inside the method body. The Parser is used to parse the syntax structure of the code text to be supplemented "var demo = {name: "John", age: 25, add", and the first code metadata of the object element ("demo") is obtained as follows: "name": "demo", "signature": "var demo", "full_name": "project.path.demo", "fields", and the first code metadata of the function element ("add") is obtained as follows: "methods": {"add": {"method_name": "add", "signature": "add: function()"}}. Based on the target position information, the position information of each code element recorded in the basic syntax tree is traversed. The target code elements are determined to be object elements ("demo") and function elements ("add"). Based on the element names of the target code elements, a reference dictionary pre-recorded with query information is queried to obtain target query information: module path + module name + demo + add.
[0179] Obtain target location information of the code elements to be processed in the code text to be processed; perform grammatical structure analysis on the code in the code text to be processed to obtain a basic syntax tree for the code text to be processed; parse the basic syntax tree to obtain first code metadata for each code element; based on the target location information, traverse each code element recorded in the basic syntax tree to determine the target code element; and determine target query information based on the element information of the target code element. Using the target location information, traverse each code element in the parsed basic syntax tree to determine the target code element, and then accurately determine the corresponding target query information, providing an accurate query basis for subsequent queries to obtain second code metadata.
[0180] In an optional embodiment of the present disclosure, before determining the target query information based on the element information of the target code element, the following specific steps are further included:
[0181] Based on the element identifier of the target code element, query the first reference dictionary to obtain the element information of the target code element, wherein the first reference dictionary is updated in the process of traversing the basic syntax tree, and the first reference dictionary records the correspondence between the element identifier and the element information of each code element;
[0182] Correspondingly, determining target query information based on target element information of the target code element includes the following specific steps:
[0183] Based on the element information of the target code element, the second reference dictionary is queried to obtain the storage path of the target code element as target query information, wherein the second reference dictionary records the correspondence between the element information of each code element and the storage path.
[0184] The element identifier is the node identifier of the code element in the basic syntax tree. For example, if the target location information determines that the node is node 5, the element identifier is node 5.
[0185] The first reference dictionary is a pre-built table that records the correspondence between the element identifiers and element information of each code element. It is a dynamic information table that is continuously updated during the traversal of the basic syntax tree. The first reference dictionary can also record the location information of each node (element identifier). For example, for node 1, the first reference dictionary records: object element - demo; location information: starting row and column number [0, 1]; ending row and column number [0, 8]; scope: inside the method body, inside the class, method definition, class definition, etc.
[0186] The second reference dictionary is a pre-built record of the correspondence between the element information and storage path of each code element, where the storage path is used as query information to determine the target query information. The element information and storage path of each code element are recorded in the form of key-value pairs, realizing the corresponding query between the element information and the storage path. For example, for variable element-a, the second reference dictionary records: element information: object element-demo+function element-add+variable element-a, storage path: module path+module name+demo+add+a.
[0187] FIG4 shows a schematic diagram of a first reference dictionary and a second reference dictionary in a code text processing method provided by an embodiment of the present disclosure, as shown in FIG4 :
[0188] For the abstract syntax tree of the code text in Figure 3, the first reference dictionary is updated during the traversal of the abstract syntax tree. The first reference dictionary contains the following records: Node 1: Object element - demo; Position information: Starting row and column numbers [0, 1]; Ending row and column numbers [0, 8]; Scope: Object definition. Node 2: Function element - add; Position information: Starting row and column numbers [0, 2]; Ending row and column numbers [0, 4]; Scope: Internal module. Node 3: Function element - mul; Position information: Starting row and column numbers [0, 5]; Ending row and column numbers [0, 7]; Scope: Internal module. Node 4: Variable element - a; Position information: Starting row and column numbers [0, 2]; Ending row and column numbers [0, 3]; Scope: Method definition. Node 5: Variable element - b; Position information: Starting row and column numbers [0, 2]; Ending row and column numbers [0, 3]; Scope: Method definition. Node 6: Variable element-a; Location information: Starting row and column numbers [0, 5]; Ending row and column numbers [0, 6]; Scope: Method definition. Node 7: Variable element-b; Location information: Starting row and column numbers [0, 5]; Ending row and column numbers [0, 6]; Scope: Method definition.
[0189] The second reference dictionary records the following in the form of key-value pairs: (element information: object element-demo; storage path: module path + module name + demo), (element information: object element-demo+function element-add; storage path: module path + module name + demo+add), (element information: object element-demo+function element-mul; storage path: module path + module name + demo+mul), (element information: object element-demo+function element-add+variable element-a; storage path: module path + module name + demo+add+a), (element information: object element-demo+function element-add+variable element-b; storage path: module path + module name + demo+add+b), (element information: object element-demo+function element-mul+variable element-a; storage path: module path + module name + demo+mul+b), (element information: object element-demo+function element-mul+variable element-b; storage path: module path + module name + demo+mul+b).
[0190] Through the second reference dictionary, we can understand that for code elements with the same name (for example, variable element a corresponding to node 4 and variable element a corresponding to node 6), due to the clear and identifiable element information (object element-demo + function element-add + variable element-a, and object element-demo + function element-mul + variable element-a), the code elements can be accurately distinguished, ensuring the accuracy of the determined target query information, obtaining valid second code metadata, avoiding the introduction of invalid reference code and the resulting model illusion in the code generation model, and achieving highly accurate code text processing. This is a method for obtaining code metadata with a full-path unique identifier.
[0191] It should be noted that, since the first reference dictionary is updated during the traversal process, when duplicate element information is encountered, the duplicate element information is merged and the element information with updated position information is used to determine the target query information.
[0192] Exemplarily, based on the element identification of the target code element: node 1 and node 2, the first reference dictionary Refmap is queried to obtain the element names of the target code elements (object element ("demo") and function element ("add")), and based on the element names of the target code elements, the second reference dictionary importRefmap is queried to obtain the storage path of the target code elements as the target query information: module path + module name + demo + add.
[0193] Based on the element identifier of the target code element, the first reference dictionary is queried to obtain the element information of the target code element. The first reference dictionary is updated during the process of traversing the basic syntax tree and records the correspondence between the element identifier and element information of each code element. Based on the element information of the target code element, the second reference dictionary is queried to obtain the storage path of the target code element as the target query information. The second reference dictionary records the correspondence between the element information and storage path of each code element. This implements path query with a full path unique identifier, ensures accurate acquisition of valid second code metadata, avoids the introduction of invalid reference code, and avoids model illusions in the code generation model, thus achieving highly accurate code text processing.
[0194] In an optional embodiment of the present disclosure, after searching the first reference dictionary based on the element identifier of the target code element, the following specific steps are further included:
[0195] In the case where the element information of the target code element is found, the position information of the target code element in the first reference dictionary is updated using the target position information.
[0196] The first reference dictionary is updated during the process of traversing the base syntax tree. Therefore, if the target location information is the location information of the currently written code element, as the writing process progresses, if the location information of the target code element is recorded in the first reference dictionary, it is necessary to dynamically update the location information using the target location information to ensure the accuracy of the information recorded in the first reference dictionary. For example, for node 1, the first reference dictionary already records node 1: object element - demo; location information: starting row and column number [0, 1]; ending row and column number [0, 8]; range: inside the object body. If the target location information is row 9, the updated element information of node 1 is: object element - demo; location information: starting row and column number [0, 1]; ending row and column number [0, 9]; range: inside the object body.
[0197] Exemplarily, when the element names of the target code elements (object element ("demo") and function element ("add")) are queried, the target position information (line 9) is used to update the position information of the target code elements in the first reference dictionary: starting row and column number [0,1]; ending row and column number [0,9]; range: inside the object body.
[0198] When the element information of the target code element is found, the target position information is used to update the position information of the target code element in the first reference dictionary, thereby ensuring the accuracy of the element information and position information recorded in the first reference dictionary and the accuracy of subsequent queries.
[0199] In an optional embodiment of the present disclosure, the method further includes the following specific steps:
[0200] In the case that the element information of the target code element is not found, the element identifier, element information and position information of the target code element are recorded in the first reference dictionary.
[0201] The first reference dictionary is updated during the process of traversing the basic syntax tree. Therefore, if the target location information is the location information of the code element currently being written, as the writing process progresses, if the location information of the target code element is not recorded in the first reference dictionary, it indicates that the code element being written is a new code element, and the element identifier, element information, and location information of the target code element need to be recorded in the first reference dictionary to ensure the integrity of the information recorded in the first reference dictionary. For example, for node 1, node 7 is not recorded in the first reference dictionary. The element identifier (node 7), element information (variable element-b), and location information (starting row and column number [0,5]; ending row and column number [0,6]; range: method definition) are recorded in the first reference dictionary.
[0202] Exemplarily, when the element name of the target code element (variable element (b)) is not found, the element identifier (node 7), element information (variable element-b) and position information (starting row and column number [0,5]; ending row and column number [0,6]; range: method definition) are recorded in the first reference dictionary.
[0203] If the element information of the target code element is not found, the element identifier, element information and location information of the target code element are recorded in the first reference dictionary, thereby ensuring the integrity of the element identifier, element information and location information recorded in the first reference dictionary and the accuracy of subsequent queries.
[0204] In an optional embodiment of the present disclosure, the second code metadata is at least one;
[0205] Correspondingly, step 110 includes the following specific steps:
[0206] determining a weight of the at least one second code metadata based on position information of the target code element corresponding to the at least one second code metadata;
[0207] Based on the weight of each second code metadata, each second code metadata is placed into a code text sequence;
[0208] The first code metadata and the code text sequence are input into a code generation model to generate a target code text.
[0209] Because the code generation model is trained on sample code text by the text processing model, the length of text that the text generation model can process is limited. When acquiring a large amount of secondary code metadata, not all of it can be fed into the code generation model, requiring selection. The location information recorded in the first reference dictionary is dynamically updated. Similar code elements written before and after can all serve as secondary code metadata. Newly written code metadata is prioritized, giving it a higher weight.
[0210] The weight of the second code metadata is the sorting weight in the text sequence. The higher the weight, the higher the priority in the text sequence. For example, for the second code metadata: "methods":{"add":{"method_name":"add","signature":"add:function()"}} and "methods":{"mul":{"method_name":"mul","signature":"mul:function()"}}, the former has a higher weight than the latter. The code text sequence is "methods":{"add":{"method_name":"add","signature":"add:function()"}}, and the separator is "methods":{"mul":{"method_name":"mul","signature":"mul:function()"}}.
[0211] In an optional embodiment of the present disclosure, before determining the weight of at least one second code metadata based on the position information of the target code element corresponding to the at least one second code metadata, the following specific steps are further included:
[0212] Obtain target position information of the code element to be processed in the code text to be processed;
[0213] Correspondingly, determining the weight of at least one second code metadata based on the position information of the target code element corresponding to the at least one second code metadata includes the following specific steps:
[0214] determining a position distance between the target code element and the code element to be processed based on the target position information and position information of the target code element corresponding to the at least one second code metadata;
[0215] The weight of at least one second code metadata is determined based on the position distance between the target code element and the code element to be processed.
[0216] Optionally, when the length of the code text sequence reaches a preset threshold, step 108 is stopped.
[0217] In addition, if the second code metadata is in a high-speed read / write medium, a fixed value is added to the weight. Moreover, different weights are set according to the code language and different ranges.
[0218] Optionally, the weights may also be calculated using a preset algorithm, which includes but is not limited to:
[0219] Method 1:
[0220] When obtaining the project file of the target project, a code reference sequence in the project file is constructed and indexed using an n-gram method. When obtaining and weighting the second code metadata, code reference sequence information within a certain range before the target location information is obtained and predicted using an n-gram algorithm to obtain the second code metadata. The weight of the code reference sequence information in the second code metadata is determined as the weight.
[0221] Method 2:
[0222] By constructing sample data, a deep learning model is trained based on multiple feature dimensions, such as the location, scope, and grammatical structure of the sample code metadata, to generate a ranking model for sorting the code metadata. Before inserting each piece of second code metadata into the code text sequence, the ranking model is fed with information on all the acquired code metadata, including its location, scope, and grammatical structure. A probability value is output for each piece of code metadata, and a weight is determined based on the probability value before each piece of second code metadata is inserted into the code text sequence.
[0223] Exemplarily, based on the position information (line 9 and line 5) and target position information (line 14) of the target code elements (function element ("add") and function element ("mul") corresponding to the two second code metadata, the distance between the two target code elements is: 5 and 9. Based on the distance of the two target code elements, the weights of the two second code metadata are determined to be: 1 / 5 and 1 / 9. Based on the weights of the two second code metadata (1 / 5 is greater than 1 / 9), the two second code metadata are put into the code text sequence: "methods":{"add":{"method_name":"add","signature":"add:function()"}}, separator, "methods":{"mul":{"method_name":"mul","signature":"mul:function()"}}, the first code metadata and the code text sequence are input into the large language model to generate supplementary code text: "var demo={name:"John",age:25,add:function(){console.log(this.name);},};demo.add".
[0224] Based on the location information of the target code element corresponding to at least one second code metadata, a weight for at least one second code metadata is determined; based on the weight of each second code metadata, each second code metadata is placed into a code text sequence; and the first code metadata and the code text sequence are input into a code generation model to generate the target code text. Using this location information, the second code metadata are rationally sorted and a code text sequence is constructed. When the text sequence constraints of the code generation model are met, the second code metadata that is more likely to be referenced is prioritized, thereby improving the accuracy of code text processing.
[0225] In an optional embodiment of the present disclosure, before step 106, the following specific steps are further included:
[0226] Identify the code language of the code text to be processed;
[0227] Correspondingly, step 106 includes the following specific steps:
[0228] Determining target query information under the code language according to the code language and element information of each code element;
[0229] The second code metadata corresponding to the target query information is searched from the target database of the code language.
[0230] Different coding languages correspond to different element information, which in turn determines different query information. For details, see the description of query information in step 104. Furthermore, different coding languages also have different code metadata. Therefore, it is necessary to determine the target query information for a given coding language based on the coding language and the element information of each code element, and then search the target database for the code language for the second code metadata corresponding to the target query information.
[0231] For example, the code text to be supplemented is: "var demo = {name: "John", age: 25, add", which identifies the code language of the code text to be supplemented as JavaScript. According to the code language JavaScript and the element names of the target code elements (object element ("demo") and function element ("add") in each code element, a reference dictionary pre-recorded with the correspondence between JavaScript element information and query information is queried to obtain the target query information: module path + module name + demo + add. Based on the key-value pairs of the correspondence between the reference query information and the reference code metadata recorded in the JavaScript database, the target query information "module path + module name + demo + add" is used as the key to search for the second code metadata of the corresponding value: {"name":"dem o","signature":"vardemo","full_name":"project.path.demo","fields":{"name":{"field_name":"name","field_value":"John","signature":"name:'John' "},"age":{"field_name":"age","field_value":"25","signature":"age:25"}},"methods":{"add":{"method_name":"add","signature":"add:function()"}}}.
[0232] The code language of the code text to be processed is identified; based on the code language and the element information of each code element, the target query information of the code language is determined; and the second code metadata corresponding to the target query information is searched from the target database of the code language. By identifying the code language, the accuracy of the determined target query information and the validity of the second code metadata obtained from the query are guaranteed, thus ensuring the accuracy of the code text processing.
[0233] FIG5 shows a flowchart of updating a code text sequence in a code text processing method provided by an embodiment of the present disclosure, as shown in FIG5 :
[0234] When a developer opens, closes, or modifies a project file in a target project, the code metadata needs to be updated accordingly. The updated code metadata may be code metadata that will be needed in the near future. The updated code metadata is added to the dynamic file queue (located in high-speed read-write media such as memory or cache) to improve the efficiency of subsequent code text processing.
[0235] Specifically: open project file / close project file / modify project file; request local file change interface; parse the changed project file to obtain code metadata of each code element; add code metadata to the dynamic file queue according to the code language; determine whether the queue length exceeds the threshold; if not, end directly; if so, remove the code metadata in the queue in first-in-first-out order and end.
[0236] FIG6 shows a flowchart of real-time analysis of code text in a code text processing method provided by one embodiment of the present disclosure, as shown in FIG6 :
[0237] When developers are writing code text, the code generation interface in the code text processing process is called in real time, and the code text to be processed of the target project being written is passed to the process to implement the following steps:
[0238] First, obtain the code text to be processed of the target project; request the local code generation interface.
[0239] Next, the code text to be processed is parsed to obtain the first code metadata of each code element, and target query information is determined based on the element information of each code element; and second code metadata corresponding to the target query information is searched from the target database.
[0240] At the same time, the code text to be processed is preprocessed to obtain code context information.
[0241] Finally, based on the first code metadata, the second code metadata and the code context information, the code generation model is called to generate the target code text.
[0242] FIG7 shows a flowchart of code text parsing in a code text processing method provided by an embodiment of the present disclosure, as shown in FIG7 :
[0243] A feasible embodiment of the code parsing in FIG6 is as follows:
[0244] Obtain the target position information of the code text to be processed and the code elements to be processed in the code text to be processed; perform grammatical structure analysis on the code in the code text to be processed to obtain the basic grammar tree of the code text to be processed; parse the basic grammar tree to obtain the first code metadata of each code element; based on the target position information, traverse each code element recorded in the basic grammar tree to determine the target code element; based on the element identifier of the target code element, query the first reference dictionary and record the correspondence between the element identifier and the element information of each code element; determine whether the element information of the target code element is queried; if not, return the element identifier and metadata of the target code element to the first reference dictionary; The element information and position information of the target code element are recorded in the first reference dictionary; if yes, the position information of the target code element in the first reference dictionary is updated using the target position information; based on the element information of the target code element, the second reference dictionary is queried to obtain the storage path of the target code element as the target query information, and the correspondence between the element information and the storage path of each code element is recorded; from the target database, the second code metadata corresponding to the target query information is searched, and the weight of the second code metadata is determined based on the position information of the target code element corresponding to the second code metadata; based on the weight of each second code metadata, each second code metadata is placed into the code text sequence.
[0245] FIG8 shows a front-end schematic diagram of a code text processing method provided by an embodiment of the present disclosure, as shown in FIG8 :
[0246] The front-end interface of the integrated development environment includes code parsing controls, code generation interface (start or close), project file directory of the target project, and code text writing area.
[0247] The project file directory of the target project has a hierarchical structure as follows: Project Engineering - Project File 1 - Module 1.1 and Module 1, 2; Project Engineering - Project File 2.
[0248] As shown in the figure above: the code text editing area has the following code text to be added: "var demo = {name: "John", age: 25, add", and the current cursor position stays on the second line.
[0249] As shown in the figure below, by executing the above embodiment, a supplementary code text is obtained: "var demo = {name: "John", age: 25, add: function () {console.log (this.name);},}; demo.add". The supplementary code text is used to supplement the code text to be supplemented.
[0250] 9 , which shows a flowchart of a code supplement method provided by an embodiment of the present disclosure. The method is applied to a cloud-side device and includes the following specific steps:
[0251] Step 902: The code text to be supplemented of the target item is input by the receiving device.
[0252] Step 904: Parse the code text to be supplemented to obtain first code metadata of each code element, wherein each code element has corresponding element information.
[0253] Step 906: Determine target query information based on the element information of each code element.
[0254] Step 908: Search the target database for the second code metadata corresponding to the target query information, wherein the target database records the correspondence between the reference query information and the reference code metadata, and the reference query information is constructed based on the element information of each code element in the project file of the target project.
[0255] Step 910: Generate supplementary code text based on the first code metadata and the second code metadata using a code generation model, wherein the code generation model is obtained by training a text processing model based on sample code text.
[0256] Step 912: Send the supplementary code text to the end-side device, so that the end-side device supplements the code text to be supplemented by using the supplementary code text.
[0257] The embodiment of the present disclosure is applied to a network cloud device where the server of a web page, application or mini-program with code text processing function is located. It is a virtual device, and a code generation function model with code text processing function is deployed on the cloud-side device. The end-side device is a physical device where the client of the web page, application or mini-program with code text processing function logged in by the user is located. The cloud-side device and the end-side device are connected through a network transmission channel for data transmission. The computing power performance and storage performance of the cloud-side device are higher than those of the end-side device.
[0258] The embodiment of the present disclosure and the embodiment of the specification of FIG1 are based on the same inventive concept. The specific methods of steps 904 to 910 refer to the contents of steps 104 to 110 in the embodiment of the specification of FIG1 , which will not be repeated here.
[0259] In an embodiment of the present disclosure, a code text to be supplemented of a target project is received input by an end-side device; the code text to be supplemented is parsed to obtain first code metadata of each code element, wherein each code element has corresponding element information; target query information is determined based on the element information of each code element; second code metadata corresponding to the target query information is searched from a target database, wherein the target database records a correspondence between reference query information and reference code metadata, and the reference query information is constructed based on the element information of each code element in a project file of the target project; based on the first code metadata and the second code metadata, a code generation model is used to generate supplementary code text, wherein the code generation model is obtained by training a text processing model based on sample code text; and the supplementary code text is sent to the end-side device, so that the end-side device uses the supplementary code text to supplement the code text to be supplemented. By pre-parsing the project files of the target project, reference query information is constructed based on the element information of each code element, and stored in the target database corresponding to the reference code metadata. During the code text processing, the first code metadata of each code element is obtained by parsing the code text to be processed, and then the target query information is determined based on the element information of each code element. Using this clear code reference relationship, the target database is queried to obtain valid second code metadata, which is used as a reference code to guide the code generation model and generate highly accurate supplementary code text, thereby improving the accuracy of code supplementation. At the same time, it is implemented on cloud-side devices with high computing performance and high storage performance, thereby improving the efficiency and accuracy of code supplementation.
[0260] In an optional embodiment of the present disclosure, after step 912, the following specific steps are further included:
[0261] Supplementary feedback information sent by the receiving end-side device, wherein the supplementary feedback information is information providing feedback on the supplementary code text;
[0262] Based on the supplemental feedback information, the parameters of the code generation model are adjusted.
[0263] Supplementary feedback information refers to feedback on supplementary code text, including but not limited to: code syntax errors, code naming errors, and incomplete code supplements.
[0264] Exemplarily, the receiving-side device provides supplementary feedback information for supplementing the code text: code syntax errors, code naming errors, and incomplete code supplementation. Based on the supplementary feedback information, the parameters of the large language model are adjusted.
[0265] In the disclosed embodiment, feedback adjustment of the parameters of the code generation model is completed in an interactive manner, thereby specifically improving the effect of code supplementation.
[0266] The following further illustrates the code text processing method provided by the present disclosure using the application of the code text processing method in an integrated development environment as an example, in conjunction with FIG10 . FIG10 shows a flowchart of a process of a code text processing method applied to an integrated development environment provided by one embodiment of the present disclosure, including the following specific steps:
[0267] Step 1002: When the integrated development environment is started, the project file of the target project is obtained.
[0268] Step 1004: Filter invalid project files.
[0269] Step 1006: Perform syntax structure analysis on the code in the project file to obtain a project syntax tree of the project file, parse the project syntax tree, and obtain reference code metadata of each code element.
[0270] Step 1008: Determine the structural information between the code elements according to the element information of each code element, and construct the storage path of the reference code metadata of each code element as reference query information based on the structural information between the code elements.
[0271] Step 1010: Build a target database based on the correspondence between the reference query information and the reference code metadata.
[0272] Step 1012: Obtain the target position information of the code text to be processed and the code element to be processed where the current cursor is located.
[0273] Step 1014: Perform grammatical structure analysis on the code in the code text to be processed to obtain a basic syntax tree of the code text to be processed, parse the basic syntax tree, and obtain first code metadata of each code element.
[0274] Step 1016: Based on the target position information, traverse each code element recorded in the basic syntax tree to determine the target code element, and based on the element identifier of the target code element, query the first reference dictionary to obtain the element information of the target code element, and when the element information of the target code element is queried, use the target position information to update the position information of the target code element in the first reference dictionary, or when the element information of the target code element is not queried, record the element identifier, element information and position information of the target code element in the first reference dictionary.
[0275] Step 1018: Based on the element information of the target code element, query the second reference dictionary to obtain the storage path of the target code element as target query information.
[0276] Step 1020: searching the target database for the second code metadata corresponding to the target query information, and merging duplicate second code metadata based on the position information of the target code elements corresponding to the second code metadata.
[0277] Step 1022: Based on the position information of the target code element corresponding to each second code metadata, determine the weight of each second code metadata, and based on the weight of each second code metadata, place each second code metadata into the code text sequence.
[0278] Step 1024: Input the first code metadata and the code text sequence into a code generation model to generate a target code text;
[0279] Step 1026: The target code text is used to supplement the code text to be processed, and rendered on the front-end interface of the integrated development environment.
[0280] By parsing the entire target project, analyzing the reference relationship between codes, constructing highly identifiable query information, and querying the code metadata actually used from the target database, the recall rate and hit rate of code references are effectively improved, and effective second code metadata is obtained as a reference code to guide the code generation model, overcoming the model hallucination problem, generating highly accurate target code text, and improving the accuracy of code text processing.
[0281] Corresponding to the above method embodiment, the present disclosure also provides an embodiment of a code text processing device. FIG11 shows a schematic diagram of the structure of a code text processing device provided by an embodiment of the present disclosure. As shown in FIG11, the device includes:
[0282] An acquisition module 1102 is configured to acquire the code text to be processed of the target project;
[0283] A first parsing module 1104 is configured to parse the code text to be processed and obtain first code metadata of each code element, wherein each code element has corresponding element information;
[0284] A first determining module 1106 is configured to determine target query information based on element information of each code element;
[0285] A first search module 1108 is configured to search a target database for second code metadata corresponding to the target query information, wherein the target database records a correspondence between reference query information and reference code metadata, and the reference query information is constructed based on element information of each code element in a project file of the target project;
[0286] The first generation module 1110 is configured to generate target code text based on the first code metadata and the second code metadata using a code generation model, wherein the code generation model is obtained by training a text processing model based on sample code text.
[0287] Optionally, the device also includes: a construction module, configured to obtain a project file of a target project; parse the project file to obtain reference code metadata of each code element; construct reference query information based on the element information of each code element; and construct a target database based on the correspondence between the reference query information and the reference code metadata.
[0288] Optionally, the construction module is further configured to: perform syntax structure analysis on the code in the project file to obtain a project syntax tree of the project file; and parse the project syntax tree to obtain reference code metadata of each code element.
[0289] Optionally, the construction module is further configured to: determine the structural information between each code element according to the element information of each code element; and construct the storage path of the reference code metadata of each code element as reference query information based on the structural information between each code element.
[0290] Optionally, the device also includes: a position information acquisition module, configured to obtain target position information of the code element to be processed in the code text to be processed; correspondingly, the first parsing module 1104 is further configured to: perform grammatical structure analysis on the code in the code text to be processed to obtain a basic syntax tree of the code text to be processed; parse the basic syntax tree to obtain first code metadata of each code element; correspondingly, the first determination module 1106 is further configured to: based on the target position information, traverse each code element recorded in the basic syntax tree to determine the target code element; and determine the target query information based on the element information of the target code element.
[0291] Optionally, the device also includes: a query module, configured to query a first reference dictionary based on the element identifier of the target code element to obtain the element information of the target code element, wherein the first reference dictionary is updated in the process of traversing the basic syntax tree, and the first reference dictionary records the correspondence between the element identifier and the element information of each code element; correspondingly, the first determination module 1106 is further configured to: query a second reference dictionary based on the element information of the target code element to obtain the storage path of the target code element as the target query information, wherein the second reference dictionary records the correspondence between the element information and the storage path of each code element.
[0292] Optionally, the apparatus further comprises: an updating module configured to update the position information of the target code element in the first reference dictionary using the target position information when the element information of the target code element is found in the query.
[0293] Optionally, the apparatus further comprises: a dictionary recording module configured to record the element identifier, element information and position information of the target code element into the first reference dictionary if the element information of the target code element is not found.
[0294] Optionally, there is at least one second code metadata; correspondingly, the first generation module 1110 is further configured to: determine the weight of at least one second code metadata based on the position information of the target code element corresponding to the at least one second code metadata; based on the weight of each second code metadata, place each second code metadata into a code text sequence; input the first code metadata and the code text sequence into the code generation model to generate the target code text.
[0295] Optionally, the device further includes: a position information acquisition module configured to acquire target position information of a code element to be processed in the code text to be processed; correspondingly, the first generation module 1110 is further configured to: determine the position distance between the target code element and the code element to be processed based on the target position information and the position information of the target code element corresponding to at least one second code metadata; and determine the weight of at least one second code metadata based on the position distance between the target code element and the code element to be processed.
[0296] Optionally, the device further includes: a code language identification module configured to identify the code language of the code text to be processed; correspondingly, the first determination module 1106 is further configured to: determine target query information under the code language based on the code language and element information of each code element; and search for second code metadata corresponding to the target query information from a target database of the code language.
[0297] In the disclosed embodiment, by pre-parsing the project file of the target project, reference query information is constructed based on the element information of each code element, and stored in the target database corresponding to the reference code metadata. During the code text processing, the first code metadata of each code element is obtained by parsing the code text to be processed, and then the target query information is determined based on the element information of each code element. By utilizing this clear code reference relationship, the target database is queried to obtain valid second code metadata, which is used as a reference code to guide the code generation model and generate highly accurate target code text, thereby improving the accuracy of code text processing.
[0298] The above is a schematic diagram of a code text processing device according to this embodiment. It should be noted that the technical solution of this code text processing device and the technical solution of the aforementioned code text processing method are based on the same concept. For details not described in detail in the technical solution of the code text processing device, please refer to the description of the technical solution of the aforementioned code text processing method.
[0299] Corresponding to the above method embodiment, the present disclosure also provides a code supplement device embodiment. Figure 12 shows a schematic diagram of the structure of a code supplement device provided by one embodiment of the present disclosure. As shown in Figure 12, the device is applied to a cloud-side device and includes:
[0300] The receiving module 1202 is configured to receive the code text to be supplemented of the target item input by the terminal side device;
[0301] The second parsing module 1204 is configured to parse the code text to be supplemented and obtain first code metadata of each code element, wherein each code element has corresponding element information;
[0302] The second determining module 1206 is configured to determine target query information based on element information of each code element;
[0303] A second search module 1208 is configured to search for second code metadata corresponding to the target query information from a target database, wherein the target database records a correspondence between reference query information and reference code metadata, and the reference query information is constructed based on element information of each code element in a project file of the target project;
[0304] The second generation module 1210 is configured to generate supplementary code text based on the first code metadata and the second code metadata using a code generation model, wherein the code generation model is trained on a text processing model based on the sample code text;
[0305] The sending module 1212 is configured to send the supplementary code text to the terminal side device, so that the terminal side device supplements the code text to be supplemented by using the supplementary code text.
[0306] In the disclosed embodiment, by pre-parsing the project file of the target project, reference query information is constructed based on the element information of each code element, and stored in the target database corresponding to the reference code metadata. During the code text processing, the first code metadata of each code element is obtained by parsing the code text to be processed, and then the target query information is determined based on the element information of each code element. By utilizing this clear code reference relationship, the target database is queried to obtain valid second code metadata, which is used as a reference code to guide the code generation model and generate highly accurate supplementary code text, thereby improving the accuracy of code supplementation. At the same time, it is implemented on a cloud-side device with high computing performance and high storage performance, thereby improving the efficiency and accuracy of code supplementation.
[0307] The above is a schematic scheme of a code supplement device of this embodiment. It should be noted that the technical scheme of the code supplement device and the technical scheme of the above-mentioned code supplement method are of the same concept. For details not described in detail in the technical scheme of the code supplement device, please refer to the description of the technical scheme of the above-mentioned code supplement method.
[0308] Figure 13 shows a block diagram of a computing device according to an embodiment of the present disclosure. Components of the computing device 1300 include, but are not limited to, a memory 1310 and a processor 1320. The processor 1320 is connected to the memory 1310 via a bus 1330, and a database 1350 is used to store data.
[0309] The computing device 1300 also includes an access device 1340 that enables the computing device 1300 to communicate via one or more networks 1360. Examples of such networks include a Public Switched Telephone Network (PSTN), a Local Area Network (LAN), a Wide Area Network (WAN), a Personal Area Network (PAN), or a combination of communication networks such as the Internet. The access device 1340 may include one or more of any type of network interface (e.g., a Network Interface Controller (NIC)) whether wired or wireless, such as an IEEE 802.11 Wireless Local Area Network (WLAN) wireless interface, a Worldwide Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, or a Near Field Communication (NFC).
[0310] In one embodiment of the present disclosure, the aforementioned components of the computing device 1300 and other components not shown in FIG13 may also be connected to each other, for example, via a bus. It should be understood that the block diagram of the computing device structure shown in FIG13 is for illustrative purposes only and does not limit the scope of the present disclosure. Those skilled in the art may add or replace other components as needed.
[0311] Computing device 1300 can be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (e.g., a tablet computer, personal digital assistant, laptop computer, notebook computer, netbook computer, etc.), a mobile phone (e.g., a smartphone), a wearable computing device (e.g., a smartwatch, smart glasses, etc.), or other types of mobile devices, or a stationary computing device such as a desktop computer or personal computer (PC). Computing device 1300 can also be a mobile or stationary server.
[0312] The processor 1320 is configured to execute the following computer-executable instructions, which, when executed by the processor, implement the steps of the above-mentioned code text processing method or code supplementation method.
[0313] The above is a schematic diagram of a computing device according to this embodiment. It should be noted that the technical solution of this computing device is based on the same concept as the technical solutions of the aforementioned code text processing method and code supplementation method. For details not described in detail in the technical solution of the computing device, please refer to the description of the technical solutions of the aforementioned code text processing method or code supplementation method.
[0314] An embodiment of the present disclosure further provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the steps of the above-mentioned code text processing method or code supplementation method.
[0315] The above is a schematic diagram of a computer-readable storage medium according to this embodiment. It should be noted that the technical solution of this storage medium is based on the same concept as the technical solutions of the aforementioned code text processing method and code supplementation method. For details not described in detail in the technical solution of the storage medium, please refer to the description of the technical solutions of the aforementioned code text processing method or code supplementation method.
[0316] An embodiment of the present disclosure further provides a computer program, wherein when the computer program is executed in a computer, the computer is caused to execute the steps of the above-mentioned code text processing method or code supplementation method.
[0317] The above is a schematic diagram of a computer program according to this embodiment. It should be noted that the technical solution of this computer program is based on the same concept as the technical solutions of the aforementioned code text processing method and code supplementation method. For details not described in detail in the technical solution of the computer program, please refer to the description of the technical solutions of the aforementioned code text processing method or code supplementation method.
[0318] The foregoing description describes specific embodiments of the present disclosure. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0319] The computer instructions include computer program code, which may be in source code form, object code form, executable file, or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal, and software distribution medium. It should be noted that the content contained in the computer-readable medium may be appropriately increased or decreased according to the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media does not include electric carrier signals and telecommunication signals.
[0320] It should be noted that for the sake of simplicity, the aforementioned method embodiments are described as a series of actions. However, those skilled in the art should be aware that the embodiments of the present disclosure are not limited by the order of the actions described, as certain steps may be performed in other orders or simultaneously according to the embodiments of the present disclosure. Furthermore, those skilled in the art should also be aware that the embodiments described in this specification are preferred embodiments, and the actions and modules involved are not necessarily required for the embodiments of the present disclosure.
[0321] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.
[0322] The preferred embodiments of the present disclosure disclosed above are only used to help illustrate the present disclosure. The optional embodiments do not describe all details in detail, nor do they limit the invention to only the specific embodiments described. Obviously, many modifications and variations can be made based on the content of the embodiments of the present disclosure. The present disclosure selects and specifically describes these embodiments in order to better explain the principles and practical applications of the embodiments of the present disclosure, so that those skilled in the art can better understand and utilize the present disclosure. The present disclosure is limited only by the claims and their full scope and equivalents.
Claims
1. A code text processing method, comprising: Get the code text to be processed of the target project; Parsing the code text to be processed to obtain first code metadata of each code element, wherein each code element has corresponding element information; Determining target query information according to element information of each code element; Searching for second code metadata corresponding to the target query information from a target database, wherein the target database records a correspondence between reference query information and reference code metadata, and the reference query information is constructed based on element information of each code element in a project file of the target project; Based on the first code metadata and the second code metadata, a target code text is generated using a code generation model, wherein the code generation model is obtained by training a text processing model based on sample code text.
2. The method according to claim 1, before searching the target database for the second code metadata corresponding to the target query information, further comprises: Get the project file of the target project; Parsing the project file to obtain reference code metadata of each code element; Constructing reference query information according to element information of each code element; Based on the correspondence between the reference query information and the reference code metadata, a target database is constructed.
3. According to the method of claim 2, the step of parsing the project file to obtain reference code metadata of each code element comprises: Performing grammatical structure analysis on the code in the project file to obtain a project syntax tree of the project file; The project syntax tree is parsed to obtain reference code metadata of each code element.
4. The method according to claim 2, wherein the step of constructing reference query information according to the element information of each code element comprises: Determine structural information between the code elements according to the element information of the code elements; Based on the structural information between the code elements, a storage path of the reference code metadata of the code elements is constructed as reference query information.
5. The method according to any one of claims 1 to 4, before parsing the code text to be processed to obtain the first code metadata of each code element, further comprising: Obtain target position information of the code element to be processed in the code text to be processed; The parsing of the code text to be processed to obtain first code metadata of each code element includes: Performing grammatical structure analysis on the code in the code text to be processed to obtain a basic grammar tree of the code text to be processed; Parsing the basic syntax tree to obtain first code metadata of each code element; The determining target query information according to the element information of each code element includes: Based on the target position information, traverse the code elements recorded in the basic syntax tree to determine the target code element; Based on the element information of the target code element, target query information is determined.
6. The method according to claim 5, before determining the target query information based on the element information of the target code element, further comprising: Based on the element identifier of the target code element, query a first reference dictionary to obtain element information of the target code element, wherein the first reference dictionary is updated in the process of traversing the basic syntax tree, and the first reference dictionary records the correspondence between the element identifier and the element information of each code element; The determining target query information based on the target element information of the target code element includes: Based on the element information of the target code element, a second reference dictionary is queried to obtain the storage path of the target code element as target query information, wherein the second reference dictionary records the correspondence between the element information of each code element and the storage path.
7. The method according to claim 6, after searching the first reference dictionary based on the element identification of the target code element, further comprises: When the element information of the target code element is found, the position information of the target code element in the first reference dictionary is updated using the target position information.
8. The method according to claim 6, further comprising: In the case that the element information of the target code element is not found, the element identifier, element information and position information of the target code element are recorded in the first reference dictionary.
9. The method according to any one of claims 1 to 8, wherein the second code metadata is at least one; The step of generating a target code text based on the first code metadata and the second code metadata by using a code generation model includes: Determining a weight of the at least one second code metadata based on the position information of the target code element corresponding to the at least one second code metadata; Based on the weight of each second code metadata, placing each second code metadata into a code text sequence; The first code metadata and the code text sequence are input into a code generation model to generate a target code text.
10. The method according to claim 9, before determining the weight of the at least one second code metadata based on the position information of the target code element corresponding to the at least one second code metadata, further comprising: Obtain target position information of the code element to be processed in the code text to be processed; The determining the weight of the at least one second code metadata based on the position information of the target code element corresponding to the at least one second code metadata comprises: Determine a position distance between the target code element and the code element to be processed based on the target position information and the position information of the target code element corresponding to the at least one second code metadata; Based on the position distance between the target code element and the to-be-processed code element, a weight of the at least one second code metadata is determined.
11. The method according to any one of claims 1 to 10, before determining the target query information according to the element information of each code element, further comprising: Identify the code language of the code text to be processed; The determining target query information according to the element information of each code element includes: Determining target query information under the code language according to the code language and element information of each code element; The target database of the code language is searched for second code metadata corresponding to the target query information.
12. A code supplement method, applied to a cloud-side device, comprising: The code text to be supplemented of the target item input by the receiving device; Parsing the code text to be supplemented to obtain first code metadata of each code element, wherein each code element has corresponding element information; Determining target query information according to element information of each code element; Searching for second code metadata corresponding to the target query information from a target database, wherein the target database records a correspondence between reference query information and reference code metadata, and the reference query information is constructed based on element information of each code element in a project file of the target project; Based on the first code metadata and the second code metadata, using a code generation model, generating supplementary code text, wherein the code generation model is obtained by training a text processing model based on sample code text; The supplementary code text is sent to the end-side device, so that the end-side device supplements the code text to be supplemented by using the supplementary code text.
13. The method according to claim 12, after sending the supplementary code text to the end-side device, further comprises: Receiving supplementary feedback information sent by the terminal side device, wherein the supplementary feedback information is information for providing feedback on the supplementary code text; Based on the supplemental feedback information, parameters of the code generation model are adjusted.
14. A computing device comprising: Memory and processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the method described in any one of claims 1 to 13 are implemented.
15. A computer-readable storage medium storing computer-executable instructions, wherein the computer-executable instructions, when executed by a processor, implement the steps of the method according to any one of claims 1 to 13.
16. A computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 13.
Citation Information
Patent Citations
Determination method and device for dependency information of database storage process and electronic equipment
CN116069808A
Code generation method and device
CN116719520A
Code text processing method, code supplementing method and computing equipment
CN117806601A
Source code generation, completion, checking, correction
US20150135166A1
Code processing method, apparatus, device, and medium
WO2022089188A1
Cited By
Code generation model training method, device and equipment
CN120873595A