Code completion method and device, computer equipment and storage medium

By creating knowledge data and dependency data of code files in the code repository and training the code completion model, the problem of insufficient accuracy of the code completion model in the existing technology is solved, and higher code generation accuracy is achieved.

CN120371270APending Publication Date: 2025-07-25TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410107835.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-25
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The existing code completion model is relatively simple in the training process, which leads to insufficient accuracy and cannot effectively improve the accuracy of code generation.

Method used

By obtaining the code files in the code repository, creating the knowledge and dependencies of the code file, including code elements and their inclusion relationships and dependencies between the files, training the code completion model to improve its performance.

Benefits of technology

The accuracy of the code completion model in the completion code task is improved, so that it can learn more fully the knowledge in the code repository and improve the quality of generated code.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120371270A_ABST
    Figure CN120371270A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a code completion method and device, computer equipment and a storage medium, and belongs to the technical field of computers. The method comprises the steps of obtaining at least two code files in a code warehouse; based on the at least two code files, knowledge data and dependency data of the at least two code files are created, the knowledge data comprise code elements in the code files and inclusion relations between the code elements, and the dependency data represent dependency relations between the at least two code files; based on the knowledge data and the dependency data of the at least two code files, a code completion model is trained, and the code completion model is used for completing codes. According to the method and the device, the code completion model can more fully learn knowledge in a code warehouse, the performance of the code completion model is improved, and the accuracy of the code completion model on a code completion task is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present application relate to the field of computer technologies, and particularly to a code completion method, apparatus, computer device, and storage medium. Background Art

[0002] With the development of computer technologies and the wide application of artificial intelligence models, a code completion solution has been proposed currently. It can complete other codes based on existing codes through a code completion model, so as to obtain complete codes, without the need for technicians to write complete codes, improving the code generation speed.

[0003] To ensure the performance of the code completion model, one or more code files can be obtained, and the code completion model can be trained based on the one or more code files, so that the code completion model has the ability to complete other codes based on existing codes. However, the above training process is relatively simple, and even after training, the accuracy of the code completion model is still insufficient. Summary of the Invention

[0004] Embodiments of the present application provide a code completion method, apparatus, computer device, and storage medium, which improve the performance of the code completion model and the accuracy of the code completion model in the code completion task. The technical solutions are as follows:

[0005] On the one hand, a code completion method is provided. The method includes:

[0006] Obtain at least two code files in a code repository;

[0007] Based on the at least two code files, create knowledge data and dependency data of the at least two code files. The knowledge data includes code elements in the code files and the inclusion relationships between the code elements, and the dependency data represents the dependency relationships between the at least two code files;

[0008] Train a code completion model based on the knowledge data and dependency data of the at least two code files. The code completion model is used to complete codes.

[0009] On the other hand, a code completion apparatus is provided. The apparatus includes:

[0010] A file acquisition module, configured to obtain at least two code files in a code repository;

[0011] A data creation module, configured to create knowledge data and dependency data of the at least two code files based on the at least two code files. The knowledge data includes code elements in the code files and the inclusion relationships between the code elements, and the dependency data represents the dependency relationships between the at least two code files;

[0012] A code completion module, configured to train a code completion model based on knowledge data and dependency data of at least two of the code files, where the code completion model is used to complete code.

[0013] In a possible implementation, the data creation module includes:

[0014] A knowledge data creation unit, configured to, for each of the code files: obtain first description data, where the first description data includes at least two code element tags, and the positions of at least two of the code element tags represent an inclusion relationship between code elements corresponding to at least two of the code element tags; obtain the code elements included in the code file; add the obtained code elements to a first target position in the first description data to obtain the knowledge data of the code file, where the first target position is the code element position corresponding to the code element tag matching the code element.

[0015] In a possible implementation, the knowledge data creation unit is configured to convert the code file into a syntax tree, where the syntax tree represents the syntax structure of the code file; and obtain the code elements from the syntax tree.

[0016] In a possible implementation, the data creation module includes:

[0017] A dependency data creation unit, configured to obtain second description data, where the second description data includes at least two code file tags, and the sorting of at least two of the code file tags represents a dependency relationship between code files corresponding to at least two of the code file tags; determine the sorting of at least two of the code files based on other code files that each code file depends on, where the sorting represents a dependency relationship between at least two of the code files; add the code contents of at least two of the code files to second target positions in the second description data to obtain the dependency data, where the second target positions are the code content positions corresponding to the code file tags with the same sorting order as the code files.

[0018] In a possible implementation, the dependency data creation unit is configured to create a dependency relationship graph based on other code files that each code file depends on, where the dependency relationship graph represents a dependency relationship between at least two of the code files; and process the dependency relationship graph using a topological sorting algorithm to obtain the sorting of at least two of the code files.

[0019] In a possible implementation manner, the dependent data creation unit is configured to, for each of the code files: split the code in the code file to obtain a plurality of elements arranged in sequence; add the plurality of elements to the second target position corresponding to the code file in the second description data.

[0020] In a possible implementation manner, the code completion module is configured to train the code completion model based on the knowledge data of at least two of the code files; after training the code completion model based on the knowledge data of at least two of the code files, train the code completion model based on the dependent data.

[0021] In a possible implementation manner, the apparatus further includes:

[0022] A repository screening module, configured to screen code repositories that meet the target quality conditions from a plurality of code repositories, where the target quality conditions include at least one of the following conditions:

[0023] The annotation times of the code repository are not less than the target times, where the annotation times are the times of annotating the code repository as a quality code repository;

[0024] The code repository does not include a first code file, where the first code file is a code file that is not written manually;

[0025] The code repository does not include a second code file, where the second code file is a code file with security vulnerabilities;

[0026] The code repository does not include a third code file, where the third code file is a code file that causes device operation problems;

[0027] The code repository does not include a fourth code file, where the fourth code file is a code file with a risk of information leakage.

[0028] In a possible implementation manner, the code completion module is further configured to, in response to a code completion instruction, obtain a first code to be completed; complete the first code through the code completion model to obtain a second completed code.

[0029] On the other hand, a computer device is provided, where the computer device includes a processor and a memory, and at least one computer program is stored in the memory, and the at least one computer program is loaded and executed by the processor to implement the operations performed by the code completion method as described in the above aspect.

[0030] On the other hand, a computer-readable storage medium is provided. At least one computer program is stored in the computer-readable storage medium and is loaded and executed by a processor to implement the operations performed by the code completion method as described in the above aspect.

[0031] On the other hand, a computer program product is provided, including a computer program that is loaded and executed by a processor to implement the operations performed by the code completion method as described in the above aspect.

[0032] The solution provided by the embodiments of the present application creates knowledge data and dependency data for at least two code files based on at least two code files in a code repository. The knowledge data of the code file can reflect the characteristics of code elements in the code file, and the dependency data can reflect the dependency relationships between different code files. Training a code completion model based on the knowledge data and dependency data of at least two code files can enable the code completion model to learn the knowledge in the code repository more fully, improve the performance of the code completion model, and increase the accuracy of the code completion model in the code completion task. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0034] Figure 1 is a schematic diagram of an implementation environment provided by an embodiment of the present application;

[0035] Figure 2 is a flowchart of a code completion method provided by an embodiment of the present application;

[0036] Figure 3 is a flowchart of another code completion method provided by an embodiment of the present application;

[0037] Figure 4 is a schematic flowchart of a process for screening a code repository provided by an embodiment of the present application;

[0038] Figure 5 is a schematic structural diagram of a code completion device provided by an embodiment of the present application;

[0039] Figure 6 is a schematic structural diagram of another code completion device provided by an embodiment of the present application;

[0040] Figure 7 is a schematic structural diagram of a terminal provided by an embodiment of the present application;

[0041] Figure 8 It is a schematic structural diagram of a server provided by an embodiment of the present application. Detailed implementation manners

[0042] To make the objectives, technical solutions and advantages of the embodiments of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the accompanying drawings.

[0043] It can be understood that the terms "first", "second", etc. used in the present application may be used herein to describe various concepts, but unless otherwise specified, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the present application, the first code file may be referred to as the second code file, and similarly, the second code file may be referred to as the first code file.

[0044] Among them, at least two means two or more. For example, at least two code files may be two code files, three code files, etc., any integer greater than or equal to two. Each refers to each of at least two. For example, each code file refers to each of the at least two code files. If the at least two code files are 3 code files, then each code file refers to each of the 3 code files.

[0045] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data for analysis, stored data, displayed data, input data, etc.) and signals (including but not limited to signals transmitted between user terminals and other devices, etc.) involved in the present application are all fully authorized by users or relevant parties, and the collection, use and processing of relevant data need to comply with relevant laws, regulations and standards of relevant countries and regions.

[0046] Artificial Intelligence (AI) is a theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable the machines to have the functions of perception, reasoning and decision-making.

[0047] Artificial intelligence technology is an interdisciplinary subject that involves a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, pre-trained model technology, operation / interaction systems, mechatronics, etc. Among them, the pre-trained model, also known as the large model or the foundation model, can be widely applied to downstream tasks in various directions of artificial intelligence after fine-tuning. The software technologies of artificial intelligence mainly include computer vision technology, speech processing technology, natural language processing technology, as well as machine learning / deep learning, autonomous driving, intelligent transportation, and several other major directions.

[0048] With the research and progress of artificial intelligence technology, artificial intelligence technology has been studied and applied in multiple fields. For example, common applications include smart homes, smart wearable devices, virtual assistants, smart speakers, intelligent marketing, driverless, autonomous driving, drones, digital twins, virtual humans, robots, artificial intelligence-generated content, conversational interactions, intelligent healthcare, intelligent customer service, game AI, etc. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.

[0049] The pre-trained model (Pre-Training Model, PTM), also known as the foundation model or the large model, refers to a deep neural network (Deep Neural Network, DNN) with a large number of parameters. It is trained on a large amount of unlabeled data, and the function approximation ability of the large-parameter DNN is used to enable the PTM to extract common features from the data. Through techniques such as fine-tuning, parameter-efficient fine-tuning (Parameter-Efficient Fine-Tuning, PEFT), and prompt-tuning, it is suitable for downstream tasks. Therefore, the pre-trained model can achieve ideal results in few-shot or zero-shot scenarios. PTM can be divided into language models, vision models, speech models, multi-modal models, etc. according to the data modalities processed. Among them, the multi-modal model refers to a model that establishes feature representations of two or more data modalities. The pre-trained model is an important tool for outputting artificial intelligence-generated content and can also be used as a general interface connecting multiple specific task models.

[0050] Adaptive computing refers to automatically adjusting the computational amount and precision of the model according to different input data, in order to improve the computational efficiency of the model while maintaining the model's precision. Adaptive computing can flexibly adjust the computational amount and precision of the model on different input data, thereby better balancing the computational efficiency and precision of the model.

[0051] The solution provided in the embodiments of this application relates to artificial intelligence technology, and is specifically described through the following embodiments:

[0052] The code completion method provided by the embodiments of the present application is used in a computer device. Optionally, the computer device is a terminal or a server. Optionally, the terminal is a smart phone, a computer, a laptop, a desktop computer, a smart speaker, a smart watch, a smart voice interaction device, a smart home appliance, a vehicle terminal, an aircraft, etc., but is not limited thereto. Optionally, the server is an independent physical server, or a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. The embodiments of the present application can be applied to various scenarios, including but not limited to cloud technology, artificial intelligence, intelligent transportation, assisted driving, etc.

[0053] In a possible implementation manner, the computer program involved in the embodiments of the present application can be deployed to be executed on a computer device, or on multiple computer devices located at one place, or on multiple computer devices distributed at multiple places and interconnected through a communication network. The multiple computer devices distributed at multiple places and interconnected through a communication network can form a blockchain system.

[0054] In a possible implementation manner, the computer device in the embodiments of the present application is a node in a blockchain system. The node can store a code completion model or a code file in the blockchain, and then the node or other device corresponding nodes in the blockchain can query the code completion model or the code file by accessing the blockchain.

[0055] Figure 1 It is a schematic diagram of an implementation environment provided by the embodiments of the present application. Refer to Figure 1 , this implementation environment includes: a terminal 101 and a server 102, and the terminal 101 and the server 102 are connected through a wired network or a wireless network.

[0056] Based on the code files in the code repository, the server 102 creates knowledge data and dependency data of at least two code files, and thus trains a code completion model based on the knowledge data and dependency data of at least two code files. After a technician writes part of the code in the terminal 101 and triggers a code completion instruction, the written code is sent to the server 102. The server 102 completes the code through the code completion model, sends the completed code to the terminal 101, and the terminal 101 displays the completed code. In this way, the technician only needs to write part of the code instead of writing all the code, which improves the efficiency of generating code.

[0057] The embodiments of the present application are applied to any scenario of developing code.

[0058] For example, in the scenario of developing a game application, by adopting the solution provided by the embodiments of the present application, a code completion model can be trained based on a code repository. After a technician writes part of the code of the game application, the complete code after completion can be obtained through the code completion model.

[0059] Or, in the scenario of developing an instant messaging application, by adopting the solution provided by the embodiments of the present application, a code completion model can be trained based on a code repository. After a technician writes part of the code of the instant messaging application, the complete code after completion can be obtained through the code completion model.

[0060] Figure 2 It is a flowchart of a code completion method provided by the embodiments of the present application. The embodiments of the present application are executed by a computer device, and the computer device is a terminal or a server. Refer to Figure 2 and the method includes:

[0061] 201. The computer device obtains at least two code files in the code repository.

[0062] Among them, the code repository includes at least two code files, and the code files contain code. In a possible implementation manner, the code files include complete code, that is, the code in the code files can implement a function independently without depending on other code files. Or, the code files include incomplete code, and the code files depend on other code files, that is, the code in the code files needs to reference the code in other code files. For example, the code in the code files includes instructions for calling other code files or information such as interfaces provided by other code files. The code in the code files cannot implement a function independently and needs to run together with the code in other dependent code files to implement the function.

[0063] In addition, in addition to code files, the code repository may also include files in other formats. Then, in a possible implementation manner, the computer device filters out code files from the code repository. For example, the computer device pre-determines the target format of the code files, and the suffix name of each file in the code repository represents the format of the file. Then, the suffix names of each file in the code repository are traversed, and the files in the target format are selected, and other format files are no longer selected. Such as the target format includes Java format, exe (executable) format or other formats, etc.

[0064] 202. The computer device creates knowledge data and dependency data of at least two code files based on the at least two code files.

[0065] In the embodiments of the present application, considering that directly using code files as training samples for the code completion model and training the code completion model based on the code files will result in the code completion model not learning the knowledge in the code files sufficiently, thereby leading to poor training effects. Therefore, the computer device first creates knowledge data and dependency data for at least two code files based on the at least two code files, and then trains the code completion model based on the created knowledge data and dependency data. For example, a knowledge graph is created based on at least two code files, and the knowledge graph includes the knowledge data and dependency data of the at least two code files.

[0066] Among them, the knowledge data of the code file includes code elements in the code file and the inclusion relationships between the code elements. Any code file includes at least two code elements, and there may be inclusion relationships between the code elements. For example, some code elements may further include other code elements. In a possible implementation manner, the code in the code file is segmented, and each part obtained by segmentation is used as a code element, or the content representing the name of a certain object in the code file is determined as a code element, that is, the code element is the name of a certain object in the code, and other content related to the object but not belonging to the object name does not need to be considered, such as the operator executed on the object, the call instruction indicating a function to be called, etc.

[0067] For example, in a code file, one or more namespaces are first defined. Each namespace includes one or more classes. Each class includes one or more functions and one or more variables. Each function includes a function description, a function name, a parameter list, and a return value description. The function description can be used as a comment of the function. The parameter list includes at least one parameter in the function. Each parameter includes a variable type and a variable name. The return value description includes a return value type and a return value name. Then the code elements in the knowledge data include: namespaces, classes, functions, variables, function descriptions, function names, parameter lists, return value descriptions, variable types and variable names of parameters, return value descriptions, return value types, and return value names, etc. And the knowledge data also includes the inclusion relationships between the above code elements, such as a namespace includes classes, and a class includes functions, etc. According to the inclusion relationships between the above code elements, it can be known which code elements are included inside each code element.

[0068] In a possible implementation manner, in addition to the code elements in the code file and the inclusion relationships between the code elements, the knowledge data of the code file further includes at least one of the file path of the code file and the syntax tree of the code file.

[0069] In addition, any code file may contain information about other code files, indicating that the operation of this code file depends on other code files. Therefore, dependency data is created based on at least two code files, and the dependency data represents the dependency relationship between at least two code files.

[0070] 203. The computer device trains a code completion model based on the knowledge data and dependency data of at least two code files, and the code completion model is used to complete code.

[0071] Among them, the code completion model is a large language model or other types of models. A large language model is a deep learning model trained based on a large amount of text data, which can generate natural language text, deeply understand the meaning of text, and process various natural language tasks. The computer device uses the knowledge data and dependency data of at least two code files as training samples. In the pre-training stage of the code completion model, the knowledge data and dependency data of at least two code files are input into the code completion model, and the code completion model learns the knowledge data and dependency data of at least two code files, and trains the code completion model one or more times to improve the performance of the code completion model.

[0072] The trained code completion model can be used to complete code. Completing code means completing code according to the code context. For example, after part of the code of the bubble sort algorithm is written, the code completion model can complete other code (italicized code):

[0073]

[0074]

[0075] Compared with directly training a code completion model based on code files, in the embodiments of the present application, the computer device sorts out the code files to obtain structured knowledge data and dependency data. The knowledge data of the code file can reflect the characteristics of code elements in the code file, and the dependency data can reflect the dependency relationship between different code files. Training the code completion model based on the structured knowledge data and dependency data, rather than purely training the code completion model based on code, can enable the code completion model to learn the knowledge in the code repository more fully.

[0076] The solution provided by the embodiments of this application creates knowledge data and dependency data for at least two code files based on the at least two code files in the code repository. The knowledge data of the code file can reflect the characteristics of the code elements in the code file, and the dependency data can reflect the dependency relationships between different code files. Training a code completion model based on the knowledge data and dependency data of at least two code files can enable the code completion model to learn the knowledge in the code repository more fully, improve the performance of the code completion model, and increase the accuracy of the code completion model in the code completion task.

[0077] Based on the above Figure 2 On the basis of the embodiment shown, the embodiments of this application also provide another code completion method. Figure 3 FIG. is a flowchart of another code completion method provided by the embodiments of this application. The embodiments of this application are executed by a computer device, which is a terminal or a server. Refer to Figure 3 and the method includes:

[0078] 301. The computer device obtains at least two code files in the code repository.

[0079] Among them, the process of obtaining the code file is the same as step 201 above and will not be elaborated here.

[0080] In the embodiments of this application, the computer device can access multiple code repositories and obtain the code files in the multiple code repositories. In order to ensure the code quality, the computer device can only use high-quality code repositories and filter out low-quality code repositories.

[0081] Among them, the code repository can be a code repository on Github (a hosting platform for open source and private software projects) or a code repository provided by other platforms, etc.

[0082] In a possible implementation, the computer device screens the code repositories that meet the target quality conditions from multiple code repositories, where the target quality conditions are determined by the computer device according to the quality requirements for the code repositories. The code repositories that meet the target quality conditions are the code repositories whose quality meets the requirements of the computer device. Among them, the target quality conditions include at least one of the following conditions:

[0083] 1. The annotation times of the code repository are not less than the target times, and the annotation times are the times of annotating the code repository as a quality code repository.

[0084] For any code repository, if any user recognizes that the quality of the code repository meets the requirements, the code repository can be annotated as a quality code repository, that is, a high-quality code repository. Then the annotation times of the code repository can represent the degree of recognition of the code repository by users and can represent whether the quality of the code repository meets the requirements.

[0085] Therefore, if the annotation times of the code repository are not less than the target times, the code repository can be used to train the code completion model. If the annotation times of the code repository are less than the target times, it indicates that the quality of the code repository is poor, and the code repository will be filtered out and no longer used to train the code completion model.

[0086] In a possible implementation, the operation of labeling a code repository as a quality code repository may include the operation of giving a like to the code repository, in which case the annotation times are the number of likes. Or, the operation of labeling a code repository as a quality code repository may include the operation of adding a star to the code repository, in which case the annotation times are the number of stars obtained by the code repository, etc.

[0087] 2. The code repository does not include a first code file, where the first code file is a code file that is not manually written.

[0088] A manually written code file is a code file written by a technician based on their own experience and knowledge, and has relatively high quality. A code file that is not manually written, that is, an automatically generated code file, may be a code file generated by an artificial intelligence model or a code file written using a code writing tool, etc. Such code files have strong randomness and it is difficult to guarantee their quality. Therefore, the computer device requires that the code repository does not include the first code file. That is, if the code repository does not include the first code file, the code repository can be used to train the code completion model. If the code repository includes the first code file, the code repository will be filtered out and no longer used to train the code completion model.

[0089] Exemplarily, a code file that is not manually written contains a non-manual label, which indicates that the code file is not manually written. For example, the non-manual label is a "automatically generated code" label, etc. By identifying the label contained in the code file, it can be determined whether the code file is a file that is not manually written.

[0090] 3. The code repository does not include a second code file, where the second code file is a code file with security vulnerabilities.

[0091] The second code file has security vulnerabilities and will cause security problems when running, such as the theft of accounts or passwords. To prevent the code completion model from learning incorrect knowledge, the computer device requires that the code repository does not include the second code file. That is, if the code repository does not include the second code file, the code repository can be used to train the code completion model. If the code repository includes the second code file, the code repository will be filtered out and no longer used to train the code completion model.

[0092] Exemplarily, the code file with security vulnerabilities contains a security vulnerability label, and this security vulnerability label indicates that the code file has security vulnerabilities. Then, by identifying the label contained in the code file, it can be determined whether the code file has security vulnerabilities. Alternatively, the computer device pre-determines the conditions that the code file with security vulnerabilities satisfies, and based on this condition, it detects the code file to determine whether the code file has security vulnerabilities.

[0093] 4. The code repository does not include a third code file, where the third code file is a code file that causes problems in the device operation.

[0094] When the third code file runs, it will cause problems in the device operation, such as the device freezing or lagging, etc. To prevent the code completion model from learning incorrect knowledge, the computer device requires that the code repository does not include the third code file. That is, if the code repository does not include the third code file, then this code repository can be used to train the code completion model; while if the code repository includes the third code file, then this code repository is filtered out and is no longer used to train the code completion model.

[0095] Exemplarily, the code file that causes problems in the device operation contains a running problem label, and this running problem label indicates that the code file will cause problems in the device operation. Then, by identifying the label contained in the code file, it can be determined whether the code file will cause problems in the device operation. Alternatively, the computer device pre-determines the conditions that the code file that causes problems in the device operation satisfies, and based on this condition, it detects the code file to determine whether the code file will cause problems in the device operation.

[0096] 5. The code repository does not include a fourth code file, where the fourth code file is a code file with a risk of information leakage.

[0097] The fourth code file has a risk of information leakage and may leak information during operation, such as leaking the password of the account, the personal identity information of the account, etc. To prevent the code completion model from learning incorrect knowledge, the computer device requires that the code repository does not include the fourth code file. That is, if the code repository does not include the fourth code file, then this code repository can be used to train the code completion model; while if the code repository includes the fourth code file, then this code repository is filtered out and is no longer used to train the code completion model.

[0098] Exemplarily, the code file with a risk of information leakage contains an information leakage label, and this information leakage label indicates that the code file has a risk of information leakage. Then, by identifying the label contained in the code file, it can be determined whether the code file has a risk of information leakage. Alternatively, the computer device pre-determines the conditions that the code file with a risk of information leakage satisfies, and based on this condition, it detects the code file to determine whether the code file has a risk of information leakage.

[0099] In a possible implementation, after the computer device determines the code repository to be used, it formats the code repository so that the format of the code in the processed code repository is unified into a fixed format, facilitating the subsequent creation of knowledge data and dependency data based on the fixed format of the code files.

[0100] The above several target quality conditions can be combined in any way. For example, see Figure 4 The embodiment of the present application provides an exemplary code repository screening process. The screening process includes: First, perform a quality analysis on the code repository to detect whether the annotation times of the code repository are not less than the target times and do not include code files that are not manually written. If the quality meets the requirements, perform a vulnerability detection on the code repository to detect whether there are security vulnerabilities in the code files in the code repository. If there are security vulnerabilities in the code files in the code repository, filter out the code repository. If there are no security vulnerabilities in the code files in the code repository, perform a problem detection on the code repository to detect whether the code files in the code repository cause device operation problems. If the code files in the code repository cause device operation problems, filter out the code repository. If the code files in the code repository do not cause device operation problems, perform an information security detection on the code repository to detect whether there is a risk of information leakage in the code files in the code repository. If there is a risk of information leakage in the code files in the code repository, filter out the code repository. If there is no risk of information leakage in the code files in the code repository, perform a formatting process on the code repository to obtain a code repository with a standardized format, and this code repository is a high-quality code repository.

[0101] 302. For each code file, the computer device obtains first description data.

[0102] Among them, the first description data is used to describe the style of the knowledge data of the code file and can be regarded as a kind of template knowledge data. By adding the relevant knowledge of the code file to the first description data, the knowledge data of the code file can be obtained.

[0103] The first description data includes at least two types of code element tags. Each type of code element tag represents a code element, and different types of code element tags represent different types of code elements. For example, the code element tag <namespace>represents a namespace, while the code element tag <function>Represents a function. And the positions of at least two code element tags represent the inclusion relationship between the code elements corresponding to at least two code element tags. For example, if code element tag 1 is on the line above code element tag 2, and the indentation distance of code element tag 2 is greater than that of code element tag 1, it means that the code element corresponding to code element tag 1 includes the code element corresponding to code element tag 2. Or, the code element tags corresponding to a code element include a start tag and an end tag, and the position between the start tag and the end tag is used to describe the code element. Then, the position between the start tag and the end tag is the code element position corresponding to the code element tag and is used to store the code element. So, if the start tag and end tag of code element 2 are both between the start tag and end tag of code element 1, it means that code element 1 includes code element 2.

[0104] In addition, the first description data can be a file in XML (eXtensible Markup Language) format or a file in other formats, and the embodiments of the present application do not limit this.

[0105] 303. The computer device obtains the code elements included in the code file.

[0106] The code elements in the code file are the same as the code elements in step 202 above and will not be elaborated here.

[0107] In a possible implementation, the computer device converts the code file into a syntax tree and obtains the code elements from the syntax tree. Exemplarily, a code parser is used to perform syntax analysis on the code file to obtain the syntax tree of the code file, such as an AST (Abstract Syntax Tree). Since the syntax tree represents the syntax structure of the code file and can clearly and intuitively represent each code element and the relationship between each code element, obtaining code elements from the syntax tree can speed up the processing speed.

[0108] 304. The computer device adds the obtained code elements to the first target position in the first description data to obtain the knowledge data of the code file, and the first target position is the code element position corresponding to the code element tag that matches the code element.

[0109] The first description data includes at least two code element tags, and each code element tag corresponds to a code element position. This code element position is used to add the code element corresponding to the code element tag. Therefore, after the computer device obtains the code elements from the code file, it adds the code element to the code element position corresponding to the code element tag that matches the code element, that is, the first target position corresponding to the code element.

[0110] In a possible implementation, different code element tags are located on different lines. In each line, the position after the position of the code element tag is the position of the code element corresponding to the code element tag. Alternatively, each type of code element tag includes two code element tags. The first code element tag is the start tag, and the second code element tag is the end tag. These two code element tags can be on the same line or different lines, and the position between these two code element tags is the position of the code element corresponding to the code element tag, which is used to store the code element corresponding to these two code element tags. These two code element tags can be the same tag, or among these two code element tags, the second code element tag includes the first code element tag and an end identifier, indicating that the second code element tag is the end tag corresponding to the first code element tag. Among them, the end identifier can be a certain punctuation mark or a certain letter, etc.

[0111] Among them, in the first description data, the first target positions corresponding to different code elements are different. By adding each code element in the code file to the corresponding position after the first target position respectively, the knowledge data of the code file can be obtained. In a possible implementation, the computer device obtains one code element from the code file each time, adds the code element to the corresponding first target position, and then obtains another code element from the code file and adds the other code element to the corresponding first target position until each code element in the code file is added. After that, the first description data is determined as the knowledge data of the code file.

[0112] For example, the knowledge data of the code file is as follows:

[0113]

[0114]

[0115] In the above knowledge data, <repoknowledge>Start tag for knowledge data,< / repoknowledge> is the end tag of the knowledge data, <repoknowledge>and< / repoknowledge> and what is between them is the knowledge data, <file>Start tag for code file,< / file> is the end tag of the code file, <file>and< / file> Between them is the knowledge data of a code file, and the tags of other code elements are the same. From the above knowledge data, it can be seen that the path of the code file is repo1 / file1.java. The code file contains the namespace namespace_A. The namespace namespace_A contains class_A. Class_A includes a function. The function is used to obtain the student name. The function name is "getStudentName". In the parameter list of the function, the type of the first parameter is int, and the value is id. The return value type of the function is String, and the return value name is name.

[0116] 305. The computer device obtains the second description data.

[0117] Among them, the second description data is used to describe the style of the dependency data and can be regarded as a kind of template dependency data. By adding the relevant knowledge of the code file on the basis of the second description data, the dependency data between at least two code files can be obtained.

[0118] The second description data includes at least two code file tags. Each code file tag represents a code file, and the sorting of at least two code file tags represents the dependency relationship between the code files corresponding to at least two code file tags. For example, if the code file tag 1 is before the code file tag 2, it means that the code file corresponding to the code file tag 2 depends on the code file corresponding to the code file tag 1.

[0119] In addition, the second description data can be a file in XML format or other formats, and the embodiments of the present application do not limit this.

[0120] In a possible implementation manner, if the first description data and the second description data are located in the same description file, after adding the relevant knowledge of the code file to the description file based on the code file, the description file can include the knowledge data and dependency data of the code file. Exemplarily, in the same description file, if the first description data is before the second description data, when training the code completion model based on this file, the code completion model can be trained in the order of knowledge data first and dependency data second.

[0121] For example, the description file is as follows:

[0122]

[0123] Among them, <repo>Start tag for knowledge,< / repo> is the end tag of the knowledge, <repo>and< / repo> between them is the content of the knowledge, <repoknowledge>Start tag for knowledge data,< / repoknowledge> is the end tag of the knowledge data, <repoknowledge>and< / repoknowledge> between them is the knowledge data, <repodepedencydata>Start tag for dependency data,< / repodepedencydata> is the end tag for dependent data, <repodepedencydata>and< / repodepedencydata> The data in between is the dependent data.

[0124] 306. The computer device determines the sorting of at least two code files based on the other code files that each code file depends on, and this sorting represents the dependency relationship between the at least two code files.

[0125] In a possible implementation, the computer device creates a dependency graph based on the other code files that each code file in the code repository depends on. The dependency graph represents the dependency relationship between at least two code files, and a topological sorting algorithm is used to process the dependency graph to obtain the sorting of at least two code files.

[0126] For example, nodes corresponding to each code file are created, and according to the other code files that the code file depends on, a directed line is added between the node corresponding to the code file and the nodes corresponding to the other code files it depends on, thereby obtaining a dependency graph. The directed line points from the node corresponding to the code file to the nodes corresponding to the other code files that the code file depends on, or the directed line points from the node corresponding to the code file to the nodes corresponding to the other code files that depend on the code file. And the topological sorting algorithm is used to perform topological sorting on the dependency graph, so as to arrange the code files corresponding to each node in the dependency graph into a linear sequence, thereby determining the sorting of at least two code files.

[0127] 307. The computer device adds the code contents of at least two code files to the second target positions in the second description data respectively to obtain dependent data, and the second target positions are the code content positions corresponding to the code file labels with the same sorting positions as the code files.

[0128] The second description data includes at least two code file labels, and each code file label corresponds to a code content position, which is used to add the code content in the code file corresponding to the code file label. Moreover, the at least two code file labels in the second description data are arranged in order, and the at least two code files in the code repository are also arranged in order. Then, in order to reflect the dependency relationship between the code files in the second description data, the code content of the code file is added to the second target position corresponding to the code file in the second description data, so that in the obtained dependent data, the positions of the code contents in the at least two code files match the dependency relationship between the at least two code files, so that the dependency relationship between the code contents can be determined according to the positions of the code contents.

[0129] In a possible implementation, different code file tags are located on different lines. In each line, the position after the position of the code file tag is the position of the code content corresponding to the code file tag. Alternatively, each type of code file tag includes two code file tags. The first code file tag is the start tag, and the second code file tag is the end tag. These two code file tags can be on the same line or different lines, and the position between these two code file tags is the position of the code content corresponding to the code file tag, which is used to store the code content of the code file corresponding to these two code file tags. These two code file tags can be the same tag, or among these two code file tags, the second code file tag includes the first code file tag and an end identifier, indicating that the second code file tag is the end tag corresponding to the first code file tag. Among them, the end identifier can be a certain punctuation mark or a certain letter, etc.

[0130] Among them, in the second description data, the second target positions corresponding to different code files are different. By adding the code content of the code files to the positions after the corresponding second target positions respectively, dependent data can be obtained. In a possible implementation, the computer device adds the code content of the code file ranked first to the second target position corresponding to the first code file tag in the second description data according to the sorting of at least two code files, and then adds the code content of the code file ranked second to the second target position corresponding to the second code file tag in the second description data, until the code content of the code file ranked last is added to the second target position corresponding to the last code file tag in the second description data, and then determines the second description data as the dependent data.

[0131] In a possible implementation, adding the code content of at least two code files to the second target positions in the second description data to obtain dependent data includes: for each code file, splitting the code in the code file to obtain multiple elements arranged in order, and adding the multiple elements to the second target position corresponding to the code file in the second description data.

[0132] Among them, the code content of the code file is different from the style of the original code in the code file. The code has been split in the code content, and the obtained elements are independent of each other, which is convenient for the code completion model to learn one or more elements separately each time training is performed, ensuring the diversity of learned knowledge and improving the performance of the code completion model.

[0133] Exemplarily, the code in the code file is split at the positions of spaces and punctuation marks, and each part obtained by the splitting is used as an element, thereby obtaining multiple elements. These multiple elements are added together to the second target position, indicating that these multiple elements constitute a piece of code.

[0134] For example, the dependency data is as follows:

[0135]

[0136] In the above dependency data, <repodepedencydata>Start tag for dependency data,< / repodepedencydata> is the end tag of the dependency data, <repodepedencydata>and< / repodepedencydata> and the data between them is the dependency data. <filetokens>Indicates a code file, and the code files arranged later depend on the code files arranged earlier. <path>Start tag for file path,< / path> Is the end tag for the file path. <tokens>Is an element label, indicating an element separated out in the code file, then multiple <tokens>Composed the code content of the code file.

[0137] 308. The computer device trains a code completion model based on the knowledge data of at least two code files. After training the code completion model based on the knowledge data of at least two code files, it trains the code completion model based on the dependency data.

[0138] Since the knowledge data of the code file only gives the defined code elements and the structural relationships between the code elements, but cannot reflect what kind of processing these code elements need to perform. And the dependency data reflects the code content of each code file and the dependency relationships between these code contents. Especially in the code repository, the dependency chain is very long and the dependency relationships are relatively complex. Therefore, it is necessary to first train the code completion model based on the knowledge data of at least two code files so that the code completion model can first learn the defined code elements and the structural relationships between the code elements, and then train the code completion model based on the dependency data so that the code completion model can learn the processing logic and dependency relationships of the learned code elements. In this way, the learning of the code completion model is more solid and the training effect is better.

[0139] In a possible implementation, training the code completion model based on the knowledge data of the code file includes: dividing the code elements of any code file in the knowledge data into first code elements and second code elements, removing the second code elements from the knowledge data, passing the knowledge data after removing the second code elements through the code completion model to predict the third code elements that need to be supplemented in the knowledge data, and adjusting the model parameters of the code completion model based on the difference information between the second code elements and the third code elements to improve the accuracy of the code completion model in completing code elements.

[0140] In another possible implementation, training the code completion model based on the dependency data includes: dividing the code content of any code file in the dependency data into first code content and second code content, removing the second code content from the dependency data, passing the dependency data after removing the second code content through the code completion model to predict the third code content that needs to be supplemented in the dependency data, and adjusting the model parameters of the code completion model based on the difference information between the second code element content and the third code content to improve the accuracy of the code completion model in completing code content.

[0141] In another possible implementation, considering that the training process of the code completion model will include multiple iterative trainings, then in each iterative training, it will first train the code completion model based on the knowledge data of at least two code files, and then train the code completion model based on the dependency data, so that the code completion model can fully learn the knowledge in the knowledge data and the dependency data.

[0142] It should be noted that, in the embodiment of the present application, it is taken as an example that steps 305-307 are executed after steps 302-304, that is, after creating the knowledge data of at least two code files, the dependency data is created. And it is taken as an example that step 308 is executed after steps 302-307, that is, after creating the knowledge data and dependency data of at least two code files, the code completion model is trained.

[0143] In another embodiment, steps 305-307 may also be executed before steps 302-304, or steps 305-307 and steps 302-304 may be executed in parallel. And after executing steps 302-304, regardless of whether steps 305-307 have been executed, the code completion model can be trained based on the knowledge data of at least two code files. In the case where steps 305-307 have been executed and the code completion model has been trained based on the knowledge data of at least two code files, the code completion model can be trained based on the dependency data. The embodiment of the present application does not limit the timing relationship between the above steps.

[0144] In another embodiment, before training the code completion model based on the knowledge data and dependency data of at least two code files, the code completion model is obtained first. The code completion model is an initialized model, or it can also be a model that has been trained. For example, other devices send a code completion model that has been trained once or multiple times to the computer device, and the computer device trains the code completion model using the method of the embodiment of the present application. Or, the computer device first trains the code completion model based on the code files in the code repository, and then continues to train the code completion model using the method provided by the embodiment of the present application after training to improve the accuracy of the code completion model.

[0145] 309. The computer device responds to the code completion instruction, obtains the first code to be completed, and completes the first code through the code completion model to obtain the completed second code.

[0146] Among them, the first code is a part of the code for implementing any target function. After the first code is completed through the code completion model, the obtained second code includes the first code and the code supplemented by the code completion model, and the second code can implement the target function of the first code.

[0147] In a possible implementation manner, the code completion model is used to process the prompt word, and the prompt word represents the task issued to the code completion model. The computer device creates a prompt word, and the prompt word includes the first code and the task text, and the task text is used to indicate to complete the first code. The prompt word is input into the code completion model, and the code completion model processes the prompt word, and then the first code can be completed according to the indication of the task text to obtain the second code. For example, the prompt word is as follows:

[0148]

[0149] In a possible implementation, the computer device is a server, and the terminal is installed with an application for generating code. The application is associated with the server and is an IDE (Integrated Development Environment) plugin, such as Vs code (Visual studio Code, a software developer tool) or Jetbrains series (a series of tools developed by the company named Jetbrains), etc. After writing code in the application, a technician triggers a code completion instruction, and sends the code completion instruction to the server through the application. The code completion instruction carries the first code written by the technician. In response to the code completion instruction, the server obtains the first code, completes the first code through a code completion model to obtain a second code, and returns the second code to the terminal. The terminal displays the second code through the application, thereby assisting developers in writing code and improving the coding efficiency of developers. Among them, the code completion instruction can be triggered by an operation of clicking the enter key, or an operation of clicking the completion button, or triggered by other operations.

[0150] In some other embodiments, the computer device is a server for implementing a question-and-answer function. The code completion model is a question-and-answer model, which has the function of completing code and may also have other question-and-answer functions. The terminal displays a question-and-answer window, inputs the written first code, and in response to the sending operation of the first code, displays a question including the first code in the question-and-answer window. The question is used to indicate the completion of the first code. The terminal sends a question-and-answer request to the server, and the question-and-answer request includes the question. After receiving the question, the server performs semantic analysis on the question through the code completion model, determines that the task to be executed is to complete the first code, so it will complete the first code to obtain a second code, generate an answer including the second code, and return it to the terminal. The terminal displays the answer in the question-and-answer window. Exemplarily, the question includes the first code and a task text, and the task text is used to indicate the completion of the first code. The server forms a prompt word with the first code and the task text, inputs the prompt word into the code completion model, and processes the prompt word through the code completion model, then the first code can be completed according to the indication of the task text to obtain a second code. Among them, the question-and-answer window can be displayed through the application for generating code. For example, the terminal displays a code writing window through the application, and in response to the opening operation of the sidebar, the question-and-answer window is displayed in the sidebar area of the current code writing window, and enters the chat mode, so as to display questions and answers through the question-and-answer window.

[0151] The solution provided by the embodiments of this application creates the knowledge data and dependency data of at least two code files based on at least two code files in a code repository. The knowledge data of a code file can reflect the characteristics of code elements in the code file, and the dependency data can reflect the dependency relationships between different code files. Training a code completion model based on the knowledge data and dependency data of at least two code files enables the code completion model to learn the knowledge in the code repository more fully, improves the performance of the code completion model, and increases the accuracy of the code completion model in the code completion task. Therefore, when using the code completion model to complete the first code, the accuracy of the completed second code is ensured.

[0152] In addition, the embodiments of this application set the first description data, and then create the knowledge data of the code file based on the first description data. Since the code element tags are already set in the first description data and the positions of the code element tags can represent the inclusion relationships between code elements, only by adding the code elements in the code file to the appropriate positions in the first description data can the knowledge data of the code file be obtained, saving workload and thus improving the operation efficiency.

[0153] In addition, convert the code file into a syntax tree, which can accurately represent the syntax structure of the code file. Then, code elements can be conveniently and quickly obtained from the syntax tree, avoiding interference from other content in the code file and improving the operation accuracy and efficiency.

[0154] In addition, the embodiments of this application set the second description data, and then create the dependency data based on the second description data. Since the code file tags are already set in the second description data and the sorting of the code file tags can represent the dependency relationships between code files, only by adding the code content in the code file to the appropriate positions in the second description data can the dependency data be obtained, saving workload and thus improving the operation efficiency.

[0155] In addition, create a dependency graph based on the other code files that each code file depends on, and use a topological sorting algorithm to process the dependency graph to obtain the sorting of at least two code files, so that the sorting of at least two code files can accurately represent the dependency relationships between at least two code files, improving the operation accuracy.

[0156] In addition, split the code in the code file to obtain multiple elements arranged in order, and add the multiple elements as the code content of the code file to the second description data to obtain the dependency data. Since the elements obtained by splitting are independent of each other, it is convenient for the code completion model to learn one or more elements separately, ensuring the diversity of the learned knowledge and improving the performance of the code completion model.

[0157] In addition, first train a code completion model based on the knowledge data of at least two code files so that the code completion model first learns the defined code elements and the structural relationships between the code elements, and then train the code completion model based on the dependency data so that the code completion model learns the processing logic and dependency relationships of the learned code elements. In this way, the learning of the code completion model is more solid and the training effect is better.

[0158] In addition, from multiple code repositories, screen the code repositories that meet the target quality conditions, and use the code repositories that meet the target quality conditions to train the code completion model, avoiding the code completion model from learning incorrect knowledge and improving the accuracy of the code completion model.

[0159] Through experimental verification using a code repository, extract a part of the code files from the code repository, remove part of the code in the extracted code files, and use the remaining code as the code to be completed, and the removed code as the code to be supplemented for the code to be completed, thereby constituting a verification dataset. The verification dataset includes at least one verification data, and each verification data includes the code to be completed and the code to be supplemented. And the verification dataset can include various completion types such as function signature completion, function parameter completion, function body completion, control flow statement completion, assignment statement completion, and function call statement completion. And use the remaining part of the code files in the code repository as the training dataset.

[0160] For the initial code completion model 1, train the code completion model 1 based on the code files in the training dataset to obtain the code completion model 2. And create knowledge data and dependency data based on the code files in the training dataset, and then train the code completion model 1 based on the knowledge data and dependency data to obtain the code completion model 3. Then, through the code completion model 1, the code completion model 2, and the code completion model 3, verify the verification dataset respectively, and determine the accuracy of each code completion model as shown in Table 1 below.

[0161] Table 1

[0162] Code completion model Code completion model 1 Code completion model 2 Code completion model 3 Accuracy 27.5% 37.5% 62.5%

[0163] As can be seen from Table 1, although compared with the code completion model 1, the accuracy of the code completion model 2 that directly learns the code files in the code repository has increased, but the increase range is not large. While the accuracy of the code completion model 3 that has learned the knowledge data and dependency data has increased significantly.

[0164] Figure 5 It is a schematic structural diagram of a code completion device provided by an embodiment of the present application. Refer to Figure 5 and this device includes:

[0165] A file acquisition module 501, configured to acquire at least two code files from a code repository;

[0166] A data creation module 502, configured to create knowledge data and dependency data for at least two code files based on the at least two code files, where the knowledge data includes code elements in the code files and the inclusion relationships between the code elements, and the dependency data represents the dependency relationships between the at least two code files;

[0167] A code completion module 503, configured to train a code completion model based on the knowledge data and dependency data of at least two code files, where the code completion model is used to complete code.

[0168] In a possible implementation, refer to Figure 6 , the data creation module 502 includes:

[0169] A knowledge data creation unit 5021, configured to, for each code file: acquire first description data, where the first description data includes at least two code element tags, and the positions of the at least two code element tags represent the inclusion relationships between the code elements corresponding to the at least two code element tags; acquire the code elements included in the code file; add the acquired code elements to a first target position in the first description data to obtain the knowledge data of the code file, where the first target position is the code element position corresponding to the code element tag that matches the code element.

[0170] In a possible implementation, refer to Figure 6 , the knowledge data creation unit 5021 is configured to convert the code file into a syntax tree, where the syntax tree represents the syntax structure of the code file; and acquire code elements from the syntax tree.

[0171] In a possible implementation, refer to Figure 6 , the data creation module 502 includes:

[0172] A dependency data creation unit 5022, configured to acquire second description data, where the second description data includes at least two code file tags, and the sorting of the at least two code file tags represents the dependency relationships between the code files corresponding to the at least two code file tags; determine the sorting of the at least two code files based on the other code files that each code file depends on, where the sorting represents the dependency relationships between the at least two code files; add the code contents of the at least two code files to second target positions in the second description data to obtain the dependency data, where the second target positions are the code content positions corresponding to the code file tags with the same sorting order as the code files.

[0173] In a possible implementation manner, the dependency data creation unit 5022 is configured to create a dependency relationship graph based on other code files upon which each code file depends, where the dependency relationship graph represents the dependency relationship between at least two code files; and process the dependency relationship graph by using a topological sorting algorithm to obtain the sorting of at least two code files.

[0174] In a possible implementation manner, for each code file, the dependency data creation unit 5022 is configured to: split the code in the code file to obtain a plurality of elements arranged in sequence; and add the plurality of elements to a second target position corresponding to the code file in the second description data.

[0175] In a possible implementation manner, the code completion module 503 is configured to train a code completion model based on the knowledge data of at least two code files; and after training the code completion model based on the knowledge data of at least two code files, train the code completion model based on the dependency data.

[0176] In a possible implementation manner, refer to Figure 6 , the apparatus further includes:

[0177] A repository screening module 504, configured to screen code repositories that meet target quality conditions from a plurality of code repositories, where the target quality conditions include at least one of the following conditions:

[0178] The annotation times of the code repository are not less than the target times, where the annotation times are the times of annotating the code repository as a quality code repository;

[0179] The code repository does not include a first code file, where the first code file is a code file not written manually;

[0180] The code repository does not include a second code file, where the second code file is a code file with security vulnerabilities;

[0181] The code repository does not include a third code file, where the third code file is a code file that causes device operation problems;

[0182] The code repository does not include a fourth code file, where the fourth code file is a code file with information leakage risks.

[0183] In a possible implementation manner, the code completion module 503 is further configured to, in response to a code completion instruction, obtain a first code to be completed; and complete the first code through the code completion model to obtain a second completed code.

[0184] It should be noted that: the code completion device provided in the above embodiments is only illustrated by dividing the above-mentioned functional modules. In practical applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the computer device is divided into different functional modules to complete all or part of the functions described above. In addition, the code completion device provided in the above embodiments and the embodiments of the code completion method belong to the same concept. For the specific implementation process, please refer to the method embodiments, which will not be elaborated here.

[0185] An embodiment of the present application also provides a computer device, which includes a processor and a memory. At least one computer program is stored in the memory, and the at least one computer program is loaded and executed by the processor to implement the operations performed in the code completion method of the above embodiments.

[0186] Optionally, the computer device is provided as a terminal. Figure 7 The structural schematic diagram of a terminal 700 provided by an exemplary embodiment of the present application is shown.

[0187] The terminal 700 includes: a processor 701 and a memory 702.

[0188] The processor 701 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 701 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field Programmable Gate Array), and PLA (Programmable Logic Array). The processor 701 may also include a main processor and a coprocessor. The main processor is a processor used to process data in the wake state, also known as the CPU (Central Processing Unit); the coprocessor is a low-power processor used to process data in the standby state. In a possible implementation manner, the processor 701 may be integrated with a GPU (Graphics Processing Unit), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 701 may also include an AI (Artificial Intelligence) processor, and the AI processor is used to process computing operations related to machine learning.

[0189] The memory 702 may include one or more computer-readable storage media, which may be non-transitory. The memory 702 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash storage devices. In a possible implementation, the non-transitory computer-readable storage medium in the memory 702 is used to store at least one computer program, and the at least one computer program is used to be possessed by the processor 701 to implement the code completion method provided in the method embodiments of this application.

[0190] In a possible implementation, the terminal 700 may also optionally include: a peripheral device interface 703 and at least one peripheral device. The processor 701, the memory 702, and the peripheral device interface 703 may be connected by a bus or signal lines. Each peripheral device may be connected to the peripheral device interface 703 through a bus, signal lines, or a circuit board. Optionally, the peripheral device includes at least one of a radio frequency circuit 704, a display screen 705, a camera assembly 706, and a power supply 707.

[0191] The peripheral device interface 703 can be used to connect at least one peripheral device related to I / O (Input / Output) to the processor 701 and the memory 702. In a possible implementation, the processor 701, the memory 702, and the peripheral device interface 703 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 701, the memory 702, and the peripheral device interface 703 can be implemented on a separate chip or circuit board, and this embodiment does not limit this.

[0192] The radio frequency circuit 704 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The radio frequency circuit 704 communicates with a communication network and other communication devices through electromagnetic signals. The radio frequency circuit 704 converts an electrical signal into an electromagnetic signal for transmission, or converts the received electromagnetic signal into an electrical signal. Optionally, the radio frequency circuit 704 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a subscriber identity module card, and so on. The radio frequency circuit 704 can communicate with other devices through at least one wireless communication protocol. The wireless communication protocol includes but is not limited to: a metropolitan area network, each generation of mobile communication networks (2G, 3G, 4G, and 5G), a wireless local area network, and / or a WiFi (Wireless Fidelity) network. In a possible implementation, the radio frequency circuit 704 may also include a circuit related to NFC (Near Field Communication), and this application does not limit this.

[0193] The display screen 705 is used to display the UI (User Interface). The UI may include graphics, text, icons, videos, and any combination thereof. When the display screen 705 is a touch display screen, the display screen 705 also has the ability to collect touch signals on or above the surface of the display screen 705. The touch signals can be input as control signals to the processor 701 for processing. At this time, the display screen 705 can also be used to provide virtual buttons and / or virtual keyboards, also known as soft buttons and / or soft keyboards. In a possible implementation, there can be one display screen 705, which is set on the front panel of the terminal 700; in some other embodiments, there can be at least two display screens 705, which are respectively set on different surfaces of the terminal 700 or in a folding design; in some other embodiments, the display screen 705 can be a flexible display screen, which is set on the curved surface or folding surface of the terminal 700. Even, the display screen 705 can also be set to an irregular non-rectangular shape, that is, a special-shaped screen. The display screen 705 can be prepared using materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).

[0194] The camera module 706 is used to collect images or videos. Optionally, the camera module 706 includes a front camera and a rear camera. The front camera is set on the front panel of the terminal 700, and the rear camera is set on the back of the terminal 700. In a possible implementation, there are at least two rear cameras, which are respectively any one of a main camera, a depth-of-field camera, a wide-angle camera, and a telephoto camera, so as to realize the function of background blurring by fusing the main camera and the depth-of-field camera, the function of panoramic shooting by fusing the main camera and the wide-angle camera, and the VR (Virtual Reality) shooting function or other fusion shooting functions. In a possible implementation, the camera module 706 can also include a flash. The flash can be a single-color-temperature flash or a two-color-temperature flash. The two-color-temperature flash refers to the combination of a warm-light flash and a cold-light flash, which can be used for light compensation under different color temperatures.

[0195] The power supply 707 is used to supply power to each component in the terminal 700. The power supply 707 can be alternating current, direct current, a disposable battery, or a rechargeable battery. When the power supply 707 includes a rechargeable battery, the rechargeable battery can support wired charging or wireless charging. The rechargeable battery can also be used to support fast charging technology.

[0196] Those skilled in the art can understand, Figure 7 The structure shown does not constitute a limitation on the terminal 700, and may include more or fewer components than shown, or combine certain components, or adopt a different component arrangement.

[0197] Optionally, the computer device is provided as a server. Figure 8 It is a schematic structural diagram of a server provided by an embodiment of the present application. The server 800 may vary greatly due to different configurations or performances, and may include one or more processors (Central Processing Units, CPUs) 801 and one or more memories 802. Among them, at least one computer program is stored in the memory 802, and the at least one computer program is loaded and executed by the processor 801 to implement the methods provided by the above various method embodiments. Of course, the server may also have components such as wired or wireless network interfaces, keyboards, and input / output interfaces for input / output. The server may also include other components for implementing device functions, which will not be elaborated here.

[0198] An embodiment of the present application also provides a computer-readable storage medium, in which at least one computer program is stored, and the at least one computer program is loaded and executed by a processor to implement the operations performed by the code completion method in the above embodiment.

[0199] An embodiment of the present application also provides a computer program product, including a computer program, and the computer program is loaded and executed by a processor to implement the operations performed by the code completion method in the above embodiment.

[0200] Those of ordinary skill in the art can understand that all or part of the steps for implementing the above embodiments can be completed by hardware, or can be completed by a program instructing relevant hardware. The program can be stored in a computer-readable storage medium, and the above-mentioned storage medium can be a read-only memory, a disk, an optical disc, etc.

[0201] The above are only optional embodiments of the embodiments of the present application, and are not intended to limit the embodiments of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the embodiments of the present application shall be included in the protection scope of the present application.< / tokens> < / tokens> < / filetokens> < / function> < / namespace>

Claims

1. A code completion method, characterized in that, The method includes: Obtaining at least two code files in a code repository; Based on the at least two code files, creating knowledge data and dependency data for the at least two code files, where the knowledge data includes code elements in the code files and the inclusion relationships between the code elements, and the dependency data represents the dependency relationships between the at least two code files; Training a code completion model based on the knowledge data and dependency data of the at least two code files, where the code completion model is used to complete code.

2. The method according to claim 1, characterized in that, Creating knowledge data for the at least two code files based on the at least two code files, including: For each of the code files: Obtaining first description data, where the first description data includes at least two code element tags, and the positions of the at least two code element tags represent the inclusion relationships between the code elements corresponding to the at least two code element tags; Obtaining the code elements included in the code file; Adding the obtained code elements to a first target position in the first description data to obtain the knowledge data of the code file, where the first target position is the code element position corresponding to the code element tag that matches the code element.

3. The method according to claim 2, characterized in that, The obtaining the code elements included in the code file includes: Converting the code file into a syntax tree, where the syntax tree represents the syntax structure of the code file; Obtaining the code elements from the syntax tree.

4. The method according to claim 2, wherein Creating the dependency data based on the at least two code files includes: Obtaining second description data, where the second description data includes at least two code file tags, and the sorting of the at least two code file tags represents the dependency relationships between the code files corresponding to the at least two code file tags; Determining the sorting of the at least two code files based on the other code files that each code file depends on, where the sorting represents the dependency relationships between the at least two code files; Adding the code contents of the at least two code files to second target positions in the second description data to obtain the dependency data, where the second target positions are the code content positions corresponding to the code file tags with the same sorting order as the code files.

5. The method according to claim 4, characterized in that, The determining the sorting of the at least two code files based on the other code files that each code file depends on includes: Creating a dependency graph based on the other code files that each code file depends on, where the dependency graph represents the dependency relationships between the at least two code files; Processing the dependency graph using a topological sorting algorithm to obtain the sorting of the at least two code files.

6. The method according to claim 4, wherein The adding the code contents of the at least two code files to second target positions in the second description data to obtain the dependency data includes: For each of the code files: Splitting the code in the code file to obtain a plurality of elements arranged in order; Adding the plurality of elements to the second target position corresponding to the code file in the second description data.

7. The method according to claim 1, wherein Training a code completion model based on the knowledge data and dependency data of at least two of the code files includes: Training the code completion model based on the knowledge data of at least two of the code files; After training the code completion model based on the knowledge data of at least two of the code files, training the code completion model based on the dependency data.

8. The method according to claim 1, characterized in that, The method further includes: Screening code repositories that meet the target quality conditions from multiple code repositories, where the target quality conditions include at least one of the following conditions: The number of times the code repository is annotated is not less than the target number of times, and the number of times the code repository is annotated is the number of times the code repository is annotated as a quality code repository; The code repository does not include a first code file, and the first code file is a code file that is not written manually; The code repository does not include a second code file, and the second code file is a code file with security vulnerabilities; The code repository does not include a third code file, and the third code file is a code file that causes device operation problems; The code repository does not include a fourth code file, and the fourth code file is a code file with a risk of information leakage.

9. The method according to any one of claims 1-8, characterized in that, The method further includes: In response to a code completion instruction, obtaining a first code to be completed; Completing the first code through the code completion model to obtain a second completed code.

10. A code completion device, characterized in that, The device includes: A file acquisition module for acquiring at least two code files in a code repository; A data creation module for creating knowledge data and dependency data of at least two of the code files based on at least two of the code files, where the knowledge data includes code elements in the code files and the inclusion relationships between the code elements, and the dependency data represents the dependency relationships between at least two of the code files; A code completion module for training a code completion model based on the knowledge data and dependency data of at least two of the code files, and the code completion model is used to complete code.

11. A computer device, characterized in that, The computer device includes a processor and a memory, and at least one computer program is stored in the memory. The at least one computer program is loaded and executed by the processor to implement the operations performed by the code completion method according to any one of claims 1 to 9.

12. A computer-readable storage medium, characterized in that, At least one computer program is stored in the computer-readable storage medium. The at least one computer program is loaded and executed by the processor to implement the operations performed by the code completion method according to any one of claims 1 to 9.

13. A computer program product, comprising a computer program, characterized in that, The computer program is loaded and executed by the processor to implement the operations performed by the code completion method according to any one of claims 1 to 9.