Code editing method and device, electronic equipment, medium and program product
By obtaining the target code and its most recently edited code snippet pair, predicting based on the code editing model and annotating the next edited code snippet, the problem of high inference learning cost of code editing model is solved, and more efficient code editing prediction is achieved.
Patent Information
- Application Number
- CN202510192019.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-20
- Publication Date
- 2025-05-27
AI Technical Summary
The reasoning learning cost of code editing models is high and lacks the ability to refuse to know, which leads to excessive amount of information processed by the model when actually developing predictions, which increases the difficulty of learning.
By obtaining the object code and its most recently edited code snippet pair, predict and annotate the next edited code snippet based on the code editing model. This method uses code snippet pairs for training, and the training data of specific constructs is better derivatable and has higher processing efficiency.
It reduces the inference learning cost of code editing models, improves the efficiency and accuracy of the model in code editing prediction, reduces the amount of information processed by the model, and reduces the difficulty of learning.
Smart Images

Figure CN120045172A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of large model technology, and in particular to a code editing method, device, electronic device, medium and program product. Background Art
[0002] Code editing recommendations are designed to provide intelligent assistance to developers in the process of modifying code. In the process of developers modifying code, they can identify potential modification points by analyzing the structure and context of the code and give corresponding code modification suggestions. Code editing recommendations can not only speed up development efficiency, but also reduce modification omissions and ensure code consistency.
[0003] In related technologies, a general large model is usually used to implement code editing. Since the general large model has a large number of parameters, the inference cost is relatively high. On the other hand, since the model has no rejection ability, a large number of edit points do not need to trigger recommendations during actual development prediction. When the entire submitted code data is input as an edit behavior to the large model, it may contain a large number of code changes, which will cause the model to process a large amount of information and increase the difficulty of learning. Summary of the invention
[0004] In view of this, the present disclosure provides a code editing method, device, electronic device, medium and program product to solve the problem of high reasoning learning cost of code editing models.
[0005] In a first aspect, the present disclosure provides a code editing method, the method comprising:
[0006] Obtaining a target code and a pair of code snippets last edited in the target code, wherein the pair of code snippets last edited includes a code snippet before the last edit and a code snippet after the last edit;
[0007] Based on the code editing model, the target code and the most recently edited code snippet pair, the code snippet to be edited next in the target code is predicted and annotated, the code snippet to be edited next includes the code snippet before the next edit and the code snippet after the next edit, the code editing model is obtained based on target sample training, and the target sample is determined based on the code snippet pair.
[0008] In a second aspect, the present disclosure provides a code editing device, the device comprising:
[0009] A code acquisition module, used to acquire a target code and a pair of code snippets last edited in the target code, wherein the pair of code snippets last edited includes a code snippet before the last edit and a code snippet after the last edit;
[0010] A code editing prediction module is used to predict and mark the code snippet to be edited next in the target code based on a code editing model, the target code and the most recently edited code snippet pair, wherein the code snippet to be edited next includes a code snippet before the next edit and a code snippet after the next edit, wherein the code editing model is obtained based on target sample training, and the target sample is determined based on the code snippet pair.
[0011] In a third aspect, the present disclosure provides an electronic device, comprising: a memory and a processor, the memory and the processor are communicatively connected to each other, computer instructions are stored in the memory, and the processor executes the code editing method of the first aspect or any corresponding embodiment thereof by executing the computer instructions.
[0012] In a fourth aspect, the present disclosure provides a computer-readable storage medium having computer instructions stored thereon, the computer instructions being used to enable a computer to execute the code editing method of the first aspect or any corresponding embodiment thereof.
[0013] In a fifth aspect, the present disclosure provides a computer program product, including computer instructions, which are used to enable a computer to execute the code editing method of the first aspect or any corresponding embodiment thereof.
[0014] The code editing method provided by the disclosed embodiment includes obtaining a target code and the most recent code snippet pair in the target code, including the code snippet before the most recent edit and the code snippet after the edit; based on a code editing model, the target code and the most recent edited code snippet pair, predicting and marking the code snippet to be edited next time in the target code, the code snippet to be edited next time includes the code snippet before the next edit and the code snippet after the edit, the code editing model is obtained through training based on a target sample, and the target sample is determined based on the code snippet pair. The method predicts code editing through the trained code editing model, and the code editing model is trained using code snippet pairs, and the training data of a specific structure is more derivable and more efficient to process. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] In order to more clearly illustrate the specific embodiments of the present disclosure or the technical solutions in the prior art, the drawings required for use in the specific embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0016] Figure 1 is a flowchart of a code editing method according to an embodiment of the present disclosure;
[0017] Figure 2 is a flowchart of a method for determining a target sample according to an embodiment of the present disclosure;
[0018] Figure 3 is a schematic diagram of a method for determining a target sample according to an embodiment of the present disclosure;
[0019] Figure 4 is a schematic diagram of a training process of a code editing model according to an embodiment of the present disclosure;
[0020] Figure 5 is a structural block diagram of a code editing device according to an embodiment of the present disclosure;
[0021] Figure 6 It is a schematic diagram of the hardware structure of the electronic device according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0022] In order to make the purpose, technical solution and advantages of the embodiments of the present disclosure clearer, the technical solution in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are part of the embodiments of the present disclosure, rather than all the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present disclosure.
[0023] It is understandable that before using the technical solutions disclosed in the embodiments of the present disclosure, the types, scope of use, usage scenarios, etc. of the personal information involved in the present disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.
[0024] For example, in response to receiving an active request from a user, a prompt message is sent to the user to clearly prompt the user that the operation requested to be performed will require obtaining and using the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, application, server, or storage medium that performs the operation of the technical solution of the present disclosure according to the prompt message.
[0025] As an optional but non-limiting implementation, in response to receiving an active request from the user, the prompt information may be sent to the user in the form of a pop-up window, in which the prompt information may be presented in text form. In addition, the pop-up window may also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0026] It is understandable that the above notification and the process of obtaining user authorization are merely illustrative and do not constitute a limitation on the implementation of the present disclosure. Other methods that meet the relevant laws and regulations may also be applied to the implementation of the present disclosure.
[0027] It is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and relevant provisions.
[0028] Code editing recommendation is a function that assists developers in modifying related codes during the process of modifying codes. By analyzing the structure and context of the code, as well as the user's current editing behavior, the large model can identify potential modification points and provide modification suggestions. This function can speed up development efficiency, reduce modification omissions, and ensure code consistency. In related technologies, code editing often relies on large general models, but this brings high reasoning costs, and the model uses a large amount of data in the process of processing code, which increases the complexity of the learning task and affects processing efficiency. Based on this, the embodiment of the present disclosure provides a code editing method that can be used in code editing scenarios. The code language can be SQL code or others, and there is no specific limitation.
[0029] According to an embodiment of the present disclosure, a code editing method embodiment is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0030] In this embodiment, a code editing method is provided, which can be used in terminals such as computers and tablet computers. Figure 1 is a flowchart of a code editing method according to an embodiment of the present disclosure, such as Figure 1 As shown, the process includes the following steps:
[0031] Step S101, obtaining a target code and a pair of code snippets most recently edited in the target code.
[0032] The most recently edited code snippet pair includes a code snippet before the most recently edited code snippet and a code snippet after the most recently edited code snippet.
[0033] Get a specific piece of code from the version control system, code editor, or code submitted by the user. The target code can be the complete code of the entire project, which is the complete code in the editor every time the user triggers a model call during the editing process.
[0034] During the development process, the user may need to edit the target code. The most recently edited code snippet pair corresponds to the most recently edited code operation. Any code that was modified in the target code during the most recent code editing and the content before the code was modified, that is, the code snippet pair includes the code snippet before the most recent edit and the code snippet after the most recent edit. In order to avoid positioning errors, the code snippet includes the modified code line and the context of the code line. The context range can be limited according to needs. For example, the code snippet can be limited to include the modified code line and the two lines of code above and below the code line. The form of the code snippet pair is not limited. For any code snippet pair, the code snippet before editing and the code snippet after editing will be clearly stated.
[0035] Step S102 : predicting and marking the next edited code snippet in the target code based on the code editing model, the target code and the pair of the most recently edited code snippets.
[0036] The code snippet to be edited next time includes the code snippet before the next editing and the code snippet after the next editing. The code editing model is obtained by training based on target samples, and the target samples are determined based on the code snippet pairs.
[0037] The code editing model is a pre-trained model that can learn patterns from historical code edits and predict future edits. The target samples used for code editing model training are determined based on a large number of code snippet pairs that show the changes in the code between different versions.
[0038] The code editing model infers the next editing behavior based on the input target code and the most recently edited code snippet pair, and then outputs the next editing behavior. The specific manifestation of outputting the next editing behavior is to mark the next edited code snippet pair in the target code, including the predicted code snippet before the next edit and the code snippet after the next edit.
[0039] The format of the code snippet pair for the most recent edit is the same as that for the next edit. Specifically, in units of lines, if the user leaves the current line or moves the cursor, it is considered the end of an editing behavior, such as modifying variable names, modifying expressions, adding fields to a group by statement, etc.
[0040] As an example, the embodiment applicable to the present disclosure is a scenario where the user modifies the code. At this time, the code has been written and the user updates it according to the needs. When the user edits any part of the code, the model automatically triggers the code editing function, marks other places that may need to be modified, and gives modification prompts. In this process, after the editor captures the user's editing action, it can automatically input the complete code to the model, and the model outputs the code snippet for the next edit.
[0041] The code editing method provided by the disclosed embodiment includes obtaining a target code and the most recent code snippet pair in the target code, including the code snippet before the most recent edit and the code snippet after the edit; based on a code editing model, the target code and the most recent edited code snippet pair, predicting and marking the code snippet to be edited next time in the target code, the code snippet to be edited next time includes the code snippet before the next edit and the code snippet after the edit, the code editing model is obtained through training based on a target sample, and the target sample is determined based on the code snippet pair. The method predicts code editing through the trained code editing model, and the code editing model is trained using code snippet pairs, and the training data of a specific structure is more derivable and more efficient to process.
[0042] In this embodiment, a method for determining a target sample is provided, which can be used in terminals such as computers and tablet computers. Figure 2 is a flow chart of a method for determining a target sample according to an embodiment of the present disclosure, such as Figure 2 As shown, the process includes the following steps:
[0043] Step S201: extracting code editing samples based on project logs.
[0044] The code editing sample includes at least one set of code snippet pairs. The project log records various editing behaviors in the development process, such as code addition, deletion, and modification. The extracted code editing sample includes at least one set of code snippet pairs, and each set of code snippet pairs includes a code snippet before editing and a code snippet after editing.
[0045] In software development, the differences between two versions of code files are compared to understand the differences between the two versions of code. This comparison can be called "diff". In this step, the specific differences can be extracted from the two versions of the code and presented in a form that is easy to view. The form and the code snippet are not limited to the method of extracting code editing samples. As an example, the dynamic programming method can be used to calculate the longest common subsequence method to calculate the difference between the two versions of the code, so as to obtain the line position of the edited code and the specific content of the code. In addition, the line position also includes one or more lines above and below the edited code, which is convenient for locating the specific position.
[0046] Step S202 , obtaining similarities between pairs of code snippets of adjacent editing behaviors, and grouping the pairs of code snippets based on the similarities to obtain grouping results.
[0047] The similarity between adjacent editing behaviors is measured, and the code snippet pairs are grouped according to the similarity. The similarity between code snippet pairs is mainly to compare the similarity between the difference snippets corresponding to each code snippet pair. The difference snippet corresponding to each code snippet pair is the difference between the code snippet before and after editing in one editing. The similarity includes completely consistent or different, and different includes not completely consistent and completely different.
[0048] Furthermore, the above step S202 includes:
[0049] Step S2021: determining a difference code corresponding to the code snippet pair based on the code snippet before editing and the code snippet after editing in the code snippet pair.
[0050] The difference code corresponding to each pair of code snippets is extracted. The difference code is the code portion introduced or deleted by the editing behavior, reflecting the specific changes between two versions of any code after being edited.
[0051] Step S2022: determine whether the difference codes corresponding to the code snippet pairs of adjacent editing behaviors are the same, so as to determine the similarity between the code snippet pairs.
[0052] After the difference codes are extracted, the similarity between the difference codes corresponding to the code fragments of adjacent editing behaviors is determined, and the similarity is determined based on the content of the difference codes.
[0053] Step S2023: if the difference codes corresponding to the code segment pairs of adjacent editing behaviors are the same, the code segment pairs of adjacent editing behaviors are divided into a first preset sample group.
[0054] If the content of the difference code is exactly the same, it means that the difference codes corresponding to the code fragment pairs of adjacent editing behaviors are the same. It should be noted that when comparing the difference codes, it is only necessary to compare whether the characters are consistent, and there is no need to consider the position of the difference code. The code fragment pairs with exactly the same difference codes are divided into a first preset sample group, and there are multiple groups of code fragment pairs with exactly the same difference codes in the first preset sample group, and each group of code fragment pairs with exactly the same difference codes can be two or more.
[0055] Step S2024: if the edit distance between the difference codes corresponding to the code snippet pairs of adjacent editing behaviors is less than a preset threshold, the code snippet pairs of adjacent editing behaviors are divided into a second preset sample group.
[0056] When judging, the difference codes can be determined to be the same by calculating the edit distance. The edit distance refers to the minimum number of editing operations (such as inserting, deleting, and replacing characters) required to convert one string into another. If there are code snippet pairs whose corresponding difference codes are not exactly the same, but the edit distance is lower than the preset threshold, it can be considered that the corresponding code snippet pairs are related to a certain extent, with high similarity but not exactly the same. The code snippet pairs with high difference code similarity but not exactly the same are divided into the second preset sample group. There is a certain amount of noise in the second preset sample group, which can cover more relevant situations.
[0057] Step S203: determining the degree of association between the code snippet pairs based on the grouping result to determine the target sample.
[0058] In the above grouping process, the code snippet pairs are divided based on the exactness of the difference codes, the similarity of the edit distance, or other semantic similarity. The degree of association in this step indicates the logical, functional, or editing behavior association between the code snippet pairs, and the degree of association is specifically evaluated based on the grouping results. A large model can be used as a discriminator to determine whether there is an association between the last edited code snippet pair and the next edited code snippet pair. The target sample refers to the data set selected from the code snippet pairs for subsequent model training.
[0059] In some optional implementations, the above step S203 includes:
[0060] Step S2031 , determining the association degree between the code snippet pairs in the second preset sample group, and determining the associated code snippet pairs as first target samples, and determining the unassociated code snippet pairs as second target samples.
[0061] There are multiple groups of code snippet pairs in the second preset sample group, each group of code snippet pairs may include two or more code snippet pairs, and the difference codes of each group of code snippet pairs may be similar but not identical or completely different. Based on the large model discriminator, the degree of association between the code snippet pairs in each group in the second preset sample group is measured, and the associated code snippet pairs are determined as the first target samples. The existence of association means that there is a certain degree of similarity between the difference codes corresponding to the code snippet pairs, but they are not completely identical. The code snippet pairs that are not associated are determined as the second target samples. The code snippet pairs that are not associated are samples that have not passed the discriminator. The similarity of the difference codes corresponding to these code snippet pairs is low as judged by the edit distance.
[0062] Step S2032: determining a trigger training sample based on the first target sample and the code snippet pairs in the first preset sample group.
[0063] The difference codes of each code segment pair in the first preset sample group are completely the same, so the code segment pairs in the first preset sample group can be considered to be related. The difference codes of each code segment pair in the first target sample are similar, and can be considered to have a certain correlation.
[0064] Specifically, step S2032 includes: assembling the first target sample and the code snippet pairs in the first preset sample group based on a preset format to determine a trigger training sample.
[0065] All relevant code snippet pairs can be assembled according to a preset format to obtain trigger training samples. The assembled trigger training samples include background code (background_code), the most recent edit (last_edit), and the expected next edit response (next_edit in response) to provide clear structured input for model training.
[0066] Step S2033: determining a rejection training sample based on the second target sample, and determining a target sample according to the triggering training sample and the rejection training sample.
[0067] The rejection training samples are composed of the second target samples, which are composed of pairs of code snippets that fail to pass the evaluation of the association degree discriminator and are determined to be unrelated. There is usually no significant similarity between the code snippet pairs in logic, function or editing behavior.
[0068] Unrelated code snippet pairs are assembled according to a preset format, and the assembled content is used as a rejection training sample to train the model to identify unrelated editing behaviors, or as a negative sample to enhance the generalization ability of the model. The target sample consists of trigger training samples and rejection training samples.
[0069] like Figure 3 The schematic diagram of the method for determining the target sample shown in the figure is a process for constructing the target sample for training the code editing model based on the commit log. The commit log is a file in the version control system that records the history of code changes. Whenever a user makes a change to the code base and commits it, a commit log is generated. Code snippet pairs are extracted from the log, and the code snippet pairs include the code snippet before editing and the code snippet after editing in one edit. The code snippet pairs are grouped according to the similarity between the difference codes, and a correlation discriminator is further used to determine whether there is a correlation between the code snippet pairs in adjacent editing behaviors. The code snippet pairs that are not correlated and have not passed the discriminator are used as rejection training samples, and the code snippets that are correlated are used as trigger training samples.
[0070] The target sample determination method provided in this embodiment generates training data by an automated training data construction method, and the training data is more derivable, which is conducive to model convergence. On the other hand, by adding rejection training samples, the model can learn more information about data distribution, which helps the model make more reasonable judgments when facing new data or unknown categories, thereby improving its generalization ability.
[0071] In some optional implementations, the method for determining the code editing model includes: training the code editing model based on the target sample to determine the code editing model, wherein the number of trigger training samples and rejection training samples in the target sample is set according to a preset ratio.
[0072] The training process of the code editing model is as follows Figure 4 As shown, the training process includes large model pre-training, continued pre-training, task pre-training and fine-tuning. Among them, large model pre-training is to obtain a base large language model, and continued pre-training is based on the large language model using domain corpus for training, so that the model can learn specific programming language syntax. Continued pre-training is an unsupervised learning process. Taking SQL language as an example, a large amount of high-quality SQL code can be used for training. Based on the pre-trained model and a large amount of SQL code, the model is continued to be pre-trained using the next word prediction method to obtain a pre-trained model in this field. This stage aims to let the model learn the SQL programming language syntax, and the SQL code running online can be selected as the training data set for the continued pre-training stage.
[0073] Task pre-training is the continued pre-training based on prompts. In this stage, prompt-based corpus is used for training. Specifically, the trigger training samples and rejection training samples in the above implementation mode can be used to form target samples. The model is pre-trained using the next word prediction method. During the process, the data ratio of trigger training samples and rejection training samples can be adjusted to ensure that the comprehensive ability of the model reaches the optimal level.
[0074] Finally, fine-tuning is performed. During the fine-tuning stage, labeled data, that is, high-quality manually annotated data, is used for training to further improve model performance.
[0075] As a specific application example of an embodiment of the present invention, a code editing method is provided, which can be used to edit SQL code. Specifically, during the code editing process, the code is displayed on the page. After editing one of the codes (the most recent edit), the code editing prediction of the code editing model is triggered, and other codes associated with the most recently edited code are highlighted, and the recommended modified content is displayed in the blank space. The user can choose how to modify according to the needs. In this scenario, if there are multiple identical codes that need to be modified, after modifying one, all other codes that need to be modified and the modified content are highlighted to avoid omissions.
[0076] In this embodiment, a code editing device is also provided, which is used to implement the above-mentioned embodiments and preferred implementation modes, and the descriptions that have been made will not be repeated. As used below, the term "module" can implement a combination of software and / or hardware of a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, the implementation of hardware, or a combination of software and hardware, is also possible and conceivable.
[0077] This embodiment provides a code editing device, such as Figure 5 As shown, including:
[0078] A code acquisition module 501 is used to acquire a target code and a pair of code snippets last edited in the target code, wherein the pair of code snippets last edited includes a code snippet before the last edit and a code snippet after the last edit;
[0079] The code editing prediction module 502 is used to predict and mark the code snippet to be edited next in the target code based on the code editing model, the target code and the most recently edited code snippet pair, wherein the code snippet to be edited next includes the code snippet before the next edit and the code snippet after the next edit, wherein the code editing model is obtained based on target sample training, and the target sample is determined based on the code snippet pair.
[0080] In some optional embodiments, the device includes a target sample determination module, including:
[0081] A sample extraction unit, configured to extract a code editing sample based on a project log, wherein the code editing sample includes at least one set of code snippet pairs;
[0082] A similarity acquisition unit, used to acquire the similarity between the code snippet pairs of adjacent editing behaviors, and group the code snippet pairs based on the similarity to obtain a grouping result;
[0083] The target sample determining unit is used to determine the degree of association between the code snippet pairs based on the grouping result to determine the target sample.
[0084] In some optional implementations, the similarity acquisition unit includes:
[0085] a difference code determination subunit, configured to determine a difference code corresponding to the code fragment pair based on the code fragment before editing and the code fragment after editing in the code fragment pair;
[0086] The similarity determination subunit is used to determine whether the difference codes corresponding to the code fragment pairs of adjacent editing behaviors are the same, so as to determine the similarity between the code fragment pairs.
[0087] In some optional implementations, the similarity acquisition unit includes:
[0088] A first grouping subunit, configured to group the code fragment pairs of adjacent editing behaviors into a first preset sample group if the difference codes corresponding to the code fragment pairs of adjacent editing behaviors are the same;
[0089] The second grouping subunit is configured to group the pair of code segments of the adjacent editing behaviors into a second preset sample group if the editing distance between the difference codes corresponding to the pair of code segments of the adjacent editing behaviors is less than a preset threshold.
[0090] In some optional implementations, the target sample determination unit includes:
[0091] a correlation degree determination subunit, configured to determine the correlation degree between the code snippet pairs in the second preset sample group, and determine the associated code snippet pairs as first target samples, and determine the unassociated code snippet pairs as second target samples;
[0092] a trigger training sample determination subunit, configured to determine a trigger training sample based on the first target sample and the code snippet pairs in the first preset sample group;
[0093] The rejection training sample determination subunit is configured to determine a rejection training sample based on the second target sample, and determine a target sample according to the trigger training sample and the rejection training sample.
[0094] In some optional implementations, the trigger training sample determining subunit is used to assemble the first target sample and the code snippet pairs in the first preset sample group based on a preset format to determine the trigger training sample.
[0095] In some optional embodiments, the device further comprises:
[0096] The model training module is used to train the code editing model based on the target sample to determine the code editing model, and the number of trigger training samples and rejection training samples in the target sample is set according to a preset ratio.
[0097] The further functional description of each of the above modules and units is the same as that of the above corresponding embodiments and will not be repeated here.
[0098] The code editing device in this embodiment is presented in the form of a functional unit, where the unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and memory that executes one or more software or fixed programs, and / or other devices that can provide the above functions.
[0099] The present disclosure also provides an electronic device having the above Figure 5 The code editing device shown.
[0100] See also Figure 6 , Figure 6 is a schematic diagram of the structure of an electronic device provided by an optional embodiment of the present disclosure, such as Figure 6 As shown, the electronic device includes: one or more processors 10, a memory 20, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. The various components are connected to each other using different buses for communication, and can be installed on a common mainboard or installed in other ways as needed. The processor can process instructions executed in the electronic device, including instructions stored in or on the memory to display graphical information of the GUI on an external input / output device (such as a display device coupled to the interface). In some optional embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple electronic devices can be connected, and each device provides some necessary operations (for example, as a server array, a group of blade servers, or a multi-processor system). Figure 6 A processor 10 is taken as an example.
[0101] The processor 10 may be a central processing unit, a network processor or a combination thereof. The processor 10 may further include a hardware chip. The hardware chip may be a dedicated integrated circuit, a programmable logic device or a combination thereof. The programmable logic device may be a complex programmable logic device, a field programmable gate array, a general purpose array logic or any combination thereof.
[0102] The memory 20 stores instructions executable by at least one processor 10, so that the at least one processor 10 executes the method shown in the above embodiment.
[0103] The memory 20 may include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function; the data storage area may store data created according to the use of the electronic device, etc. In addition, the memory 20 may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some optional embodiments, the memory 20 may optionally include a memory remotely arranged relative to the processor 10, and these remote memories may be connected to the electronic device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0104] The memory 20 may include a volatile memory, such as a random access memory; the memory may also include a non-volatile memory, such as a flash memory, a hard disk or a solid state drive; the memory 20 may also include a combination of the above types of memory.
[0105] The electronic device further comprises a communication interface 30 for the electronic device to communicate with other devices or a communication network.
[0106] The embodiments of the present disclosure also provide a computer-readable storage medium. The above-mentioned method according to the embodiments of the present disclosure can be implemented in hardware, firmware, or can be implemented as a computer code that can be recorded in a storage medium, or can be implemented as a computer code that is originally stored in a remote storage medium or a non-temporary machine-readable storage medium and will be stored in a local storage medium and downloaded through a network, so that the method described herein can be stored in such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only storage memory, a random access memory, a flash memory, a hard disk or a solid-state drive, etc.; further, the storage medium can also include a combination of the above-mentioned types of memory. It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component that can store or receive software or computer code. When the software or computer code is accessed and executed by a computer, a processor, or hardware, the method shown in the above embodiment is implemented.
[0107] Although the embodiments of the present disclosure have been described in conjunction with the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present disclosure, and such modifications and variations are all within the scope defined by the appended claims.
Claims
1. A code editing method, characterized in that: The method comprises: Obtaining a target code and a pair of code snippets last edited in the target code, wherein the pair of code snippets last edited includes a code snippet before the last edit and a code snippet after the last edit; Based on the code editing model, the target code and the most recently edited code snippet pair, the code snippet to be edited next in the target code is predicted and annotated. The code snippet to be edited next includes the code snippet before the next edit and the code snippet after the next edit. The code editing model is obtained by training based on the target sample, and the target sample is determined based on the code snippet pair.
2. The method according to claim 1, characterized in that The method for determining the target sample includes: Extracting a code editing sample based on a project log, wherein the code editing sample includes at least one set of code snippet pairs; Acquire the similarity between the code snippet pairs of adjacent editing behaviors, and group the code snippet pairs based on the similarity to obtain a grouping result; The association degree between the code snippet pairs is determined based on the grouping result to determine a target sample.
3. The method according to claim 2, characterized in that The obtaining the similarity between the code snippet pairs of adjacent editing behaviors includes: Determine a difference code corresponding to the code snippet pair based on the code snippet before editing and the code snippet after editing in the code snippet pair; It is determined whether the difference codes corresponding to the pair of code snippets of adjacent editing behaviors are the same, so as to determine the similarity between the pair of code snippets.
4. The method according to claim 3, characterized in that The grouping the code snippet pairs based on the similarity to obtain a grouping result includes: If the difference codes corresponding to the code fragment pairs of the adjacent editing behaviors are the same, classifying the code fragment pairs of the adjacent editing behaviors into a first preset sample group; If the edit distance between the difference codes corresponding to the pair of code segments of the adjacent editing behaviors is less than a preset threshold, the pair of code segments of the adjacent editing behaviors is divided into a second preset sample group.
5. The method according to claim 4, characterized in that The determining the degree of association between the code snippet pairs based on the grouping results to determine the target sample includes: Determining the degree of association between the code snippet pairs in the second preset sample group, and determining the associated code snippet pairs as first target samples, and determining the unassociated code snippet pairs as second target samples; Determine a trigger training sample based on the first target sample and the code snippet pairs in the first preset sample group; A rejection training sample is determined based on the second target sample, and a target sample is determined according to the trigger training sample and the rejection training sample.
6. The method according to claim 5, characterized in that The determining a trigger training sample based on the first target sample and the code snippet pair in the first preset sample group includes: The first target sample and the code snippet pairs in the first preset sample group are assembled based on a preset format to determine a triggering training sample.
7. The method according to claim 5, characterized in that The method for determining the code editing model includes: The code editing model is trained based on the target sample to determine the code editing model, and the number of the trigger training samples and the rejection training samples in the target sample is set according to a preset ratio.
8. A code editing device, characterized in that: The device comprises: A code acquisition module, used to acquire a target code and a pair of code snippets last edited in the target code, wherein the pair of code snippets last edited includes a code snippet before the last edit and a code snippet after the last edit; A code editing prediction module is used to predict and mark the code snippet to be edited next in the target code based on a code editing model, the target code and the most recently edited code snippet pair, wherein the code snippet to be edited next includes a code snippet before the next edit and a code snippet after the next edit, wherein the code editing model is obtained based on target sample training, and the target sample is determined based on the code snippet pair.
9. An electronic device, characterized in that: include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the code editing method according to any one of claims 1 to 7 by executing the computer instructions.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the code editing method according to any one of claims 1 to 7.
11. A computer program product, characterized in that The method comprises computer instructions, wherein the computer instructions are used to cause a computer to execute the code editing method according to any one of claims 1 to 7.