Method, device and equipment for identifying error code suggestions
By using a unified differential format file and a comparison method to generate candidate code suggestions in the code review scenario of AI big model, identifying and intercepting error code suggestions, the false positive phenomenon is solved and the reliability and credibility of code review is improved.
Patent Information
- Application Number
- CN202510025254.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-07
- Publication Date
- 2025-05-30
AI Technical Summary
In the code review scenario of AI big model, false positives are more serious, resulting in incorrect code suggestions being displayed to users, affecting the reliability and credibility of code review.
By obtaining the unified difference format file, input it into a pre-trained large model to generate candidate code suggestions, and compare each set of candidate code suggestions with each paragraph in the unified difference format file to identify and intercept the wrong code suggestions.
It improves the false positive interception rate in code detection, reduces the display of error code suggestions, and enhances the reliability and credibility of code reviews.
Smart Images

Figure CN120066930A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and more particularly, to a method, apparatus, and device for identifying error code suggestions. Background Art
[0002] During the software development process, an AI large model is often used to conduct code reviews on the modified code and display the review results to the user. However, in the large-scale use of the AI large model, there will be hallucination problems, resulting in false alarms. For example, incorrect code suggestions are displayed to the user. Due to the particularity of the code review scenario of the AI large model, there are higher requirements for reliability and credibility, and it is difficult to tolerate the hallucination problems of the AI large model.
[0003] In the prior art, methods such as adding longer contexts to the large model, using the reflection mechanism of the large model, and using vector recall to achieve interception are used to intercept incorrect code suggestions. However, adding longer contexts will affect the concentration of the large model, and the false alarm interception rates of the reflection mechanism of the large model and using vector recall to achieve interception are also relatively low.
[0004] In summary, how to improve the false alarm interception rate during code detection is a problem that needs to be solved currently. Summary of the Invention
[0005] In view of this, embodiments of the present invention provide a method, apparatus, and device for identifying error code suggestions, which can intercept incorrect code suggestions during code detection, thereby improving the false alarm interception rate.
[0006] In a first aspect, an embodiment of the present invention provides a method for identifying error code suggestions, the method comprising:
[0007] Obtain a unified diff file, where the unified diff file represents the differences between a first file and a second file, and the first file and the second file each include multiple pieces of code;
[0008] Input the unified diff file into a pre-trained large model to generate at least one set of candidate code suggestions, where each set of candidate code suggestions includes multiple pieces of code;
[0009] Compare each set of candidate code suggestions with each paragraph in the unified diff file one by one, where each paragraph includes the modified code and the corresponding context of a set number of lines;
[0010] In response to any set of candidate code suggestions successfully matching any paragraph in the unified diff file, determine the any set of candidate code suggestions as incorrect code suggestions.
[0011] Optionally, the method further includes:
[0012] Deleting the error code suggestions from the candidate code suggestions to generate target code suggestions.
[0013] Optionally, the method further includes:
[0014] Outputting the target code suggestions.
[0015] Optionally, the method further includes:
[0016] Obtaining a first file, where the first file is an original code file.
[0017] Optionally, the method further includes:
[0018] Obtaining a second file, where the second file is a code file generated after modifying the first file.
[0019] Optionally, the method further includes:
[0020] Generating the unified diff file according to the first file and the second file.
[0021] Optionally, the step of comparing each group of the candidate code suggestions with each paragraph in the unified diff file specifically includes:
[0022] Comparing each line of code in each group of the candidate code suggestions with each line of code in each paragraph except for the set modification method.
[0023] Optionally, the set modification method is deletion.
[0024] In a second aspect, an embodiment of the present invention provides a device for identifying error code suggestions, where the device includes:
[0025] An obtaining unit, configured to obtain a unified diff file, where the unified diff file represents the difference between a first file and a second file, and the first file and the second file each include multiple lines of code;
[0026] A generating unit, configured to input the unified diff file into a pre-trained large model to generate at least one group of candidate code suggestions, where each group of the candidate code suggestions includes multiple lines of code;
[0027] A comparing unit, configured to compare each group of the candidate code suggestions with each paragraph in the unified diff file one by one, where each paragraph includes the modified code and the corresponding context of the set number of lines;
[0028] A determination unit, in response to any group of the candidate code suggestions being successfully compared with any paragraph in the unified diff format file, is configured to determine the any group of the candidate code suggestions as error code suggestions.
[0029] Optionally, the generating unit is further configured to:
[0030] Delete the error code suggestions from the candidate code suggestions to generate target code suggestions.
[0031] Optionally, the apparatus further includes:
[0032] An output unit, configured to output the target code suggestions.
[0033] Optionally, the obtaining unit is further configured to:
[0034] Obtain a first file, where the first file is an original code file.
[0035] Optionally, the obtaining unit is further configured to:
[0036] Obtain a second file, where the second file is a code file generated after modifying the first file.
[0037] Optionally, the generating unit is further configured to:
[0038] Generate the unified diff format file according to the first file and the second file.
[0039] Optionally, the comparing unit is specifically configured to:
[0040] Compare each line of code in each group of the candidate code suggestions with each line of code in each paragraph except for the set modification method one by one.
[0041] Optionally, the set modification method is deletion.
[0042] In a third aspect, an embodiment of the present invention provides an electronic device, including a memory and a processor, where the memory is configured to store one or more computer program instructions, and where the one or more computer program instructions are executed by the processor to implement the method as described in any item of the first aspect or any possible implementation of the first aspect.
[0043] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium, on which computer program instructions are stored, and the computer program instructions, when executed by a processor, implement the method as described in any item of the first aspect or any possible implementation of the first aspect.
[0044] In an embodiment of the present invention, by obtaining a unified diff file, where the unified diff file represents the differences between a first file and a second file, and the first file and the second file each include multiple lines of code; inputting the unified diff file into a pre-trained large model to generate at least one set of candidate code suggestions, where each set of candidate code suggestions includes multiple lines of code; comparing each set of candidate code suggestions one by one with each paragraph in the unified diff file, where each paragraph includes the modified code and the corresponding context of a set number of lines; in response to any set of candidate code suggestions successfully matching any paragraph in the unified diff file, determining the any set of candidate code suggestions as incorrect code suggestions. Through the above method, incorrect code suggestions can be intercepted, thereby improving the false alarm interception rate. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Through the following description of the embodiments of the present invention with reference to the accompanying drawings, the above and other objects, features, and advantages of the present invention will become clearer. In the drawings:
[0046] Figure 1 is a flowchart of a method for identifying incorrect code suggestions in an embodiment of the present invention;
[0047] Figure 2 is a flowchart of another method for identifying incorrect code suggestions in an embodiment of the present invention;
[0048] Figure 3 is a flowchart of yet another method for identifying incorrect code suggestions in an embodiment of the present invention;
[0049] Figure 4 is a schematic diagram of a device for identifying incorrect code suggestions in an embodiment of the present invention;
[0050] Figure 5 is a schematic diagram of an electronic device in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0051] The following is a description of the present application based on embodiments, but the present application is not limited to these embodiments. In the following detailed description of the present application, some specific details are described in detail. Those skilled in the art can fully understand the present application without the description of these details. To avoid obscuring the essence of the present application, well-known methods, processes, procedures, elements, and circuits are not described in detail.
[0052] In addition, those of ordinary skill in the art should understand that the accompanying drawings provided herein are for illustrative purposes only, and the drawings are not necessarily drawn to scale.
[0053] Unless the context clearly requires otherwise, words such as "including" and "comprising" in the entire application document shall be interpreted in an inclusive sense rather than an exclusive or exhaustive sense; that is, it is the meaning of "including but not limited to".
[0054] In the description of this application, it should be understood that terms such as "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance. In addition, in the description of this application, unless otherwise specified, "a plurality of" means two or more.
[0055] In the prior art, first, a Unified Diff text is determined based on the source code and the modified code, and then the code review result is generated according to the above Unified Diff text; specifically, the Unified Diff text is input into a large model to generate the above code review result, where the code review result includes a problem description and a solution, relevant files and relative paths, the method fragment where the problem is located, and code suggestions for fixing the problem; the user adjusts the modified code again according to the above code suggestions, but some of the above code suggestions are incorrect, that is, false positives occur. The false positive means that the large model responds with a code error when detecting the Unified Diff text and gives code suggestions, but in fact, the code has no error, resulting in a relatively high false positive rate in the overall code review. The false positive rate refers to the sum of the probabilities of false positives or the occurrence of ambiguous or neutral code suggestions; showing the false positive code suggestions to the user will cause an additional processing burden to the user.
[0056] In order to reduce the false positive rate and improve the false positive interception rate, generally three methods are adopted, as follows: Method 1: The method of adding longer context to the large model; specifically, longer context is added in the stage of finding problems. The number of context lines to be displayed for the modified code in the common Unified Diff text is generally default set to three lines. In order to improve the understanding ability of the large model, the number of context lines is increased. However, under the current basic ability background of the large model, longer context will significantly reduce the concentration of the large model, resulting in it being prone to missing many local code problems; Method 2: Use the reflection mechanism of the large model to intercept incorrect code suggestions. For example, in the multi-round reflection module of the Chain of Thought (CoT), letting the large model reflect can effectively reduce its hallucination problem, but the false positive interception rate is also low; Method 3: Use vector recall to implement interception of incorrect code suggestions, but using vector recall is more suitable for solving the problem of prediction interception based on user behavior, and it will bring certain side effects to the interception of specific features, resulting in a relatively low false positive interception rate. Therefore, how to improve the false positive interception rate during code detection is a problem that needs to be solved currently.
[0057] In an embodiment of the present invention, the large model may also be referred to as an Artificial Intelligence (AI) model or a Large Language Model (LLM). Among them, the large language model is a deep learning model based on the transformer architecture, capable of processing and generating natural language text. It is usually trained on a large amount of text data, has the ability to understand and generate language, and is widely used in dialogue systems, text generation, and other natural language processing tasks.
[0058] In an embodiment of the present invention, by using historical data on the backhaul line, that is, obtaining historical review results, analyzing and summarizing them, it is determined that there are two types of false alarms with a relatively high proportion. Type 1: Local logical modifications, for example, removing an if statement, removing a case in a switch statement, changing to call another method, deleting part of the logic, etc.; Type 2: Refactoring, for example, modifying the signature at the definition or reference, such as adding a method, deleting a parameter, changing to reference another parameter, renaming a method, etc.; The common feature of the above two types of false alarms is the lack of integrity of the code context and the lack of requirements background. The code suggestions given for the above false alarms are called error code suggestions. The error code suggestions for the above common features are collectively referred to as rollback codes, that is, it is recommended that the user change the modified code back. By identifying the error code suggestions that give rollback codes, the false alarm rate can be effectively reduced. On this basis, in an embodiment of the present invention, a method for identifying error code suggestions is proposed, specifically as Figure 1 shown, the method includes:
[0059] Step S101, obtain a unified diff format file.
[0060] Among them, the unified diff format file represents the difference between a first file and a second file, and the first file and the second file each include multiple pieces of code;
[0061] In a possible implementation manner, the Unified Diff is a variant of the context format, which omits redundant context lines and has a more compact format. When generating the Unified Diff, a parameter lines is set to indicate the number of context lines to be displayed for the modified code, generally defaulting to three lines. Specifically, the GNU diff can be used to generate the Unified Diff format. Among them, the GNU diff is a command-line tool widely used in UNIX and UNIX-like systems for comparing the differences between two files or directories.
[0062] In a possible implementation, after the user modifies the first file to generate a second file, the Unified Diff file is then generated according to GNU diff. In the Unified Diff file, in addition to showing the modified code, the N lines of code above and below the modified code are also shown, where N is a positive integer greater than or equal to 0. Suppose N is equal to 3. An example of the Unified Diff file is as follows:
[0063] ---
[0064] a / force-base / src / main / java / com / alibaba / force / base / mr / llm / LlmCodeReviewExecutor.java+++
[0065] b / force-base / src / main / java / com / alibaba / force / base / mr / llm / LlmCodeReviewExecutor.java@@-45,6+45,7@@
[0066] import java.util.concurrent.ScheduledFuture;
[0067] import java.util.concurrent.TimeUnit;
[0068] import java.util.concurrent.TimeoutException;
[0069] +import java.util.concurrent.atomic.AtomicBoolean;
[0070] import java.util.stream.Collectors;
[0071] import lombok.extern.slf4j.Slf4j;
[0072] import org.apache.commons.lang3.StringUtils;
[0073] @@-90,6+91,7@@public abstract class LlmCodeReviewExecutor{
[0074] protected final int fastFinished(Long id,Long projectId,List<Future <integer>>futures){ScheduledExecutorService scheduledExecutorService = Executors.newScheduledThreadPool(1);
[0075] +AtomicBoolean isFinished = new AtomicBoolean(false);
[0076] ScheduledFuture<?> statusCheckFuture =
[0077] scheduledExecutorService.scheduleAtFixedRate(
[0078] ()->{
[0079] @@-100,6+102,7@@public abstract class LlmCodeReviewExecutor{
[0080] log.info("Immediately finish LlmCodeReview,id:{}",id);
[0081] / / Clear the pending AI comments generated this time
[0082] codeReviewCommentService.clearPendingAiComments(projectId,id);
[0083] +isFinished.set(true);
[0084] scheduledExecutorService.shutdownNow();
[0085] }
[0086] },
[0087] @@-107,12+110,14@@public abstract class LlmCodeReviewExecutor{
[0088] 1,
[0089] TimeUnit.SECONDS);
[0090] - int result = getLlmGeneratedCommentsCount(futures);
[0091] + if (isFinished.get()) {
[0092] + return 0;
[0093] +}
[0094] statusCheckFuture.cancel(true);
[0095] scheduledExecutorService.shutdownNow();
[0096] - return result;
[0097] + return getLlmGeneratedCommentsCount(futures);
[0098] }
[0099] / ** Determine whether the current CodeReview needs to immediately end the intelligent review * /
[0100] In the above Unified Diff file, it has a specific format. Specifically, the "---" indicates the file name before the change, and the "+++" indicates the file name after the change; a line starting and ending with "@" indicates the line range where the change occurs. For example, "@-1,7+1,7@" means there is a difference between 7 consecutive lines starting from line 1 in the first file and 7 consecutive lines starting from line 1 in the second file; each line is preceded by an identifier, which is used to indicate that the line is deleted in the original file (indicated by "-"), added in the modified file (indicated by "+"), or unchanged (indicated by a space). The above Unified Diff file is only for illustrative purposes and is specifically determined according to the actual situation.
[0101] Step S102: Input the unified difference format file into a pre-trained large model to generate at least one set of candidate code suggestions.
[0102] Among them, each set of the candidate code suggestions includes multiple pieces of code.
[0103] In a possible implementation, the above Unified Diff file is input into a pre-trained large model to generate multiple review results. Each review result has a specific format, including suggestion content, relevant file, existing code, and code suggestion. Among them, the code suggestion is the candidate code suggestion, and the candidate code suggestion includes multiple pieces of code. An example 2 of the candidate code suggestion is as follows:
[0104] - int result = getLlmGeneratedCommentsCount(futures);
[0105] statusCheckFuture.cancel(true);
[0106] scheduledExecutorService.shutdownNow();
[0107] - return result;
[0108] Step S103: Compare each group of the candidate code suggestions with each paragraph in the unified diff format file one by one.
[0109] Among them, each paragraph includes the modified code and the context of the corresponding set number of lines.
[0110] In a possible implementation, the above Example 1 can be divided into 4 paragraphs, specifically as follows:
[0111] The first paragraph:
[0112] @@ -45,6 +45,7 @@
[0113] import java.util.concurrent.ScheduledFuture;
[0114] import java.util.concurrent.TimeUnit;
[0115] import java.util.concurrent.TimeoutException;
[0116] + import java.util.concurrent.atomic.AtomicBoolean;
[0117] import java.util.stream.Collectors;
[0118] import lombok.extern.slf4j.Slf4j;
[0119] import org.apache.commons.lang3.StringUtils;
[0120] Second paragraph:
[0121] @@-90,6+91,7@@public abstract class LlmCodeReviewExecutor{
[0122] protected final int fastFinished(Long id,Long projectId,List<Future <integer>>
[0123] futures){
[0124] ScheduledExecutorService scheduledExecutorService =
[0125] Executors.newScheduledThreadPool(1);
[0126] +AtomicBoolean isFinished = new AtomicBoolean(false);
[0127] ScheduledFuture<?> statusCheckFuture =
[0128] scheduledExecutorService.scheduleAtFixedRate(
[0129] ()->{
[0130] The third paragraph:
[0131] @@-100,6+102,7@@public abstract class LlmCodeReviewExecutor{
[0132] log.info("Immediately finish LlmCodeReview,id:{}",id);
[0133] / / Clear the pending AI comments generated this time
[0134] codeReviewCommentService.clearPendingAiComments(projectId,id);
[0135] +isFinished.set(true);
[0136] scheduledExecutorService.shutdownNow();
[0137] }
[0138] },
[0139] The fourth paragraph:
[0140] @@-107,12+110,14@@public abstract class LlmCodeReviewExecutor{
[0141] 1,
[0142] TimeUnit.SECONDS);
[0143] -int result = getLlmGeneratedCommentsCount(futures);
[0144] +if(isFinished.get()){
[0145] +return 0;
[0146] +}
[0147] statusCheckFuture.cancel(true);
[0148] scheduledExecutorService.shutdownNow();
[0149] -return result;
[0150] +return getLlmGeneratedCommentsCount(futures);
[0151] }
[0152] The line starting and ending with "@@" above indicates the range of lines where changes occur, that is, a paragraph.
[0153] Before making the comparison, first locate the Unified Diff file to be compared based on the relevant file and the raw diff. For example, the located Unified Diff file is the file shown in Example 2, and then compare it with the above four paragraphs respectively according to the candidate code suggestions in Example 2; specifically, each line of code in each group of the candidate code suggestions is compared one by one with each line of code in each of the paragraphs except for the set modification method, and the set modification method is deletion; the above comparison process can be called the block-level sliding window method.
[0154] In a possible implementation, the candidate code suggestions in Example 2 include 6 lines, two of which are blank lines. The blank lines are also compared. Each of the above 6 lines is compared with each line of 4 paragraphs in the Unified Diff file. After one line is successfully compared, the comparison of the next line continues. If a line comparison fails, the comparison ends.
[0155] Step S104, in response to any group of the candidate code suggestions being successfully compared with any paragraph in the unified difference format file, determine the any group of the candidate code suggestions as error code suggestions.
[0156] In a possible implementation, the candidate code suggestions in the above Example 2 match successfully with the fourth paragraph in Example 1, that is, the fourth paragraph includes all the codes in Example 2. It can be inferred that the candidate code suggestions may be rollback codes, and then determine the candidate code suggestions as error code suggestions.
[0157] In a possible implementation, after the above step S104, there are also other steps, specifically as Figure 2 shown, the method includes:
[0158] Step S105, delete the error code suggestions in the candidate code suggestions to generate target code suggestions.
[0159] Step S106, output the target code suggestions.
[0160] In a possible implementation, other steps may also be included before the above step S101, specifically as Figure 3 described, including:
[0161] Step S107, obtain a first file.
[0162] Wherein, the first file is an original code file.
[0163] Step S108, obtain a second file.
[0164] Wherein, the second file is a code file generated after modifying the first file.
[0165] Step S109, generate the unified difference format file according to the first file and the second file.
[0166] Through the above embodiments, the matching algorithm of the code block-level sliding window based on the Unified Diff file does not have the unpredictability of the large model, can stably identify false positives of specific features (such as rollback codes), can intercept error code suggestions, and thus improve the false positive interception rate.
[0167] In an embodiment of the present invention, an apparatus for identifying error code suggestions is provided, as Figure 4 shown, specifically including: an acquisition unit 401, a generation unit 402, a comparison unit 403, and a determination unit 404. Among them, the acquisition unit 401 is used to acquire a unified difference format file, where the unified difference format file represents the differences between a first file and a second file, and the first file and the second file respectively include multiple pieces of code; the generation unit 402 is used to input the unified difference format file into a pre-trained large model to generate at least one group of candidate code suggestions, where each group of the candidate code suggestions includes multiple pieces of code; the comparison unit 403 is used to compare each group of the candidate code suggestions with each paragraph in the unified difference format file one by one, where each paragraph includes the modified code and the corresponding context of the set number of lines; the determination unit 404, in response to any group of the candidate code suggestions being successfully compared with any paragraph in the unified difference format file, is used to determine the any group of the candidate code suggestions as error code suggestions.
[0168] Further, the generation unit is further used to:
[0169] Delete the error code suggestions from the candidate code suggestions to generate target code suggestions.
[0170] Further, the apparatus further includes:
[0171] An output unit, which is used to output the target code suggestions.
[0172] Further, the acquisition unit is further used to:
[0173] Acquire a first file, where the first file is an original code file.
[0174] Further, the acquisition unit is further used to:
[0175] Acquire a second file, where the second file is a code file generated after modifying the first file.
[0176] Further, the generation unit is further used to:
[0177] Generate the unified difference format file according to the first file and the second file.
[0178] Further, the comparison unit is specifically used to:
[0179] Compare each line of code in each group of the candidate code suggestions with each line of code in each paragraph except for the set modification method one by one.
[0180] Further, the setting modification method is deletion.
[0181] Figure 5 is a schematic structural diagram of the electronic device in the embodiment of the present invention. As Figure 5 shown, it includes a general computer hardware structure, which at least includes a processor 501 and a memory 502. The processor 501 and the memory 502 are connected through a bus 503. The memory 502 is adapted to store instructions or programs executable by the processor 501. The processor 501 can be an independent microprocessor or a set of one or more microprocessors. Thus, the processor 501 executes the instructions stored in the memory 502 to execute the method flow of the embodiment of the present invention as described above to implement the processing of data and the control of other devices. The bus 503 connects the above-mentioned multiple components together, and at the same time connects the above-mentioned components to a display controller 504, a display device, and an input / output (I / O) device 505. The input / output (I / O) device 505 can be a mouse, a keyboard, a modem, a network interface, a touch input device, a somatosensory input device, a printer, and other devices well known in the art. Typically, the input / output (I / O) device 505 is connected to the system through an input / output (I / O) controller 506.
[0182] Among them, the instructions stored in the memory 502 are executed by at least one processor 501 to implement: obtaining a unified diff file, where the unified diff file represents the difference between a first file and a second file, and the first file and the second file respectively include multiple pieces of code; inputting the unified diff file into a pre-trained large model to generate at least one set of candidate code suggestions, where each set of the candidate code suggestions includes multiple pieces of code; comparing each set of the candidate code suggestions with each paragraph in the unified diff file one by one, where each paragraph includes the modified code and the corresponding context of the set number of lines; in response to any set of the candidate code suggestions being successfully compared with any paragraph in the unified diff file, determining the any set of the candidate code suggestions as an incorrect code suggestion.
[0183] Specifically, the electronic device includes: one or more processors 501 and a memory 502, Figure 5 Taking one processor 501 as an example. The processor 501 and the memory 502 can be connected through a bus or other means, Figure 5 Take the bus connection as an example. As a non-volatile computer-readable storage medium, the memory 502 can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules. By running the non-volatile software programs, instructions, and modules stored in the memory 502, the processor 501 executes various functional applications and data processing of the device, that is, implements the method for determining and identifying error code suggestions as described above.
[0184] The memory 502 may include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store an option list, etc. In addition, the memory 502 may include high-speed random access memory and may also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some embodiments, the memory 502 may optionally include a memory remotely set relative to the processor 501, and these remote memories can be connected to an external device through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0185] One or more modules are stored in the memory 502 and, when executed by one or more processors 501, execute the method for identifying error code suggestions in any of the above method embodiments.
[0186] As those skilled in the art will realize, various aspects of the embodiments of the present invention can be implemented as a system, a method, or a computer program product. Therefore, various aspects of the embodiments of the present invention can take the following forms: a completely hardware implementation, a completely software implementation (including firmware, resident software, microcode, etc.), or an implementation that combines software aspects with hardware aspects, which is generally referred to herein as "circuit", "module", or "system". In addition, various aspects of the embodiments of the present invention can take the following form: a computer program product implemented in one or more computer-readable media, with computer-readable program code implemented thereon.
[0187] Any combination of one or more computer-readable media may be utilized. A computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium may be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium would include the following: an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In the context of embodiments of the present invention, a computer-readable storage medium may be any tangible medium that can contain, or store a program for use by or in connection with an instruction execution system, apparatus, or device.
[0188] A computer-readable signal medium may include a propagated data signal with computer-readable program code embodied therein, either as in a baseband or as part of a carrier wave. Such a propagated signal may take any of a variety of forms, including but not limited to: electromagnetic, optical, or any suitable combination thereof. A computer-readable signal medium may be any computer-readable medium that is not a computer-readable storage medium and that can communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device.
[0189] Any suitable medium may be used to transmit the program code embodied on a computer-readable medium, including but not limited to wireless, wired, fiber optic cable, RF, etc., or any suitable combination of the foregoing.
[0190] The computer program code for performing operations for aspects of embodiments of the present invention may be written in any combination of one or more programming languages, including: object-oriented programming languages such as Java, Smalltalk, C++; and conventional procedural programming languages such as the "C" programming language or similar programming languages. The program code may execute entirely on the user's computer as a stand-alone software package, partly on the user's computer and partly on a remote computer, partly on the user's computer and partly on a remote computer or server, or entirely on the remote computer or server. In the latter case, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection may be made to an external computer (e.g., through the Internet using an Internet service provider).
[0191] The flowchart illustrations and / or block diagrams of the methods, apparatuses (systems) and computer program products according to the embodiments of the present invention described above depict various aspects of the embodiments of the present invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and the combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device to produce a machine, such that the instructions (executed via the processor of the computer or other programmable data processing device) create a means for implementing the functions / actions specified in the flowchart and / or block diagram block or blocks.
[0192] These computer program instructions can also be stored in a computer-readable medium that can direct a computer, other programmable data processing device, or other apparatus to operate in a particular manner, such that the instructions stored in the computer-readable medium produce an article of manufacture including instructions for implementing the functions / actions specified in the flowchart and / or block diagram block or blocks.
[0193] The computer program instructions can also be loaded onto a computer, other programmable data processing device, or other apparatus, so as to perform a series of operational steps on the computer, other programmable device, or other apparatus to produce a computer-implemented process, such that the instructions executed on the computer or other programmable device provide a process for implementing the functions / actions specified in the flowchart and / or block diagram block or blocks.
[0194] The foregoing are only the preferred embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.
[0195] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or fully authorized by all parties. Moreover, the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of the relevant countries and regions, and corresponding operation entrances are provided for the user to select authorization or rejection. If the user refuses to process personal information other than the necessary information required for the basic functions, it will not affect the user's use of the basic functions.< / integer> < / integer>
Claims
1. A method for identifying error code suggestions, characterized in that, The method comprises: Obtain a unified difference format file, wherein the unified difference format file represents the difference between a first file and a second file, and the first file and the second file respectively include a plurality of codes; Inputting the unified difference format file into a pre-trained large model to generate at least one group of candidate code suggestions, wherein each group of candidate code suggestions includes multiple codes; Compare each group of candidate code suggestions with each paragraph in the unified difference format file one by one, wherein each paragraph includes the modified code and a corresponding set number of lines of context; In response to any group of the candidate code suggestions being successfully compared with any paragraph in the unified difference format file, the any group of the candidate code suggestions is determined as an erroneous code suggestion.
2. The method according to claim 1, characterized in that The method further comprises: The erroneous code suggestion is deleted from the candidate code suggestions to generate a target code suggestion.
3. The method according to claim 2, characterized in that The method further comprises: The target code suggestion is output.
4. The method according to claim 1, characterized in that: The method further comprises: A first file is obtained, wherein the first file is an original code file.
5. The method according to claim 4, characterized in that The method further comprises: A second file is obtained, wherein the second file is a code file generated after the first file is modified.
6. The method according to claim 5, characterized in that The method further comprises: The unified difference format file is generated according to the first file and the second file.
7. The method according to claim 1, characterized in that The step of comparing each group of candidate code suggestions with each paragraph in the unified difference format file one by one specifically includes: Each line of code in each group of candidate code suggestions is compared one by one with each line of code in each paragraph except for the set modification mode.
8. The method according to claim 1, characterized in that The setting modification method is deletion.
9. A device for identifying error code suggestions, characterized in that The device comprises: An acquiring unit, configured to acquire a unified difference format file, wherein the unified difference format file indicates a difference between a first file and a second file, and the first file and the second file respectively include a plurality of codes; A generating unit, configured to input the unified difference format file into a pre-trained large model to generate at least one group of candidate code suggestions, wherein each group of candidate code suggestions includes a plurality of codes; A comparison unit, configured to compare each group of candidate code suggestions with each paragraph in the unified difference format file one by one, wherein each paragraph includes the modified code and a corresponding set number of lines of context; A determination unit is configured to determine any group of the candidate code suggestions as an erroneous code suggestion in response to a successful comparison between any group of the candidate code suggestions and any paragraph in the unified difference format file.
10. An electronic device comprising a memory and a processor, characterized in that: The memory is used to store one or more computer program instructions, wherein the one or more computer program instructions are executed by the processor to implement the method according to any one of claims 1 to 8.
11. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.