Code task processing, code error correction and code processing model training method

By adding region identifiers to the original code file and modification plan, the efficiency problem of users needing to manually integrate modification is solved, and the method of quickly generating target modification files is realized, which significantly improves the speed and efficiency of code generation.

CN119645488BActive Publication Date: 2025-05-13ALIBABA (CHINA) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510158648.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-13
Publication Date
2025-05-13
Estimated Expiration
2045-02-13

AI Technical Summary

Technical Problem

In the prior art, after receiving the code modification plan, the user needs to manually operate to integrate the modification into the original code file, resulting in additional workload and efficiency reduction.

Method used

By adding region identifiers to the original code file and the original modification scheme, the code processing model can accurately locate the code content, thereby generating target modification files, reducing the length of the model output text, and improving generation speed and efficiency.

Benefits of technology

It significantly speeds up the code generation speed, improves the generation efficiency of target modified files, and allows users to obtain complete and runnable modified code files more quickly.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119645488B_ABST
    Figure CN119645488B_ABST
Patent Text Reader

Abstract

The embodiments of this specification provide a method for code task processing, code error correction, and code processing model training, wherein the code task processing method includes: obtaining an original code file and an original modification plan of the original code file; adding a region identifier to the original code file and the original modification plan to obtain a code file to be processed and a modification plan to be processed; inputting the code file to be processed and the modification plan to be processed into a code processing model to obtain a region identifier sequence; and generating a target modification file of the original code file according to the region identifier sequence. By introducing the region identifier, the mode of model code generation is switched to a mode of generating only the region identifier, which greatly reduces the length of the model output text and speeds up the generation speed. Since each region identifier is associated with the original code file and the original modification plan, the target modification file can be quickly generated based on the region identifier sequence, thereby improving the generation efficiency of the target modification file.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of this specification relate to the field of artificial intelligence technology, and in particular to code task processing, code error correction, and code processing model training methods. Background Art

[0002] With the development of computer technology, large models have begun to shine, demonstrating extraordinary capabilities in language understanding, generation, interaction, and reasoning, and are widely used in natural language processing fields such as dialogue, translation, and code processing.

[0003] Taking the field of code processing as an example, large models can be used to automatically generate code modification plans to reduce labor costs. However, after obtaining the code modification plan, users often need to perform further manual operations to integrate these modifications into the original code file to obtain a complete and executable modified file. This process adds extra workload and reduces overall efficiency. Summary of the invention

[0004] In view of this, an embodiment of the present specification provides a code task processing method. One or more embodiments of the present specification also relate to a code error correction method, a code processing model training method, an information processing method based on a code processing model, a task platform, a code task processing device, a code error correction device, a code processing model training device, an information processing device based on a code processing model, a computing device, an electronic device, a computer-readable storage medium, and a computer program product to solve the technical defects existing in the prior art.

[0005] According to a first aspect of an embodiment of the present specification, a code task processing method is provided, including: obtaining an original code file and an original modification plan of the original code file; adding a region identifier to the original code file and the original modification plan to obtain a code file to be processed and a modification plan to be processed; inputting the code file to be processed and the modification plan to be processed into a code processing model to obtain a region identifier sequence; and generating a target modification file of the original code file according to the region identifier sequence.

[0006] In the code task processing method provided by one or more embodiments of the present specification, since the original modification plan includes the modification points for the original code file, the code content of the target modification file comes from the original code file or the original modification plan. On this basis, by adding region identifiers to the original code file and the original modification plan, the code processing model can accurately locate the location of the code content, thereby switching the model code generation mode to a mode that only generates region identifiers, greatly reducing the length of the model output text and significantly speeding up the generation speed. Moreover, since each region identifier is associated with the original code file and the original modification plan, after obtaining the region identifier sequence, the target modification file can be quickly generated, thereby improving the generation efficiency of the target modification file. BRIEF DESCRIPTION OF THE DRAWINGS

[0007] Figure 1 is a flow chart of a code task processing method provided by an embodiment of this specification;

[0008] Figure 2 is an architecture diagram of a code task processing system provided by an embodiment of this specification;

[0009] Figure 3 is a flow chart of a code error correction method provided by an embodiment of this specification;

[0010] Figure 4 is a flowchart of a code processing model training method provided by an embodiment of this specification;

[0011] Figure 5 is a processing flow chart of a code task processing method provided by an embodiment of this specification;

[0012] Figure 6 is a flowchart of an information processing method based on a code processing model provided by an embodiment of this specification;

[0013] Figure 7 It is a structural diagram of a task platform provided by an embodiment of this specification;

[0014] Figure 8 It is a structural diagram of a code task processing device provided by an embodiment of this specification;

[0015] Fig. 9 It is a structural schematic diagram of a code error correction device provided by an embodiment of this specification;

[0016] Fig.10 It is a structural diagram of a code processing model training device provided by an embodiment of this specification;

[0017] Fig.11 is a schematic diagram of the structure of an information processing device based on a code processing model provided by an embodiment of this specification;

[0018] Fig.12 is a structural block diagram of a computing device provided by one embodiment of this specification;

[0019] Fig.13 It is a structural block diagram of an electronic device provided by an embodiment of this specification. DETAILED DESCRIPTION

[0020] Many specific details are described in the following description to facilitate a full understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar generalizations without violating the connotation of this specification, so this specification is not limited to the specific implementation disclosed below.

[0021] The terms used in one or more embodiments of this specification are only for the purpose of describing specific embodiments, and are not intended to limit one or more embodiments of this specification. The singular forms of "a", "said" and "the" used in one or more embodiments of this specification and the appended claims are also intended to include plural forms, unless the context clearly indicates other meanings. It should also be understood that the term "and / or" used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.

[0022] It should be understood that although the terms first, second, etc. may be used to describe various information in one or more embodiments of this specification, this information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of one or more embodiments of this specification, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".

[0023] In addition, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0024] In one or more embodiments of this specification, a large model refers to a deep learning model with large-scale model parameters, which usually contains hundreds of millions, tens of billions, hundreds of billions, trillions, or even more than 10 trillion model parameters. A large model can also be called a foundation model / foundation model. The large model is pre-trained with large-scale unlabeled corpus to produce a pre-trained model with more than 100 million parameters. This model can adapt to a wide range of downstream tasks, and the model has good generalization ability, such as a large-scale language model (LLM, Large Language Model), a multi-modal pre-training model, etc.

[0025] When the big model is used in practice, only a small number of samples are needed to fine-tune the pre-trained model and it can be applied to different tasks. The big model can be widely used in natural language processing (NLP), computer vision and other fields. Specifically, it can be applied to computer vision tasks such as visual question answering (VQA), image description (IC, Image Caption), image generation, as well as natural language processing tasks such as text-based sentiment classification, text summary generation, and machine translation. The main application scenarios of the big model include digital assistants, intelligent robots, search, online education, office software, e-commerce, intelligent design, etc.

[0026] First, the terms involved in one or more embodiments of this specification are explained.

[0027] Intelligent Coding Assistant: Utilizes big models and artificial intelligence technology to deeply understand the developer's programming intent and contextual information, and provides functions such as code completion, code generation, code error detection and correction, and code optimization to help developers reduce repetitive work and achieve significant improvements in programming efficiency and code quality.

[0028] Post-modification code generation: Generate a complete post-modification code file based on the pre-modification code file and the code modification plan, thereby realizing the implementation from the modification plan to the specific code.

[0029] Supervised Fine-Tuning (SFT): is a further training method based on a pre-trained model. In this method, the model is trained on a dataset containing input and expected output pairs so that the model can learn how to generate answers closer to human level. Supervised fine-tuning is usually used to adapt the model to data in a specific task or domain, thereby improving the model's performance on these tasks.

[0030] Code modification suggestion generation: By analyzing the code files and providing suggestions on requirements implementation, error correction, code refactoring, and security improvement based on the user's input problems, the development team can improve the code and enhance the quality and maintainability of the entire code project. For reasons such as readability, modification suggestions often only include the modified fragments, and other code that does not need to be modified is often omitted.

[0031] In code project development scenarios, users often need to modify existing project codes based on development requirements raised by product managers or defect issues raised by other users. This process is a core step in software development, which places very high demands on developers' project understanding and coding skills, and is also the most time-consuming and labor-intensive step in the project development process. With the rapid development of artificial intelligence technology and big model technology, in recent years, more and more work has begun to try to use big models to automatically generate code modification plans to reduce labor costs. However, after obtaining the code modification plan, users need to manually confirm whether to adopt the code modification plan. After confirming to adopt the code modification plan, users cannot directly obtain the executable modified file, and further manual operations are required to integrate these modifications into the original code file to obtain a complete and executable modified file. This process adds extra workload and reduces overall efficiency.

[0032] After analysis, the modified code generation task has a significant feature: in the code modification plan, the modification points have been proposed. Therefore, when generating the modified code, the model only needs to be generated according to the modification plan. Therefore, in the embodiments of this specification, it is believed that the code content of the modified code can only come from the pre-modification code or the code modification plan. In addition, since the code files are often relatively long, it often takes a long time to generate the modified code (a code file of 100 lines takes 1 minute), which affects the user experience. Based on this, the embodiments of this specification propose a code task processing solution to obtain the original code file and the original modification plan of the original code file; add a region identifier to the original code file and the original modification plan to obtain the code file to be processed and the modification plan to be processed; input the code file to be processed and the modification plan to be processed into the code processing model to obtain a region identifier sequence; according to the region identifier sequence, generate a target modification file of the original code file. By utilizing the particularity of the modified code generation task and adding region identifiers to the original code file and the original modification plan, the code processing model can accurately locate the location of the code content, thereby switching the model code generation mode to a mode that only generates region identifiers, greatly reducing the length of the model output text and significantly speeding up the generation speed. In addition, since each region identifier is associated with the original code file and the original modification plan, after obtaining the region identifier sequence, the target modification file can be quickly generated, thereby improving the generation efficiency of the target modification file.

[0033] In this specification, a code task processing method is provided. This specification also involves a code error correction method, a code processing model training method, an information processing method based on a code processing model, a task platform, a code task processing device, a code error correction device, a code processing model training device, an information processing device based on a code processing model, a computing device, an electronic device, a computer-readable storage medium and a computer program product, which are described in detail one by one in the following embodiments.

[0034] See also Figure 1 , Figure 1 A flowchart of a code task processing method provided by an embodiment of the present specification is shown, which specifically includes the following steps:

[0035] Step 102: Obtain an original code file and an original modification plan of the original code file.

[0036] It should be noted that the original code file refers to the source code file that has not been modified or processed. The original code file contains all the logic and structure of the program. It is a collection of codes written by the developer and used to implement specific functions or solve specific problems. Therefore, the original code file can be understood as the code file before modification. The original modification plan refers to the modification suggestions made for the original code file. Therefore, the original modification plan can be understood as the original modification suggestion. The specific modification content included in the original modification plan is related to the code task, and the code task includes but is not limited to code completion tasks, code generation tasks, code error detection and correction tasks, code optimization tasks, etc. For example, in the case where the code task is a code error detection and correction task, the modification content included in the original modification plan is the accurate modification code corresponding to the error code in the original code file. In the case where the code task is a code optimization task, the modification content included in the original modification plan is the optimized code corresponding to the unoptimized code in the original code file.

[0037] For example, assume that the original code file is as follows:

[0038] def calculate_total(prices, tax_rate):

[0039] total = sum(prices)

[0040] tax = total * tax_rate

[0041] return total + tax# Error: The total amount after tax is not processed and displayed correctly

[0042] # This is a blank line

[0043] items = [10, 20, 30]

[0044] tax_rate = 0.08

[0045] print("Total:", calculate_total(items, tax_rate))

[0046] It should be noted that the purpose of the original code file is to calculate the total amount of a set of commodity prices, add the corresponding tax amount, and finally output the total amount including tax. The original code file defines a function called calculate_total, which receives two parameters: prices, which is a price list containing multiple items (such as the commodity price is [10,20, 30]). tax_rate, a floating point number, represents the tax rate (such as the tax rate of the commodity is 8%). Use the built-in sum function in Python to sum all the prices in the price list prices to get the total amount of the commodity prices and assign it to the variable total. According to the total amount total and the given tax rate tax_rate, calculate the tax amount to be added and assign it to the variable tax. The function returns the result of the total amount total plus the tax amount tax. This return value represents the total amount of all commodity prices plus tax. However, in the original code file, although the calculate_total function calculates the total amount including tax, it only returns the value of the total amount plus the tax amount in the end, and it does not clearly state that this is the total amount including tax when output, which is easy to cause confusion.

[0047] Therefore, in response to the above problems, the original modification plan of the original code file introduces the total_with_tax variable to clearly indicate the total amount including tax, and ensures that the function returns this value. In addition, the format string is used to retain two decimal places to improve the clarity and accuracy of the output. The original modification plan is as follows:

[0048] tax_amount = total * tax_rate

[0049] total_with_tax = total + tax_amount

[0050] return total_with_tax# Modification: Clearly return the total amount after tax

[0051] # This is a blank line

[0052] print(f"Total with Tax: {calculate_total(items, tax_rate):.2f}")# Clearly display the total amount after tax

[0053] In an optional embodiment of the present specification, since the original code file usually includes complex code content, if the original modification plan only includes the modification content, the code processing model cannot accurately locate the code content corresponding to the modification content in the original code file, which will affect the accuracy of the model generation result. Therefore, the original modification plan can also include the context code of the modification content, so as to facilitate the code processing model to locate the modification position. Citing the example of the above original modification plan, the original modification plan including the modification content context is as follows:

[0054] total = sum(prices)

[0055] tax_amount = total * tax_rate

[0056] total_with_tax = total + tax_amount

[0057] return total_with_tax# Modification: Clearly return the total amount after tax

[0058] # This is a blank line

[0059] items = [10, 20, 30]

[0060] tax_rate = 0.08

[0061] print(f"Total with Tax: {calculate_total(items, tax_rate):.2f}")# Clearly display the total amount after tax

[0062] In practical applications, there are many ways to obtain the original code file and the original modification scheme of the original code file, which can be selected according to the actual situation, and the embodiments of this specification do not limit this. In one possible implementation of this specification, the original code file and the original modification scheme of the original code file can be read from other data acquisition devices or databases. In another possible implementation of this specification, the original code file and the original modification scheme of the original code file sent by the user through the terminal device can be received.

[0063] Step 104: Add the region identifier to the original code file and the original modification plan to obtain the code file to be processed and the modification plan to be processed.

[0064] It should be noted that the region identifier refers to a mark inserted in the original code file and the original modification plan. The region identifier is used to indicate the region position where the code content is located in the original code file and the region position where the modified content is located in the original modification plan. The region identifier can be used to provide the code processing model with clear boundaries of the code content. The region position supported by the region identifier can be a function position, a class position, a code block position, a line position, etc. In other words, the region identifier includes but is not limited to a function identifier, a class identifier, a code block identifier, a line identifier, etc. Preferably, the region identifier can be a line identifier, which is used to clearly indicate the specific line position of the code.

[0065] In practical applications, there are many ways to add a region identifier to the original code file and the original modification plan to obtain the code file to be processed and the modification plan to be processed, and the embodiments of this specification do not limit this. In one possible implementation of this specification, the location of adding the region identifier can be determined in the original code file according to the region location indicated by the region identifier, and the location of adding the region identifier can be determined in the original modification plan, and then based on the addition location determined in the original code file and the original modification plan, the region identifier is added to the original code file and the original modification plan to obtain the code file to be processed and the modification plan to be processed. In another possible implementation of this specification, the original code file and the original modification plan can be split according to the region location indicated by the region identifier to obtain the original code culture after the split and the original modification plan after the split, and the region identifier is directly added to the original code file after the split and the original modification plan after the split to obtain the code file to be processed and the modification plan to be processed. During the split, if the region identifier is a line identifier, the original code file after the split and the original modification plan after the split are separate code lines; if the region identifier is a code block identifier, the original code file after the split and the original modification plan after the split are separate code blocks.

[0066] In an optional embodiment of the present specification, in order to distinguish the code content in the original code file from the modified content in the original modification scheme, the region identifiers added in the original code file and the original modification scheme may be different, that is, the region identifier includes a first region identifier and a second region identifier, the first region identifier is different from the second region identifier, and the first region identifier is used to identify the location of the code content in the original code file. The second region identifier is used to identify the location of the modified content in the original modification scheme. The above-mentioned adding region identifiers to the original code file and the original modification scheme to obtain the code file to be processed and the modification scheme to be processed may include the following steps: adding the first region identifier to the original code file to obtain the code file to be processed, and adding the second region identifier to the original modification scheme to obtain the modification scheme to be processed.

[0067] Furthermore, taking the area identifier as a row identifier (including the first row identifier and the second row identifier) ​​as an example, when adding the area identifier to the original code file and the original modification plan to obtain the code file to be processed and the modification plan to be processed, the first row identifier can be used to identify the specific position of each first code line to obtain the code file to be processed, and the second row identifier can be used to identify the specific position of each second code line to obtain the modification plan to be processed.

[0068] In an optional embodiment of the present specification, the region identifier includes a first line identifier and a second line identifier, the original code file includes at least one first code line, and the original modification plan includes at least one second code line; the above adding the region identifier to the original code file and the original modification plan to obtain the code file to be processed and the modification plan to be processed may include the following steps:

[0069] A first line identifier is added to the first code line to obtain a code file to be processed, and a second line identifier is added to the second code line to obtain a modification plan to be processed, wherein the first line identifier is used to identify the position of the first code line, and the second line identifier is used to identify the position of the second code line.

[0070] It should be noted that the first line identifier refers to a mark that identifies the specific location of each first code line in the original code file, such as "@P line number @". The first line identifier corresponds to the first code line one by one. The second line identifier refers to a mark that identifies the specific location of each second code line in the original modification plan, such as "@M line number @". The second line identifier corresponds to the second code line one by one. The first line identifier and the second line identifier are different.

[0071] In actual applications, when adding a first line identifier to a first code line, that is, when using the first line identifier to identify the specific location of each first code line, the first line identifier can be added before the first code line, or after the first code line. When adding a second line identifier to a second code line, that is, when using the second line identifier to identify the specific location of each second code line, the second line identifier can be added before the second code line, or after the second code line. The locations for adding the first line identifier and the second line identifier are selected according to actual conditions, and the embodiments of this specification do not impose any restrictions on this.

[0072] The solution of the embodiment of this specification uses line identifiers to ensure that each modification suggestion can accurately correspond to a specific line of code in the original code file, provide clear guidance for the code processing model, and improve the accuracy of code processing.

[0073] In an optional embodiment of the present specification, the above-mentioned adding a first line identifier to the first code line to obtain a code file to be processed, and adding a second line identifier to the second code line to obtain a modification plan to be processed may include the following steps:

[0074] A first line of identifier is added before the first code line to obtain a code file to be processed, and a second line of identifier is added before the second code line to obtain a modification plan to be processed.

[0075] Exemplarily, the original code file in the above example is referenced, and the first line identifier "@P line number @" is added before the first code line included in the original code file, and the code file to be processed is obtained as follows:

[0076] @P0@ def calculate_total(prices, tax_rate):

[0077] @P1@total = sum(prices)

[0078] @P2@tax = total * tax_rate

[0079] @P3@return total + tax# Error: The total amount after tax is not processed and displayed correctly

[0080] @P4@# This is a blank line

[0081] @P5@ items = [10, 20, 30]

[0082] @P6@ tax_rate = 0.08

[0083] @P7@ print("Total:", calculate_total(items, tax_rate))

[0084] Referring to the original modification scheme including the modification content context in the above example, a second line identifier "@M line number@" is added before the second code line included in the original modification scheme, and the modification scheme to be processed is obtained as follows:

[0085] @M0@total = sum(prices)

[0086] @M1@tax_amount = total * tax_rate

[0087] @M2@total_with_tax = total + tax_amount

[0088] @M3@return total_with_tax# Modification: Clearly return the total amount after tax

[0089] @M4@# This is a blank line

[0090] @M5@ items = [10, 20, 30]

[0091] @M6@ tax_rate = 0.08

[0092] @M7@ print(f"Total with Tax: {calculate_total(items, tax_rate):.2f}")# Clearly display the total amount after tax

[0093] By applying the solution of the embodiment of this specification, a first line identifier is added before the first code line, and a second line identifier is added before the second code line, so that the code processing model can quickly determine the specific code line corresponding to each modification suggestion in the original code file, thereby improving code processing efficiency.

[0094] Since the generation speed of deep learning models is often slow, for example, the size of models that perform well in the modified code generation task is often 32B or larger, and its output speed is about 50 token / s, while a 100-line code file is often 1500 tokens long, resulting in a long waiting time for users and a poor experience. In order to solve this problem, in an optional embodiment of this specification, the code snippets that have been modified compared to the original code file in the original modification scheme can be retained, and the code snippets that have not been modified can be abbreviated to improve the processing speed of the model, that is, the above-mentioned second line identifier is added to the second code line. Before obtaining the modification scheme to be processed, the following steps can also be included:

[0095] From the second code line, filter out the unmodified code line that is the same as the first code line;

[0096] Using the unmodified identifier, replace the unmodified code line to obtain a replaced second code line;

[0097] Adding a second line identifier to the second code line and obtaining a pending modification plan may include the following steps:

[0098] A second line identifier is added to the replaced second code line to obtain a modification plan to be processed.

[0099] It should be noted that an unmodified code line refers to a second code line in at least one second code line included in the original modification scheme that is identical to any first code line in the original code file. An unmodified identifier refers to a special symbol or label used to replace an unmodified code line in the original modification scheme, such as " / / ... existing code ...".

[0100] In actual applications, there are many ways to filter out the unmodified code lines that are identical to the first code lines from the second code lines, and the specific selection is based on the actual situation. The embodiments of this specification do not impose any restrictions on this. In one possible implementation of this specification, the original code file (including at least one first code line) and the original modification plan (including at least one second code line) can be compared line by line to find the unmodified code lines that are identical to the first code lines. In another possible implementation of this specification, a text comparison tool (such as diff, git diff), a string comparison function in a programming language, or a dedicated code comparison library can be used to filter out the unmodified code lines that are identical to the first code lines from the second code lines.

[0101] In an optional embodiment of the present specification, if the original modification plan only includes the modified content, the code processing model cannot accurately locate the code content corresponding to the modified content in the original code file, which will affect the accuracy of the model generation results. Therefore, when replacing the unmodified code line with the unmodified identifier, at least one second code line in the unmodified code line that is adjacent to the modified code line can be retained without replacement.

[0102] Exemplarily, referring to the original code file in the above example and the original modification scheme including the modification content context, the unmodified code lines identical to the first code line are screened out from the original modification scheme, including the following four lines:

[0103] total = sum(prices)

[0104] # This is a blank line

[0105] items = [10, 20, 30]

[0106] tax_rate = 0.08

[0107] Furthermore, when the unmodified code line is replaced with the unmodified identifier, "total = sum(prices)" and "tax_rate = 0.08" adjacent to the modified code line are retained, and the second code line after replacement is obtained as follows:

[0108] total = sum(prices)

[0109] tax_amount = total * tax_rate

[0110] total_with_tax = total + tax_amount

[0111] return total_with_tax# Modification: Clearly return the total amount after tax

[0112] / / ... existing code ...

[0113] tax_rate = 0.08

[0114] print(f"Total with Tax: {calculate_total(items, tax_rate):.2f}")# Clearly display the total amount after tax

[0115] Finally, add the second line of identifiers before the second line of code after replacement to obtain the pending modification plan as follows:

[0116] @M0@total = sum(prices)

[0117] @M1@tax_amount = total * tax_rate

[0118] @M2@total_with_tax = total + tax_amount

[0119] @M3@return total_with_tax# Modification: Clearly return the total amount after tax

[0120] @M4@ / / ... existing code ...

[0121] @M5@ tax_rate = 0.08

[0122] @M6@ print(f"Total with Tax: {calculate_total(items, tax_rate):.2f}")# Clearly display the total amount after tax

[0123] By applying the solution of the embodiment of this specification, an unmodified identifier is used to replace an unmodified code line in the second code line that is identical to the first code line, to obtain a replaced second code line, and a second line identifier is added to the replaced second code line to obtain a pending modification solution, thereby shortening the length of the pending modification solution and improving the efficiency of code task processing.

[0124] In an optional embodiment of the present specification, after adding the region identifier to the original code file and the original modification plan to obtain the code file to be processed and the modification plan to be processed, the following steps may also be included:

[0125] Based on the content correspondence between the region identifier and the original code file, and the content correspondence between the region identifier and the original modification scheme, identifier mapping information is constructed.

[0126] It should be noted that the content correspondence between the region identifier and the original code file can be obtained based on the correspondence between the region identifier and the first code line in the code file to be processed. The content correspondence between the region identifier and the original modification scheme can be obtained based on the correspondence between the region identifier and the second code line in the modification scheme to be processed. The identifier mapping information is used to record the mapping relationship between the region identifier and the first code line, and the region identifier and the second code line. The process of constructing the identifier information can be regarded as a forward mapping process from code content to region identifier. Optionally, after constructing the identifier mapping information, the identifier mapping information can be stored in the information library.

[0127] Exemplarily, taking "@P1@total = sum(prices)" in the code file to be processed as an example, it can be determined that there is a corresponding relationship between the regional identifier "@P1@" and the first code line "total = sum(prices)". Taking the modification scheme to be processed "@M1@tax_amount = total * tax_rate" as an example, it can be determined that there is a corresponding relationship between the regional identifier "@M1@" and the second code line "tax_amount = total * tax_rate". Then the identifier mapping information is: "@P1@" can be mapped to "total = sum(prices)", and "@M1@" can be mapped to "tax_amount = total * tax_rate".

[0128] By applying the solution of the embodiment of this specification, by constructing the identifier mapping information, it is ensured that the target modified file can be quickly obtained based on the target region identifier mapping included in the region identifier sequence.

[0129] Step 106: Input the code file to be processed and the modification plan to be processed into the code processing model to obtain a region identifier sequence.

[0130] It should be noted that the code processing model refers to a deep learning model with the ability to predict region identifiers. The code processing model is trained based on sample code files, sample modification plans, and sample identifier sequences. The specific training method of the code processing model can be found in Figure 4 A code processing model training method is given. The region identifier sequence includes multiple target region identifiers. The identification granularity of the target region identifier corresponds to the identification granularity of the region identifier. If the region identifier added in the original code file and the original modification plan is a line identifier, the target region identifier here is the target line identifier; if the region identifier added in the original code file and the original modification plan is a code block identifier, the target region identifier here is the target code block identifier. Exemplarily, taking the original code file in the above example and the modification plan to be processed obtained by adding the second line identifier before the second code line after replacement as an example, the code file to be processed and the modification plan to be processed are input into the code processing model, and the region identifier sequence is obtained as follows: @P0@@M0@@M1@@M2@@M3@@P4@@P5@@P6@@M7@.

[0131] Step 108: Generate a target modified file of the original code file according to the region identifier sequence.

[0132] It should be noted that the target modified file refers to the code file obtained by modifying the original code file based on the original modification scheme. The process of generating the target modified file of the original code file according to the region identifier sequence can be regarded as a reverse mapping process from mapping the target region identifier to the code content.

[0133] In practical applications, there are many ways to generate a target modification file of an original code file based on a region identifier sequence, and the specific selection is based on the actual situation. The embodiments of this specification do not impose any restrictions on this. In one possible implementation of this specification, the target modification file of the original code file can be mapped according to the region identifier sequence and the identifier mapping information. In another possible implementation of this specification, the code content corresponding to each target region identifier in the region identifier sequence can be searched in the code file to be processed and the modification scheme to be processed, and the found code content can be spliced ​​according to the arrangement of the region identifier sequence to obtain the target modification file.

[0134] By applying the solution of the embodiments of the present specification, by adding region identifiers to the original code file and the original modification plan, the code processing model can accurately locate the position of the code content, thereby switching the model code generation mode to a mode of only generating region identifiers, greatly reducing the length of the model output text and significantly speeding up the generation speed. Moreover, since each region identifier is associated with the original code file and the original modification plan, after obtaining the region identifier sequence, the target modification file can be quickly generated, thereby improving the generation efficiency of the target modification file.

[0135] In an optional embodiment of the present specification, the above-mentioned generating a target modified file of the original code file according to the region identifier sequence may include the following steps:

[0136] Obtaining identifier mapping information, wherein the identifier mapping information is used to describe the content correspondence between the region identifier and the original code file and the original modification plan;

[0137] Generate a target modified file of the original code file according to the region identifier sequence and the identifier mapping information.

[0138] In practical applications, there are many ways to obtain identifier mapping information, which are selected according to actual conditions, and the embodiments of this specification do not impose any restrictions on this. In one possible implementation of this specification, the identifier mapping information can be read from the information library. In another possible implementation of this specification, the identifier mapping information can be constructed based on the content correspondence between the region identifier and the original code file, and the content correspondence between the region identifier and the original modification plan.

[0139] Furthermore, since the identifier mapping information is used to describe the content correspondence between the region identifier and the original code file and the original modification plan, after obtaining the identifier mapping information, the code content in the original code file and the original modification plan that has no content correspondence with the target region identifiers in the region identifier sequence can be deleted, or only the code content that has a content correspondence with the target region identifiers in the region identifier sequence can be retained to obtain the target modification file.

[0140] Exemplarily, referring to the region identifier sequence in the above example, a target modified file of the original code file is generated according to the region identifier sequence as follows:

[0141] def calculate_total(prices, tax_rate):

[0142] total = sum(prices)

[0143] tax_amount = total * tax_rate

[0144] total_with_tax = total + tax_amount

[0145] return total_with_tax# Modification: Clearly return the total amount after tax

[0146] # This is a blank line

[0147] items = [10, 20, 30]

[0148] tax_rate = 0.08

[0149] print(f"Total with Tax: {calculate_total(items, tax_rate):.2f}")# Clearly display the total amount after tax

[0150] By applying the solution of the embodiment of this specification, a target modified file of an original code file is generated according to the region identifier sequence and the identifier mapping information, so that the target modified file can be generated efficiently and accurately.

[0151] Considering that the model parameters of the code processing model are relatively large and the computing resources of the terminal device are limited, the code task processing method proposed in the embodiment of this specification can be applied to Figure 2 The code task processing system shown is not limited to this. Figure 2 , Figure 2 An architecture diagram of a code task processing system provided by an embodiment of the present specification is shown, and the code task processing system may include a terminal device 202 and a server 204;

[0152] The terminal device 202 is used to send the original code file and the original modification plan of the original code file to the server 204;

[0153] The server 204 is used to add a region identifier to the original code file and the original modification plan to obtain a code file to be processed and a modification plan to be processed; input the code file to be processed and the modification plan to be processed into a code processing model to obtain a region identifier sequence; generate a target modification file of the original code file according to the region identifier sequence; and send the target modification file to the terminal device 202;

[0154] The terminal device 202 is also used to receive the target modification file sent by the server 204.

[0155] like Figure 2As shown, the code processing model is deployed in the server 204, and the server 204 can be connected to one or more terminal devices 202 through a local area network connection, a wide area network connection, an Internet connection, or other types of data networks. The terminal device 202 may include, but is not limited to: a smart phone, a tablet computer, a laptop computer, a PDA, a personal computer, a smart home device, a vehicle-mounted device, etc. The terminal device 202 can also interact with the user through a graphical user interface to implement the call to the code processing model, and then implement the code task processing method provided in the embodiment of this specification. The code task processing method provided in the embodiment of this specification is generally executed by the server, but in other embodiments of this specification, when the operating resources of the terminal device can meet the deployment and operating conditions of the code processing model, the terminal device can also have similar functions with the server, thereby executing the code task processing method provided in the embodiment of this specification. In other embodiments, the code task processing method provided in the embodiment of this specification can also be jointly executed by the terminal device and the server.

[0156] The following combination Figure 3 , taking the application of the code task processing method provided in this specification in the code error correction scenario as an example, the code task processing method is further described. Figure 3 A flowchart of a code error correction method provided by an embodiment of this specification is shown, which specifically includes the following steps:

[0157] Step 302: Obtain an original code file and an error correction solution for the original code file.

[0158] Step 304: Add the region identifier to the original code file and the error correction solution to obtain the code file to be processed and the error correction solution to be processed.

[0159] Step 306: Input the code file to be processed and the error correction solution to be processed into the code processing model to obtain a region identifier sequence.

[0160] Step 308: Generate an error correction code file of the original code file according to the region identifier sequence.

[0161] It should be noted that the error correction scheme refers to the detailed modification suggestions and steps proposed for the errors or problems existing in the original code file. The error correction scheme includes the accurate modification code corresponding to the error code in the original code file. The error correction code file refers to the code file obtained after the errors in the original code file are modified based on the error correction scheme. The error correction code file does not include the errors proposed in the error correction scheme. The implementation method of steps 302 to 308 can refer to the implementation method of the above-mentioned steps 102 to 108, and will not be repeated in the embodiments of this specification.

[0162] By applying the solution of the embodiments of the present specification, by adding region identifiers to the original code file and the error correction solution, the code processing model can accurately locate the position of the code content, thereby switching the model code generation mode to a mode of only generating region identifiers, which greatly reduces the length of the model output text and significantly speeds up the generation speed. In addition, since each region identifier is associated with the original code file and the error correction solution, after obtaining the region identifier sequence, the error correction code file can be quickly generated, thereby improving the generation efficiency of the error correction code file.

[0163] The following combination Figure 4 , the training method of the code processing model used in the code task processing method and the code error correction method provided in this specification is described. Among them, Figure 4 A flowchart of a code processing model training method provided by an embodiment of this specification is shown, which specifically includes the following steps:

[0164] Step 402: Obtain a sample set, wherein the sample set includes a sample code file, a sample modification scheme, and a sample identifier sequence.

[0165] Step 404: Add the sample region identifier to the sample code file and the sample modification plan to obtain the sample processing file and the sample processing plan.

[0166] Step 406: Input the sample processing file and the sample processing plan into the code processing model to obtain a predicted identifier sequence.

[0167] Step 408: Train the code processing model according to the predicted identifier sequence and the sample identifier sequence to obtain a trained code processing model.

[0168] It should be noted that the training method of the code processing model is supervised fine-tuning, that is, the sample set includes the real processing label (sample identifier sequence) of the task processing model. The sample identifier sequence is the generation target of the code processing model, which is used to guide the training process of the code processing model. The code processing model trained according to the prediction identifier sequence and the sample identifier sequence can be a large model of the code field or an untrained neural network model. The definition of "sample code file, sample modification scheme and prediction identifier sequence" can refer to the definition of "original code file, original modification scheme and region identifier sequence" mentioned above, and the implementation method of step 404 and step 406 can refer to the implementation method of step 102 and step 106 mentioned above, and the embodiment of this specification will not be repeated. The sample identifier sequence includes multiple sample region identifiers. The prediction identifier sequence includes multiple prediction region identifiers. There are many ways to obtain a sample set, which are selected according to the actual situation, and the embodiment of this specification does not limit this. In one possible implementation of this specification, a sample set including a sample code file, a sample modification scheme and a sample identifier sequence can be read from other data acquisition devices or databases. In another possible implementation of the present specification, a sample set including a sample code file, a sample modification scheme, and a sample identifier sequence sent by a user through a terminal device may be received.

[0169] In practical applications, when the code processing model is trained according to the prediction identifier sequence and the sample identifier sequence, the loss value can be calculated according to the prediction identifier sequence and the sample identifier sequence, and the model parameters of the code processing model can be adjusted according to the loss value until the training process meets the preset stop condition, and the trained code processing model is obtained. Among them, there are many functions for calculating the loss value, such as the cross entropy loss function, the L1 norm loss function, the maximum loss function, the mean square error loss function, the logarithmic loss function, etc., which are selected according to the actual situation, and the embodiments of this specification do not make any restrictions on this. The preset stop conditions include but are not limited to the loss value being less than or equal to the preset threshold and the number of iterations reaching the preset number of iterations, wherein the preset threshold and the preset number of iterations are selected according to the actual situation, and the embodiments of this specification do not make any restrictions on this.

[0170] In a possible implementation of the present specification, after calculating the loss value, the loss value is compared with a preset threshold. Specifically, if the loss value is greater than the preset threshold, it means that the difference between the predicted identifier sequence and the sample identifier sequence is large, and the code processing model has poor prediction ability for the predicted identifier sequence. At this time, the model parameters of the code processing model can be adjusted, and the step of inputting the sample processing file and the sample processing scheme into the code processing model to obtain the predicted identifier sequence is returned to continue training the code processing model until the loss value is less than or equal to the preset threshold, indicating that the difference between the predicted identifier sequence and the sample identifier sequence is small, and the preset stop condition is reached, and the trained code processing model is obtained.

[0171] In another possible implementation of the present specification, in addition to comparing the size relationship between the loss value and the preset threshold, the number of iterations can also be combined to determine whether the current code processing model has been trained. Specifically, if the loss value is greater than the preset threshold, the model parameters of the code processing model are adjusted, and the step of inputting the sample processing file and the sample processing scheme into the code processing model to obtain the predicted identifier sequence is returned to continue training the code processing model until the preset number of iterations is reached, and the iteration is stopped to obtain a trained code processing model.

[0172] It is worth noting that when the code processing model is trained based on the sample set, the sample region identifier can be given to the code processing model as a special tag of the model to provide additional information to the code processing model to help the code processing model understand the specific format of the input and output. In addition, since the sample region identifier is given to the code processing model as a special tag of the model, the code processing model can recognize and correctly process the sample region identifier, ensuring that each region identifier is converted to only one token instead of attempting further segmentation. Specifically, when the sample region identifier is given to the code processing model as a special tag of the model, the tokenizer of the code processing model can be updated.

[0173] By applying the solution of the embodiment of this specification, the code processing model is trained according to the predicted identifier sequence and the sample identifier sequence, and the code processing model is continuously trained when the preset stop condition is not met until the preset stop condition is met, and the code processing model is obtained after the training is completed. By continuously adjusting the model parameters of the code processing model, the code processing model finally obtained can have accurate identifier sequence prediction capabilities.

[0174] In an optional embodiment of the present specification, the obtaining of the sample set may include the following steps:

[0175] Obtain a sample code file and a sample modification file of the sample code file;

[0176] Generate a sample modification plan according to the sample code file and the sample modification file;

[0177] A sample identifier sequence is generated based on the sample modification file and the sample region identifier.

[0178] It should be noted that the sample modification file refers to a code file obtained by modifying the sample code file. The sample modification plan refers to a modification suggestion generated for the sample code file based on the sample modification file. The sample modification file can be obtained by modifying the sample code file according to the sample modification plan. The process of generating the sample modification plan based on the sample code file and the sample modification file can be understood as the process of generating the code modification suggestion.

[0179] In practical applications, there are many ways to obtain sample code files and sample modification files of sample code files, and the embodiments of this specification do not limit this. In one possible implementation of this specification, sample code files and sample modification files of sample code files can be read from other databases or data acquisition devices. In another possible implementation of this specification, sample code files and sample modification files of sample code files sent by users through terminal devices can be received.

[0180] Furthermore, there are multiple ways to generate a sample modification scheme based on the sample code file and the sample modification file, and the embodiments of this specification do not limit this. In one possible implementation of this specification, the differences between the sample code file and the sample modification file can be compared, and the sample modification scheme can be written based on the differences. In another possible implementation of this specification, a modification scheme generation model can be used to generate a sample modification scheme based on the sample code file and the sample modification file.

[0181] By applying the solution of the embodiment of this specification, a sample modification plan is generated according to the sample code file and the sample modification file, which ensures the accuracy of the sample modification plan, further improves the accuracy of the sample identifier sequence, and provides precise guidance for the code processing model training process.

[0182] In an optional embodiment of the present specification, the above-mentioned generation of the sample modification scheme according to the sample code file and the sample modification file may include the following steps:

[0183] The sample code file and the sample modification file are input into the modification solution generation model to obtain the sample modification solution.

[0184] It should be noted that the modification scheme generation model refers to a deep learning model with the ability to generate modification schemes. The modification scheme generation model can be a large model in the code field, or it can be a deep learning model trained based on training modification schemes, training code files, and training modification files. The training process of the modification scheme generation model is supervised fine-tuning, and the training process of the above-mentioned code processing model can be referred to, and the embodiments of this specification will not be repeated.

[0185] By applying the solution of the embodiment of this specification and utilizing the code processing capability of the modification solution generation model, a sample modification solution is generated according to a sample code file and a sample modification file, thereby improving the accuracy of the sample modification solution.

[0186] In practical applications, there are many ways to generate a sample identifier sequence based on the sample modification file and the sample region identifier, which are selected according to the actual situation, and the embodiments of this specification do not limit this. In a possible implementation of this specification, sample code lines with high similarity to each modification line in the sample modification file can be screened out from the sample processing file and the sample processing scheme, and the sample region identifier included in the sample code line can be used to replace the corresponding modification line in the sample modification file to obtain a sample identifier sequence.

[0187] In another possible implementation of the present specification, the sample region identifier includes a first sample identifier and a second sample identifier, the first sample identifier is used to identify the position of the first sample code line in the sample code file, and the second sample identifier is used to identify the position of the second sample code line in the sample modification solution; the above-mentioned generating a sample identifier sequence according to the sample modification file and the sample region identifier may include the following steps:

[0188] Filter out a first sample modification line and a second sample modification line from the sample modification file, wherein the first sample modification line is the same as the first sample code line, and the second sample modification line is the same as the second sample code line;

[0189] The first sample modification line is replaced with the first sample identifier, and the second sample modification line is replaced with the second sample identifier to obtain a sample identifier sequence.

[0190] Exemplarily, the target modification file in the above example is used as a sample modification file, the original code file in the above example is used as a sample code file, and the original modification scheme including the modification content context in the above example is used as a sample modification scheme. The five first sample modification lines selected from the sample modification file are as follows:

[0191] def calculate_total(prices, tax_rate):

[0192] total = sum(prices)

[0193] # This is a blank line

[0194] items = [10, 20, 30]

[0195] tax_rate = 0.08

[0196] The four second sample modification lines selected from the sample modification file are as follows:

[0197] tax_amount = total * tax_rate

[0198] total_with_tax = total + tax_amount

[0199] return total_with_tax# Modification: Clearly return the total amount after tax

[0200] print(f"Total with Tax: {calculate_total(items, tax_rate):.2f}")# Clearly display the total amount after tax

[0201] Assume that sample area identifiers are added in the sample code file and the sample modification plan, and in the process of obtaining the sample processing file and the sample processing plan, the first sample identifier added to the first sample code line “def calculate_total(prices, tax_rate):” is “@P0@”, the first sample identifier added to the first sample code line “total = sum(prices)” is “@M0@”, the first sample identifier added to the first sample code line “ ” is “@P4@”, the first sample identifier added to the first sample code line “items = [10, 20, 30]” is “@P5@”, the first sample identifier added to the first sample code line “tax_rate = 0.08” is “@P6@”, the second sample identifier added to the second sample code line “tax_amount =total * tax_rate” is “@M1@”, the second sample identifier added to the second sample code line “total_with_tax= total + tax_amount” is “@M2@”, the first sample identifier added to the second sample code line “returntotal_with_tax# Modification: The second sample identifier added for "Explicitly return the total amount after tax" is "@M3@", and the second sample identifier added for the second sample code line "print(f"Total with Tax: {calculate_total(items, tax_rate):.2f}")# Explicitly display the total amount after tax" is "@M7@". Therefore, the first sample modification line is replaced with the first sample identifier, and the second sample modification line is replaced with the second sample identifier, and the sample identifier sequence is: @P0@@M0@@M1@@M2@@M3@@P4@@P5@@P6@@M7@.

[0202] By applying the solution of the embodiment of this specification, the first sample modification line in the sample modification file is replaced with the first sample identifier, and the second sample modification line in the sample modification file is replaced with the second sample identifier, thereby ensuring the accuracy of the sample identifier sequence.

[0203] See also Figure 5 , Figure 5 A flowchart of a processing process of a code task processing method provided by an embodiment of the present specification is shown. The processing process of the code task processing method can be divided into a code processing model training phase and a code task processing phase. Next, the implementation process of the above two phases is described in detail.

[0204] Code processing model training phase: obtain sample code files and sample modification files of sample code files; input the sample code files and sample modification files into the modification scheme generation model to obtain the sample modification scheme; add sample area identifiers to the sample code files and sample modification schemes to obtain sample processing files and sample processing schemes; use the sample area identifiers to represent the sample modification files to obtain a sample identifier sequence; input the sample processing files and sample processing schemes into the code processing model to obtain a predicted identifier sequence; train the code processing model based on the predicted identifier sequence and the sample identifier sequence to obtain a trained code processing model.

[0205] Code task processing stage: obtain the original code file and the original modification plan of the original code file; add a first line identifier before each first code line included in the original code file to obtain the code file to be processed, and add a second line identifier before each second code line included in the original modification plan to obtain the modification plan to be processed, wherein the first line identifier and the second line identifier are different, the first line identifier is used to identify the position of the first code line in the original code file (such as the line number), and the second area identifier is used to identify the line number of the second code line in the original modification plan (such as the line number). Input the code file to be processed and the modification plan to be processed into the trained code processing model to obtain a line identifier sequence; generate the target modification file of the original code file according to the line identifier sequence.

[0206] By applying the solution of the embodiments of the present specification, by adding line identifiers in the original code file and the original modification plan, the code processing model can accurately locate the line where the code content is located, thereby switching the model code generation mode to a mode of only generating line identifiers, greatly reducing the length of the model output text and significantly speeding up the generation speed. Moreover, since each line identifier is associated with the content in the original code file and the original modification plan, after obtaining the line identifier sequence, the target modification file can be quickly mapped through the line identifiers, thereby improving the generation efficiency of the target modification file.

[0207] See also Figure 6 , Figure 6 A flowchart of an information processing method based on a code processing model provided by an embodiment of the present specification is shown. The information processing method based on the code processing model is applied to a task platform and specifically includes the following steps:

[0208] Step 602: Receive a model request sent by a terminal device, wherein the model request includes a scenario identifier of a target code scenario, scenario input data of the target code scenario, and at least one of a model specification parameter.

[0209] Step 604: Based on the model request, determine a target code processing model from a plurality of code processing models, wherein the plurality of code processing models are trained using a code processing model training method.

[0210] It should be noted that the target code processing model is a code processing model suitable for the target code scenario. Based on the model request, there are multiple ways to determine the target code processing model from multiple code processing models. In one possible implementation of this specification, based on the model request, the corresponding target code processing model can be found from at least one code processing model included in the model library; in another possible implementation, based on the model request, the target code processing model can be trained and obtained; in another optional method of this specification, the target code processing model can be constructed based on the model request.

[0211] For example, based on the scenario identification of the target code scenario, at least one pre-trained code processing model can be searched from the model library, and then based on the model specification parameters, an initial code processing model can be obtained by screening at least one code processing model, and then based on the scenario input data of the target code scenario, the initial code processing model obtained by screening can be trained to obtain a target code processing model suitable for user needs. Figure 4 The code processing model is obtained by training using the training method shown, and will not be described in detail in the embodiments of this specification.

[0212] The solution of the embodiment of this specification is applied to obtain the target code processing model according to user needs, realize personalized model service, provide users with an efficient, flexible and easy-to-use model service method, and improve user experience.

[0213] In an optional embodiment of the present specification, the model request includes a scenario identifier of a target code scenario; and determining the target code processing model from a plurality of code processing models based on the model request may include the following steps:

[0214] Based on the scenario identifier of the target code scenario, a target code processing model suitable for the target code scenario is searched from a model library, wherein the model library stores a plurality of code processing models suitable for different code scenarios.

[0215] It should be noted that scenario identification refers to a unique or specific label used to distinguish different code task scenarios. The model library is a database for storing and managing various pre-trained deep learning models. Multiple code processing models adapted to different code task scenarios cover different code application scenarios and requirements. The model library allows users to select appropriate models according to their needs, or directly call the model for code task processing through the application programming interface.

[0216] The various models stored in the model library that are suitable for different code task scenarios are optimized for specific application environments. Figure 4 The model obtained by training the code processing model shown in the training method is as follows: For example, based on the scenario identifier "code optimization" of the target code scenario, a target code processing model suitable for the code optimization scenario can be searched from the model library.

[0217] By applying the solution of the embodiments of this specification, based on scenario requirements, the target code processing model suitable for the scenario is accurately found through scenario identification, so that code task processing is more accurate and more in line with the scenario, thereby improving user experience and code task processing quality.

[0218] In an optional embodiment of the present specification, the model request includes scenario input data of the target code scenario; and the above-mentioned determining the target code processing model from multiple code processing models based on the model request may include the following steps:

[0219] Determine an initial code processing model suitable for a target code scenario from among multiple code processing models;

[0220] Based on the scenario input data of the target code scenario, the initial code processing model is trained to obtain the target code processing model.

[0221] It should be noted that the initial code processing model refers to a model among multiple code processing models that is applicable to the target code scenario. The initial code processing model may not only be applicable to the target code scenario, but also to other code scenarios. It is a general code processing model that can be applied to different code scenarios. The initial code processing model can be used to process target code tasks, but the effect may not be very good. At this time, the parameters of the initial code processing model can be adjusted based on the scenario input data of the target code scenario. For example, the initial code processing model can be optimized based on the scenario input data of the code optimization scenario to obtain a target code processing model suitable for the code optimization scenario. The scenario input data of the target code scenario can be understood as a sample set of sample code tasks in the target code scenario (including sample code files, sample modification plans, and sample identifier sequences). The initial code processing model is trained based on the scenario input data of the target code scenario. The method of obtaining the target code processing model can refer to Figure 4 The training method of the code processing model shown in the present specification will not be described in detail in the embodiments.

[0222] By applying the solution of the embodiments of this specification, based on scenario requirements, a general initial code processing model is further trained through scenario input data to obtain a target code processing model adapted to the scenario, so that the target code processing model is more in line with the scenario, thereby improving the user experience and the processing quality of code tasks.

[0223] In an optional embodiment of the present specification, the model request includes model specification parameters; and determining the target code processing model from the plurality of code processing models based on the model request may include the following steps:

[0224] Based on the model specification parameters, a corresponding target code processing model is searched from a model library, wherein the model library stores a plurality of code processing models with different model specification parameters.

[0225] It should be noted that model specification parameters refer to various parameters that define the structure and behavior of the model. These parameters can be roughly divided into two categories: model parameters (learnable parameters) and hyperparameters. Model parameters refer to parameters that are automatically adjusted by the back-propagation algorithm during model training, including but not limited to weight matrices (weights) and biases (biases). For example, in a simple fully connected layer, the weight matrix is ​​a two-dimensional tensor that connects the neurons in the input layer and the output layer; the bias term is a one-dimensional vector that provides an additional offset value for each output neuron. Hyperparameters refer to parameters set before starting model training to control the learning process and architecture of the model. Hyperparameters include but are not limited to learning rate (Learning Rate) and the number of neurons per layer (Number of Neurons per Layer), which should be selected according to actual conditions.

[0226] By applying the solution of the embodiments of this specification, the corresponding target code processing model can be accurately found based on the model specification parameters, thereby ensuring the efficient and stable operation of the target code processing model and improving the user experience.

[0227] In an optional embodiment of the present specification, after determining the target code processing model from multiple code processing models based on the model request, the following steps may also be included:

[0228] Deploy the target code processing model, and build a code task processing interface based on the target code processing model to enable the terminal device to schedule the target code processing model to execute the target code task.

[0229] It should be noted that the code task processing interface is an interactive programming interface for the terminal device to schedule the target code processing model to perform target code task processing, which is usually provided in the form of an application programming interface. Through the code task processing interface, the user can input the task data of the target code task (including the original code file and the original modification plan of the original code file) to perform code processing.

[0230] In practical applications, there are many ways to deploy the target code processing model. In one possible implementation, the target code processing model can be deployed on a cloud-side device using the infrastructure provided by a cloud service provider. In another possible implementation, the target code processing model can be deployed on an edge device using a lightweight framework. For example, the target code processing model can be deployed on a distributed system, and based on the target code processing model, a code task processing interface can be constructed and provided to the terminal device, so that the terminal device can schedule the target code processing model to execute the target code task.

[0231] By applying the solution provided in the embodiments of this specification, deploying the target code processing model, and building a code task processing interface based on the target code processing model, the terminal device can efficiently call the target code processing model and improve the processing quality and response speed of the target code task.

[0232] See also Figure 7 , Figure 7 A schematic diagram of the structure of a task platform provided by an embodiment of the present specification is shown, wherein the task platform 700 includes a request interface 702 and a response unit 704;

[0233] The request interface 702 is used to receive a model request sent by a terminal device, wherein the model request includes a scenario identifier of a target code scenario, scenario input data of the target code scenario, and at least one of a model specification parameter;

[0234] The response unit 704 is used to determine a target code processing model from a plurality of code processing models based on the model request, wherein the plurality of code processing models are trained using a code processing model training method.

[0235] In an optional embodiment of the present specification, the task platform further includes a code task processing interface, and the code task processing interface is constructed based on the target code processing model;

[0236] The code task processing interface is used for terminal devices to schedule and execute target code tasks.

[0237] By applying the solution of the embodiments of this specification, the task platform adapts to user needs to obtain the target code processing model, realizes personalized model services, provides users with an efficient, flexible and easy-to-use model service platform, and improves user experience.

[0238] The above is a schematic scheme of a task platform of this embodiment. It should be noted that the technical scheme of the task platform and the technical scheme of the information processing method based on the code processing model belong to the same concept. For details not described in detail in the technical scheme of the task platform, please refer to the description of the technical scheme of the information processing method based on the code processing model.

[0239] Corresponding to the above-mentioned code task processing method embodiment, this specification also provides a code task processing device embodiment, Figure 8 FIG. 1 shows a schematic diagram of the structure of a code task processing device provided by an embodiment of the present specification. Figure 8 As shown, the device comprises:

[0240] A first acquisition module 802 is configured to acquire an original code file and an original modification plan of the original code file;

[0241] A first adding module 804 is configured to add a region identifier to an original code file and an original modification plan to obtain a code file to be processed and a modification plan to be processed;

[0242] A first input module 806 is configured to input the code file to be processed and the modification plan to be processed into the code processing model to obtain a region identifier sequence;

[0243] The first generating module 808 is configured to generate a target modified file of the original code file according to the region identifier sequence.

[0244] Optionally, the area identifier includes a first line identifier and a second line identifier, the original code file includes at least one first code line, and the original modification plan includes at least one second code line; the first adding module 804 is further configured to add the first line identifier to the first code line to obtain the code file to be processed, and to add the second line identifier to the second code line to obtain the modification plan to be processed, wherein the first line identifier is used to identify the position of the first code line, and the second line identifier is used to identify the position of the second code line.

[0245] Optionally, the first adding module 804 is further configured to add a first line identifier before the first code line to obtain a code file to be processed, and add a second line identifier before the second code line to obtain a modification plan to be processed.

[0246] Optionally, the device also includes: a screening module, configured to screen out unmodified code lines that are identical to the first code lines from the second code lines; replace the unmodified code lines with unmodified identifiers to obtain replaced second code lines; and a first adding module 804, further configured to add a second line identifier to the replaced second code lines to obtain a modification plan to be processed.

[0247] Optionally, the first generation module 808 is further configured to obtain identifier mapping information, wherein the identifier mapping information is used to describe the content correspondence between the region identifier and the original code file and the original modification plan respectively; and generate a target modification file of the original code file according to the region identifier sequence and the identifier mapping information.

[0248] By applying the solution of the embodiments of the present specification, by adding region identifiers to the original code file and the original modification plan, the code processing model can accurately locate the position of the code content, thereby switching the model code generation mode to a mode of only generating region identifiers, greatly reducing the length of the model output text and significantly speeding up the generation speed. Moreover, since each region identifier is associated with the original code file and the original modification plan, after obtaining the region identifier sequence, the target modification file can be quickly generated, thereby improving the generation efficiency of the target modification file.

[0249] The above is a schematic scheme of a code task processing device of this embodiment. It should be noted that the technical scheme of the code task processing device and the technical scheme of the above-mentioned code task processing method belong to the same concept, and the details not described in detail in the technical scheme of the code task processing device can be referred to the description of the technical scheme of the above-mentioned code task processing method.

[0250] Corresponding to the above-mentioned code error correction method embodiment, this specification also provides a code error correction device embodiment, Fig. 9 FIG. 2 shows a schematic diagram of the structure of a code error correction device provided by an embodiment of the present specification. Fig. 9 As shown, the device comprises:

[0251] A second acquisition module 902 is configured to acquire an original code file and an error correction solution for the original code file;

[0252] The second adding module 904 is configured to add the region identifier to the original code file and the error correction solution to obtain the code file to be processed and the error correction solution to be processed;

[0253] The second input module 906 is configured to input the code file to be processed and the error correction solution to be processed into the code processing model to obtain a region identifier sequence;

[0254] The second generating module 908 is configured to generate an error correction code file of the original code file according to the region identifier sequence.

[0255] By applying the solution of the embodiments of the present specification, by adding region identifiers to the original code file and the error correction solution, the code processing model can accurately locate the position of the code content, thereby switching the model code generation mode to a mode of only generating region identifiers, which greatly reduces the length of the model output text and significantly speeds up the generation speed. In addition, since each region identifier is associated with the original code file and the error correction solution, after obtaining the region identifier sequence, the error correction code file can be quickly generated, thereby improving the generation efficiency of the error correction code file.

[0256] The above is a schematic scheme of a code error correction device of this embodiment. It should be noted that the technical scheme of the code error correction device and the technical scheme of the above-mentioned code error correction method belong to the same concept, and the details not described in detail in the technical scheme of the code error correction device can be referred to the description of the technical scheme of the above-mentioned code error correction method.

[0257] Corresponding to the above-mentioned code processing model training method embodiment, this specification also provides a code processing model training device embodiment, Fig.10 FIG. 1 is a schematic diagram showing the structure of a code processing model training device provided by an embodiment of the present specification. Fig.10 As shown, the device comprises:

[0258] The third acquisition module 1002 is configured to acquire a sample set, wherein the sample set includes a sample code file, a sample modification scheme, and a sample identifier sequence;

[0259] The third adding module 1004 is configured to add the sample region identifier to the sample code file and the sample modification scheme to obtain the sample processing file and the sample processing scheme;

[0260] The third input module 1006 is configured to input the sample processing file and the sample processing plan into the code processing model to obtain a predicted identifier sequence;

[0261] The training module 1008 is configured to train the code processing model according to the prediction identifier sequence and the sample identifier sequence to obtain a trained code processing model.

[0262] Optionally, the third acquisition module 1002 is further configured to acquire a sample code file and a sample modification file of the sample code file; generate a sample modification scheme according to the sample code file and the sample modification file; and generate a sample identifier sequence according to the sample modification file and the sample region identifier.

[0263] Optionally, the third acquisition module 1002 is further configured to input the sample code file and the sample modification file into a modification solution generation model to obtain a sample modification solution.

[0264] Optionally, the sample area identifier includes a first sample identifier and a second sample identifier, the first sample identifier is used to identify the position of the first sample code line in the sample code file, and the second sample identifier is used to identify the position of the second sample code line in the sample modification scheme; the third acquisition module 1002 is further configured to filter out the first sample modification line and the second sample modification line from the sample modification file, wherein the first sample modification line is the same as the first sample code line, and the second sample modification line is the same as the second sample code line; the first sample modification line is replaced by the first sample identifier, and the second sample modification line is replaced by the second sample identifier to obtain a sample identifier sequence.

[0265] By applying the solution of the embodiment of this specification, the code processing model is trained according to the predicted identifier sequence and the sample identifier sequence, and the code processing model is continuously trained when the preset stop condition is not met until the preset stop condition is met, and the code processing model is obtained after the training is completed. By continuously adjusting the model parameters of the code processing model, the code processing model finally obtained can have accurate identifier sequence prediction capabilities.

[0266] The above is a schematic scheme of a code processing model training device of this embodiment. It should be noted that the technical scheme of the code processing model training device and the technical scheme of the above-mentioned code processing model training method belong to the same concept, and the details not described in detail in the technical scheme of the code processing model training device can be referred to the description of the technical scheme of the above-mentioned code processing model training method.

[0267] Corresponding to the above-mentioned information processing method embodiment based on the code processing model, this specification also provides an information processing device embodiment based on the code processing model. Fig.11 FIG. 1 shows a schematic diagram of the structure of an information processing device based on a code processing model provided by an embodiment of the present specification. Fig.11 As shown, the device is applied to a task platform and includes:

[0268] The receiving module 1102 is configured to receive a model request sent by a terminal device, wherein the model request includes a scenario identifier of a target code scenario, scenario input data of the target code scenario, and at least one of a model specification parameter;

[0269] The determination module 1104 is configured to determine a target code processing model from a plurality of code processing models based on the model request, wherein the plurality of code processing models are trained using a code processing model training method.

[0270] Optionally, the model request includes a scenario identifier of the target code scenario; the determination module 1104 is further configured to search for a target code processing model suitable for the target code scenario from a model library based on the scenario identifier of the target code scenario, wherein the model library stores multiple code processing models suitable for different code scenarios.

[0271] Optionally, the model request includes scenario input data of the target code scenario; the determination module 1104 is further configured to determine an initial code processing model suitable for the target code scenario from multiple code processing models; based on the scenario input data of the target code scenario, the initial code processing model is trained to obtain the target code processing model.

[0272] Optionally, the model request includes model specification parameters; the determination module 1104 is further configured to search for a corresponding target code processing model from a model library based on the model specification parameters, wherein the model library stores a plurality of code processing models with different model specification parameters.

[0273] Optionally, the apparatus further includes: a deployment module configured to deploy a target code processing model, and construct a code task processing interface based on the target code processing model, so that the terminal device schedules the target code processing model to execute the target code task.

[0274] The solution of the embodiment of this specification is applied to obtain the target code processing model according to user needs, realize personalized model service, provide users with an efficient, flexible and easy-to-use model service method, and improve user experience.

[0275] The above is a schematic scheme of an information processing device based on a code processing model of this embodiment. It should be noted that the technical scheme of the information processing device based on the code processing model and the technical scheme of the information processing method based on the code processing model belong to the same concept, and the details not described in detail in the technical scheme of the information processing device based on the code processing model can be referred to the description of the technical scheme of the information processing method based on the code processing model.

[0276] Fig.12 The structural block diagram of a computing device 1200 provided by an embodiment of the present specification is shown. The computing device 1200 includes: a memory 1210 and a processor 1220; the memory 1210 is used to store computer programs / instructions, and the processor 1220 is used to execute computer programs / instructions. When the computer program / instructions are executed by the processor 1220, the steps of the above-mentioned code task processing method, code error correction method, code processing model training method, or information processing method based on code processing model are implemented.

[0277] In one or more embodiments of the present specification, the computing device may be understood as an integrated intelligent terminal, including but not limited to a server, a desktop computer, a personal computer (PC), an all-in-one model machine, a mobile phone, a tablet computer or other portable intelligent terminal, etc. The computing device may be pre-installed with the model described in the above embodiments of the present application.

[0278] Specifically, the computing device can pre-set multiple types of models, including but not limited to models in the fields of natural language processing, visual processing, speech processing, code processing, multimodal task processing, etc., so as to provide a variety of model choices. In different product forms, the computing device can support one or more model usage methods, including but not limited to model training, model calling, model fine-tuning, model deployment, model reasoning and application, etc. In some product forms, the computing device also supports model management, including but not limited to multi-type model management (supporting the management of multiple types of models such as discriminants and generative models), model version control (supporting the control of different model versions), model evaluation (based on model evaluation tools, evaluating the performance and effect of the model), etc. In other product forms, the computing device can also create applications based on models, provide application programming interface (API, Application Programming Interface) calling capabilities, and can call models to the created applications through API interfaces, while providing application management tools to achieve management and monitoring of applications.

[0279] Furthermore, the computing device can also include data management (supporting the creation and management of model tuning data sets), a training center (providing rich training resources to help users learn and master artificial intelligence technology), and basic management and control capabilities (providing enterprise-level basic management and control capabilities to ensure the security and efficient operation of the system). Through the above functions, a comprehensive, integrated artificial intelligence development, training, deployment and application device is provided.

[0280] Fig.13 A structural block diagram of an electronic device 1300 provided according to an embodiment of the present specification is shown.

[0281] The memory 1310 and the processor 1320 are connected via a bus 1330;

[0282] The memory 1310 is used to store computer programs / instructions, and the processor 1320 is used to execute computer programs / instructions. When the computer program / instructions are executed by the processor 1320, the steps of the above-mentioned code task processing method or code error correction method or code processing model training method or information processing method based on the code processing model are implemented.

[0283] Specifically, the components of the electronic device 1300 include, but are not limited to, a memory 1310 and a processor 1320. The processor 1320 and the memory 1310 may be connected via a bus 1330. The electronic device 1300 may further include an access device 1340, which enables the electronic device 1300 to communicate with a database 1350 storing data via one or more networks 1360. Examples of these networks include a public switched telephone network (PSTN), a local area network (LAN), a wide area network (WAN), a personal area network (PAN), or a combination of communication networks such as the Internet. The access device 1340 may include one or more of any type of network interface, wired or wireless (e.g., a network interface card (NIC)), such as an IEEE802.11 wireless local area network (WLAN) wireless interface, a World Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, a Near Field Communication (NFC) interface, and the like.

[0284] In one embodiment of the present specification, the above components of the electronic device 1300 and Fig.13 Other components not shown in the figure may also be connected to each other, such as through a bus. It should be understood that Fig.13 The electronic device structure block diagram shown is only for the purpose of illustration, and is not intended to limit the scope of this specification. Those skilled in the art can add or replace other components as needed.

[0285] The electronic device 1300 may be any type of stationary or mobile electronic device, including a mobile computer or mobile computing device (e.g., a tablet computer, a personal digital assistant, a laptop computer, a notebook computer, a netbook, etc.), a mobile phone (e.g., a smart phone), a wearable electronic device (e.g., a smart watch, smart glasses, etc.), or other types of mobile devices, or a stationary electronic device such as a desktop computer or a PC. The electronic device 1300 may also be a mobile or stationary electronic device.

[0286] The above is a schematic scheme of an electronic device of this embodiment. It should be noted that the technical scheme of the electronic device and the technical schemes of the above-mentioned code task processing method, code error correction method, code processing model training method and information processing method based on the code processing model belong to the same concept. For details not described in detail in the technical scheme of the electronic device, please refer to the description of the technical scheme of the above-mentioned code task processing method or code error correction method or code processing model training method or information processing method based on the code processing model.

[0287] An embodiment of the present specification also provides a computer-readable storage medium storing a computer program / instruction, which, when executed by a processor, implements the steps of the above-mentioned code task processing method or code error correction method or code processing model training method or information processing method based on a code processing model.

[0288] The above is a schematic scheme of a computer-readable storage medium of this embodiment. It should be noted that the technical scheme of the storage medium and the technical scheme of the above-mentioned code task processing method, code error correction method, code processing model training method and information processing method based on the code processing model belong to the same concept, and the details not described in detail in the technical scheme of the storage medium can be referred to the description of the technical scheme of the above-mentioned code task processing method or code error correction method or code processing model training method or information processing method based on the code processing model.

[0289] An embodiment of the present specification also provides a computer program product, including a computer program / instruction, which, when executed by a processor, implements the steps of the above-mentioned code task processing method or code error correction method or code processing model training method or information processing method based on a code processing model.

[0290] The above is a schematic scheme of a computer program product of this embodiment. It should be noted that the technical scheme of the computer program product and the technical scheme of the above-mentioned code task processing method, code error correction method, code processing model training method and information processing method based on code processing model belong to the same concept, and the details not described in detail in the technical scheme of the computer program product can be referred to the description of the technical scheme of the above-mentioned code task processing method or code error correction method or code processing model training method or information processing method based on code processing model.

[0291] The above is a description of a specific embodiment of the specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0292] The computer instructions include computer program codes, which may be in source code form, object code form, executable files or some intermediate forms, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium, etc. It should be noted that the content contained in the computer-readable medium may be appropriately increased or decreased according to the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media do not include electric carrier signals and telecommunication signals.

[0293] It should be noted that, for the above-mentioned method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the embodiments of this specification are not limited by the order of the actions described, because according to the embodiments of this specification, some steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the embodiments of this specification.

[0294] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0295] The preferred embodiments of this specification disclosed above are only used to help explain this specification. The optional embodiments do not describe all the details in detail, nor do they limit the invention to only the specific implementation methods described. Obviously, many modifications and changes can be made according to the content of the embodiments of this specification. This specification selects and specifically describes these embodiments in order to better explain the principles and practical applications of the embodiments of this specification, so that technicians in the relevant technical field can well understand and use this specification. This specification is only limited by the claims and their full scope and equivalents.

Claims

1. A code task processing method, comprising: Obtaining an original code file and an original modification plan of the original code file; Adding a region identifier to the original code file and the original modification scheme to obtain a code file to be processed and a modification scheme to be processed, wherein the region identifier is used to indicate the region position where the code content in the original code file is located and the region position where the modification content in the original modification scheme is located; Inputting the code file to be processed and the modification plan to be processed into a code processing model to obtain a region identifier sequence; A target modified file of the original code file is generated according to the region identifier sequence.

2. The method according to claim 1, wherein the region identifier comprises a first line identifier and a second line identifier, the original code file comprises at least one first code line, and the original modification scheme comprises at least one second code line; The adding of the region identifier in the original code file and the original modification scheme to obtain the code file to be processed and the modification scheme to be processed comprises: Add the first line identifier to the first code line to obtain the code file to be processed, and add the second line identifier to the second code line to obtain the modification plan to be processed, wherein the first line identifier is used to identify the position of the first code line, and the second line identifier is used to identify the position of the second code line.

3. The method according to claim 2, wherein the adding the first line identifier to the first code line to obtain the code file to be processed, and the adding the second line identifier to the second code line to obtain the modification solution to be processed comprises: The first line identifier is added before the first code line to obtain the code file to be processed, and the second line identifier is added before the second code line to obtain the modification plan to be processed.

4. The method according to claim 2, before adding the second line identifier to the second code line and obtaining the pending modification solution, further comprises: Filter out, from the second code lines, unmodified code lines that are identical to the first code lines; Replacing the unmodified code line with an unmodified identifier to obtain a replaced second code line, wherein the unmodified identifier refers to a symbol or label used to replace the unmodified code line; The adding the second line identifier to the second code line to obtain the modification solution to be processed includes: The second line identifier is added to the replaced second code line to obtain the modification solution to be processed.

5. The method according to claim 1, wherein generating a target modified file of the original code file according to the region identifier sequence comprises: Acquire identifier mapping information, wherein the identifier mapping information is used to describe the content correspondence between the region identifier and the original code file and the original modification scheme respectively; A target modified file of the original code file is generated according to the region identifier sequence and the identifier mapping information.

6. A method for correcting code errors, comprising: Obtaining an original code file and an error correction solution for the original code file; Adding a region identifier to the original code file and the error correction scheme to obtain a code file to be processed and an error correction scheme to be processed, wherein the region identifier is used to indicate the region position where the code content in the original code file is located and the region position where the modified content in the error correction scheme is located; Inputting the code file to be processed and the error correction solution to be processed into a code processing model to obtain a region identifier sequence; An error correction code file of the original code file is generated according to the region identifier sequence.

7. A code processing model training method, comprising: Acquire a sample set, wherein the sample set includes a sample code file, a sample modification scheme, and a sample identifier sequence, wherein the sample modification scheme is a sample modification scheme of the sample code file, and the sample identifier sequence is a generation target of a code processing model; Adding a sample area identifier to the sample code file and the sample modification scheme to obtain a sample processing file and a sample processing scheme, wherein the sample area identifier is used to indicate the area position where the code content in the sample code file is located and the area position where the modified content in the sample modification scheme is located; Inputting the sample processing file and the sample processing scheme into the code processing model to obtain a predicted identifier sequence; The code processing model is trained according to the predicted identifier sequence and the sample identifier sequence to obtain a trained code processing model, wherein the trained code processing model is used in the processing process of the method as claimed in any one of claims 1 to 6.

8. The method according to claim 7, wherein obtaining the sample set comprises: Obtaining the sample code file and a sample modified file of the sample code file; Generate the sample modification plan according to the sample code file and the sample modification file; The sample identifier sequence is generated according to the sample modification file and the sample region identifier.

9. The method according to claim 8, wherein generating the sample modification scheme according to the sample code file and the sample modification file comprises: The sample code file and the sample modification file are input into a modification solution generation model to obtain the sample modification solution.

10. The method according to claim 8, wherein the sample region identifier comprises a first sample identifier and a second sample identifier, the first sample identifier is used to identify the position of a first sample code line in the sample code file, and the second sample identifier is used to identify the position of a second sample code line in the sample modification scheme; The step of generating the sample identifier sequence according to the sample modification file and the sample region identifier comprises: Filter out a first sample modification line and a second sample modification line from the sample modification file, wherein the first sample modification line is the same as the first sample code line, and the second sample modification line is the same as the second sample code line; The first sample modification line is replaced by the first sample identifier, and the second sample modification line is replaced by the second sample identifier to obtain the sample identifier sequence.

11. An information processing method based on a code processing model, applied to a task platform, comprising: Receiving a model request sent by a terminal device, wherein the model request includes a scenario identifier of a target code scenario, scenario input data of the target code scenario, and at least one of a model specification parameter; Based on the model request, a target code processing model is determined from a plurality of code processing models, wherein the plurality of code processing models are trained using the code processing model training method according to any one of claims 7 to 10.

12. The method according to claim 11, wherein the model request includes a scene identification of a target code scene; The step of determining a target code processing model from a plurality of code processing models based on the model request includes: Based on the scenario identifier of the target code scenario, a target code processing model adapted to the target code scenario is searched from a model library, wherein the model library stores a plurality of code processing models adapted to different code scenarios.

13. The method of claim 11, wherein the model request comprises scenario input data of a target code scenario; The step of determining a target code processing model from a plurality of code processing models based on the model request includes: Determining an initial code processing model adapted to the target code scenario from a plurality of code processing models; Based on the scenario input data of the target code scenario, the initial code processing model is trained to obtain a target code processing model.

14. The method of claim 11, wherein the model request comprises model specification parameters; The step of determining a target code processing model from a plurality of code processing models based on the model request includes: Based on the model specification parameters, a corresponding target code processing model is searched from a model library, wherein the model library stores a plurality of code processing models with different model specification parameters.

15. The method according to any one of claims 11 to 14, after determining the target code processing model from a plurality of code processing models based on the model request, further comprising: The target code processing model is deployed, and based on the target code processing model, a code task processing interface is constructed so that the terminal device schedules the target code processing model to execute the target code task.

16. A task platform, comprising a request interface and a response unit; The request interface is used to receive a model request sent by a terminal device, wherein: The model request includes a scenario identifier of a target code scenario, scenario input data of the target code scenario, and at least one of a model specification parameter; The response unit is used to determine a target code processing model from a plurality of code processing models based on the model request, wherein the plurality of code processing models are trained using the code processing model training method according to any one of claims 7 to 10.

17. The task platform according to claim 16, further comprising a code task processing interface, wherein the code task processing interface is constructed based on the target code processing model; The code task processing interface is used for the terminal device to schedule and execute the target code task.

18. A computing device comprising: Memory and processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions. When the computer program / instructions are executed by the processor, the steps of the method described in any one of claims 1 to 15 are implemented.

19. An electronic device comprising: A memory and a processor, wherein the memory and the processor are connected via a bus; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions. When the computer program / instructions are executed by the processor, the steps of the method described in any one of claims 1 to 15 are implemented.

20. A computer-readable storage medium storing a computer program / instruction, wherein the computer program / instruction, when executed by a processor, implements the steps of the method according to any one of claims 1 to 15.

21. A computer program product, comprising a computer program / instruction, which, when executed by a processor, implements the steps of the method according to any one of claims 1 to 15.

Citation Information

Patent Citations

  • Source code problem repairing method and device based on deep learning model

    CN118363845A

  • Code fixing method and apparatus, and device and medium

    US20240385945A1