Code extraction method, computing device, storage medium and computer program product
By dynamically adjusting the window size and re-acquisitioning the code segment, the problem of incomplete extraction of code segments in the code editor is solved, and the integrity and independent operation ability of the code segment are achieved.
Patent Information
- Application Number
- CN202510229326.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-05-30
AI Technical Summary
The code segment extracted in the code editor may be incomplete and lack the necessary context information, resulting in the inability to run independently.
By dynamically adjusting the window size, gradually optimize the obtained code segment. After each adjustment of the window, re-get the code segment in the window and analyze it as a new code segment to be detected until the termination condition is met.
Ensure that the final obtained object code segment is complete and can be run and understood independently.
Smart Images

Figure CN120066571A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present application relate to the field of artificial intelligence technology, and in particular, to a code extraction method, a computing device, a computer-readable storage medium, and a computer program product. Background Art
[0002] In the process of software development, code extraction is a crucial technology. Especially in large code libraries, developers often need to quickly understand, reuse, or modify specific parts of the code.
[0003] In related technologies, the selection function of the code editor is usually used to implement code extraction based on the text range selected by the user. The user can directly use the mouse or keyboard to select a piece of code in the code editor, and then the code editor can extract the code segment selected by the user.
[0004] The inventor found that during the implementation of the concept of the present application, the code segment extracted by the user in the code editor may be incomplete, such as lacking necessary context information, etc., resulting in the inability of the extracted code segment to run independently. Summary of the Invention
[0005] Embodiments of the present application provide a code extraction method, a computing device, a computer-readable storage medium, and a computer program product.
[0006] In a first aspect, an embodiment of the present application provides a code extraction method, including:
[0007] In response to a code selection operation triggered in a code editor, obtain a first window formed by expanding a predetermined number of lines centered on the selected position, and obtain a first code segment within the first window;
[0008] Use the first code segment as a code segment to be detected, and input the code segment to be detected into a large model to use the large model to detect whether the integrity of the code segment to be detected meets a preset condition. If the integrity of the code segment to be detected meets the preset condition, determine the code segment to be detected as a target code segment. If the code segment to be detected does not meet the preset condition, generate window adjustment prompt information;
[0009] Adjust the first window to a second window according to the window adjustment prompt information;
[0010] Obtain a second code segment within the second window;
[0011] Use the second code segment as a code segment to be detected, and return to the step of inputting the code segment to be detected into the large model to continue execution.
[0012] Second aspect, an embodiment of the present application provides a code extraction device, including:
[0013] A first code segment acquisition module, configured to, in response to a code selection operation triggered in a code editor, acquire a first window formed by expanding a predetermined number of lines centered on the selected position, and acquire a first code segment within the first window;
[0014] A code segment input module, configured to use the first code segment as a code segment to be detected, and input the code segment to be detected into a large model, so as to use the large model to detect whether the integrity of the code segment to be detected meets a preset condition. If the integrity of the code segment to be detected meets the preset condition, determine the code segment to be detected as a target code segment. If the code segment to be detected does not meet the preset condition, generate window adjustment prompt information;
[0015] A window adjustment module, configured to adjust the first window to a second window according to the window adjustment prompt information;
[0016] A second code segment acquisition module, configured to acquire a second code segment within the second window;
[0017] An iteration module, configured to use the second code segment as a code segment to be detected, and return to the step of inputting the code segment to be detected into the large model and continue to execute.
[0018] Third aspect, an embodiment of the present application provides a computing device, including a processing component and a storage component;
[0019] The storage component stores one or more computer instructions; the one or more computer instructions are used to be called and executed by the processing component to implement the code extraction method provided by the embodiment of the present application.
[0020] Fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processing component, the code extraction method provided by the embodiment of the present application is implemented.
[0021] Fifth aspect, an embodiment of the present application provides a computer program product, including a computer program / instructions. When the computer program / instructions are executed by a processing component, the code extraction method provided by the embodiment of the present application is implemented.
[0022] Embodiments of the present application can gradually optimize the extracted code segment by dynamically adjusting the window size. After each window adjustment, the code segment within the window can be re-acquired and used as a new code segment to be detected for analysis. This process can be iterative until the newly acquired code segment to be detected meets the termination condition, so that the finally extracted target code segment has integrity.
[0023] These aspects or other aspects of the present application will be more clearly understood in the following description of the embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0025] Figure 1 A flowchart showing a code extraction method provided by an embodiment of the present application;
[0026] Figure 2 A block diagram showing a code extraction device provided by an embodiment of the present application;
[0027] Figure 3 A block diagram showing a computing device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0028] In order to enable those skilled in the art to better understand the solution of the present application, the following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the drawings in the embodiments of the present application.
[0029] In some processes described in the specification and claims of the present application and the above drawings, there are multiple operations that appear in a specific order. However, it should be clearly understood that these operations may not be executed in the order in which they appear herein or may be executed in parallel. The operation numbers such as 101, 102, etc. are only used to distinguish different operations, and the numbers themselves do not represent any execution order. In addition, these processes may include more or fewer operations, and these operations may be executed in sequence or in parallel. It should be noted that the descriptions such as "first" and "second" in this article are used to distinguish different messages, devices, modules, etc., and do not represent a sequence, nor do they limit that "first" and "second" are of different types.
[0030] In the process of software development, code extraction is a crucial technology. Especially in large code libraries, developers often need to quickly understand, reuse, or modify specific parts of the code.
[0031] In the related art, usually, the selection function of the code editor is used to implement code extraction based on the text range selected by the user. The user can directly use the mouse or keyboard in the code editor to select a section of code, and then the code editor can extract the selected code segment.
[0032] In the process of implementing the concept of this application, the inventor found that the code segments extracted by users in a code editor may be incomplete. For example, they may lack necessary context information, etc., resulting in the inability of the extracted code segments to run independently.
[0033] To solve the technical problems existing in the related art, the embodiments of this application provide a code extraction method. The method can gradually optimize the extracted code segments by dynamically adjusting the window size. After each adjustment of the window, the code segments within the window can be re-obtained and used as new code segments to be detected for analysis. This process can be iterative until the newly obtained code segments to be detected meet the termination conditions, so that the finally extracted target code segments are complete.
[0034] Next, the technical solutions in the embodiments of this application will be described clearly and completely with reference to the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of this application.
[0035] Figure 1 The flowchart of a code extraction method provided by the embodiments of this application is shown. This code extraction method can be executed by a plug-in module. In the embodiments of this application, the plug-in module can be directly integrated into the code editor as a part of its built-in functions. However, it is not limited to this. The plug-in module can also be designed as an independent application or service and communicate with the code editor through a standardized interface. In actual applications, the plug-in module can select its implementation method according to actual needs.
[0036] As Figure 1 shown, the method can specifically include the following steps:
[0037] 101: In response to a code selection operation triggered in the code editor, obtain a first window formed by expanding a predetermined number of lines centered on the selected position, and obtain a first code segment within the first window.
[0038] Among them, the plug-in module can listen to user operations in the code editor in real time, such as cursor movement, code selection, etc. When the user triggers a code selection operation, for example, by dragging the mouse to select a section of code or by using a shortcut key to select a certain line of code, the plug-in module can capture the event and record the selected position. Among them, the selected position can include a single cursor position or a selected area.
[0039] After obtaining the selected position of the user, the plugin module can use the selected position as the center and expand a predetermined number of lines upstream and downstream to generate a first window. Among them, the predetermined number of lines to be expanded can be configured according to actual needs, and this predetermined number of lines can be implemented as a fixed value or dynamically adjusted according to code complexity. For example, if the user selects the tenth line of code and sets the number of expanded lines to 5 lines, the range of the first window is from the 5th to the 15th line.
[0040] After determining the window range of the first window, the plugin module can extract the code content within the first window from the code file as the first code segment.
[0041] The first code segment included in the first window may, for example, lack necessary context information (such as variable definitions, function declarations, etc.) and is incomplete, resulting in the code segment being unable to run or be understood independently. To ensure the integrity of the finally extracted code segment, the embodiments of this application introduce a large model to detect the integrity of the first code segment extracted from the first window.
[0042] 102: Use the first code segment as the code segment to be detected, and input the code segment to be detected into the large model to use the large model to detect whether the integrity of the code segment to be detected meets the preset conditions. If the integrity of the code segment to be detected meets the preset conditions, determine the code segment to be detected as the target code segment. If the code segment to be detected does not meet the preset conditions, generate window adjustment prompt information.
[0043] Among them, the large model can be a small pre-trained language model that has been fine-tuned. Through fine-tuning, the large model can significantly improve its code semantic understanding ability, enhance its context association ability, and make it better adapt to specific programming languages. The fine-tuned model can more deeply understand the syntax structure of programming languages, variable naming rules, function call relationships, etc., so as to more accurately parse the semantics of the code. At the same time, the fine-tuned model can better identify the logical relationships between code segments, provide support for context extraction, and enable it to more accurately grasp the overall logic when processing complex code. In addition, fine-tuning can also enable the model to adapt to the characteristics of specific programming languages such as Swift and Objective-C, thereby improving the accuracy of code analysis and ensuring that the model performs well in different programming language environments.
[0044] After obtaining the first code segment, the plugin module can input the first code segment into the large model, and the large model can detect the integrity of the first code segment according to the preset conditions. Among them, the preset conditions can be one or a set of criteria for evaluating the integrity of code segments.
[0045] If the large model detects that the code to be detected meets the preset conditions, it can be used as the target code segment, and the target code segment can be the finally extracted code segment. If the large model detects that the code segment to be detected does not meet the preset conditions, it can generate window adjustment prompt information to guide subsequent window adjustment.
[0046] 103: Adjust the first window to the second window according to the window adjustment prompt information.
[0047] 104: Obtain the second code segment within the second window.
[0048] According to the window adjustment prompt information, the plug-in module can dynamically adjust the window size to adjust the first window to the second window to re-extract the second code segment within the second window. Compared with the first code segment, the second code segment may contain more context information or fix certain problems.
[0049] 105: Use the second code segment as the code segment to be detected and return to continue executing the step of inputting the code segment to be detected into the large model.
[0050] After obtaining the second code segment, the second code segment can be re-input into the large model for integrity detection, and the large model can re-determine whether the updated code segment to be detected meets the preset conditions. If the second code segment still does not meet the preset conditions, the plug-in module can generate new window adjustment prompt information according to the detection results. The window adjustment prompt information can be used to guide the next window adjustment strategy.
[0051] According to the new prompt information, the plug-in module can adjust the window size again, re-extract the code content (such as the third code segment), and input it into the large model as the new code segment to be detected for analysis.
[0052] In the embodiments of the present application, the code segment obtained by extraction can be gradually optimized by dynamically adjusting the window size. After each window adjustment, the code segment within the window can be re-obtained and used as the new code segment to be detected for analysis. This process can be iterative until the newly obtained code segment to be detected meets the termination conditions, so that the finally extracted target code segment has integrity.
[0053] In some embodiments, the preset conditions include multiple sub-conditions, and each sub-condition has a corresponding window adjustment strategy.
[0054] In some embodiments, adjusting the first window to the second window according to the window adjustment prompt information can be specifically implemented as:
[0055] Determine the target sub-condition that the code segment to be detected indicated by the window prompt information does not meet;
[0056] Adjust the first window to the second window by using a target window adjustment strategy corresponding to a target sub - condition.
[0057] As described above, the large - model can determine whether the code segment to be detected meets the integrity requirements through preset conditions. Among them, the preset conditions can include multiple sub - conditions, and each sub - condition corresponds to a specific window adjustment strategy. In this way, after receiving the output of the large - model, the plug - in module can take targeted adjustment measures according to the specific problems existing in the code segment to be detected, so as to more efficiently optimize the code segment extraction process.
[0058] When it is necessary to adjust the first window to the second window according to the window adjustment prompt information, the plug - in module can first analyze the window adjustment prompt information to clarify the specific target sub - condition that the code segment to be detected does not meet. For example, the prompt information may indicate problems such as undefined variables, low semantic scores, or truncated logical blocks in the code segment. Each problem corresponds to a specific target sub - condition, and each target sub - condition is associated with a specific window adjustment strategy. Next, the plug - in module can use the target window adjustment strategy corresponding to the target sub - condition to adjust the first window to generate the second window.
[0059] The design of multiple sub - conditions of the preset conditions and the corresponding window adjustment strategies can provide clear guidance for the dynamic window adjustment of the plug - in module. By identifying the target sub - conditions and matching the corresponding adjustment strategies, the plug - in module realizes precise and efficient context window optimization, thus improving the accuracy and rationality of code excerpt.
[0060] In a possible implementation manner of the present application, the multiple sub - conditions of the preset conditions can include semantic score threshold conditions, dependency - missing conditions, structure - break conditions, etc. Among them, each sub - condition can be configured with a corresponding window adjustment strategy. For example, if the window adjustment prompt information indicates that there are undefined variables in the code segment, the plug - in module can adopt a "dependency repair" window adjustment strategy; if the window adjustment prompt information shows a low semantic score, the plug - in module can adopt a "semantic completion" window adjustment strategy; if the window adjustment prompt information indicates that the logical block is truncated, the plug - in module can adopt a "structure alignment" window adjustment strategy. In this way, the plug - in module can adopt the most appropriate adjustment method for different problems, gradually optimize the window range, and finally extract a complete and non - redundant code segment.
[0061] The judgment method of each sub - condition and the corresponding window adjustment strategy will be described in detail below.
[0062] In some embodiments, using the large - model to detect whether the integrity of the code segment to be detected meets the preset conditions can be specifically implemented as:
[0063] In the case where an undefined variable is detected in the code segment to be detected, it is determined that the code segment to be detected does not meet the dependency integrity sub-condition.
[0064] During the code analysis process, a complete code snippet needs to meet the dependency integrity condition. Dependency integrity can mean that all elements such as variables, functions, or classes referenced in the code snippet have clear definitions or declarations. If a variable is used in the code snippet but its definition is not included within the current code scope, then this variable is considered an "undefined variable". This situation can cause the code snippet to be unable to run or be understood independently, thus breaking the logical structure and semantic integrity of the code.
[0065] Specifically, when the large model performs semantic analysis on the code, it will check whether all references in the code segment to be detected have corresponding definitions. For example, if the variable x appears in the code snippet but the plugin module cannot find the definition location of x within the current code scope, then it will be determined that there is a dependency missing problem in this code snippet. At this time, the large model can classify this problem as "the dependency integrity sub-condition is not met".
[0066] The dependency integrity sub-condition is part of the preset conditions and is used to evaluate whether the code snippet meets the extraction requirements. Once an undefined variable is detected, the large model can determine that the code segment to be detected does not meet the dependency integrity sub-condition and generate corresponding window adjustment prompt information to guide the subsequent window adjustment strategy.
[0067] In some embodiments, using the target window adjustment strategy corresponding to the target sub-condition, adjusting the first window to the second window can be specifically implemented as:
[0068] Trace back to the variable definition location in the upstream code;
[0069] Merge the variable definition code segment at the variable definition location into the first window to obtain the second window.
[0070] In the case where an undefined variable is detected in the code segment to be detected, the plugin module can trace back to the variable definition location in the upstream code. After finding the variable definition location, the plugin module can merge the variable definition code at this location into the first window to generate the second window.
[0071] The second window not only contains the code content in the first window but also adds the code segment at the variable definition location. Thus, the second code segment included in the second window can meet the dependency integrity sub-condition.
[0072] In some embodiments, using the large model to detect whether the integrity of the code segment to be detected meets the preset conditions can be specifically implemented as:
[0073] When it is detected that the semantic score of the code segment to be detected is lower than the preset semantic threshold, it is determined that the code segment to be detected does not meet the semantic integrity sub-condition.
[0074] In an embodiment of the present application, the preset condition can also be a semantic integrity sub-condition. The core of the semantic integrity sub-condition lies in evaluating whether the semantic score of the code segment to be detected reaches a certain standard (such as a preset semantic threshold). If the semantic score is lower than the preset semantic threshold, it is determined that the code segment does not meet the semantic integrity sub-condition.
[0075] Among them, the semantic score can be a quantitative index used to measure the semantic integrity of the code segment to be detected, which can reflect the logical coherence and syntactic correctness of the code segment.
[0076] After inputting the code segment to be detected into the large model, the large model can calculate the semantic score according to its understanding ability of the code. For example, assuming that the preset semantic threshold is 0.95 and the semantic score of the code segment to be detected is 0.85, the large model will determine that the code segment to be detected does not meet the semantic integrity sub-condition and generate corresponding window adjustment prompt information. The window adjustment prompt information can be used to guide subsequent window adjustment strategies. For example, in the case of a low semantic score, the window adjustment strategy can include expanding the window to add more context content, thereby improving the semantic score.
[0077] In some embodiments, using the window adjustment strategy corresponding to the type of prompt information, adjusting the first window to the second window can be specifically implemented as:
[0078] Extend the window boundaries of the first window by a predetermined number of code lines to the upstream code and / or downstream code to generate the second window.
[0079] Among them, the window can be extended to the upstream code, downstream code, or in both directions simultaneously. Among them, the upstream code can refer to the code content before the cursor position, and the downstream code refers to the code content after the cursor position.
[0080] The plug-in module can extend the window boundary according to a predetermined number of code lines (such as ±5 lines or ±10 lines). For example, assuming that the range of the first window is from line 10 to line 20, the plug-in module can extend the window to from line 5 to line 25 to form the second window. This extension method ensures that the second code segment included in the second window can cover more context content, thereby improving the semantic score and making the second code segment meet the semantic integrity sub-condition.
[0081] Among them, the number of extended lines can be dynamically adjusted according to the actual situation. For example, for simple code segments, the number of extended lines can be reduced; while for complex code segments, the number of extended lines can be increased.
[0082] In some embodiments, using a large model to detect whether the integrity of the code segment to be detected meets the preset conditions can be specifically implemented as follows:
[0083] When it is detected that the starting position or the ending position of the code segment to be detected truncates a code logic block, it is determined that the code segment to be detected does not meet the structural integrity sub-condition.
[0084] Among them, a code logic block can refer to a code range with clear logical meaning, such as a complete function body, a loop body, a conditional branch (such as an if statement), or a class definition. A logic block is usually delimited by specific syntax structures, such as curly braces {}, indentation levels, or other language-specific markers.
[0085] If the starting or ending position of a code snippet truncates a certain logic block, it will cause the code snippet to be unable to run or be understood independently. For example, a function call may depend on certain variables or operations within the function body. If the function body is truncated, then the extracted code snippet will lose its meaning.
[0086] A large model can perform semantic analysis on the code snippet to identify the scope of the logic block. For example, the model can parse the Abstract Syntax Tree (AST) of the code and locate the starting and ending positions of each logic block. By comparing the starting and ending positions of the code snippet with the boundaries of the logic block, it is determined whether there is a truncation problem.
[0087] If the starting position of the code segment to be detected is in the middle of a certain logic block (i.e., it does not start from the starting position of the logic block), or the ending position is in the middle of the logic block (i.e., it does not reach the ending position of the logic block), then it is determined that the code segment to be detected truncates the logic block. For example:
[0088] Suppose the range of a certain function body is from line 5 to line 15, and the range of the code segment to be detected is from line 8 to line 12, then it can be considered that the code segment to be detected truncates the function body. Similarly, if the range of a certain loop body is from line 20 to line 30, and the range of the code segment to be detected is from line 25 to line 35, then it can also be considered that the code segment to be detected truncates the loop body.
[0089] In some embodiments, using a window adjustment strategy corresponding to the type of prompt information, adjusting the first window to the second window can be specifically implemented as follows:
[0090] Align the window boundary of the first window with the logic block boundary of the code logic block to adjust the first window to the second window.
[0091] The plug-in module can parse the abstract syntax tree of the code file through a static analysis tool (such as SwiftSyntax) to locate the start and end positions of each logical block. Then, the plug-in module can adjust the range of the first window according to the boundaries of the logical blocks. For example: if the start position of the first window is in the middle of a function body, the window start position is adjusted to the start position of the function body; if the end position of the first window is in the middle of a loop body, the window end position is adjusted to the end position of the loop body.
[0092] Thus, the second code segment included in the second window generated after adjustment can meet the structural integrity sub-condition.
[0093] In some embodiments, using a large model to detect whether the integrity of the code segment to be detected meets a preset condition includes:
[0094] In the case where the score of the code segment to be detected is higher than the preset semantic threshold and the number of lines of the code segment to be detected is higher than the preset number of code lines, it is determined that the code segment to be detected does not meet the conciseness sub-condition;
[0095] Using a target window adjustment strategy corresponding to the target sub-condition to adjust the first window to the second window includes:
[0096] Shrink the first window to obtain the second window.
[0097] Among them, the conciseness sub-condition can be used to evaluate whether the code segment to be detected is both semantically complete and not overly verbose. If the semantic score of the code segment is high (indicating good semantic integrity), but the number of lines is too large (indicating that it may contain redundant content), the large model can determine that the code segment to be detected does not meet the conciseness sub-condition.
[0098] The large model can preset a preset semantic threshold (such as 0.95) and a preset number of code lines (such as 20 lines). These two parameters are used to comprehensively judge whether the code segment meets the conciseness sub-condition: if the semantic score of the code segment to be detected is higher than the preset semantic threshold, it indicates good semantic integrity; if the number of lines of the code segment to be detected is higher than the preset number of code lines, it indicates that it may contain redundant content.
[0099] By shrinking the window, redundant content in the code segment to be detected can be removed while retaining necessary context information. The plug-in module can gradually narrow the window range according to the semantic score gradient direction and code structure characteristics of the code segment to be detected. For example: if the start part of the code segment contains content irrelevant to the current logic, the window can be shrunk upstream; if the end part of the code segment contains redundant content, the window can be shrunk downstream.
[0100] After reducing the first window, the second window obtained has less redundant content compared to the first window, but still maintains the semantic integrity of the code snippet. For example, assume the range of the first window is from line 1 to line 30, and the content from line 21 to line 30 is irrelevant to the current logic. Then, the window range can be reduced to from line 1 to line 20 to form the second window.
[0101] In some embodiments, when it is determined that the code segment to be detected does not meet multiple target sub-conditions according to the window adjustment prompt information, using the target window adjustment strategy corresponding to the target sub-condition, the adjustment of the first window to the second window can be specifically implemented as:
[0102] Determine the adjustment priorities of multiple target adjustment strategies respectively;
[0103] Execute multiple target adjustment strategies in sequence according to the adjustment priorities.
[0104] In an actual application scenario, there may be multiple problems in the code segment to be detected. For example: there is an undefined variable x, indicating a missing dependency; the semantic score is 0.85, lower than the preset threshold of 0.95, indicating insufficient semantic integrity; the window boundary truncates a function body, indicating insufficient structural integrity. These problems correspond to different target sub-conditions respectively, and different window adjustment strategies need to be adopted to solve them.
[0105] When multiple target sub-conditions exist simultaneously, directly executing the adjustment strategies randomly or disorderly may lead to low efficiency or conflicts. Therefore, it is necessary to assign a priority to each window adjustment strategy to ensure that the most critical problems are solved first.
[0106] For example, dependency repair can be set as the highest priority, semantic completion as the second highest priority, and structural alignment as the lowest priority.
[0107] The plugin module can execute multiple target adjustment strategies in sequence according to the priority order to gradually optimize the window range. For example: First, for the problem of missing dependencies, extend the window upstream to merge the definition location of the undefined variable x into the first window; Second, for the problem of low semantic score, extend the window to add more context; Finally, for the problem of logical block truncation, adjust the window boundary to align with the start or end position of the logical block.
[0108] After each adjustment, the plugin module can re-extract the new code segment and input it into the large model for analysis again. If there are still problems unsolved, continue to execute the remaining adjustment strategies until all problems are solved or the termination condition is reached.
[0109] During code analysis and processing, it is usually necessary to extract the context related to the target code segment from the entire code file. In some embodiments, the method further includes:
[0110] Input the abstract syntax tree of the code file and the target code segment into a large model to use the large model to output code nodes associated with the target code segment based on the abstract syntax tree;
[0111] Obtain the associated code segment corresponding to the code node from the code editor;
[0112] Output the associated code segment.
[0113] Among them, the abstract syntax tree is a tree-shaped representation of the code structure, which can clearly show elements such as variables, functions, and classes in the code and their mutual relationships. For example, a function call may depend on a variable definition or another function declaration, and these dependencies can all be parsed through the AST.
[0114] After inputting the target code segment and the abstract syntax tree into the large model, the large model can utilize its semantic understanding ability to analyze the relationship between the target code segment and other code nodes based on the AST. For example, the large model can identify variables referenced, functions called, or classes depended on in the target code segment, thereby identifying the code nodes associated with the target code segment and the location information of these code nodes.
[0115] Thus, the plugin module can obtain the associated code segment corresponding to the code node from the code editor according to the location information.
[0116] During the code extraction process, after the target code segment is extracted, it needs to be further verified to ensure its syntactic correctness, logical integrity, and semantic consistency. This verification process can be completed by using a large model. Combining with preset verification rules, a comprehensive quality assessment is performed on the target code segment, and the verification result is output.
[0117] In some embodiments, the method may further include:
[0118] Input the target code segment into the large model to use the large model to verify the target code segment according to the preset verification rules and output the verification result.
[0119] The large model can analyze the target code segment according to the preset verification rules. The preset verification rules may include the following aspects, for example:
[0120] Syntactic correctness: Verify whether the target code segment conforms to the syntax specifications of the programming language (such as Swift or Objective-C). For example, check whether the parentheses match and whether the variable declarations are legal.
[0121] Logical integrity: Check whether the target code segment contains complete statements, expressions, or logical blocks. For example, whether a function call includes all necessary arguments.
[0122] Semantic consistency: Ensure that the target code segment is semantically consistent with other parts of the file where it is located. For example, whether the use of variables conforms to their defined types and scopes.
[0123] Correspondingly, the verification results can include syntax correctness scores, logical integrity scores, and semantic consistency scores.
[0124] After outputting the verification results, developers can quickly locate and fix code problems with the help of the verification results, and use them as a basis for scoring to quantify the quality of the target code segment.
[0125] Figure 2 The block diagram of a code extraction device provided in an embodiment of the present application is shown, as Figure 2 shown, the device may specifically include:
[0126] The first code segment acquisition module 201 is configured to, in response to a code selection operation triggered in a code editor, acquire a first window formed by expanding a predetermined number of lines centered on the selected position, and acquire a first code segment within the first window;
[0127] The code segment input module 202 is configured to use the first code segment as the code segment to be detected, and input the code segment to be detected into the large model to detect whether the integrity of the code segment to be detected meets a preset condition. If the integrity of the code segment to be detected meets the preset condition, the code segment to be detected is determined as the target code segment. If the code segment to be detected does not meet the preset condition, a window adjustment prompt message is generated;
[0128] The window adjustment module 203 is configured to adjust the first window to a second window according to the window adjustment prompt message;
[0129] The second code segment acquisition module 204 is configured to acquire a second code segment within the second window;
[0130] The iteration module 205 is configured to use the second code segment as the code segment to be detected, and return to the step of inputting the code segment to be detected into the large model to continue execution.
[0131] In some embodiments, the preset condition includes multiple sub-conditions, and each sub-condition has a corresponding window adjustment strategy.
[0132] In some embodiments, the window adjustment module 203 is specifically configured to:
[0133] Determine the target sub-condition that the code segment to be detected indicated by the window prompt message does not meet;
[0134] Adjust the first window to the second window by using a target window adjustment strategy corresponding to a target sub - condition.
[0135] In some embodiments, the code segment input module 202 is specifically configured to:
[0136] When it is detected that there is an undefined variable in the code segment to be detected, determine that the code segment to be detected does not meet the dependency - complete sub - condition;
[0137] In some embodiments, the window adjustment module 203 is specifically configured to:
[0138] Trace back to the variable definition position in the upstream code;
[0139] Merge the variable definition code segment at the variable definition position into the first window to obtain the second window.
[0140] In some embodiments, the code segment input module 202 is specifically configured to:
[0141] When it is detected that the semantic score of the code segment to be detected is lower than a preset semantic threshold, determine that the code segment to be detected does not meet the semantic - complete sub - condition;
[0142] In some embodiments, the window adjustment module 203 is specifically configured to:
[0143] Expand the window boundary of the first window by a predetermined number of code lines in the upstream code and / or downstream code to generate the second window.
[0144] In some embodiments, the code segment input module 202 is specifically configured to:
[0145] When it is detected that the start position or the end position of the code segment to be detected truncates a code logic block, determine that the code segment to be detected does not meet the structure - complete sub - condition;
[0146] In some embodiments, the window adjustment module 203 is specifically configured to:
[0147] Align the window boundary of the first window with the logic block boundary of the code logic block to adjust the first window to the second window.
[0148] In some embodiments, when it is determined that the code segment to be detected does not meet multiple target sub - conditions according to the window adjustment prompt information, the window adjustment module 203 is specifically configured to:
[0149] Respectively determine the adjustment priorities of multiple target adjustment strategies;
[0150] Execute multiple target adjustment strategies in sequence according to the adjustment priorities.
[0151] In some embodiments, the device may further include:
[0152] A syntax tree input module for inputting the abstract syntax tree of a code file and a target code segment into a large model, so as to utilize the large model to output code nodes associated with the target code segment based on the abstract syntax tree;
[0153] An associated code segment determination module for obtaining the associated code segment corresponding to the code node from a code editor;
[0154] An associated code segment output module for outputting the associated code segment.
[0155] In some embodiments, the apparatus may further include:
[0156] A verification module for inputting the target code segment into the large model, so as to utilize the large model to verify the target code segment according to a preset verification rule and output a verification result.
[0157] Figure 2 The described code extraction apparatus may execute Figure 1 The code extraction method described in the illustrated embodiment, and its implementation principle and technical effects will not be elaborated. For the code extraction apparatus in the above embodiments, the specific manners in which each module and unit perform operations have been described in detail in the embodiments related to the method, and will not be elaborated here.
[0158] An embodiment of the present application further provides a computing device, as Figure 3 shown, the device may include a storage component and a processing component;
[0159] The storage component stores one or more computer instructions, wherein the one or more computer instructions are called and executed by the processing component to implement the code extraction method provided by the embodiment of the present application.
[0160] Of course, the computing device may necessarily further include other components, such as an input / output interface, a display component, a communication component, etc.
[0161] The input / output interface provides an interface between the processing component and a peripheral interface module, and the above peripheral interface module may be an output device, an input device, etc. The communication component is configured to facilitate communication between the computing device and other devices in a wired or wireless manner, etc.
[0162] Among them, the processing component may include one or more processors to execute computer instructions to complete all or part of the steps in the above-mentioned method. Of course, the processing component may also be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors or other electronic components for executing the above-mentioned method.
[0163] The storage component is configured to store various types of data to support the operation of the terminal. The storage component can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.
[0164] The display component can be an electroluminescent (EL) element, a liquid crystal display or a microdisplay with a similar structure, or a retina direct display or a similar laser scanning display.
[0165] The embodiment of the present application also provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a computer, it can implement the Figure 1 code extraction method of the above-mentioned embodiment. The computer-readable medium can be included in the electronic device described in the above-mentioned embodiment; or it can exist alone without being assembled into the electronic device.
[0166] The embodiment of the present application also provides a computer program product, which includes a computer program carried on a computer-readable storage medium, and when the computer program is executed by a computer, it can implement the code extraction method as described in the above-mentioned Figure 1 embodiment. In such an embodiment, the computer program can be downloaded and installed from the network and / or installed from a removable medium. When the computer program is executed by the processor, it executes various functions defined in the system of the present application.
[0167] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the above-described system, device and unit can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0168] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative labor.
[0169] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0170] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application and are not intended to limit them. Although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features. These modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A code extraction method, characterized in that: include: In response to a code selection operation triggered in the code editor, a first window formed by extending a predetermined number of lines with the selected position as the center is obtained, and a first code segment in the first window is obtained; The first code segment is used as a code segment to be detected, and the code segment to be detected is input into a large model, so as to use the large model to detect whether the integrity of the code segment to be detected meets a preset condition; if the integrity of the code segment to be detected meets the preset condition, the code segment to be detected is determined as a target code segment; if the code segment to be detected does not meet the preset condition, a window adjustment prompt message is generated; According to the window adjustment prompt information, adjusting the first window to a second window; Obtaining a second code segment in the second window; The second code segment is used as the code segment to be detected, and the process returns to the step of inputting the code segment to be detected into the large model for further execution.
2. The method according to claim 1, characterized in that: The preset condition includes a plurality of sub-conditions, each of which has a corresponding window adjustment strategy; The adjusting the first window to the second window according to the window adjustment prompt information comprises: Determine a target sub-condition that the code segment to be detected does not satisfy, as indicated by the window prompt information; The first window is adjusted to a second window by using a target window adjustment strategy corresponding to the target sub-condition.
3. The method according to claim 2, characterized in that The step of using the large model to detect whether the integrity of the code segment to be detected meets a preset condition includes: In the case where it is detected that there is an undefined variable in the code segment to be detected, determining that the code segment to be detected does not satisfy the dependency complete sub-condition; The adjusting the first window to the second window by using the target window adjustment strategy corresponding to the target sub-condition includes: Trace the upstream code back to the variable definition location; The variable definition code segment at the variable definition position is merged into the first window to obtain the second window.
4. The method according to claim 2, characterized in that: The step of using the large model to detect whether the integrity of the code segment to be detected meets a preset condition includes: In the case where it is detected that the semantic score of the code segment to be detected is lower than a preset semantic threshold, determining that the code segment to be detected does not meet the semantic integrity sub-condition; The adjusting the first window to the second window by using the window adjustment strategy corresponding to the prompt information type includes: The window boundary of the first window is extended toward the upstream code and / or the downstream code by a predetermined code line to generate the second window.
5. The method according to claim 2, characterized in that: The step of using the large model to detect whether the integrity of the code segment to be detected meets a preset condition includes: In the case where it is detected that the start position or the end position of the code segment to be detected truncates the code logic block, determining that the code segment to be detected does not meet the structural integrity sub-condition; The adjusting the first window to the second window by using the window adjustment strategy corresponding to the prompt information type includes: The window boundary of the first window is aligned with the logic block boundary of the code logic block to adjust the first window to a second window.
6. The method according to claim 2, characterized in that When it is determined according to the window adjustment prompt information that the code segment to be detected does not satisfy multiple target sub-conditions, adjusting the first window to the second window by using a target window adjustment strategy corresponding to the target sub-conditions includes: Determine the adjustment priorities of multiple target adjustment strategies respectively; The multiple target adjustment strategies are executed in sequence according to the adjustment priorities.
7. The method according to claim 1, characterized in that The method further comprises: Inputting the abstract syntax tree of the code file and the target code segment into the large model, so as to output a code node associated with the target code segment based on the abstract syntax tree by using the large model; Obtaining an associated code segment corresponding to the code node from the code editor; The associated code segment is output.
8. The method according to claim 1, characterized in that The method further comprises: The target code segment is input into the large model, so as to use the large model to verify the target code segment according to preset verification rules and output the verification result.
9. A computing device, characterized in that including a processing component and a storage component; The storage component stores one or more computer instructions; the one or more computer instructions are used to be called and executed by the processing component to implement the code extraction method as described in any one of claims 1 to 8.
10. A computer-readable storage medium, characterized in that: A computer program is stored thereon, and when the computer program is executed by the processing component, the code extraction method according to any one of claims 1 to 8 is implemented.
11. A computer program product, characterized in that The method comprises a computer program / instruction, wherein when the computer program / instruction is executed by a processing component, the code extraction method according to any one of claims 1 to 8 is implemented.