Method and device for obtaining code detection result, equipment, storage medium and product
By introducing advanced language models and prediction models into code detection, analyzing code differences and generating relevant rule objects, the problem of limited accuracy of rule matching in existing code detection methods is solved, and a higher code detection recall and accuracy is achieved.
Patent Information
- Application Number
- CN202510252119.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-04
- Publication Date
- 2025-06-03
AI Technical Summary
In existing code detection methods, non-custom rules have strong dependence on model capabilities, limited accuracy of rule matching, lack context fusion and poor interpretability, resulting in inaccurate code detection.
By introducing advanced language models, the object code with different differences in the program code to be detected is analyzed, the rule objects of the detection code are filtered out using the prediction model, and prompt words are generated based on the rule objects and the object code, and the object code detection results are generated in combination with the language model.
Ensure the comprehensiveness of code detection, improve the recall and accuracy of code detection, and reduce dependence on custom rules.
Smart Images

Figure CN120086111A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technology, and particularly to a method, apparatus, device, storage medium and product for obtaining code detection results. Background Art
[0002] Code detection is a crucial link in the software development process. Its core objective is to discover potential problems, ensure code compliance, improve software quality and promote team collaboration by analyzing code changes.
[0003] With the development of technology, the currently commonly used code detection method is code detection based on LLM (Large Language Models). It can automatically generate code detection opinions and provide specific optimization suggestions according to project rules and context information. Among them, the rules in the code detection process are the key information for detecting code, which can be divided into two categories: non-custom rules (these rules are pre-known and can be collected in advance and used for training models. For example, general coding specifications, common security vulnerability detection rules, industry standards, etc.) and custom rules (these rules are customized for specific projects, teams or scenarios).
[0004] However, non-custom rules are highly dependent on the model's capabilities, have limited accuracy in rule matching, lack context integration and poor interpretability; custom rules, since they do not require pre-model training, can combine context information in specific modules, functions or business logics to improve the accuracy of rule matching. However, the current custom rules rely on traditional recall methods based on fixed rule sets and simple keyword matching, resulting in many important custom rules being ignored during the code detection process, affecting code quality and project consistency, and making code detection inaccurate. Summary of the Invention
[0005] In view of this, the present disclosure provides a method, apparatus, device, storage medium and product for obtaining code detection results to solve the problem of inaccurate code detection.
[0006] In a first aspect, the present disclosure provides a method for obtaining code detection results, the method comprising:
[0007] Obtain the program code to be detected;
[0008] Obtain the target code with differences in the program code;
[0009] Process the target code based on the target model to obtain the target code detection result of the program code. The target model is a model obtained by optimizing the initial model. The initial model includes a language model and a prediction model. The prediction model is used to obtain the rule object of the detection code based on the target code, and the language model is used to obtain the program code detection result based on the prompt word. The prompt word is generated based on the rule object and the target code.
[0010] In the embodiments of the present disclosure, by introducing an advanced language model, analyze the target code with differences in the program code to be detected, and the prediction model screens out the rule object of the detection code. Then, based on the prompt word generated by the rule object and the target code and combined with the language model, generate the target code detection result. Since the rule object obtained by the prediction model is most relevant to the target code, the comprehensiveness of code detection is ensured, and the recall rate and accuracy of code detection are improved.
[0011] In a second aspect, the present disclosure provides an apparatus for obtaining a code detection result. The apparatus includes:
[0012] A first acquisition module, configured to acquire the program code to be detected;
[0013] A second acquisition module, configured to acquire the target code with differences in the program code;
[0014] A obtaining module, configured to process the target code based on the target model to obtain the target code detection result of the program code. The target model is a model obtained by optimizing the initial model. The initial model includes a language model and a prediction model. The prediction model is used to obtain the rule object of the detection code based on the target code, and the language model is used to obtain the program code detection result based on the prompt word. The prompt word is generated based on the rule object and the target code.
[0015] In a third aspect, the present disclosure provides a computer device, including: a memory and a processor. The memory and the processor are communicatively connected to each other. The memory stores computer instructions, and the processor executes the computer instructions to execute the method for obtaining the code detection result in the first aspect or any corresponding embodiment thereof.
[0016] In a fourth aspect, the present disclosure provides a computer-readable storage medium, on which computer instructions are stored. The computer instructions are used to cause a computer to execute the method for obtaining the code detection result in the first aspect or any corresponding embodiment thereof.
[0017] In a fifth aspect, the present disclosure provides a computer program product, including computer instructions, which are used to cause a computer to execute the method for obtaining the code detection result in the first aspect or any corresponding embodiment thereof. Description of the Drawings
[0018] To more clearly illustrate the specific embodiments of the present disclosure or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present disclosure. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0019] Figure 1 It is a schematic flowchart of a method for obtaining code detection results according to an embodiment of the present disclosure;
[0020] Figure 2 It is a schematic flowchart of another method for obtaining code detection results according to an embodiment of the present disclosure;
[0021] Figure 3 It is a schematic flowchart of yet another method for obtaining code detection results according to an embodiment of the present disclosure;
[0022] Figure 4 It is a structural block diagram of a device for obtaining code detection results according to an embodiment of the present disclosure;
[0023] Figure 5 It is a schematic hardware structure diagram of a computer device according to an embodiment of the present disclosure. Specific Embodiments
[0024] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present disclosure with reference to the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are some, but not all, of the embodiments of the present disclosure. Based on the embodiments of the present disclosure, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present disclosure.
[0025] Code detection is a crucial part of the software development process. With the development of technology, language models have powerful natural language understanding and generation capabilities, and can process structured and unstructured data such as code, comments, context, etc. This makes code detection based on language models possible, which can not only automatically generate code detection results, but also provide specific optimization suggestions according to project rules and context information. Among them, the current code detection solutions based on language models mainly focus on non-custom rules. Through prompt engineering, a unified template is constructed, taking code changes, context, and rule descriptions as inputs, and directly invoking the generation ability of the language model to output code detection opinions. However, it is only applicable to non-custom models and highly dependent on the model's capabilities. Based on this, a large amount of internal data or open-source data is used to perform supervised fine-tuning on the language model to make it perform better in specific code detection scenarios. However, it requires a large amount of labeled data and cannot cover custom rules.
[0026] Currently, to address the deficiencies of strong dependence on model capabilities for non-custom rules, limited accuracy of rule matching, lack of context integration, and poor interpretability, a code detection solution with custom rules is proposed. However, the current custom rules are limited to being applied in specific modules, functions, or business logics. The lack of context awareness can lead to these rules being either misapplied or omitted. At the same time, the current custom rules rely on traditional recall methods based on fixed rule sets and simple keyword matching, resulting in inaccurate final code detection.
[0027] To solve the above problems, according to an embodiment of the present disclosure, a method embodiment for obtaining code detection results is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0028] In this embodiment, a method for obtaining code detection results is provided, as Figure 1 shown. Figure 1 FIG. is a flowchart of the method for obtaining code detection results according to an embodiment of the present disclosure. This method is applied to the server side that can run the target model. The method process includes the following steps:
[0029] Step S101, obtain the program code to be detected.
[0030] Optionally, in the embodiments of the present disclosure, the program code to be detected can be obtained from the local development environment. Specifically: (1) Determine the code storage location: First, it is necessary to clarify the storage path of the code in the local computer. If development is carried out using an integrated development environment, the code is usually stored in the project folder. (2) Open the project file: In the integrated development environment, select "Open Project" through the "File" menu or directly select the recently used project on the welcome interface to find and open the project containing the code to be detected. It is also possible to directly enter the code storage directory in the file explorer and find the main code file of the project, such as the.java file in Java, the.py file in Python, etc. (3) Browse and select the code: In the project view of the integrated development environment, the project structure can be expanded to browse through each folder and file to find the specific code file and related classes, functions, etc. that need to be detected.
[0031] The program code to be detected can also be obtained from the code hosting platform. Specifically: (1) Log in to the code hosting platform: Common code hosting platforms include Platform A, Platform B, Platform C, etc. Open the browser, enter the platform website, and log in using the account. (2) Find the target project: Enter the project name or relevant keywords in the search bar of the platform, or find the project containing the code to be detected by browsing the project list, organizing projects, etc. (3) Obtain the program code to be detected.
[0032] The program code to be detected can also be obtained from the server. Specifically: (1) Establish a connection: If the code is stored on the server, a suitable tool is needed to establish a connection with the server. (2) Locate the code path: After logging in to the server, according to the file system structure and code deployment on the server, use the command-line tool or server management interface to find the directory where the code is located. (3) Obtain the program code to be detected. The embodiments of the present disclosure include, but are not limited to, the above methods for obtaining program code.
[0033] Step S102, obtain the target code with differences in the program code.
[0034] Optionally, the target code with differences can be obtained by comparing within the local repository. Specifically: Using the command line, (1) View the differences between the working area and the staging area: Open the terminal in the project root directory and use the diff command, which will display the code differences that have been modified in the working area but not yet added to the staging area. (2) View the differences between the staging area and the latest commit: Use the diff --staged or diff --cached command to view the code changes that have been added to the staging area but not yet committed. (3) Compare the differences between two commits: Use diff <commit1> <commit2>Command, wherein <commit1>And <commit2>Is the submitted hash value, so that the specific code changes between these two submissions can be seen.
[0035] It is also possible to obtain the target code with differences by comparing the remote repository with the local repository. Specifically: (1) Pull the update of the remote repository: Use the fetch command in the terminal to obtain the latest commit information from the remote repository, but do not merge it into the local branch. (2) Compare the differences: Use diff <local_branch> <remote branch>Command to compare the differences between the local branch and the remote branch and find the modified parts of the code.
[0036] It can be understood that the parts of the program code where there are changes, additions, or deletions are called the differential code that differs from the original code, and here it is called the target code.
[0037] Step S103, process the target code based on the target model to obtain the target code detection result of the program code. Among them, the target model is a model obtained by optimizing the initial model. The initial model includes a language model and a prediction model. The prediction model is used to obtain the rule object of the detection code based on the target code, and the language model is used to obtain the program code detection result based on the prompt word. The prompt word is generated based on the rule object and the target code.
[0038] Optionally, in the embodiments of the present disclosure, after obtaining the target code with differences in the program code, the target code is processed based on the target model obtained by optimizing the initial model, and the target code detection result of the program code is output.
[0039] Among them, the initial model consists of two parts, namely a language model and a prediction model. Among them, the prediction model is used to obtain the rule object of the detection code based on the target code. Here, the rule object is usually some code specifications that are most relevant to the target code. For example, in Python development, the PEP 8 coding specification stipulates rules such as code indentation and naming styles; developers may also have their own specific rules, such as the maximum number of lines of a function limit, the annotation specification that must be added, etc. Using the rule object to detect the target code improves the accuracy and comprehensiveness of the detection.
[0040] Generate a prompt word based on the rule object and the target code, and the language model obtains the program code detection result based on the prompt word. For example, the prompt word:
[0041] You are a code detection expert and need to detect the following changed code:
[0042] {Target code}
[0043] For the code to be detected, it is necessary to determine whether there are the following problems:
[0044] {Rule object}
[0045] When you are fully confident that the code within the detection range has the above problems, return the detection result in Chinese. The specific content includes:
[0046] ##line
[0047] The line number, that is, the starting position of the marked problem code;
[0048] ##issue_type
[0049] The type of the issue, i.e., the rule ID in the above issues;
[0050] ##issue
[0051] Specific issue description, including what the problem is, why there is this problem, and what consequences it may cause; for the same type of issues, at most 1 detection result can be output, and even if the line numbers are different, they should be merged.
[0052] Based on the above prompt words, the language model only needs to output the detection result of the target code. Based on the above prompt words, the detection result of the target code output can be: (1) Problem description: The generation model will clearly describe the problems existing in the code differences. (2) Reasons for the problem: Syntax errors, potential logical errors, violations of coding specifications. (3) Impact analysis: Analyze the impact of the problem on the local part of the code, such as the impact on the current function or class; consider the impact of the problem on the entire project, including potential impacts on other modules and system functions.
[0053] In addition, repair suggestions can also be output to help developers solve the problems, etc.
[0054] In the embodiments of the present disclosure, by introducing an advanced language model, the target code with differences in the program code to be detected is analyzed, the rule object of the detection code is screened out by the prediction model, and then based on the prompt words generated by the rule object and the target code combined with the language model, the detection result of the target code is generated. Since the rule object obtained by the prediction model is most relevant to the target code, the comprehensiveness of code detection is ensured, and the recall rate and accuracy of code detection are improved.
[0055] In some alternative embodiments, as Figure 2 shown, before step S103, the method further includes:
[0056] Step a1, obtaining training code samples.
[0057] Optionally, in the embodiments of the present disclosure, it is necessary to first train the initial models (i.e., the language model and the prediction model), and obtain the target model through the optimization and adjustment of the initial models, and then use the target model in the subsequent application stage of code detection.
[0058] Therefore, in the training stage, it is necessary to obtain training code samples. It should be noted that the way to obtain training code samples here can be to obtain them from the public code dataset platform, or from professional technical communities and forums, etc. The embodiments of the present disclosure include but are not limited to the above ways of obtaining training code samples.
[0059] Step a2: Obtain the code information with differences in the training code samples.
[0060] Optionally, after obtaining the training code samples, the fetch command can be used to display the code information with differences in the training code samples. For the specific method of obtaining the code information with differences, reference can be made to the steps of step S102 in the above embodiments, which will not be elaborated here.
[0061] In addition, the code information here can include the code itself with differences, or some associated information related to the code itself. Currently, it is necessary to perform syntax parsing on the code information to generate a corresponding abstract syntax tree and extract syntax structures and node information.
[0062] Step a3: Obtain feature vectors based on the embedded sub-model of the language model. Among them, the feature vectors include the first feature vector of the code information and the second feature vector of the information describing the preset rules.
[0063] Optionally, the language model includes an embedded sub-model (embedding model), which can generate high-dimensional code difference representation vectors to achieve the conversion of vector representations.
[0064] At this time, the embedded sub-model of the language model can be used to convert the code information (including syntax structures and node information in the abstract syntax tree, etc.) into the first feature vector, and convert the information describing the preset rules into the second feature vector. Among them, the preset rules refer to custom rules, and the information describing the custom rules can be rich metadata attached to the rules, including the applicable programming language, framework, rule category (security, performance, code style, etc.), applicable scope (variable level, function level, module level, project level), priority, etc. These metadata help to accurately locate the applicable scenarios of the rules and improve the accuracy of rule matching.
[0065] It should be noted that the description text of the rules (including custom rules) is encoded into the second feature vector. In this way, the semantic information of the rules is effectively captured, which is convenient for efficient comparison and matching with the first feature vector of the code differences.
[0066] Here, the above first feature vector and second feature vector are collectively referred to as the feature vectors obtained by the embedded sub-model.
[0067] Step a4: Input the feature vectors into the prediction model, and select the target rule object from the rule library containing the preset rules, where the relevance of the target rule object to the code information is greater than the correlation threshold.
[0068] Optionally, the obtained feature vectors are input into the constructed prediction model, and the prediction model intelligently predicts which rule categories or scopes may not be applicable.
[0069] Among them, through learning a large number of diverse codes and rules (including historical custom rules), the prediction model has strong generalization ability and can adapt to code changes under different projects, languages, and frameworks, and can also accurately predict codes and custom rules not encountered before.
[0070] After that, the prediction model selects target rule objects with a relevance greater than the relevance threshold (such as 90%) to the code information from the rule library containing preset rules. Only the selected target rule objects are the rules most relevant to the code information. Then, inapplicable rules can be excluded, greatly narrowing the scope of candidate rules, reducing the burden of subsequent calculations, and improving the efficiency and accuracy of rule recall.
[0071] Step a5: Generate a prompt word based on the target rule object and the feature vector, and input the prompt word into the language model to obtain the code detection result of the training code sample.
[0072] Optionally, after obtaining the selected target rule objects, generate a prompt word based on the target rule object and the feature vector, such as formulating some fixed templates, and filling the information in the feature vector and the target rule object into the templates to obtain the prompt word. Then input the prompt word into the language model, and the language model generates the code detection result of the training code sample based on the prompt word.
[0073] Step a6: Adjust the embedded sub-model and the prediction model of the language model based on the code detection result to obtain the target model.
[0074] Optionally, in the embodiments of the present disclosure, the embedded sub-model and the prediction model can be adjusted according to the code detection result. For example, a user interface is provided to allow the detector to give feedback on the applicability and effectiveness of each rule. When the code detection result is inconsistent with the result desired by the user, and the received feedback from the user is "applicable", "not applicable", "needs improvement", etc., the model parameters of the embedded sub-model and the prediction model can be adjusted, and then the finally trained target model can be obtained.
[0075] In the embodiments of the present disclosure, according to the specific code information and the corresponding selected target rule objects, targeted detection results are directly output, reducing manual intervention and improving detection efficiency; at the same time, based on the collected feedback data, the embedded sub-model and the prediction model are continuously trained and optimized to improve the judgment accuracy and self-adaptability of the target model.
[0076] In some alternative embodiments, the code information includes a differential code sample, and step a4 includes:
[0077] Step b1: Input the first feature vector and the second feature vector of the differential code sample into the prediction model.
[0078] Optionally, if the code information refers to a differential code sample (i.e., the code feature itself of the differential code), at this time, the first feature vector corresponding to the differential code sample and the second feature vector describing the information of the preset rules are input into the prediction model, and the prediction model is allowed to perform the screening of the custom rules.
[0079] In some alternative embodiments, the code information includes a differential code sample and the context information of the differential code sample. The above step a4 includes:
[0080] Step c1: Obtain the third feature vector of the context information based on the embedded sub-model of the language model.
[0081] Step c2: Fuse the first feature vector and the third feature vector to obtain the fourth feature vector.
[0082] Step c3: Input the fourth feature vector and the second feature vector into the prediction model.
[0083] Optionally, collect the relevant information of the file, module, and project where the differential code sample is located, including file path, module dependency relationship, project configuration, programming language and framework used, etc., to obtain the context information of the differential code sample. It should be noted that the differential code sample is constantly changing, but the context information related to the differential code sample is usually fixed. Therefore, after obtaining the differential code sample, it is very easy to obtain the context information.
[0084] If the code information refers to a differential code sample and the context information of the differential code sample, at this time, it is necessary to obtain the third feature vector of the context information based on the embedded sub-model of the language model, and then fuse the first feature vector and the third feature vector. For example, the simple splicing method: connect the two feature vectors end to end in sequence to form a longer new feature vector; the weighted summation method: assign a weight to each feature vector, and then multiply the corresponding elements by the weights and add them to obtain the fused feature vector. By adjusting the weights, the contribution degree of different feature vectors to the fusion result can be controlled. Fusion based on machine learning models: Use machine learning models (such as neural networks, support vector machines, etc.) to fuse the two feature vectors. After that, the fused fourth feature vector is obtained.
[0085] Input the fourth feature vector and the second feature vector describing the information of the preset rules into the prediction model, and let the prediction model perform the screening of the custom rules.
[0086] In the embodiments of the present disclosure, based on the differential code sample and the context information, the applicability and matching degree of the more accurate judgment rules can be obtained, and the accuracy and efficiency of the custom rule recall can be improved.
[0087] In some alternative embodiments, step a4 includes:
[0088] Step d1: Select a plurality of rule candidate objects from the rule library.
[0089] Step d2: Obtain the similarity between the feature vector and the feature vectors of the rule candidate objects to obtain the correlation score for each rule candidate object.
[0090] Step d3: Obtain the target rule object based on the correlation score.
[0091] Optionally, according to the output result of the prediction model, filter out the possibly applicable rule sets from the rule library containing custom rules to obtain a plurality of rule candidate objects. Then calculate the similarity between the feature vector and the feature vectors of the rule candidate objects to obtain the correlation score for each rule candidate object. Combine factors such as the priority of the rule candidate objects, historical application effects, and the preferences of the detectors to weight the correlation scores to obtain a comprehensive score. According to the comprehensive score, sort the recalled rule candidate objects in descending order and preferentially display the target rule object with the highest score.
[0092] To ensure that the detector can pay attention to the most relevant target rule object, a correlation score threshold can be set at this time to filter out the rule candidate objects with too low scores.
[0093] In the embodiments of the present disclosure, by performing correlation evaluation and sorting on the filtered rule candidate objects, it is ensured that the most relevant target rule objects are recalled and displayed preferentially.
[0094] In some alternative embodiments, step a6 includes:
[0095] Step e1: Obtain the feedback information input based on the detection result.
[0096] Step e2: Adjust the embedded sub-model and the prediction model based on the feedback information until the evaluation score output by the evaluation model is greater than the preset threshold, then stop the adjustment to obtain the target model.
[0097] Optionally, in addition to setting the training code samples for initial model training in the embodiments of the present disclosure, an evaluation set can also be set, and a large number of evaluation code samples are provided in this evaluation set.
[0098] By obtaining the feedback information input based on the detection result, adjusting the model parameters of the embedded sub-model and the prediction model based on the feedback information, and then inputting the evaluation code samples into the adjusted embedded sub-model and the adjusted prediction model to obtain the prediction results of the evaluation set. Then input the prediction results into the evaluation model to calculate the evaluation score.
[0099] Compare the evaluation score output by the evaluation model with a preset threshold. If the evaluation score is greater than the preset threshold, stop adjusting the embedded sub-model and the prediction model; otherwise, continue the adjustment to obtain the final target model.
[0100] In the embodiments of the present disclosure, the performance of the system is continuously optimized by collecting feedback to achieve the self-adaptation and continuous improvement of the target model.
[0101] In some alternative embodiments, the method further includes:
[0102] Step f1: Create a rule import interface.
[0103] Step f2: Create, modify, and manage preset rules based on the rule import interface, and add the preset rules to the rule library.
[0104] Optionally, provide a friendly rule import interface through which developers are allowed to create, modify, and manage preset rules according to project requirements. After the preset rules are created or modified, they are automatically incorporated into the structured rule library and support the setting of metadata (such as applicable language, applicable scope, rule category, priority, etc.), and are treated equally in terms of management and application as the system-built rules.
[0105] In the embodiments of the present disclosure, a structured rule library is established by creating a rule import interface to achieve dynamic rule screening and improve the efficiency and accuracy of rule recall.
[0106] As a specific application embodiment of the embodiments of the present disclosure, as Figure 3 shown, the specific process is as follows:
[0107] Extract code difference features, obtain context information, and metadata of preset rules;
[0108] Input the code difference features, context information, and metadata into the embedded sub-model;
[0109] Obtain a code difference representation vector, a context information representation vector, and a rule representation vector from the embedded sub-model of the language model;
[0110] Fuse the code difference representation vector and the context information representation vector to obtain a fused representation vector;
[0111] Input the fused representation vector and the rule representation vector into the prediction model for rule screening;
[0112] Output an applicable rule candidate set by the prediction model;
[0113] Calculate the similarity and sort the rule candidate objects in the fused representation vector and the rule candidate set to obtain a target rule object;
[0114] Generate a prompt based on the target rule object and the fused representation vector;
[0115] Obtain the code detection result from the language model based on the prompt;
[0116] Optimize the embedded sub-model and the prediction model of the language model based on the code detection result.
[0117] By introducing an advanced language model, the embodiments of the present disclosure perform context understanding and analysis on code differences, intelligently screen and recall relevant rule candidate objects, and dynamically rank the rule candidate objects, aiming to obtain the most relevant target rule object, improving the efficiency and accuracy of rule recall in the code detection process.
[0118] In this embodiment, a device for obtaining the code detection result is further provided. This device is used to implement the above embodiments and preferred implementation manners, and those that have been described will not be repeated. As used below, the term "module" may be a combination of software and / or hardware that implements a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.
[0119] This embodiment provides a device for obtaining the code detection result, as Figure 4 shown, including:
[0120] The first acquisition module 401 is used to acquire the program code to be detected;
[0121] The second acquisition module 402 is used to acquire the target code with differences in the program code;
[0122] The obtaining module 403 is used to process the target code based on the target model to obtain the target code detection result of the program code. The target model is a model obtained by optimizing the initial model. The initial model includes a language model and a prediction model. The prediction model is used to obtain the rule object for detecting the code based on the target code, and the language model is used to obtain the program code detection result based on the prompt. The prompt is generated based on the rule object and the target code.
[0123] In some alternative implementation manners, the device further includes:
[0124] The third acquisition module is used to acquire the training code sample before processing the target code based on the target model to obtain the target code detection result of the program code;
[0125] The fourth acquisition module is used to acquire the code information with differences in the training code sample;
[0126] A first obtaining module, configured to obtain a feature vector based on an embedded sub-model of a language model, where the feature vector includes a first feature vector of code information and a second feature vector of information describing a preset rule;
[0127] A selection module, configured to input the feature vector into a prediction model, and select a target rule object from a rule library including preset rules, where the relevance between the target rule object and the code information is greater than a relevance threshold;
[0128] A second obtaining module, configured to generate a prompt word based on the target rule object and the feature vector, and input the prompt word into the language model to obtain a code detection result for a training code sample;
[0129] A third obtaining module, configured to adjust the embedded sub-model and the prediction model of the language model based on the code detection result to obtain a target model.
[0130] In some alternative embodiments, the code information includes a differential code sample, and the selection module is configured to input the first feature vector and the second feature vector of the differential code sample into the prediction model.
[0131] The code information includes a differential code sample and context information of the differential code sample. The selection module is further configured to obtain a third feature vector of the context information based on the embedded sub-model of the language model; fuse the first feature vector and the third feature vector to obtain a fourth feature vector; and input the fourth feature vector and the second feature vector into the prediction model.
[0132] In some alternative embodiments, the selection module is further configured to select multiple rule candidate objects from the rule library; obtain the similarity between the feature vector and the feature vectors of the rule candidate objects to obtain a relevance score for each rule candidate object; and obtain the target rule object based on the relevance score.
[0133] In some alternative embodiments, the third obtaining module is configured to obtain feedback information input based on the detection result; adjust the embedded sub-model and the prediction model based on the feedback information until the evaluation score output by an evaluation model is greater than a preset threshold, and then stop the adjustment to obtain the target model.
[0134] In some alternative embodiments, the apparatus further includes:
[0135] A creation module, configured to create a rule import interface;
[0136] An addition module, configured to create, modify, and manage preset rules based on the rule import interface, and add the preset rules to the rule library.
[0137] The further function descriptions of the above-mentioned modules and units are the same as those in the corresponding foregoing embodiments, and will not be elaborated herein.
[0138] The device for obtaining the code detection result in this embodiment is presented in the form of a functional unit. Here, the unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and a memory that execute one or more software or fixed programs, and / or other devices that can provide the above functions.
[0139] This embodiment of the disclosure also provides a computer device having the above Figure 4 device for obtaining the code detection result as shown.
[0140] Please refer to Figure 5 , Figure 5 which is a schematic structural diagram of a computer device provided by an alternative embodiment of the disclosure. As shown in Figure 5 , the computer device includes: one or more processors 10, a memory 20, and an interface for connecting various components, including a high-speed interface and a low-speed interface. Each component communicates with each other using different buses and can be installed on a common motherboard or in other ways as needed. The processor can process instructions executed within the computer device, including instructions stored in the memory or on the memory to display graphical information of the GUI on an external input / output device (such as a display device coupled to the interface). In some alternative embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple computer devices can be connected, and each device provides some necessary operations (such as an array of servers, a set of blade servers, or a multi-processor system). Figure 5 In
[0141] FIG. 14, one processor 10 is taken as an example.
[0142] The processor 10 may be a central processing unit, a network processor, or a combination thereof. Among them, the processor 10 may further include a hardware chip. The above hardware chip may be an application specific integrated circuit, a programmable logic device, or a combination thereof. The above programmable logic device may be a complex programmable logic device, a field programmable gate array, a generic array logic, or any combination thereof.
[0143] The memory 20 may include a program storage area and a data storage area. Among them, the program storage area may store an operating system and application programs required for at least one function; the data storage area may store data created according to the use of the computer device, etc. In addition, the memory 20 may include a high-speed random access memory, and may also include a non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some alternative embodiments, the memory 20 may optionally include a memory remotely provided with respect to the processor 10, and these remote memories may be connected to the computer device through a network. Examples of the above-mentioned network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0144] The memory 20 may include a volatile memory, for example, a random access memory; the memory may also include a non-volatile memory, for example, a flash memory, a hard disk, or a solid-state drive; the memory 20 may further include a combination of the above types of memories.
[0145] The computer device further includes a communication interface 30 for the computer device to communicate with other devices or a communication network.
[0146] The embodiments of the present disclosure also provide a computer-readable storage medium. The methods according to the embodiments of the present disclosure may be implemented in hardware, firmware, or be implemented as computer code that can be recorded on a storage medium, or be implemented as computer code originally stored in a remote storage medium or a non-transitory machine-readable storage medium and to be downloaded through a network and stored in a local storage medium, so that the methods described herein can be stored in such software processes on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium may be a magnetic disk, an optical disk, a read-only memory, a random access memory, a flash memory, a hard disk, or a solid-state drive, etc.; further, the storage medium may also include a combination of the above types of memories. It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component that can store or receive software or computer code, and when the software or computer code is accessed and executed by the computer, the processor, or the hardware, the methods shown in the above embodiments are implemented.
[0147] A part of the present disclosure can be applied as a computer program product, such as computer program instructions, which, when executed by a computer, can invoke or provide the methods and / or technical solutions according to the present disclosure through the operations of the computer. Those skilled in the art should understand that the forms of existence of computer program instructions in a computer-readable medium include but are not limited to source files, executable files, installation package files, etc. Correspondingly, the ways in which computer program instructions are executed by a computer include but are not limited to: the computer directly executes the instructions, or the computer compiles the instructions and then executes the corresponding compiled program, or the computer reads and executes the instructions, or the computer reads and installs the instructions and then executes the corresponding installed program. Herein, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to the computer.
[0148] Although the embodiments of the present disclosure have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the present disclosure, and such modifications and variations all fall within the scope defined by the appended claims.< / remote> < / commit1>
Claims
1. A method for obtaining code detection results, characterized in that: The method comprises: Obtain the program code to be tested; Obtaining object code having differences in the program code; The target code is processed based on a target model to obtain a target code detection result of the program code, wherein the target model is a model obtained after optimizing an initial model, the initial model includes a language model and a prediction model, the prediction model is used to obtain a rule object for detecting the code based on the target code, and the language model is used to obtain a program code detection result based on a prompt word, and the prompt word is generated based on the rule object and the target code.
2. The method according to claim 1, characterized in that Before the target code is processed based on the target model to obtain the target code detection result of the program code, the method further includes: Get training code samples; Obtaining code information having differences in the training code samples; Obtaining a feature vector based on the embedded sub-model of the language model, wherein the feature vector includes a first feature vector of the code information and a second feature vector of information describing a preset rule; Inputting the feature vector into the prediction model, selecting a target rule object from a rule library containing the preset rules, wherein the relevance between the target rule object and the code information is greater than a relevance threshold; Generate a prompt word based on the target rule object and the feature vector, and input the prompt word into the language model to obtain a code detection result for the training code sample; The embedded sub-model of the language model and the prediction model are adjusted based on the code detection result to obtain the target model.
3. The method according to claim 2, characterized in that The code information includes a difference code sample, and the inputting the feature vector into the prediction model includes: The first feature vector and the second feature vector of a difference code sample are input into the prediction model.
4. The method according to claim 2, characterized in that: The code information includes a difference code sample and context information of the difference code sample, and the inputting the feature vector into the prediction model includes: Obtaining a third feature vector of the context information based on the embedded sub-model of the language model; Merging the first eigenvector and the third eigenvector to obtain a fourth eigenvector; The fourth eigenvector and the second eigenvector are input into the prediction model.
5. The method according to claim 2, characterized in that: The step of selecting a target rule object from a rule library containing the preset rule includes: Selecting a plurality of rule candidates from the rule base; Obtaining the similarity between the feature vector and the feature vector of the rule candidate object, and obtaining a relevance score for each of the rule candidate objects; The target rule object is obtained based on the relevance score.
6. The method according to claim 2, characterized in that The step of adjusting the embedded sub-model of the language model and the prediction model based on the code detection result to obtain the target model includes: Obtaining feedback information based on the detection result input; The embedded sub-model and the prediction model are adjusted based on the feedback information until the evaluation score output by the evaluation model is greater than a preset threshold, then the adjustment is stopped to obtain the target model.
7. The method according to claim 2, characterized in that The method further comprises: Create a rule import interface; The preset rules are created, modified and managed based on the rule import interface, and the preset rules are added to the rule base.
8. A device for obtaining code detection results, characterized in that: The device comprises: A first acquisition module, used for acquiring a program code to be detected; A second acquisition module is used to acquire target codes that are different from the program codes; A module is obtained, which is used to process the target code based on a target model to obtain a target code detection result of the program code, wherein the target model is a model obtained after optimizing the initial model, the initial model includes a language model and a prediction model, the prediction model is used to obtain a rule object for detecting the code based on the target code, and the language model is used to obtain a program code detection result based on a prompt word, and the prompt word is generated based on the rule object and the target code.
9. A computer device, characterized in that: include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the method for obtaining code detection results according to any one of claims 1 to 7 by executing the computer instructions.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the method for obtaining code detection results according to any one of claims 1 to 7.
11. A computer program product, characterized in that The method comprises computer instructions, wherein the computer instructions are used to cause a computer to execute the method for obtaining a code detection result according to any one of claims 1 to 7.