Code detection method and device, electronic equipment and readable storage medium
Through static detection tools, word frequency inverse document frequency characteristics, weak learners and language models, the early comprehensive problem of code quality detection in the application industry is solved, multi-level detection of code quality is realized, and detection costs and risks are reduced.
Patent Information
- Application Number
- CN202510308400.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-07-29
AI Technical Summary
During the rapid iteration of the application industry, existing test left-shift measures are difficult to fully detect code quality in the code development stage, resulting in high risk of stability problems after launch, and static detection tools lack understanding of the code context and complex maintenance.
Multi-level detection methods of static detection tools, word frequency inverse document frequency characteristics, weak learners and language models are adopted to realize multi-source detection of code files through combination of static analysis, frequency analysis and semantic analysis.
Move the test process left to the code not executed stage, reduce detection costs and risks, improve code detection depth and effect, increase detection dimensions, solve the limitations of static detection tools, and provide multi-level quality detection.
Smart Images

Figure CN120386702A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technology, and particularly to a code detection method, apparatus, electronic device, and readable storage medium. Background Art
[0002] Compared with traditional industries that pay more attention to safety and stability and have slow product updates, application industries (such as the game industry) have frequent business iterations, and it is often difficult for testing work to reach the depth and meticulousness of unit testing. In this case, it becomes more difficult to ensure the stability of application products after each iteration goes live.
[0003] To detect problems as early as possible, shorten the testing cycle, and reduce testing costs, test shifting left can be promoted. The core idea of test shifting left is that the earlier unreasonable points are found, the lower the probability of problems occurring. A series of test shifting left measures such as use case standardization, standard efficiency library, continuous integration, automated testing, and program code coverage have indeed advanced the testing work to the development stage, but it still needs to be carried out after the code is executed, and there is still a problem of relatively high risk.
[0004] It should be noted that the information disclosed in the above background art section is only used to enhance the understanding of the background of the present disclosure, and thus may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention
[0005] The present disclosure provides a code detection method, apparatus, electronic device, and readable storage medium to solve or at least partially solve the above problems, specifically as follows.
[0006] In a first aspect, the present disclosure provides a code detection method, the method comprising:
[0007] Determining a first detection result of a code file to be detected through a static detection tool based on a detection policy;
[0008] Determining the term frequency-inverse document frequency feature of the code file;
[0009] Determining a second detection result of the code file through a weak learner according to the term frequency-inverse document frequency feature of the code file;
[0010] Determining a third detection result of the code file through a first language model;
[0011] Obtaining a multi-source detection result of the code file through a second language model according to the first detection result, the second detection result, the third detection result, and the code file.
[0012] In a second aspect, the present disclosure further provides a code detection apparatus, the apparatus comprising:
[0013] The first code detection module is configured to determine a first detection result of a code file to be detected through a static detection tool based on a detection policy;
[0014] The feature determination module is configured to determine the term frequency-inverse document frequency feature of the code file;
[0015] The second code detection module is configured to determine a second detection result of the code file through a weak learner according to the term frequency-inverse document frequency feature of the code file;
[0016] The third code detection module is configured to determine a third detection result of the code file through a first language model;
[0017] The detection result fusion module is configured to obtain a multi-source detection result of the code file through a second language model according to the first detection result, the second detection result, the third detection result, and the code file.
[0018] In a third aspect, the present disclosure further provides an electronic device, including: a processor, a memory, and computer program instructions stored on the memory and executable on the processor;
[0019] When the processor executes the computer program instructions, the code detection method described in the first aspect above is implemented.
[0020] In a fourth aspect, the present disclosure further provides a computer-readable storage medium, in which computer program instructions are stored, and when the computer program instructions are executed by a processor, the code detection method described in the first aspect above is implemented.
[0021] The exemplary embodiments of the present disclosure have the following beneficial effects:
[0022] The code detection method provided by the present disclosure can perform basic static detection on the code file to be detected through a static detection tool based on a detection strategy, ensure the basic quality of the code, and obtain the first detection result of the code file; it can determine the term frequency-inverse document frequency feature of the code file, and based on the term frequency-inverse document frequency feature of the code file, perform deeper quality detection on the code file through a weak learner to obtain the second detection result of the code file; it can understand the code semantic structure through a first language model, comprehensively analyze the context and intention of the code, identify potential logical errors in the code, and obtain the third detection result of the code; finally, according to the first detection result, the second detection result, and the third detection result, obtain the multi-source detection result of the code through a second language model. The present disclosure can not only shift the testing process to the stage where the code has not been executed, reduce the detection cost and problem risk, but also utilize a three-layer bottom-up detection structure of static analysis, frequency analysis, and semantic analysis to perform multi-level quality detection on the code, improving the depth and effect of code detection; new code problems can be solved not only by maintaining and configuring static detection strategies, but also by frequency analysis and semantic analysis, without relying solely on maintaining and configuring static detection strategies to solve, increasing the ways to detect new code problems. In addition, the present disclosure applies the statistical method of term frequency-inverse document frequency to code detection, analyzes the code file from the frequency dimension, increases the analysis dimension of the code file, and improves the comprehensiveness of code detection. Description of the Drawings
[0023] Figure 1 is a flowchart of a code detection method provided by one embodiment of the present disclosure;
[0024] Figure 2 is a schematic diagram of a step framework of a code detection provided by one embodiment of the present disclosure;
[0025] Figure 3 is a schematic diagram of a front-end page for managing and displaying detection strategies provided by one embodiment of the present disclosure;
[0026] Figure 4 is a block diagram of a code detection device provided by one embodiment of the present disclosure;
[0027] Figure 5 is a schematic diagram of the logical structure of an electronic device for implementing code detection provided by one embodiment of the present disclosure. Detailed Embodiments
[0028] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present disclosure. Apparently, the described embodiments are only some, rather than all, of the embodiments of the present disclosure. The components of the embodiments of the present disclosure described and illustrated herein generally may be arranged and designed in a variety of different configurations. Therefore, the detailed description of the embodiments of the present disclosure provided herein is not intended to limit the scope of the claimed present disclosure, but merely represents selected embodiments of the present disclosure. Based on the embodiments of the present disclosure, every other embodiment obtained by a person of ordinary skill in the art without creative efforts shall fall within the protection scope of the present disclosure.
[0029] As used in this specification, the terms "a", "an", "the", and "said" are used to denote the presence of one or more elements / components / etc.; the terms "comprising" and "having" are used to mean an open inclusion and refer to the presence of additional elements / components / etc. in addition to the listed elements / components / etc.; the terms "first", "second", etc. are only used as labels and are not a limitation on the quantity of their objects.
[0030] It should be understood that in the embodiments of the present disclosure, "at least one" means one or more, and "a plurality" means two or more. "And / or" is merely a description of the association relationship of associated objects and indicates that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally means that the associated objects before and after are in an "or" relationship. "Including A, B, and / or C" means including any one, any two, or all three of A, B, and C.
[0031] It should be understood that in the embodiments of the present disclosure, "B corresponding to A", "B corresponding to A relatively", "A corresponding to B relatively", or "B corresponding to A relatively" means that B is associated with A, and B can be determined according to A. Determining B according to A does not mean determining B only according to A, and B can also be determined according to A and / or other information.
[0032] The code detection method in one of the embodiments of the present disclosure may run on a local terminal device or a server. When the code detection method runs on the server, the method may be implemented and executed based on a cloud interaction system, where the cloud interaction system includes a server and a client device.
[0033] In an alternative embodiment, various cloud applications can run under the cloud interaction system, such as cloud games. Taking cloud games as an example, cloud games refer to a game mode based on cloud computing. In the operation mode of cloud games, the running entity of the game program and the entity presenting the game screen are separated. The storage and running of the code detection method are completed on the cloud game server, and the role of the client device is to receive, send data, and present the game screen. For example, the client device can be a display device with data transmission function near the user side, such as a mobile terminal, a television, a computer, a personal digital assistant, etc.; however, the cloud game server in the cloud is responsible for information processing. When playing a game, the player operates the client device to send an operation instruction to the cloud game server. The cloud game server runs the game according to the operation instruction, encodes and compresses data such as the game screen, and returns it to the client device through the network. Finally, the client device decodes and outputs the game screen.
[0034] In an alternative embodiment, taking a game as an example, the local terminal device stores the game program and is used to present the game screen. The local terminal device is used to interact with the player through the graphical user interface, that is, conventionally, the game program is downloaded and installed on the electronic device and run. The manner in which the local terminal device provides the graphical user interface to the player can include various ways. For example, it can be rendered and displayed on the display screen of the terminal, or provided to the player through holographic projection. For example, the local terminal device can include a display screen and a processor. The display screen is used to present the graphical user interface, which includes the game screen, and the processor is used to run the game, generate the graphical user interface, and control the display of the graphical user interface on the display screen.
[0035] In a possible embodiment, the present disclosure provides a code detection method, which provides a graphical user interface through a terminal device, where the terminal device can be the aforementioned local terminal device or the client device in the aforementioned cloud interaction system.
[0036] The execution entity of this method can be a terminal device or a server. The terminal device can be a desktop computer, a laptop computer, a tablet computer, a mobile phone, etc., or other electronic devices, which are not specifically limited in this disclosure. The server is used to provide background services for the client of the application program in the terminal device. For example, the server can be the background server of the above application program. The server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms, which are not specifically limited in this disclosure.
[0037] Before elaborating on the embodiments of the present disclosure in detail, the related technologies will be further introduced first.
[0038] Currently, a series of test shift-left measures such as use case standardization, standard efficiency library, continuous integration, automated testing, and program code coverage do indeed push the testing work forward to the development stage, but it still needs to be carried out in the stage after the code is executed, and there are still problems of relatively high detection costs and problem risks.
[0039] Therefore, it can be considered to shift the testing process left to the stage when the program has just written the code but has not yet executed it. Currently, the detection method for code at this stage usually relies on static detection tools, but static detection tools detect code problems through built-in fixed policies, lack an understanding of the code context, and cannot provide deeper solutions. Secondly, static detection tools need to continuously update the built-in policies to solve new code problems, resulting in relatively complex maintenance and configuration.
[0040] Figure 1 Illustrates a code detection method provided by one embodiment of the present disclosure, as Figure 1 shown, this method includes the following steps S101 to S105.
[0041] Step S101: Determine the first detection result of the code file to be detected through a static detection tool based on a detection policy.
[0042] The following will refer to Figure 2 the step framework shown to elaborate on the embodiments of the present disclosure in detail.
[0043] In this step, the code file to be detected can be subjected to basic static detection through a static detection tool based on a detection policy to ensure the basic quality of the code file. Optionally, the code file to be detected can include multiple code files.
[0044] In an alternative embodiment, first, a plurality of static detection tools based on detection policies can be used to perform static detection on the code file to be detected respectively, obtaining a plurality of static detection results; if the plurality of static detection results indicate that the code file has an anomaly, the code anomaly types included in the plurality of static detection results can be mapped to the preset code anomaly categories in the preset code anomaly category set, obtaining the first detection result of the code file.
[0045] The first detection result may include a conclusion of no anomaly for the code file (indicating that the code file has no problem), or, the problematic code snippet in the code file and the corresponding code anomaly type of the problematic code snippet (indicating that the code file has a problem).
[0046] Different static detection tools can provide different but rich detection policy sets, which can complement each other well in detection policies. However, the result output styles and the classification differences of code anomaly types of different static detection tools are relatively large. In this embodiment, the code file is statically detected by different static detection tools respectively, and each static detection tool can output a static detection result. The preset code anomaly category set may include a plurality of preset code anomaly categories. If at least one static detection result indicates that the code file has an anomaly, that is, there is an error in the code, the code anomaly types included in each static detection result can be mapped to the corresponding preset code anomaly categories in the preset code anomaly category set, that is, the code anomaly types output by each static detection tool are converted into the corresponding preset code anomaly categories in the preset code anomaly category set, so as to more effectively unify and manage the outputs of different static detection tools.
[0047] Optionally, the static detection tools to be used can be determined according to the language type of the code file. For example, for a code file written in the Java language, the PMD (Programming Mistake Detector) tool and the CheckStyle tool can be used for static detection. For example, for a code file written in the JavaScript language, the PMD (Programming Mistake Detector) tool and the ESLint tool can be used for static detection.
[0048] Optionally, the preset code anomaly categories included in the preset code anomaly category set can be SonarLint categories (i.e., the code anomaly categories provided by the SonarLint tool).
[0049] Optionally, all detection policies can be imported into the database, and all detection policies can be divided into different policy groups, such as Figure 3As shown, on the front - end page, policy information of the detection policies included in the policy group can be displayed for the policy group. The policy information can include, for example, the name of the detection policy, the code anomaly category corresponding to the detection of the detection policy, the severity level of the code anomaly category corresponding to the detection of the detection policy, the description information of the detection policy, the recommended code corresponding to the detection policy (i.e., the correct code that is detected as having no anomaly using this detection policy), the non - recommended code corresponding to the detection policy (i.e., the incorrect code that is detected as having an anomaly using this detection policy), etc. Optionally, the user can add a new detection policy or delete an added detection policy.
[0050] Optionally, the code anomaly types included in the multiple static detection results can be mapped to the preset code anomaly categories in the preset code anomaly category set through a third - language model (i.e., LLM, Large Language Model). In this embodiment, the LLM can be used as a classification model to perform secondary processing on the results output by the static code detection method, so as to map them to the corresponding categories in the preset code anomaly category set (such as the SonarLint category). Exemplarily, if it is required to map all the code anomaly types included in the static detection results output by the ESlint tool to the SonarLint category, when using the LLM, the input prompt information can include: one is the static detection results output by the ESlint tool, and the other is all the SonarLint categories. Thus, the LLM can map all the code anomaly types included in the static detection results output by the ESlint tool to the SonarLint category based on the prompt information.
[0051] Optionally, if all the multiple static detection results indicate that the code file has no anomaly, it is determined that the first detection result of the code file is that the code file has no anomaly.
[0052] Step S102: Determine the term frequency - inverse document frequency feature of the code file.
[0053] Step S103: According to the term frequency - inverse document frequency feature of the code file, determine the second detection result of the code through a weak learner.
[0054] In steps S102 and S103, the term frequency - inverse document frequency is a statistical method used to evaluate the importance of a word in a document set. The present disclosure can apply this statistical method to code detection, analyze the code file from the frequency dimension, and perform a deeper - level code quality detection on the code file by analyzing the frequency feature of the code file.
[0055] The second detection result can include a conclusion of no anomaly for the code file (indicating that the code file has no problem), or a conclusion of having an anomaly for the code file (indicating that the code file has a problem).
[0056] In an optional embodiment, the steps of determining the term frequency-inverse document frequency feature of the code file can be implemented in the following manner, including:
[0057] Perform text segmentation on the code file through a text splitter to obtain multiple code snippets;
[0058] For the symbol units in the code snippets, determine the term frequency-inverse document frequency values of the symbol units;
[0059] Concatenate the term frequency-inverse document frequency values of the symbol units in multiple code snippets to obtain the term frequency-inverse document frequency feature of the code file.
[0060] In this embodiment, the code file can be segmented into multiple code snippets through a text splitter (such as the text splitter provided by the LangChain framework). Optionally, the text segmentation of the code file can be performed according to the logical functions in the code. For example, the function definition and initialization parts in the code file can be segmented into one code snippet, the start part of the main loop in the code file can be segmented into one code snippet, the logical part for processing dictionary items in the code file can be segmented into one code snippet, the logical part for processing list items in the code file can be segmented into one code snippet, and the return result part in the code file can be segmented into one code snippet.
[0061] Then, for the symbol units (tokens) in each code snippet, such as keywords, identifiers, constants, operators, delimiters, etc., determine their term frequency-inverse document frequency values, that is, TF-IDF (Term Frequency_Inverse Document Frequency) values. The TF-IDF value can be used to evaluate the importance of a word for a document set. In this embodiment, the TF-IDF value can be used to evaluate the importance of a token for a code file set. In this embodiment, the term frequency (TF value) refers to the frequency of a token appearing in a certain code snippet, and the inverse document frequency (IDF value) refers to the prevalence of a token in all code snippets. If the number of code snippets containing a certain token is less, the IDF value of that token is larger, indicating that the token has good category discrimination ability. The term frequency-inverse document frequency value of a token is the product of the TF value and the IDF value of that token. The formula is as follows.
[0062] W x,y = TF x,y × IDF x
[0063]
[0064] Wherein, TFx,y is the word frequency of symbol unit x in code snippet y, n x,y is the number of times symbol unit x appears in code snippet y, ∑ k n k,y is the total number of times all symbol units appear in code snippet y, IDF x is the inverse document frequency of symbol unit x in all code snippets, N is the total number of all code snippets, N x is the number of code snippets in which symbol unit x appears in all code snippets.
[0065] Exemplarily, for example, for the code snippet of the logical part of the processing dictionary item in the code file, the calculation result of its TF-IDF value can be as shown in Table 1 below.
[0066] Table 1
[0067] token TF-IDF value isinstance 2.7 dict 3.6 … …
[0068] After that, the word frequency-inverse document frequency values of the symbol units in all code snippets can be concatenated to form a feature vector, so as to obtain the word frequency-inverse document frequency feature of the code file, which is used as the input of the weak learner.
[0069] In an optional embodiment, the second detection result of the code file can be determined by a heterogeneous weak learner according to the word frequency-inverse document frequency feature of the code file.
[0070] A heterogeneous weak learner refers to that the weak learner part in an ensemble learning system (including weak learners and a meta-model) contains different types of weak learners. The weak learners in the heterogeneous weak learner not only have different learning abilities but also different types. This diversity helps to improve the prediction performance of the heterogeneous weak learner.
[0071] In this embodiment, first, the word frequency-inverse document frequency feature of the code file can be input into multiple different individual learners respectively. Using the stacking method in ensemble learning, multiple prediction results are generated in parallel by multiple different heterogeneous weak learners. These prediction results are used as the input features of the meta-model (also called the meta-learner). Optionally, the voting method can be used inside the meta-model to replace the traditional neural network layer. The meta-model can make a final prediction based on the prediction results of multiple weak learners, and the detection result with the highest number of votes is selected as the final prediction result from the prediction results of multiple weak learners. Since the three different weak learners are all classification models, the output results of the weak learners are both one of the two results: there is an anomaly or there is no anomaly. The output result of the meta-model is also one of the two results: there is an anomaly or there is no anomaly. Among them, both the weak learner and the meta-model act as binary classifiers.
[0072] Optionally, the heterogeneous weak learners may include the following weak learners: logistic regression classifier, support vector machine, and random forest. Exemplarily, the above three weak learners may output the results shown in Table 2 below for the code file.
[0073] Table 2
[0074] weak learner output result logistic regression classifier problematic support vector machine no problem random forest problematic
[0075] Step S104: Determine the third detection result of the code file through the first language model.
[0076] In this step, a language model may be used to perform semantic analysis on the code file. The language model can understand complex code semantic results, comprehensively parse the code context and intent, so as to identify potential logical errors.
[0077] The third detection result may include a conclusion of no exception for the code file (indicating that the code file has no problem), or, the problematic code snippet in the code file and the corresponding code exception type of the problematic code snippet (indicating that the code file has a problem).
[0078] In an optional embodiment, the third detection result of the code file may be determined through the first language model after pre-training and model fine-tuning.
[0079] In this embodiment, the language model is usually pre-trained using general domain data, so that although the language model performs well in the general domain, its performance in the vertical domain is not satisfactory. To enable the language model to better understand the complex code semantic structure and identify problems related to code quality, the pre-trained language model may be fine-tuned through the constructed code samples.
[0080] Optionally, the training method of the first language model may include the following steps:
[0081] Obtain the pre-trained first language model;
[0082] Construct positive code snippet samples and negative code snippet samples;
[0083] According to the positive code snippet samples and negative code snippet samples, perform model fine-tuning on the pre-trained first language model to obtain the first language model after pre-training and model fine-tuning.
[0084] Among them, the positive code snippet sample refers to a code snippet with an exception, and the negative code snippet sample refers to a code snippet without an exception.
[0085] In an optional example, the positive and negative code snippet samples can be constructed in the following way: The high-star Github code repositories (such as axios, nextjs, node) can be segmented, and the code snippets that meet the preset conditions are segmented based on keywords and used as the negative code snippet samples required for model fine-tuning; According to the SonarLint tool and / or custom detection strategies, introductions, cases and other materials, multiple corresponding positive code snippet samples in different scenarios are generated using prompt information in the LLM.
[0086] Then, the pre-trained first language model can be fine-tuned using contrastive learning. Contrastive Learning is a self-supervised learning method that aims to learn feature representations by bringing the representations of similar samples (input samples and positive samples) closer and pulling the representations of dissimilar samples (input samples and negative samples) farther apart, enabling the model to learn that the feature representations of similar samples are closer and those of dissimilar samples are farther apart. During the training process, the model does not rely on labels but learns through the similarity between samples. After fine-tuning, the first language model can have better performance in the vertical domain of code detection.
[0087] Optionally, the first language model can be fine-tuned using instructions, and the prompt information for model fine-tuning is output to the first language model. The prompt information includes a preset set of code anomaly categories, as well as the positive and negative code snippet samples (i.e., positive and negative sample pairs) corresponding to each preset code anomaly category in the preset set of code anomaly categories.
[0088] In this embodiment, the third detection result of the code file can be determined by the first language model pre-trained based on general data and fine-tuned based on the positive and negative code snippet samples. The code file to be detected, the preset code anomaly categories (such as all SonarLint categories), and the detection instructions can be input into the first language model. The first language model can detect the code file in response to the detection instructions and represent and output the detected code anomaly problems using the preset code anomaly categories.
[0089] Through model fine-tuning based on positive and negative samples, the semantics of each anomaly problem can be solidified into the first language model, improving the performance of the first language model in the vertical domain, reducing situations such as understanding errors and misalignments, and enabling it to better adapt to the application scenario of code quality detection.
[0090] Step S105: Obtain the multi-source detection result of the code file through the second language model according to the first detection result, the second detection result, the third detection result, and the code file.
[0091] In this step, the first detection result, the second detection result, and the third detection result obtained by detecting the code file in different ways can be fused through the second language model. The second language model is instructed to give the final detection result after comprehensively considering the above three detection results, and thus the multi-source detection result of the code file can be obtained. Among them, the second language model can only be pre-trained without fine-tuning.
[0092] In an alternative embodiment, the result output by static detection (i.e., the first detection result) is the code snippet without anomalies or problems and its corresponding preset code anomaly type (e.g., SonarLint category). The result output by machine learning detection (i.e., the second detection result) can include no anomalies or anomalies. The result output by language model detection (i.e., the third detection result) can include the code snippet without anomalies or problems and its corresponding preset code anomaly type (e.g., SonarLint category).
[0093] In an alternative embodiment, according to the first detection result, the second detection result, the third detection result, and the code file, obtaining the multi-source detection result of the code file through the second language model can be achieved in the following ways, including:
[0094] Determine the fourth detection result of the code file through the pre-trained second language model;
[0095] According to the first detection result, the second detection result, the third detection result, and the fourth detection result, obtain the multi-source detection result of the code file through the second language model.
[0096] In this embodiment, first, the second language model pre-trained based on general data is used to detect the code file to obtain the fourth detection result of the code file. Since the first language model is pre-trained and fine-tuned, while the second language model is only pre-trained without fine-tuning, the detection results of the first language model stronger in the vertical domain and the second language model stronger in the general domain will be different, thus enabling the detection of code to make up for deficiencies. Then, the second language model is instructed to give the final detection result after comprehensively considering the above four detection results, and thus the multi-source detection result of the code file can be obtained.
[0097] Optionally, the first language model can adopt models such as Qwen (Tongyi Qianwen), GPT (Generative Pre-trained Transformer), etc. The second language model can adopt models such as Qwen, GPT, etc. In an alternative example, the first language model can adopt the Qwen2 model, and the second language model can adopt the GPT model.
[0098] In an alternative embodiment, the code detection method may further include the following steps:
[0099] If the first detection result indicates that the code is abnormal, map the code abnormality type included in the first detection result to a preset code abnormality category in the preset code abnormality category set;
[0100] If the third detection result indicates that the code is abnormal, map the code abnormality type included in the third detection result to a preset code abnormality category in the preset code abnormality category set.
[0101] In this embodiment, the code abnormality types in the first detection result and the third detection result can both be converted into the corresponding preset code abnormality categories in the preset code abnormality category set, for example, both are converted into the corresponding SonarLint categories, so as to unify the categories of the detection results and facilitate management and processing. Optionally, this step can be executed by the above-mentioned second language model.
[0102] If the first detection result, the second detection result, and the third detection result all indicate that the code file has no abnormality, then the multi-source detection result is no abnormality.
[0103] In an alternative embodiment, the code detection method may further include the following steps:
[0104] Classify the problem code segments in the first detection result and the third detection result according to the preset code abnormality category. Optionally, this step can be executed by the above-mentioned second language model.
[0105] In this embodiment, the problem codes in the detection results obtained by performing code detection in different ways can be classified according to the preset code abnormality category, which is convenient for management and processing.
[0106] In an alternative embodiment, the code detection method may further include the following steps:
[0107] Deduplicate the first detection result, the second detection result, and the third detection result.
[0108] In this embodiment, the detection results obtained by performing code detection in different ways can be deduplicated. For example, if the same problem is detected for the same code segment p in both the first detection result and the third detection result, and the second detection result also detects a problem for the code segment p, then only the first detection result or the third detection result of the code segment p can be retained. Optionally, this step can be executed by the above-mentioned second language model.
[0109] Optionally, the multi-source detection result can be stored in a database for retrieval.
[0110] The code detection method provided by the present disclosure can perform basic static detection on the code file to be detected through a static detection tool based on a detection strategy, ensure the basic quality of the code, and obtain the first detection result of the code file; it can determine the term frequency-inverse document frequency feature of the code file, and according to the term frequency-inverse document frequency feature of the code file, perform deeper quality detection on the code file through a weak learner to obtain the second detection result of the code file; it can understand the code semantic structure through a first language model, comprehensively analyze the context and intention of the code, identify potential logical errors in the code, and obtain the third detection result of the code file; finally, according to the first detection result, the second detection result, and the third detection result, obtain the multi-source detection result of the code through a second language model. The present disclosure can not only shift the testing process to the stage where the code has not been executed, reduce the detection cost and problem risk, but also utilize a three-layer bottom-up detection structure of static analysis, frequency analysis, and semantic analysis to perform multi-level quality detection on the code, improving the depth and effect of code detection; new code problems can be solved not only by maintaining and configuring static detection strategies, but also by frequency analysis and semantic analysis, without solely relying on maintaining and configuring static detection strategies to solve, increasing the ways to detect new code problems. In addition, the present disclosure applies the statistical method of term frequency-inverse document frequency to code detection, analyzes the code file from the frequency dimension, increases the analysis dimension of the code file, and improves the comprehensiveness of code detection.
[0111] Corresponding to the code detection method provided by the embodiments of the present disclosure, the embodiments of the present disclosure also provide a code detection device. As Figure 4 shown, the device 700 includes:
[0112] A first code detection module 71, configured to determine a first detection result of a code file to be detected through a static detection tool based on a detection strategy;
[0113] A feature determination module 702, configured to determine the term frequency-inverse document frequency feature of the code file;
[0114] A second code detection module 703, configured to determine a second detection result of the code file through a weak learner according to the term frequency-inverse document frequency feature of the code file;
[0115] A third code detection module 704, configured to determine a third detection result of the code file through a first language model;
[0116] A detection result fusion module 705, configured to obtain a multi-source detection result of the code file through a second language model according to the first detection result, the second detection result, the third detection result, and the code file.
[0117] In an alternative embodiment, determining the term frequency-inverse document frequency feature of the code file includes:
[0118] Performing text segmentation on the code file through a text splitter to obtain a plurality of code snippets;
[0119] Determining the term frequency-inverse document frequency value of the symbol unit in the code snippet;
[0120] Concatenating the term frequency-inverse document frequency values of the symbol units in the plurality of code snippets to obtain the term frequency-inverse document frequency feature of the code file.
[0121] In an alternative embodiment, determining the second detection result of the code file through a weak learner according to the term frequency-inverse document frequency feature of the code file includes:
[0122] Determining the second detection result of the code file through a heterogeneous weak learner according to the term frequency-inverse document frequency feature of the code file.
[0123] In an alternative embodiment, determining the third detection result of the code file through a first language model includes:
[0124] Determining the third detection result of the code file through a first language model after pre-training and model fine-tuning;
[0125] Wherein, the training method of the first language model includes:
[0126] Obtaining a pre-trained first language model;
[0127] Constructing positive code snippet samples and negative code snippet samples;
[0128] Performing model fine-tuning on the pre-trained first language model according to the positive code snippet samples and the negative code snippet samples to obtain a first language model after pre-training and model fine-tuning.
[0129] In an alternative embodiment, determining the first detection result of a code file to be detected through a static detection tool based on a detection strategy includes:
[0130] Performing static detection on the code file to be detected through a plurality of static detection tools based on a detection strategy respectively to obtain a plurality of static detection results;
[0131] If the plurality of static detection results indicate that the code file has an anomaly, mapping the code anomaly types included in the plurality of static detection results to preset code anomaly categories in a preset code anomaly category set to obtain the first detection result of the code file.
[0132] In an alternative embodiment, mapping the code anomaly types included in the multiple static detection results to preset code anomaly categories in a preset code anomaly category set includes:
[0133] Mapping the code anomaly types included in the multiple static detection results to preset code anomaly categories in a preset code anomaly category set through a third language model.
[0134] In an alternative embodiment, obtaining the multi-source detection result of the code file through a second language model according to the first detection result, the second detection result, the third detection result, and the code file includes:
[0135] Determining a fourth detection result of the code file through a pre-trained second language model;
[0136] Merging the first detection result, the second detection result, the third detection result, and the fourth detection result through the second language model to obtain the multi-source detection result of the code file.
[0137] In an alternative embodiment, the apparatus is further configured to:
[0138] If the first detection result indicates that the code file has an anomaly, mapping the code anomaly type included in the first detection result to a preset code anomaly category in a preset code anomaly category set;
[0139] If the third detection result indicates that the code file has an anomaly, mapping the code anomaly type included in the third detection result to a preset code anomaly category in a preset code anomaly category set.
[0140] In an alternative embodiment, the apparatus is further configured to:
[0141] Classifying the problem code segments in the first detection result, the second detection result, and the third detection result according to the preset code anomaly category.
[0142] In an alternative embodiment, the apparatus is further configured to:
[0143] Removing duplicates from the first detection result, the second detection result, and the third detection result.
[0144] The code detection device provided by the present disclosure can perform basic static detection on the code file to be detected through a static detection tool based on a detection strategy, ensure the basic quality of the code, and obtain the first detection result of the code file; it can determine the term frequency-inverse document frequency feature of the code file, and based on the term frequency-inverse document frequency feature of the code file, perform deeper quality detection on the code file through a weak learner to obtain the second detection result of the code file; it can understand the code semantic structure through a first language model, comprehensively analyze the context and intention of the code, identify potential logical errors in the code, and obtain the third detection result of the code; finally, based on the first detection result, the second detection result, and the third detection result, obtain the multi-source detection result of the code through a second language model. The present disclosure can not only shift the testing process to the stage where the code has not been executed, reduce the detection cost and problem risk, but also utilize a three-layer bottom-up detection structure of static analysis, frequency analysis, and semantic analysis to perform multi-level quality detection on the code, improving the depth and effect of code detection; new code problems can be solved not only by maintaining and configuring static detection strategies, but also by frequency analysis and semantic analysis, without relying solely on maintaining and configuring static detection strategies, increasing the ways to detect new code problems. In addition, the present disclosure applies the statistical method of term frequency-inverse document frequency to code detection, analyzes the code file from the frequency dimension, increases the analysis dimension of the code file, and improves the comprehensiveness of code detection.
[0145] Next, an electronic device provided by an embodiment of the present disclosure will be introduced. Please refer to Figure 5 , Figure 5 which is a schematic structural diagram of the electronic device provided by an embodiment of the present disclosure. Among them, the code detection device described in the embodiments of the present disclosure can be deployed on the electronic device 800 to implement the functions in the embodiments of the present disclosure. Specifically, the electronic device 800 includes: a receiver 801, a transmitter 802, a processor 803, and a memory 804 (where the number of processors 803 in the electronic device 800 can be one or more, Figure 5 and one processor is taken as an example here), where the processor 803 can include an application processor 8031 and a communication processor 8032. In some embodiments of the present disclosure, the receiver 801, the transmitter 802, the processor 803, and the memory 804 can be connected through a bus or other means.
[0146] The memory 804 may include a read-only memory and a random access memory, and provide instructions and data to the processor 803. A part of the memory 804 may also include a non-volatile random access memory (NVRAM). The memory 804 stores processor and operation instructions, executable modules, or data structures, or subsets thereof, or extended sets thereof, where the operation instructions may include various operation instructions for implementing various operations.
[0147] The processor 803 controls the operation of the execution device. In a specific application, the various components of the execution device are coupled together through a bus system, where the bus system may include a power bus, a control bus, a status signal bus, etc. in addition to a data bus. However, for the sake of clarity, all kinds of buses are referred to as the bus system in the figure.
[0148] The methods disclosed in the above embodiments of the present disclosure may be applied to the processor 803 or implemented by the processor 803. The processor 803 may be an integrated circuit chip with signal processing capabilities. In the implementation process, the steps of the above methods may be completed by the integrated logic circuit in hardware or instructions in software form in the processor 803. The above-mentioned processor 803 may be a general-purpose processor, a digital signal processor (DSP), a microprocessor or a microcontroller, and may further include an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The processor 803 may implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments of the present disclosure. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present disclosure may be directly embodied as being executed and completed by a hardware decoding processor, or executed and completed by a combination of hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory 804, and the processor 803 reads the information in the memory 804 and combines its hardware to complete the steps of the above method.
[0149] The receiver 801 can be used to receive input digital or character information, and generate signal inputs related to the relevant settings and function controls of the execution device. The transmitter 802 can be used to output digital or character information through the first interface; the transmitter 802 can also be used to send instructions to the disk array through the first interface to modify the data in the disk array; the transmitter 802 may also include a display device such as a display screen.
[0150] In an embodiment of the present disclosure, the application processor 8031 in the processor 803 is used to execute the code detection method in the embodiment of the present disclosure. It should be noted that the specific manner in which the application processor 8031 executes each step is based on the same concept as each method embodiment in the present disclosure, and the technical effects brought by it are the same as those of each method embodiment in the present disclosure. For specific content, reference can be made to the description in the method embodiments shown above in the present disclosure, and details will not be repeated here.
[0151] The embodiment of the present disclosure also provides a chip for running instructions, and the chip is used to execute the technical solution of the code detection method in the above embodiment.
[0152] The embodiment of the present disclosure also provides a computer-readable storage medium, in which computer instructions are stored. When the computer instructions run on a processor, the processor is caused to execute the technical solution of the code detection method in the above embodiment.
[0153] The embodiment of the present disclosure also provides a computer program product, including a computer program, and the computer program is used to execute the technical solution of the code detection method in the above embodiment when executed by a processor.
[0154] The above computer-readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk or an optical disk. The readable storage medium can be any available medium accessible by a general or dedicated server.
[0155] It should be understood that the present disclosure is not limited to the exact structure already described and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.
[0156] Although the present disclosure is disclosed above with preferred embodiments, it is not used to limit the present disclosure. Any person skilled in the art can make possible changes and modifications without departing from the spirit and scope of the present disclosure. Therefore, the protection scope of the present disclosure should be subject to the scope defined by the claims of the present disclosure.
Claims
1. A code detection method, characterized in that, The method includes: Determining a first detection result of a code file to be detected by a static detection tool based on a detection policy; Determining the term frequency-inverse document frequency feature of the code file; Determining a second detection result of the code file by a weak learner according to the term frequency-inverse document frequency feature of the code file; Determining a third detection result of the code file by a first language model; Obtaining a multi-source detection result of the code file by a second language model according to the first detection result, the second detection result, the third detection result, and the code file.
2. The method according to claim 1, wherein The determining the term frequency-inverse document frequency feature of the code file includes: Performing text segmentation on the code file by a text splitter to obtain a plurality of code segments; Determining the term frequency-inverse document frequency value of the symbol unit in the code segment; Concatenating the term frequency-inverse document frequency values of the symbol units in the plurality of code segments to obtain the term frequency-inverse document frequency feature of the code file.
3. The method according to claim 1, wherein The determining a second detection result of the code file by a weak learner according to the term frequency-inverse document frequency feature of the code file includes: Determining a second detection result of the code file by a heterogeneous weak learner according to the term frequency-inverse document frequency feature of the code file.
4. The method according to claim 1, wherein The determining a third detection result of the code file by a first language model includes: Determining a third detection result of the code file by a first language model after pre-training and model fine-tuning; Wherein, the training method of the first language model includes: Obtaining a pre-trained first language model; Constructing positive code segment samples and negative code segment samples; Performing model fine-tuning on the pre-trained first language model according to the positive code segment samples and the negative code segment samples to obtain a first language model after pre-training and model fine-tuning.
5. The method according to claim 1, characterized in that The determining a first detection result of a code file to be detected by a static detection tool based on a detection policy includes: Performing static detection on the code file to be detected by a plurality of static detection tools based on a detection policy respectively to obtain a plurality of static detection results; If the plurality of static detection results indicate that the code file has an abnormality, mapping the code abnormality types included in the plurality of static detection results to preset code abnormality categories in a preset code abnormality category set to obtain the first detection result of the code file.
6. The method according to claim 5, characterized in that, The mapping the code abnormality types included in the plurality of static detection results to preset code abnormality categories in a preset code abnormality category set includes: Mapping the code abnormality types included in the plurality of static detection results to preset code abnormality categories in a preset code abnormality category set by a third language model.
7. The method according to claim 1, wherein The obtaining a multi-source detection result of the code file by a second language model according to the first detection result, the second detection result, the third detection result, and the code file includes: Determining a fourth detection result of the code file by a pre-trained second language model; The multi-source detection result of the code file is obtained by combining the first detection result, the second detection result, the third detection result, and the fourth detection result through the second language model.
8. The method according to claim 1, characterized in that, The method further includes: If the first detection result indicates that the code file has an anomaly, map the code anomaly type included in the first detection result to a preset code anomaly category in the preset code anomaly category set; If the third detection result indicates that the code file has an anomaly, map the code anomaly type included in the third detection result to a preset code anomaly category in the preset code anomaly category set.
9. The method according to claim 1, wherein The method further includes: Classify the problem code segments in the first detection result, the second detection result, and the third detection result according to the preset code anomaly category.
10. The method according to claim 1, characterized in that, The method further includes: Deduplicate the first detection result, the second detection result, and the third detection result.
11. A code detection device, characterized in that, The apparatus includes: A first code detection module, configured to determine a first detection result of a code file to be detected through a static detection tool based on a detection policy; A feature determination module, configured to determine the term frequency-inverse document frequency feature of the code file; A second code detection module, configured to determine a second detection result of the code file through a weak learner according to the term frequency-inverse document frequency feature of the code file; A third code detection module, configured to determine a third detection result of the code file through a first language model; A detection result fusion module, configured to obtain the multi-source detection result of the code file through a second language model according to the first detection result, the second detection result, the third detection result, and the code file.
12. An electronic device, characterized in that, It includes: A processor, a memory, and computer program instructions stored on the memory and executable on the processor; When the processor executes the computer program instructions, the code detection method according to any one of claims 1 to 10 above is implemented.
13. A computer-readable storage medium, characterized in that, Computer program instructions are stored in the computer-readable storage medium, and when the computer program instructions are executed by the processor, they are used to implement the code detection method according to any one of claims 1 to 10 above.