A code location method based on programming language migration

By building a function library and a total keyword list, a subject word self-test crawler system and an adversarial generator are designed, and a comparison learning fine-tuning large model encoder and pre-trained Softmax regression classifier are combined to solve the problem of cross-programming language code positioning and achieve efficient and accurate code positioning.

CN119045880BActive Publication Date: 2025-05-06NINGBO JINWANG INFORMATION IND CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411536704.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-31
Publication Date
2025-05-06
Estimated Expiration
2044-10-31

AI Technical Summary

Technical Problem

It is difficult for the prior art to implement code positioning across programming languages, and existing methods cannot quickly locate the relevant code for specifying functions or specifying typical processes.

Method used

By building a function library and a total keyword list, designing a subject word self-test crawler system and an adversarial generator, combining a comparative learning and fine-tuning large model encoder and pre-trained Softmax regression classifier, code positioning across programming languages ​​is achieved.

Benefits of technology

It implements the specified functions across programming languages ​​or code positioning that specify typical processes, improving the efficiency and accuracy of code positioning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119045880B_ABST
    Figure CN119045880B_ABST
Patent Text Reader

Abstract

The invention discloses a code location method based on programming language migration, which relates to computer software. Functions are refined into typical processes to construct a function library and a total keyword table, function labels are generated through quasi-transformation, a subject word self-checking crawler system is designed, a large number of accurate main and sub-language code segments are obtained based on a subject relevance discrimination method and a hash matching method to construct a main and sub-language code library, an adversarial generator is pre-trained based on an adversarial training strategy to achieve main and sub-language feature alignment, a large model encoder is fine-tuned using contrastive learning in combination with the main code library to maximize the separability of code representation, a pre-trained Softmax regression classifier is assisted to achieve function label assignment of code segments, words to be located and codes to be located are obtained, the codes to be located are split and the words to be located are identified based on a dual-channel query strategy to select positive examples, and the positive examples and code segments are input into a pre-trained and fine-tuned network to achieve code location of a specified function or a specified typical process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to computer software, and in particular to a code positioning method based on programming language migration. Background Art

[0002] When developing new required functions or maintaining existing code, developers will try to obtain code snippets corresponding to specific theme functions and understand their logical structure. When it comes to code-level analysis, the complex source code structure and the developer's poor coding style will cause the developer to spend a lot of time to locate the code snippets and sort out the implementation logic. Therefore, finding a method to quickly locate the code snippets is of great practical significance for developers.

[0003] The invention patent with the existing announcement number CN109240700B proposes a key code positioning method and system, which collects the function call relationship starting from the entry function of the interface parameter constraint code in the scenario of preset input parameters through program instrumentation, and performs key code analysis on each function based on this to locate all constraint codes related to the interface parameters.

[0004] The key code location of the existing technology is often only for a single programming language, and it is impossible to achieve code location for multiple programming languages. In addition, code location is achieved by inputting preset parameters for key code analysis, and it is impossible to quickly locate the relevant code of a specified function or a specified typical process. Summary of the invention

[0005] In view of the deficiencies in the prior art, the present invention proposes a code location method based on programming language migration to achieve code location of a specified function or a specified typical process across programming languages.

[0006] To achieve the above object, the present invention provides the following technical solutions:

[0007] A code location method based on programming language migration includes the following specific steps:

[0008] Building a function library and the total keyword list ;

[0009] Retrieve the main language code snippet through the keyword self-checking crawler system, and add functional tags to build the main code base ;

[0010] Combine the keyword self-check crawler system and sample enhancement to collaboratively obtain paralanguage code snippets to build a paracode library ;

[0011] Build an adversarial generator and pre-train it based on the adversarial training strategy;

[0012] Contrastive learning is used to fine-tune the large model encoder, and the Softmax regression classifier is further pre-trained;

[0013] Obtain the words to be located, split the code to be located based on artificial intelligence or static code analysis tools, select positive examples based on the dual-channel query strategy, and locate the code.

[0014] Further, build a function library and the total keyword list The specific steps include:

[0015] All functions that can be implemented by statistical programming languages;

[0016] Get all typical processes corresponding to a single function;

[0017] Extract keywords corresponding to a single typical process to obtain all keywords corresponding to a single function and construct a keyword table for a single function;

[0018] Store single functions and corresponding keyword tables in the function library until all single functions have been stored, the function library The build is complete;

[0019] Extract function library All keywords in the table are merged to obtain non-repeated keywords, and non-repeated keywords are rearranged to construct a total keyword table. .

[0020] Further, build the main code base The specific steps include:

[0021] Determine the main language and traverse the function library one by one For a single function in the project, the single function and the main language are delivered as keywords to the keyword self-check crawler system to build a code set corresponding to the single function;

[0022] Get the function library The keyword table corresponding to a single function in the comparison table is compared with the total keyword table. To generate a function code corresponding to a single function;

[0023] The function code corresponding to a single function is converted into a function label through a reversible transformation, and a code set corresponding to the single function is assigned;

[0024] The code sets and assigned function labels are stored sequentially until all code sets for a single function have been stored. Build completed.

[0025] Further, build a sub-code library The specific steps include:

[0026] Determine the paralanguage and deliver it as a subject word to the subject word self-check crawler system until the search stops;

[0027] Determine whether the number of sub-language code segments reaches the total sub-threshold. If so, directly build the sub-code library ;

[0028] If it is not reached, the sub-language code segments are generated through sample enhancement until the total sub-threshold is reached, and the sub-code library is constructed. .

[0029] Furthermore, the subject word self-checking crawler system includes a scheduling module, a downloading module, an extraction module, and a data pipeline module;

[0030] The scheduling module receives and determines whether the subject words have been crawled, generates crawler requests based on the uncrawled subject words, and decides whether to send the crawler requests to the download module or insert them into the Redis task queue based on whether the crawling task is being executed;

[0031] The download module uses random dynamic User-Agent and random delay to disguise crawler requests, access the website and download crawler responses;

[0032] The extraction module introduces a web crawling algorithm to extract page elements and calculate scores based on text density. It normalizes the page elements with the highest scores to generate the body text, and uses find() and find_all() commands to extract the title, content, and URL in the body text.

[0033] The data pipeline module uses the topic relevance discrimination method to filter out relevant text from the text, and removes duplicate related text through hash matching deduplication method, and uses OCR technology to recognize characters in the content to automatically generate code snippets.

[0034] Furthermore, the topic relevance determination method includes the following specific steps:

[0035] Get all the subject words and convert them into subject word vectors one by one;

[0036] The TextRank algorithm is used to extract several keywords and convert them into keyword vectors in turn;

[0037] Calculate the cosine distance between a single topic word vector and a single keyword vector, and take the mean of all cosine distances as the topic relevance;

[0038] Set a relevance threshold. If the topic relevance is greater than or equal to the relevance threshold, it is considered as relevant text.

[0039] Furthermore, the hash matching deduplication method removes duplicate related texts including the following specific steps:

[0040] Use the URLparse() command to parse the new URL, concatenate the protocol and domain name to generate the core segment;

[0041] Set a fixed-length bit array and compress the core segment into a fingerprint sequence using the SHA1 algorithm;

[0042] Input the fingerprint sequence into the Bloom filter to generate several hash values, and take the modulus of the bit array length to generate index keys;

[0043] Based on whether the index key exists in Redis, the new URL is discarded or stored.

[0044] Furthermore, building and pre-training the adversarial generator includes the following specific steps:

[0045] From the main repository and sub-repositories The same number of primary language code segments and secondary language code segments are randomly extracted from the BERT network, and intuitive features are generated as the input of the adversarial generator.

[0046] The intuitive features are input into the confusion generation network, the potential features are extracted through multiple convolutions, and unified features are generated through linear modulation and ReLU function mapping;

[0047] The unified features are input into the feature restoration network, the dimensions are expanded through deconvolution, and the reconstructed intuitive features are obtained through nonlinear transformation restoration;

[0048] The unified features are input into the domain discrimination network, the unified features are mapped to the separable space through nonlinear transformation, and the unified features are determined to come from the primary language or the secondary language through logistic regression;

[0049] Considering feature restoration loss and domain discrimination loss, we introduce balance parameters and define the loss function ;

[0050] Repeatedly input intuitive features into the adversarial generator and update the loss function based on the adversarial training strategy Until the loss function At a minimum, the main language code segment and the secondary language code segment achieve feature alignment.

[0051] Furthermore, the adversarial training strategy refers to the confusion generation network and feature restoration network to perform positive updates, and the domain discrimination network to perform bidirectional updates;

[0052] Forward update means that the hyperparameters in the confusion generation network and the feature restoration network are updated along the gradient direction of the feature restoration loss reduction;

[0053] Bidirectional updating means that the hyperparameters of the domain discrimination network are updated in the direction of the gradient that reduces the domain discrimination loss, and the hyperparameters in the confusion generation network and the feature restoration network are updated in the direction of the gradient that increases the domain discrimination loss.

[0054] During the training process, the confusion generation network and the feature restoration network will continue to compete with the domain discrimination network to improve their own capabilities. When the adversarial generator reaches a Nash equilibrium, the loss function Minimum.

[0055] Furthermore, fine-tuning the large model encoder using contrastive learning includes the following specific steps:

[0056] Select the GraphCodeBERT pre-trained model as the large model encoder;

[0057] From the main repository Randomly select two main language code segments from all code sets and extract the corresponding function labels;

[0058] All extracted main language code segments are input into the adversarial generator to obtain unified features, and the unified features are further input into the large model encoder to generate code representation;

[0059] Defining the loss function And through contrastive learning, multiple rounds of code representation optimization are performed to continuously reduce the loss function , when the loss function When is minimum, the distinguishability of the code representation is maximized, fixing the hyperparameters of the large model encoder.

[0060] Furthermore, the pre-training Softmax regression classifier includes the following specific steps:

[0061] Set up a Softmax regression model, reduce the dimension of the code representation through the fully connected layer, and map it to Within the interval, output the predicted probability vector;

[0062] Setting the multi-classification loss function To indicate the similarity between the distribution of the predicted probability vector and the actual probability vector;

[0063] The code representation is input into the Softmax regression classifier for multiple rounds, and the hyperparameters of the fully connected layer are updated using the gradient descent method so that the multi-classification loss function To achieve the minimum, fix the hyperparameters of the Softmax regression classifier.

[0064] Furthermore, code location includes the following specific steps:

[0065] Based on the dual-channel query strategy, the words to be located are identified and the main code base is Select positive examples from

[0066] Input the positive examples and several code segments obtained by splitting the code to be located into the adversarial generator in sequence to obtain the unified features of all single code segments;

[0067] The unified features of all single code segments are further input into the large model encoder, and the output code representations are sequentially input into the Softmax regression classifier to output the prediction probability vector corresponding to the single code segment;

[0068] The function label with the largest probability is selected from the predicted probability vector and assigned to a single code segment, and all single code segments consistent with the positive example function label are obtained.

[0069] Furthermore, selecting positive examples based on the dual-channel query strategy includes the following specific steps:

[0070] Identify the words to be located and judge them as specified functions or typical processes;

[0071] If it is a specified function, query the function library , get the keyword table corresponding to the specified function;

[0072] Compare with the total keyword table to obtain the function code corresponding to the specified function and convert it into a function label;

[0073] Query the main code base , randomly select a main language code segment with the same function label as a positive example;

[0074] If it is a specified typical process, obtain the keywords corresponding to the typical process;

[0075] Determine the position of the keyword in the total keyword table, generate the benchmark function code and convert it into the benchmark function label;

[0076] Get all function labels that are greater than or equal to the benchmark function label, and restore them to the corresponding function codes;

[0077] Compare with the benchmark function code, and place the function label corresponding to the function code containing the specified typical process into the label set to be located;

[0078] Select a feature tag from the set of tags to be located and A main language code segment with the same function label is randomly selected as a positive example for code location, until all function labels in the label set to be located are selected.

[0079] Compared with the prior art, the present invention has the following significant advantages:

[0080] 1. Build a keyword table from the function to the typical process, compare the total keyword table to generate function codes, and generate function labels through reversible transformation, supplemented by a dual-channel query strategy to provide two code positioning methods: designated function positioning and designated typical process positioning;

[0081] 2. Design a keyword self-check crawler system, combine it with Redis to achieve fast data storage and distributed access, introduce a unified extraction rule for web crawling algorithms, design a topic relevance determination method and a hash matching deduplication method to filter out irrelevant and duplicate web pages, and improve acquisition efficiency and accuracy. At the same time, use sample enhancement to solve the problem of insufficient crawling of low-resource paralanguage code segments by the keyword self-check crawler system, and ensure that the level of the main code base and the sub-code base can effectively support the pre-training and fine-tuning of the neural network;

[0082] 3. Design an adversarial generator and perform pre-training based on the adversarial training strategy to achieve feature alignment of the primary and secondary language code segments. Fine-tune the large model encoder through contrastive learning to assist the Softmax regression classifier in assigning accurate functional labels to the code segments. BRIEF DESCRIPTION OF THE DRAWINGS

[0083] Figure 1 A flow chart of a code location method based on programming language migration;

[0084] Figure 2 A function label flow chart corresponding to a single function is generated in the present invention;

[0085] Figure 3 A flowchart of the relevant text for removing duplicates using the hash matching deduplication method of the present invention;

[0086] Figure 4 This is a diagram of the adversarial generator model in the present invention;

[0087] Figure 5 It is a flowchart of code positioning in the present invention;

[0088] Figure 6 This is a flow chart of the dual-channel query strategy in the present invention. DETAILED DESCRIPTION

[0089] The present invention is further described in detail below in conjunction with the accompanying drawings and embodiments.

[0090] like Figure 1 As shown, the embodiment provided by the present invention: a code location method based on programming language migration includes the following specific steps:

[0091] Building a function library and the total keyword list ;

[0092] Retrieve the main language code snippet through the keyword self-checking crawler system, and add functional tags to build the main code base ;

[0093] Combine the keyword self-check crawler system and sample enhancement to collaboratively obtain paralanguage code snippets to build a paracode library ;

[0094] Construct an adversarial generator and use an adversarial training strategy to pre-train the adversarial generator to achieve alignment of primary and secondary language features;

[0095] Contrastive learning is used to fine-tune the large model encoder to maximize the distinguishability of code representation, and the code representation is further input into a Softmax regression classifier for pre-training;

[0096] Obtain the words to be located and the codes to be located, split the codes to be located, identify the words to be located based on the dual-channel query strategy, select positive examples, and perform code location.

[0097] Further, build a function library and the total keyword list The specific steps include:

[0098] All functions that can be realized by statistical programming languages, assuming that there are function, denoted as ;

[0099] Get the Features Corresponding to all typical processes, assuming that The typical process includes description, looping, comparison, sorting and calculation;

[0100] Extract the keywords corresponding to each typical process in turn. correspond A typical process, so the function Also corresponds to Keywords, combination Keywords to build functionality Corresponding keyword table ,in, For Function The corresponding Keywords;

[0101] The first Features And the corresponding keyword table Synchronous record in the function library After all functions and corresponding keyword tables are recorded, the function library The build is complete;

[0102] Extract function library All keywords in the list are merged after deduplication. Assuming that the remaining keywords after deduplication are non-repeating keywords, rearrange the non-repeating keywords in lexicographical order, and construct a total keyword table .

[0103] Further, build the main code base The specific steps include:

[0104] Determine a primary language to distinguish secondary languages. The primary language specifically refers to a single programming language with a large amount of public resources and standardized data. The secondary language generally refers to all programming languages ​​other than the selected primary language. In this embodiment, C++ is selected as the primary language.

[0105] For function libraries Middle Features , the function and the main language as the subject word, based on the continuous search function of the subject word self-checking crawler system The corresponding main language code segment until the function The number of corresponding main language code segments reaches a single main threshold When the function stops, Corresponding Main language code snippets packaged into a build code set ;

[0106] like Figure 2 As shown, initialize the generation All zero code, query function library To get the function Corresponding keyword table , compare the total keyword table , if the total keyword list Middle Keywords In the keyword table In the Bit is set to 1, thus generating the function Corresponding function code , and the function Corresponding function code Transformed into functional labels through reversible transformation , and assign code set The original value and the transformed value of the reversible transformation must correspond one to one, and there must not be a one-to-many or many-to-one situation, that is, the original value is unique and the transformed value is unique. In this embodiment, the function code is directly selected. Treat it as a binary code, and then convert it into a decimal number as a function label through base conversion. ;

[0107] Code Set and function labels Synchronously stored in the main code base Continue to search through the keyword self-check crawler system Corresponding main language snippets to build your code set , additional function tags And stored synchronously in the main code base middle;

[0108] When the function library middle The code sets and function labels corresponding to each function are stored in the main code library In the main code base Build completed.

[0109] Further, build a sub-code library The specific steps include:

[0110] Determine the required sub-language. In this embodiment, Python and Rust are selected as sub-languages.

[0111] Since there may be insufficient public resources on the Internet for paralanguages, and the standardization of their writing and annotation is low, it is difficult to retrieve the paralanguage code snippets of specified functions and attach functional tags. Therefore, only the number of paralanguages ​​collected is equal to the total paralanguage threshold. The paralanguage code segments do not need to be distinguished in terms of function. The selected paralanguage is directly used as the subject word. The subject word self-checking crawler system continuously searches for the paralanguage code segments until the search stops.

[0112] If the number of sub-language code segments is less than the total sub-threshold , then the reason for stopping is that the sub-language public code resources are insufficient, and the sub-language code segments are generated through sample enhancement until the total sub-language threshold is reached. , otherwise, the reason for stopping is determined to be the completion of the retrieval task;

[0113] Will Sub-language code segments are directly stored to build sub-code bases .

[0114] Specifically, the sample enhancement and supplementary generation of the paralanguage code segment includes the following methods:

[0115] The code back-translation method is adopted, that is, through a multi-language generation model trained on large-scale code data, the multi-language generation model includes CodeGeeX that can achieve zero-sample code translation, to translate the para-language code segment into code segments of other programming languages, and then translate it back to the para-language code segment. The para-language code segment after back-translation can keep the code segment function unchanged, but the specific code presentation form is different from the para-language code segment before back-translation;

[0116] The variable name replacement method is adopted, that is, the variable name and function name in the paralanguage code segment are renamed to ensure that the presentation form of the paralanguage code segment is changed without changing the function. The specific implementation methods include naming standardization and naming randomization. Naming standardization means that the newly appearing variable names or function names in the paralanguage code segment are named in the order of appearance. ,Name randomization refers to randomly selecting an unselected name from a pre-established name pool to replace the newly appearing variable name or function name in the ,paralanguage code segment.

[0117] Furthermore, the keyword self-check crawler system is improved based on the Scrapy crawler framework, including scheduling module, download module, extraction module and data pipeline module;

[0118] The scheduling module receives the subject words and determines whether the subject words have been crawled, discards the crawled subject words and generates crawler requests based on the uncrawled subject words, and further determines whether the crawling task is being executed. If there is no crawling task being executed, the crawler request is sent to the download module, otherwise the crawler request is inserted into the Redis task queue. The traditional Scrapy crawler framework stores crawler requests in disk space, but disk space cannot be shared by multiple hosts. Redis can share crawler requests to achieve multi-host collaborative crawling and improve crawler efficiency.

[0119] The download module uses random dynamic User-Agent and random delay to disguise crawler requests to improve the success rate of crawler systems accessing websites and downloading crawler responses. User-Agent is used to inform websites of the client's operating system, browser and other properties. If crawler requests always use a fixed User-Agent, it is easy for websites to identify and prohibit crawlers. Random dynamic User-Agent pre-configures an identity list that saves commonly used User-Agent information and randomly selects User-Agent during a single crawl. In addition, the default access interval in the Scrapy crawler framework is 2 seconds, while the interval for ordinary clients to access websites is generally random. Random delay creates a class called random delay to generate a random interval of 1-5 seconds. The crawler system simulates the random behavior of clients accessing websites based on the random interval.

[0120] The extraction module introduces the existing web crawling algorithm into the Spider module in the Scrapy crawler framework, obtains and optimizes the crawler response to generate the text, and uses the find() and find_all() commands to extract the title, content and URL in the text. Since the number and distribution of page elements on different websites are generally different, page elements include navigation links, headers, footers, advertising pop-ups and text. The Scrapy crawler framework needs to formulate one-to-one adaptive extraction rules when facing different websites. The introduced web crawling algorithm extracts page elements by constructing a DOM tree, calculates scores based on the text density of page elements, and standardizes the page elements with the highest scores to generate the text, which effectively unifies the extraction rules of different websites;

[0121] The data pipeline module uses the topic relevance discrimination method to filter out relevant texts from the acquired texts, and removes duplicate related texts through hash matching deduplication. The content in the relevant texts may include code snippets presented in text form and code snippets presented in image form. OCR technology is used to recognize characters in the content to automatically generate code snippets.

[0122] Furthermore, the topic relevance determination method includes the following specific steps:

[0123] Get all the subject words, and transform the single subject words into high-dimensional subject word vectors through the Skip-gram model.

[0124] The TextRank algorithm is used to extract several keywords from the title and content of the text. The number of extracted keywords is consistent with the number of subject words. The single keyword is converted into a high-dimensional keyword vector through the Skip-gram model.

[0125] Calculate the cosine distance between a single subject word vector and a single keyword vector in turn, and use the average of all cosine distances as the subject relevance between the subject word and the text;

[0126] Set a relevance threshold. If the topic relevance is greater than or equal to the relevance threshold, it is considered as relevant text.

[0127] like Figure 3 As shown, further, the hash matching deduplication method removes duplicate related texts including the following specific steps:

[0128] Parse the new URL using the URLparse() command in the Scrapy crawler framework, obtain the protocol and domain name in the new URL, and concatenate them to generate the core segment;

[0129] set up Dimensional bit array, and set the value of all bits to 0, and compress the core segment into a fingerprint sequence through the SHA1 algorithm;

[0130] The fingerprint sequence is input into the Bloom filter, and the Bloom filter passes Independent hash functions are used to calculate the fingerprint sequence and generate the corresponding hash values;

[0131] Will The hash values ​​correspond to the length of the bit array Take the modulus, find the corresponding position in the bit array and set the value to 1 to generate the index key;

[0132] Based on the index key, a quick query is performed in Redis to determine whether the index key exists. If so, the new URL is discarded. Otherwise, the index key corresponding to the new URL is stored in Redis.

[0133] Furthermore, the advantages of hash matching deduplication include:

[0134] Traditional deduplication solutions directly use the full URL for detection. However, the full URL consists of the protocol, domain name, port number, path, query parameters and anchor. Among them, the protocol and domain name determine the accessed web page, and the rest only confirm the port number, specific file, customized content and specific location of the accessed web page. Therefore, multiple different URLs may index the same web page, that is, the same source is different. The hash matching deduplication method pre-parses the URL to obtain the core segment, and uses the core segment for judgment to solve the repeated crawling caused by the same source is different;

[0135] The traditional solution for detecting URLs is to compare them with URLs stored in the database. However, URLs include numbers, characters, symbols, and even garbled characters. They are long and complex in structure. High computing power is required when traversing the database for comparison. The hash matching deduplication method first reduces the URL dimension through the SHA1 algorithm, and converts the complex URL comparison into a simple binary bit comparison based on the properties of the Bloom filter, greatly improving efficiency.

[0136] like Figure 4 As shown, further, constructing and pre-training the adversarial generator to achieve primary and secondary language feature alignment includes the following specific steps:

[0137] From the main repository Randomly select Main language code segment , from the secondary code base Randomly select Paralanguage code snippets , and concatenate to generate the input code segment ;

[0138] Enter the code snippet Generate intuitive features through BERT network processing As the input of the adversarial generator, intuitive features refer to the lower-level features that can be directly discovered through observation. Due to the differences between the primary language and the secondary language, the intuitive features vary greatly;

[0139] The intuitive features Input Confusion Generation Network , extracting latent features through multiple convolutions, and mapping them to a unified space through linear modulation and ReLU function to generate unified features ,Unified features refer to the essential features of the code segment, which are ,highly similar features shared by the primary language and the secondary ,language;

[0140] Unify the features Input feature restoration network , expand the unified feature dimension through multiple deconvolutions, and restore it to the space where the intuitive features are located through linear modulation and ReLU function to obtain the reconstructed intuitive features, feature restoration network The purpose is to prevent confusion in the generated network Excessive hyperlocation to the point of completely losing intuitive features;

[0141] Unify the features Input domain discrimination network , the unified features are further mapped to the separable space through the nonlinear transformation composed of linear modulation and nonlinear activation function, and the unified features are determined to be from the primary language or the secondary language through logistic regression;

[0142] Defining the loss function , the specific formula is as follows:

[0143] ,

[0144] in, , and They represent the confusion generation network , Feature Restoration Network and domain discrimination network The hyperparameters involved in is the balance parameter, is a field label used to indicate the Unifying Features From the primary or secondary language, and They are feature restoration loss and domain discrimination loss, respectively. In this embodiment, MSE and cross entropy are used respectively;

[0145] The intuitive features Repeatedly input the adversarial generator and generate the network through confusion , Feature Restoration Network and domain discrimination network After processing, the loss function is updated based on the adversarial training strategy , when the loss function At the minimum, the domain discrimination network It is impossible to distinguish whether the uniform feature comes from the primary language or the secondary language, but the feature reduction network Intuitive features can be reconstructed based on unified features, and the hyperparameters of the adversarial generator can be fixed. The main language code segment and the secondary language code segment can be aligned. That is, in the feature space where the unified features are located, the main language code segment and the secondary language code segment can be regarded as code segments of the same language.

[0146] Furthermore, adversarial training strategies refer to confusing the generated network and feature restoration network During the training process, the domain discrimination network To compete with each other, specifically, to confuse the generative network and feature restoration network Perform positive update, domain discrimination network Perform bidirectional updates;

[0147] Forward update means that when After a round of training of the adversarial generator with intuitive feature input, the feature restoration loss is obtained. , in the next round of training, the confusion generation network and feature restoration network The hyperparameters in and Will restore the loss along the feature Update in the direction of decreasing gradient;

[0148] Bidirectional updating means that when After one round of training of the adversarial generator with intuitive feature input, the domain discrimination loss is obtained. , in the next round of training, the domain discrimination network Hyperparameters Towards field judgment loss The gradient direction is updated in the decreasing direction and the balancing parameter is multiplied by The opposite number of to inform the confusion generation network and feature restoration network The hyperparameters in and Losses are determined by field Update the increasing gradient direction;

[0149] Driven by adversarial training strategies, confusing generative networks and feature restoration network Domain Discrimination Network Conduct adversarial and confuse the generated network Discriminating the network for interference domains Distinguishing between primary and secondary languages ​​will improve the ability of the superordinate to generate unified features. At the same time, the feature restoration network It will improve the ability to reconstruct intuitive features to ensure that unified features are not too high. In contrast, the domain discrimination network In order to distinguish whether the unified features come from the primary language or the secondary language, it will also improve the discrimination ability, and finally the adversarial generator will reach a Nash equilibrium. At this time, the loss function Minimum.

[0150] Furthermore, fine-tuning the large model encoder using contrastive learning includes the following specific steps:

[0151] Select the GraphCodeBERT pre-trained model as the large model encoder. Since the main language code segment or the secondary language code segment can be regarded as the same language code segment after being processed by the adversarial generator, fine-tuning the large model encoder only requires the main language code segment with the function label.

[0152] From the main repository of Code Sets Two main language code segments are randomly selected from each of the Code Sets For example, the two randomly selected main language code segments are recorded as and , and extract the corresponding functional labels , a total of Main language code segment and the corresponding function labels ;

[0153] Main language code segment Input adversarial generator to obtain unified features , further unifying the features Input large model encoder to generate code representation , the code representation dimensions are ;

[0154] Defining the loss function , the specific formula is as follows:

[0155] ,

[0156] in, is the temperature hyperparameter, is the i-th code representation and Code representation The cosine similarity loss function Describes the correlation between main language code segments with the same function label and the difference between main language code segments with different function labels;

[0157] Multiple rounds of optimization through comparative learning Code representation To continuously reduce the loss function After each round of comparative learning, the correlation between the main language code segments with the same function labels and the difference between the main language code segments with different function labels increase at the same time. When it is minimum, the cosine similarity between the code representations corresponding to the main language code segments with the same function labels will be maximized, and the cosine similarity between the code representations corresponding to the main language code segments with different function labels will be minimized. The difficulty of directly distinguishing different function labels based on code representations is minimized, the distinguishability of code representations is maximized, and the hyperparameters of the large model encoder are fixed.

[0158] Furthermore, the pre-training Softmax regression classifier includes the following specific steps:

[0159] Set up a Softmax regression model, and the code representation dimension is known to be , the function tags have Therefore, the input layer and output layer of the Softmax regression model are set and neurons, the hidden layer is a fully connected layer, and a Softmax activation function is set after the output layer for regression classification. Code representation Input Softmax regression model, and then reduce the dimension to , mapped to by the Softmax activation function In the interval, the output of the Softmax regression model is the predicted probability vector , the predicted probability vector The dimension is , among which, Elements Indicates Code representation Assign function labels The predicted probability of

[0160] Setting the multi-classification loss function To represent the predicted probability vector With the actual probability vector The similarity of the distribution is as follows:

[0161] ,

[0162] in, represents the actual probability vector The transpose of Code representation Function labels should be given , then the actual probability vector Except Elements All other values ​​except 1 are 0, so the loss function It can be further simplified to ;

[0163] Characterize the code Multiple rounds of input Softmax regression classifiers are used to update the hyperparameters of the fully connected layer using the gradient descent method so that the multi-classification loss function Continuously decreasing, multi-classification loss function The closer it is to 0, the Code representation Assign function labels The closer the predicted probability is to 1, the stronger the classification ability of the Softmax regression classifier is. When minimum, the hyperparameters of the Softmax regression classifier are fixed.

[0164] Furthermore, the code to be located is a code source file provided by the user, which can be written in a primary language or a secondary language. The code source file includes several code segments. The splitting of the code to be located can be performed by artificial intelligence or static code analysis tools. In this embodiment, Understand is used for analysis. After the code to be located is imported, the programming language can be autonomously identified, and the hierarchical structure of the code to be located can be displayed through a visual interface. At this time, the code to be located can be split into several code segments according to the hierarchical structure through automated scripts or manual operations.

[0165] like Figure 5 As shown, further, code positioning includes the following specific steps:

[0166] Based on the dual-channel query strategy, the word to be located is identified as a specified function or a specified typical process and the main code base is Select a positive example, where a positive example refers to a main language code segment whose function label matches a specified function or a specified typical process;

[0167] Input the positive examples and several code segments obtained by splitting the code to be located into the adversarial generator to obtain the unified features of all single code segments;

[0168] The unified features of all single code segments are further input into the large model encoder to obtain the code representations of all single code segments;

[0169] The code representations of all single code segments are input into the Softmax regression classifier in turn, and the prediction probability vector corresponding to the single code segment is output. The function label with the highest probability is selected from the prediction probability vector and assigned to the single code segment. All single code segments that are consistent with the function labels of the positive examples are the code segments of the required specified functions or specified typical processes.

[0170] like Figure 6 As shown, further, selecting positive examples based on the dual-channel query strategy includes the following specific steps:

[0171] Identify the words to be located and judge them as specified functions or typical processes;

[0172] If it is a specified function, query the function library , get the keyword table corresponding to the specified function;

[0173] Further compare with the total keyword table to obtain the function code corresponding to the specified function and convert it into a function label;

[0174] Query the main code base , randomly select a main language code segment with the same function label as a positive example;

[0175] If it is a specified typical process, obtain the keywords corresponding to the typical process;

[0176] Further determine the position of the keyword in the total keyword table, generate a benchmark function code, the benchmark function code has the same dimension as the function code, where the position corresponding to the keyword is 1, and the rest are 0. The benchmark function code is regarded as a binary code and converted into a benchmark function label;

[0177] Get all function labels that are greater than or equal to the benchmark function label, restore them to the corresponding function code and compare them with the benchmark function code. If the function code and the benchmark function code have 1 in the same position, put the function label corresponding to the function code into the label set to be located. Assuming that the benchmark function code is 0100, the corresponding benchmark function label is 4. The function code containing the specified typical process must have 1 in the same position as the benchmark function code, but the remaining positions can also be 1. For example, if the function code is 0111, the corresponding function label is 7, which must be greater than or equal to the benchmark function label.

[0178] Select a feature tag from the set of tags to be located and A main language code segment with the same function label is randomly selected as a positive example for code location, and the label is removed after the code location is completed. A new function label is selected from the set of labels to be located, and this step is repeated until all function labels in the set of labels to be located are selected.

[0179] The invention discloses a code location method based on programming language migration, which comprises the following steps: refining functions into typical processes to construct a function library and a total keyword table, generating function labels through quasi-transformation, designing a subject word self-checking crawler system, obtaining a large number of accurate primary and secondary language code segments based on a subject relevance discrimination method and a hash matching method to construct a primary and secondary code library, pre-training an adversarial generator based on an adversarial training strategy to achieve primary and secondary language feature alignment, fine-tuning a large model encoder using contrastive learning in combination with the main code library to maximize the separability of code representation, assisting in pre-training a Softmax regression classifier to achieve function label assignment of code segments, obtaining words to be located and codes to be located, splitting the codes to be located, identifying words to be located based on a dual-channel query strategy to select positive examples, and inputting the positive examples and code segments into a pre-trained and fine-tuned network to achieve code location of a specified function or a specified typical process.

[0180] The above is only a preferred embodiment of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions under the concept of the present invention belong to the protection scope of the present invention. It should be pointed out that for ordinary technicians in this technical field, some improvements and modifications without departing from the principle of the present invention should also be regarded as the protection scope of the present invention.

Claims

1. A code location method based on programming language migration, characterized in that: The specific steps include: Build a function library and a total keyword table based on all typical processes corresponding to a single function; Based on the subject word self-checking crawler system, the main language code segments are obtained and functional tags are added to build the main code base; The sample-enhanced auxiliary subject heading self-checking crawler system is used to obtain the paralanguage code segments to build the paracode library; Based on the adversarial training strategy, an adversarial generator is constructed and pre-trained. Contrastive learning is used to fine-tune the large model encoder to maximize the distinguishability of code representation, and the code representation is further input into the Softmax regression classifier for pre-training. Obtain the words to be located and split the codes to be located, and select positive examples based on the dual-channel query strategy for code location; The construction and pre-training of the adversarial generator based on the adversarial training strategy includes the following specific steps: Randomly extracting a primary language code segment and a secondary language code segment, and using a BERT network to generate intuitive features and input them into the adversarial generator; The confusion generation network extracts the potential features in the intuitive features through multiple convolutions, and generates unified features through nonlinear mapping; The feature restoration network expands the dimension of the unified feature through deconvolution, and obtains and reconstructs intuitive features through nonlinear transformation; The domain discrimination network maps the unified features to a separable space through nonlinear transformation and determines the source of the unified features; Introduce a balance parameter to comprehensively consider feature restoration loss and domain discrimination loss, define the loss function, and update the loss function to the minimum based on the adversarial training strategy; The method of fine-tuning the large model encoder using contrastive learning includes the following specific steps: Selecting the GraphCodeBERT pre-trained model as the large model encoder; Randomly extract two main language code segments from each code set in the main code base and extract the corresponding functional labels; Inputting the main language code segment into the adversarial generator and the large model encoder in sequence to generate a code representation; Define the loss function and optimize the code representation through contrastive learning to reduce the loss function. When the loss function is minimized, the distinguishability of the code representation is maximized, and the hyperparameters of the large model encoder are fixed.

2. A code location method based on programming language migration as claimed in claim 1, characterized in that: The adversarial training strategy includes forward updating and bidirectional updating; The forward update refers to updating the hyperparameters in the confusion generation network and the feature restoration network along the gradient direction in which the feature restoration loss decreases; The bidirectional update refers to updating the hyperparameters of the domain discrimination network along the gradient direction in which the domain discrimination loss decreases, and introducing a balance parameter so that the hyperparameters in the confusion generation network and the feature restoration network are updated along the gradient direction in which the domain discrimination loss increases.

3. A code location method based on programming language migration as claimed in claim 1, characterized in that: The dual-channel query strategy for selecting positive examples includes the following specific steps: Identify the words to be located and judge them as specified functions or typical processes; If it is a specified function, obtain the function label corresponding to the specified function, query the main code library, and randomly select the main language code segment with the same function label as a positive example; If it is a specified typical process, obtain the keywords corresponding to the typical process and generate a benchmark function label; All function labels that are greater than or equal to the benchmark function label are obtained, restored to function coding and determined whether they contain the specified typical process, and based on the function labels containing the specified typical process, the main language code segments with the same function labels are randomly selected from the main code base as positive examples.

4. A code location method based on programming language migration as claimed in claim 1, characterized in that: The construction of the function library and the total keyword table includes the following specific steps: Get all typical processes corresponding to a single function; Extracting keywords corresponding to the typical process and constructing a keyword table corresponding to the single function; Establishing a function library including all the single functions and the keyword table; All non-repeating keywords in the function library are obtained, and the non-repeating keywords are rearranged to construct a total keyword table.

5. The method for locating a code based on programming language migration as claimed in claim 1, characterized in that: The construction of the main code base includes the following specific steps: Determine the main language, traverse the single functions in the function library in turn, and deliver the single function and the main language as keywords to the keyword self-check crawler system to obtain the main language code segment; Obtaining a keyword table corresponding to a single function in the function library, and comparing the keyword table with the total keyword table to generate a function code corresponding to the single function; Convert the function code corresponding to the single function into a function label through a reversible transformation, and assign it to all main language code segments corresponding to the single function; Store all main language code segments and function labels corresponding to all single functions to build the main code base.

6. A method for locating code based on programming language migration as claimed in claim 1, characterized in that: The construction of the secondary code base includes the following specific steps: Determine the paralanguage, and deliver the paralanguage as a subject word to the subject word self-check crawler system until the search is stopped; Determine whether the number of the secondary language code segments reaches the total secondary threshold, and if so, directly build a secondary code library; If it is not reached, the para-language code segments are generated through sample enhancement until the total para-threshold is reached, and the para-code library is constructed.

7. A code location method based on programming language migration as described in any one of claims 1, 5, and 6, characterized in that: The subject word self-check crawler system includes a scheduling module, a downloading module, an extraction module and a data pipeline module; The scheduling module receives and determines whether the subject words have been crawled, generates crawler requests based on the uncrawled subject words, and decides whether to send the crawler requests to the download module or insert them into the Redis task queue based on whether the crawling task is being executed; The download module uses random dynamic User-Agent and random delay to disguise crawler requests, access the website and download crawler responses; The extraction module introduces a web crawling algorithm to extract page elements and calculate scores based on text density. It normalizes the page elements with the highest scores to generate the body text and further extracts the title, content and URL in the body text. The data pipeline module uses the topic relevance discrimination method to filter out relevant text from the text, and removes duplicate related text through hash matching deduplication method, and uses OCR technology to recognize characters in the content to automatically generate code snippets.

8. A method for locating code based on programming language migration as claimed in claim 7, characterized in that: The hash matching deduplication method for removing duplicate related texts includes the following specific steps: Parse the URL, concatenate the protocol and domain name to generate the core segment; Set a bit array and compress the core segment into a fingerprint sequence using the SHA1 algorithm; Input the fingerprint sequence into a Bloom filter to generate a number of hash values, and respectively perform modulo operation on the bit array length to generate index keys; The URL is discarded or stored based on whether the index key exists in Redis.

Citation Information

Patent Citations

  • Key Code Location Methods and Systems

    CN109240700B

  • Word representation learning method for multi-language large model

    CN116956889A

  • Systems and methods for natural language code search

    WO2023060034A1