Class imbalance processing method and system based on code entity attribute similarity
By extracting and calculating the multi-dimensional characteristic similarity of code entities, the synthesis of new positive samples is solved, and the defects of the simplified characteristics of the SMOTE method in the existing technology are solved, and the model's recognition ability and overall performance of a few class samples are significantly improved.
Patent Information
- Application Number
- CN202510052143.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-14
- Publication Date
- 2025-05-27
AI Technical Summary
The existing technology has class imbalance problems when dealing with code odor, which leads to classifiers tend to preferentially identify most class samples and ignore a few class samples. The SMOTE method simplifies code instances into numerical vectors, ignoring the multi-dimensional characteristics of the code, resulting in the synthetic samples being unable to accurately reflect the characteristics of a few class.
By extracting the dependency sets, historical change sets and code text of the code entity, the similarity between each positive sample is calculated, and the similar samples are weighted and fused, and the similar samples are determined, and a new positive sample is synthesized to generate new positive samples to balance the data distribution.
It significantly improves the ability to identify a few types of samples, can accurately reflect the characteristics of a few types, clearly distinguish the boundaries of positive and negative samples, and effectively improve model performance.
Smart Images

Figure CN120045946A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of software engineering, and in particular to a class imbalance processing method and system based on code entity attribute similarity. Background Art
[0002] In the field of software engineering, code smell refers to potential structural problems in the code. Although these problems do not directly cause software function failure, they will increase the difficulty of code maintenance, understanding and expansion. Long-term accumulation will form technical debt, affecting the maintainability, scalability and overall code quality of the system. Among them, Feature Envy is a common code smell, which means that a method is overly dependent on data or functions of other classes instead of using resources in its own class. This violates the basic principles of object-oriented design, such as encapsulation and single responsibility principle, and may lead to excessive coupling between classes, reduce the modularity of the system, and increase maintenance costs and error risks.
[0003] In the prior art, code smell detection methods are mainly divided into two categories: tool-based and machine learning-based. Tool-based methods use heuristic algorithms and preset thresholds to determine whether the code has a smell, but the effectiveness of such methods depends on the accuracy of the threshold and often has limitations when dealing with complex code smells. Machine learning-based methods can automatically learn data patterns through models to improve detection accuracy.
[0004] However, since the number of code instances with Feature Envy in actual projects is far less than that of code instances without odor, this serious class imbalance problem causes the classifier to prefer identifying majority class samples (code without odor) while ignoring minority class samples (code with odor), thus affecting the recognition ability of minority class samples.
[0005] In order to solve the class imbalance problem, researchers have proposed a variety of imbalanced learning techniques, such as cost-sensitive classifiers, single-class classifiers, and data balancing techniques, including random oversampling, random undersampling, and synthetic minority oversampling technology (SMOTE). Among them, SMOTE is one of the more commonly used methods. It selects similar instances through the KNN algorithm to generate new minority class sample points, thereby balancing the data distribution and improving the model's detection ability for minority class samples. The limitation of the SMOTE method is that it simplifies code instances into numerical vectors, ignoring the multidimensional characteristics of the code as a complex entity (such as dependencies, historical changes, and code text structure), resulting in the inability of synthetic samples to accurately reflect minority class characteristics, and may even blur the boundaries between positive and negative samples, thereby affecting model performance. Summary of the invention
[0006] In order to solve the technical problems that the existing class imbalance problem causes the classifier to prefer identifying majority class samples while ignoring minority class samples, thereby affecting the recognition ability of minority class samples, and the existing SMOTE method simplifies code instances into numerical vectors, ignoring the multidimensional characteristics of codes as complex entities, resulting in synthetic samples being unable to accurately reflect minority class characteristics and may even blur the boundaries between positive and negative samples, thereby affecting model performance, the present invention provides a class imbalance processing method and system based on code entity attribute similarity.
[0007] The technical solution provided by the embodiment of the present invention is as follows:
[0008] First aspect:
[0009] An embodiment of the present invention provides a class imbalance processing method based on code entity attribute similarity, comprising:
[0010] S1: Get sample instance;
[0011] S2: extracting a plurality of positive samples and a plurality of negative samples from the sample instance, wherein the number of the positive samples is less than the number of the negative samples;
[0012] S3: extracting the dependency set, historical change set and code text of each positive sample;
[0013] S4: Calculating the dependency similarity, historical change similarity and code text similarity between each of the positive samples according to the dependency set, the historical change set and the code text;
[0014] S5: performing weighted fusion on the dependency similarity, the historical change similarity and the code text similarity to determine the comprehensive similarity between each of the positive samples;
[0015] S6: Determine similar samples of each of the positive samples according to the comprehensive similarity;
[0016] S7: synthesize each of the positive samples with the corresponding similar samples to generate a positive sample.
[0017] Second aspect:
[0018] An embodiment of the present invention provides a class imbalance processing system based on code entity attribute similarity, comprising: a memory and one or more processors;
[0019] One or more application programs are stored in the memory, and the one or more application programs are suitable for being executed by the one or more processors to implement the above-mentioned class imbalance processing method based on code entity attribute similarity.
[0020] The beneficial effects brought about by the technical solution provided by the embodiment of the present invention include at least:
[0021] In an embodiment of the present invention, by extracting the dependency set, historical change set and code text of each positive sample, the classifier can effectively focus on minority class samples while identifying majority class samples, thereby significantly improving the recognition ability of minority class samples. By calculating the dependency similarity, historical change similarity and code text similarity between each positive sample, the code instance is not simply simplified into a numerical vector, and the multidimensional characteristics of the code as a complex entity can be fully considered. By synthesizing each positive sample with the corresponding similar sample to generate a positive sample, the minority class characteristics can be accurately reflected, the boundaries between positive and negative samples can be clearly distinguished, and the model performance can be effectively improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0023] Figure 1 A flow chart of a method for processing class imbalance based on code entity attribute similarity provided by an embodiment of the present invention;
[0024] Figure 2 A structural diagram of a class imbalance processing system based on code entity attribute similarity provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0025] The technical solution of the present invention is described below in conjunction with the accompanying drawings.
[0026] In the embodiments of the present invention, words such as "exemplarily" and "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design described as "example" in the present invention should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of the word "example" is intended to present the concept in a specific way. In addition, in the embodiments of the present invention, the meaning expressed by "and / or" can be both, or it can be either of the two.
[0027] In order to make the technical problems, technical solutions and advantages to be solved by the present invention more clear, a detailed description will be given below with reference to the accompanying drawings and specific embodiments.
[0028] Reference Manual Attached Figure 1 , shows a flow chart of a class imbalance processing method based on code entity attribute similarity provided by an embodiment of the present invention.
[0029] The embodiment of the present invention provides a class imbalance processing method based on code entity attribute similarity, which can be implemented by a class imbalance processing device based on code entity attribute similarity, and the class imbalance processing device based on code entity attribute similarity can be a terminal or a server. The processing flow of the class imbalance processing method based on code entity attribute similarity can include the following steps:
[0030] S1: Get a sample instance.
[0031] In the present invention, by acquiring sample instances, positive samples and negative samples can be distinguished, and the focus is on feature extraction and optimization of positive samples (minority classes) to facilitate subsequent processing.
[0032] S2: Extract multiple positive samples and multiple negative samples from the sample instance, where the number of positive samples is less than the number of negative samples.
[0033] Specifically, multiple positive samples and multiple negative samples are extracted from the sample instances, where positive samples refer to code instances with Feature Envy smell, and negative samples refer to code instances without the smell. Since the number of positive samples is usually much smaller than that of negative samples, there is an obvious class imbalance problem in the dataset, and subsequent sample balancing processing techniques are needed to optimize the data distribution in order to improve the model's ability to detect positive samples.
[0034] Feature Envy is a common code smell in software development, which means that a method is overly dependent on data or functions of other classes instead of making full use of its own class resources. This phenomenon violates the encapsulation and single responsibility principles in object-oriented design, and may lead to excessive coupling between classes, reducing the modularity and maintainability of the code. Feature Envy will increase the complexity and maintenance cost of the system, and is one of the issues that need to be addressed first during code refactoring.
[0035] In this invention, by distinguishing between positive and negative samples, we can accurately locate methods (positive samples) with Feature Envy odors, focus our research on solving these potential problems, and provide clear goals for subsequent optimization algorithms. The number of positive samples is usually small, while the number of negative samples is large, which directly exposes the class imbalance problem of the data set. This phenomenon provides a basis for subsequent sample balancing technology, ensuring that the model can focus on improving the recognition ability of positive samples rather than biased towards negative samples.
[0036] S3: Extract the dependency set, historical change set, and code text of each positive sample.
[0037] Specifically, the static code analysis tool Scitools Understand is used to extract the dependency set of each positive sample, including the positive sample method calls and calls, member access and references, and object instantiation. Based on the Git repository, the commit records related to the positive sample method are retrieved to extract the historical change set of each positive sample.
[0038] Among them, Scitools Understand is a professional static code analysis tool, which is widely used in software development and maintenance. It supports multiple programming languages (such as C / C++, Java, Python, etc.), and can deeply analyze the structure, dependency, complexity and quality of the code.
[0039] Among them, the Git repository is the core component for storing and managing code versions, and is the basic unit when developers use the Git version control system. It records all files, code history, change records, and branch structures of the project, and supports multi-user collaboration and code version management. Git repositories are divided into two types: local repositories and remote repositories: local repositories are stored on the developer's computer, and remote repositories are usually hosted on a cloud platform.
[0040] Here, "positive sample methods" refer to code methods (i.e. functions) that are marked as having Feature Envy smells during code analysis.
[0041] It should be noted that code text refers to the source code content of the code itself, including the combination of characters and statements that specifically implement the logic in the program.
[0042] In the present invention, the dependency set, historical change set and code text fully describe the characteristics of the positive sample method from three different dimensions: structure, historical evolution and semantics. This multi-dimensional characteristic analysis provides a rich information basis for subsequent similarity calculation and sample generation, which helps to improve the accuracy and diversity of sample synthesis. Comprehensively extracting the characteristics of positive samples can ensure that the generated new samples are consistent with the positive samples in structure, history and semantics, while avoiding repeated generation of similar samples and improving the diversity and representativeness of generated samples.
[0043] S4: Based on the dependency set, the historical change set and the code text, the dependency similarity, the historical change similarity and the code text similarity between each positive sample are calculated.
[0044] It should be noted that dependency similarity is used to measure the similarity of dependency sets between positive sample methods. When the dependency sets of two positive samples are highly similar, it means that they have a strong correlation in dependency attributes. Historical change similarity is used to evaluate the similarity of change trajectories of positive sample methods in the development history, reflecting their consistency in version evolution. Code text similarity quantifies the semantic similarity of code text through a deep learning model, accurately capturing the semantic correlation features between positive sample methods.
[0045] In the present invention, the similarity between positive samples is comprehensively measured through three dimensions: dependency, historical changes, and code text. Not only can associations be found from code structure and historical changes, but also deeper features can be captured from the code semantic level, ensuring that the similarity analysis between positive samples is more comprehensive and more accurate. Through accurate similarity calculation, the instance closest to the characteristics of each positive sample is found, which helps to generate minority class samples with real characteristics. This method can ensure the consistency of new samples in dependencies, historical trajectories, and semantic characteristics, and improve the quality and representativeness of generated data. By matching and generating high-quality similar samples, the minority class sample data set is expanded, which helps to alleviate the problem of class imbalance and significantly improve the model's ability to recognize positive samples (Feature Envy odor).
[0046] In a possible implementation, S4 specifically includes sub-steps S401 to S403:
[0047] S401: Calculate the dependency similarity between each positive sample according to the dependency set of each positive sample.
[0048] Specifically, according to the dependency set, the Jaccard similarity coefficient is used to calculate the dependency similarity between each positive sample.
[0049] Among them, the Jaccard similarity coefficient is a commonly used similarity measurement method used to compare the similarity between two sets. It expresses similarity by calculating the ratio of the size of the intersection of two sets to the size of the union. The value range of the Jaccard similarity coefficient is between 0 and 1. The closer the value is to 1, the more similar the two sets are. The value of 0 means that the two sets do not intersect at all.
[0050] In a possible implementation manner, S401 specifically includes:
[0051] According to the following formula, the dependency similarity between each positive sample is calculated:
[0052]
[0053] Among them, s 1Indicates the similarity of the dependency relationship between each positive sample, J(Dep(m i ),Dep(m k )) represents the Jaccard similarity coefficient between the dependency set of the i-th positive sample and the dependency set of the k-th positive sample, Dep(m i ) represents the dependency set of the i-th positive sample, Dep(m k ) represents the dependency set of the kth positive sample, ∩ represents the intersection operation, and ∪ represents the union operation.
[0054] In the present invention, dependency similarity measures the degree of association between positive sample methods in dependency structure, can quantify the similarity of structural characteristics between methods, and provide data support for subsequent similar sample matching and sample generation. By adopting the Jaccard similarity coefficient, the dependency sets of different positive samples can be accurately compared to reflect the degree of their shared dependencies, thereby revealing the potential correlation in the code structure. Through dependency similarity, highly similar instances between positive samples can be more accurately identified, and more realistic minority class samples can be generated in combination with other dimensional characteristics, thereby optimizing the balance of the data set and model performance.
[0055] S402: Calculate the historical change similarity between each positive sample according to the historical change set of each positive sample.
[0056] Specifically, according to the historical change set, the Jaccard similarity coefficient is used to calculate the historical change similarity between each positive sample.
[0057] In a possible implementation manner, S402 specifically includes:
[0058] According to the following formula, the historical change similarity between each positive sample is calculated:
[0059]
[0060] Among them, s 2 Indicates the historical change similarity between each positive sample, J(Hist(m i ),Hist(m k )) represents the Jaccard similarity coefficient between the historical change set of the i-th positive sample and the historical change set of the k-th positive sample, Hist(m i ) represents the historical change set of the i-th positive sample, Hist(m k ) represents the historical change set of the kth positive sample.
[0061] In the present invention, the historical change similarity reflects the similarity of the change trajectory of the positive sample in the development history. By analyzing the historical change sets of different positive samples, their common evolution laws in version control can be revealed, and the key characteristics of the historical evolution of the code can be captured. Using the Jaccard similarity coefficient to calculate the historical change set can accurately measure whether the historical change trajectory of the positive sample method is similar, providing a more reliable basis for subsequent similar sample matching and generation. As an independent dimension, historical change similarity, together with dependency similarity and code text similarity, constitutes a comprehensive feature analysis framework. It can supplement the deficiencies of structural and semantic information and further improve the accuracy of similar sample matching.
[0062] S403: Calculate the code text similarity between the positive samples according to the code text of each positive sample.
[0063] According to the code text, the deep learning model CodeBERT is used to extract the semantic embedding vector of the code text, and the code text similarity between each positive sample is calculated through the normalized Euclidean distance.
[0064] Among them, CodeBERT is a deep learning pre-training model designed for natural language and code analysis tasks. Based on the Transformer architecture, CodeBERT learns the semantic information of the code by training on large-scale code and annotation data, and supports multiple programming languages (such as Python, Java, etc.). It is widely used in code search, completion, and semantic analysis. By converting code snippets into semantic vectors, it improves code understanding and processing capabilities.
[0065] Among them, Euclidean distance is a commonly used mathematical method to measure the straight-line distance between two points, and is widely used in the fields of geometry and data analysis. The smaller the Euclidean distance, the closer the two points are. The larger the distance, the farther the two points are. Euclidean distance is widely used in cluster analysis, machine learning, image processing and other fields to measure the similarity or difference between samples.
[0066] In the present invention, by analyzing the semantic similarity of code snippets, the instances closest to the semantic features of the positive samples can be better identified, ensuring that the generated minority class samples are semantically consistent with the original samples, thereby improving the representativeness and credibility of the synthesized samples.
[0067] In a possible implementation, S403 specifically includes sub-steps S4031 to S4033:
[0068] S4031: Convert the code text of each positive sample into multiple embedding vectors through the CodeBERT model.
[0069] S4032: Calculate the similarity index between each embedding vector:
[0070]
[0071] Among them, d(X,Y) represents the similarity index between the embedding vector X and the embedding vector Y, and X n Represents the component value of the embedding vector X in the nth dimension, Y n Represents the component value of the embedding vector Y in the nth dimension, n = 1, 2, …, N, where N represents the total number of dimensions.
[0072] S4033: Normalize the similarity index to obtain the code text similarity between each positive sample:
[0073]
[0074] Among them, s 3 Indicates the code text similarity between each positive sample.
[0075] In the present invention, the deep learning model CodeBERT is used to extract the semantic embedding vector of the code text, which can accurately capture the semantic information of the positive sample code fragment, make up for the shortcomings of the traditional lexical or grammatical analysis method, and provide higher quality semantic features for subsequent sample analysis. By calculating the code text similarity by normalized Euclidean distance, the degree of semantic association between positive sample code fragments can be intuitively quantified, which helps to more accurately identify semantically similar samples. By analyzing the semantic similarity of code text, developers can find potential problems or optimization opportunities in the code, thereby improving the quality and maintainability of system code.
[0076] S5: Perform weighted fusion on dependency similarity, historical change similarity, and code text similarity to determine the comprehensive similarity between each positive sample.
[0077] In a possible implementation, S5 is specifically:
[0078] According to the following formula, the dependency similarity, historical change similarity, and code text similarity are weighted and fused to determine the comprehensive similarity between each positive sample:
[0079]
[0080] Among them, S(m i ,m k ) represents the comprehensive similarity between the i-th positive sample and the k-th positive sample, j = 1, 2, 3, α 1 Represents the weight coefficient of dependency similarity, α 2 Represents the weight coefficient of the similarity of historical changes, α 3 The weight coefficient representing the code text similarity.
[0081] In the present invention, by combining the similarities of three dimensions, namely, dependency, historical changes, and code text, the multi-dimensional characteristic correlation between positive samples can be fully reflected, avoiding the analysis bias caused by single-dimensional characteristics. The comprehensive similarity combines the structural (dependency), historical evolution (historical changes), and semantic (code text) characteristics, making the similarity calculation between positive samples more accurate and helping to find the most relevant similar samples. Through more accurate comprehensive similarity calculation, the sample generation and data balancing process can be optimized, thereby providing more accurate training data for the model and improving the classifier's ability to detect positive samples.
[0082] In a possible implementation manner, after S5 and before S6, the method further includes:
[0083] S8: Through the differential evolution algorithm, the weight coefficients of dependency similarity, historical change similarity and code text similarity are optimized respectively.
[0084] It should be noted that the Differential Evolution (DE) algorithm is a simple and efficient global optimization algorithm suitable for solving complex multivariable optimization problems. It gradually optimizes the objective function by simulating the evolutionary process in nature, including mutation, crossover, and selection. The algorithm starts with a randomly initialized population, and each individual represents a possible solution. New candidate solutions are generated through differential mutation operations and compared with the current solution, and individuals with better fitness are retained to enter the next generation. The differential evolution algorithm has the characteristics of global search capability, simple implementation, and few parameter settings. It is widely used in fields such as function optimization, machine learning, and engineering design.
[0085] Specifically, the population size of the differential evolution algorithm is set to 30, the scale factor is 0.3, the crossover rate is 0.9, and the maximum generation is 20. The F1-score maximization of the classification task is taken as the optimization goal. Finally, the optimal weight coefficient combination of the three similarities is obtained, which is used to comprehensively calculate the overall similarity between positive samples.
[0086] In the present invention, the differential evolution algorithm has a powerful global search capability, which can effectively avoid falling into the local optimum and ensure that the weight coefficients of the three similarities are optimally allocated, thereby improving the accuracy of the comprehensive similarity. By optimizing the weight coefficient, the contribution ratio of each similarity in the comprehensive calculation can be dynamically adjusted according to different data distributions and task objectives (such as classification performance), so that the results are more in line with the actual situation. Taking F1-score maximization as the optimization goal, the model performance is directly optimized to ensure that the generated comprehensive similarity can effectively support the subsequent classification model training and enhance the recognition ability of minority class positive samples.
[0087] S6: Determine similar samples of each positive sample based on the comprehensive similarity.
[0088] In a possible implementation manner, S6 specifically includes:
[0089] Determine similar samples for each positive sample according to the following formula:
[0090]
[0091] Among them, m similarity Represents the i-th positive sample m i similar samples, argmax means taking the maximum value, and M represents the positive sample set.
[0092] In the present invention, by calculating the comprehensive similarity, the dependencies between positive samples, historical changes, and the similarity of the semantics of the code text can be comprehensively measured, so as to more accurately find the samples that are most similar to the characteristics of each positive sample and avoid the errors caused by a single dimension. By matching the most similar samples for each positive sample, a reliable reference pair can be provided for subsequent sample generation, ensuring that the generated new samples can fully retain the characteristics of the positive samples and improve the representativeness and diversity of the generated samples. By selecting similar samples, minority class samples that are close to the real ones can be better generated in terms of structure, history, and semantic characteristics, further optimizing data distribution and solving the problem of insufficient minority class samples.
[0093] S7: synthesize each positive sample with the corresponding similar sample to generate a positive sample.
[0094] In a possible implementation manner, S7 specifically includes:
[0095] Generate positive samples according to the following formula:
[0096] m new =m i +λ·(m similarity -m i )
[0097]
[0098] Among them, m new represents the generated positive sample, λ represents the synthesis ratio parameter, and r represents a random number in the range of (0,1) and not equal to 0.5.
[0099] Specifically, according to the difference in feature vectors between the positive sample and the most similar sample, the generation ratio of new samples is controlled by dynamically adjusting the parameter λ to ensure the diversity and representativeness of the generated samples. The generated positive samples are then added to the training data set to expand the number of minority class samples (positive samples), thereby effectively alleviating the class imbalance problem.
[0100] In the present invention, by generating new positive samples (minority class), the number of positive samples is effectively expanded, and the distribution of positive and negative samples in the data set is balanced, thereby reducing the bias of the classification model to the majority class samples and improving the model's recognition ability for the minority class. The dynamically adjusted parameter λ controls the synthesis ratio so that the generated new samples contain both the characteristics of the original positive samples and the characteristics of similar samples, thereby generating more diverse and representative minority class samples and avoiding the generation of overly similar duplicate samples. The generated new positive samples expand the characteristic distribution of the training data and enrich the feature diversity of the minority class samples, so that the model can learn the characteristics of the positive samples more comprehensively and improve the classification performance, especially the detection ability of the minority class in actual scenarios.
[0101] The beneficial effects brought about by the technical solution provided by the embodiment of the present invention include at least:
[0102] In an embodiment of the present invention, by extracting the dependency set, historical change set and code text of each positive sample, the classifier can effectively focus on minority class samples while identifying majority class samples, thereby significantly improving the recognition ability of minority class samples. By calculating the dependency similarity, historical change similarity and code text similarity between each positive sample, the code instance is not simply simplified into a numerical vector, and the multidimensional characteristics of the code as a complex entity can be fully considered. By synthesizing each positive sample with the corresponding similar sample to generate a positive sample, the minority class characteristics can be accurately reflected, the boundaries between positive and negative samples can be clearly distinguished, and the model performance can be effectively improved.
[0103] Reference Manual Attached Figure 2 , showing a structural schematic diagram of a class imbalance processing system based on code entity attribute similarity provided by the present invention.
[0104] The present invention further provides a class imbalance processing system 30 based on code entity attribute similarity, comprising: a memory 303 and one or more processors 301 .
[0105] One or more application programs are stored in the memory 303 , and the one or more application programs are suitable for being executed by the one or more processors 301 to implement the class imbalance processing method based on code entity attribute similarity described in the method embodiment.
[0106] The class imbalance processing system 30 based on code entity attribute similarity includes: a processor 301 and a memory 303. The processor 301 and the memory 303 are connected, for example, through a bus 302.
[0107] The structure of the class imbalance processing system 30 based on code entity attribute similarity does not constitute a limitation to the embodiment of the present invention.
[0108] Processor 301 may be a CPU, a general purpose processor, a DSP, an ASIC, an FPGA or other programmable logic device, a transistor logic device, a hardware component or any combination thereof. It may implement or execute various exemplary logic blocks, modules and circuits described in conjunction with the disclosure of the present invention. Processor 301 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc.
[0109] The bus 302 may include a path to transmit information between the above components. The bus 302 may be a PCI bus or an EISA bus, etc. The bus 302 may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, only one thick line is used in the figure, but it does not mean that there is only one bus or one type of bus.
[0110] The memory 303 can be a ROM or other type of static storage device that can store static information and instructions, a RAM or other type of dynamic storage device that can store information and instructions, or an EEPROM, a CD-ROM or other optical disk storage, an optical disk storage (including a compressed optical disk, a laser disk, an optical disk, a digital versatile disk, a Blu-ray disk, etc.), a magnetic disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited to these.
[0111] It should be noted that the class imbalance processing system 30 based on code entity attribute similarity can implement the above-mentioned class imbalance processing method based on code entity attribute similarity, and can achieve the same or similar technical effects. To avoid repetition, the present invention will not go into details.
[0112] The beneficial effects brought about by the technical solution provided by the embodiment of the present invention include at least:
[0113] In an embodiment of the present invention, by extracting the dependency set, historical change set and code text of each positive sample, the classifier can effectively focus on minority class samples while identifying majority class samples, thereby significantly improving the recognition ability of minority class samples. By calculating the dependency similarity, historical change similarity and code text similarity between each positive sample, the code instance is not simply simplified into a numerical vector, and the multidimensional characteristics of the code as a complex entity can be fully considered. By synthesizing each positive sample with the corresponding similar sample to generate a positive sample, the minority class characteristics can be accurately reflected, the boundaries between positive and negative samples can be clearly distinguished, and the model performance can be effectively improved.
[0114] The present invention also provides a computer-readable storage medium having a computer program stored thereon, which can be loaded and executed by a processor to implement the class imbalance processing method based on code entity attribute similarity as described in the first aspect.
[0115] The above is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art who is familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed by the present invention, which should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention should be based on the protection scope of the claims.
[0116] There are a few points to note:
[0117] (1) The drawings of the embodiments of the present invention only relate to the structures related to the embodiments of the present invention, and other structures may refer to the general design.
[0118] (2) For the sake of clarity, in the drawings used to describe the embodiments of the present invention, the thickness of the layers or regions is exaggerated or reduced, that is, these drawings are not drawn according to the actual scale. It is understood that when an element such as a layer, film, region or substrate is referred to as being "on" or "under" another element, the element may be "directly" "on" or "under" the other element or there may be intermediate elements.
[0119] (3) In the absence of conflict, the embodiments of the present invention and the features therein may be combined with each other to obtain new embodiments.
[0120] The above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. The protection scope of the present invention shall be based on the protection scope of the claims.
Claims
1. A class imbalance processing method based on code entity attribute similarity, characterized in that: include: S1: Get sample instance; S2: extracting a plurality of positive samples and a plurality of negative samples from the sample instance, wherein the number of the positive samples is less than the number of the negative samples; S3: extracting the dependency set, historical change set and code text of each positive sample; S4: Calculating the dependency similarity, historical change similarity and code text similarity between each of the positive samples according to the dependency set, the historical change set and the code text; S5: performing weighted fusion on the dependency similarity, the historical change similarity and the code text similarity to determine the comprehensive similarity between each of the positive samples; S6: Determine similar samples of each of the positive samples according to the comprehensive similarity; S7: synthesize each of the positive samples with the corresponding similar samples to generate a positive sample.
2. The class imbalance processing method based on code entity attribute similarity according to claim 1 is characterized in that: The S4 specifically includes: S401: Calculating the dependency similarity between the positive samples according to the dependency set of the positive samples; S402: Calculating the historical change similarity between each of the positive samples according to the historical change set of each of the positive samples; S403: Calculate the code text similarity between the positive samples according to the code text of each positive sample.
3. The class imbalance processing method based on code entity attribute similarity according to claim 2 is characterized in that: The S401 is specifically as follows: According to the following formula, the dependency similarity between the positive samples is calculated: Among them, s1 represents the similarity of the dependency relationship between each positive sample, J(Dep(m i ),Dep(m k )) represents the Jaccard similarity coefficient between the dependency set of the i-th positive sample and the dependency set of the k-th positive sample, Dep(m i ) represents the dependency set of the i-th positive sample, Dep(m k ) represents the dependency set of the kth positive sample, ∩ represents the intersection operation, and ∪ represents the union operation.
4. The class imbalance processing method based on code entity attribute similarity according to claim 2 is characterized in that: The S402 is specifically as follows: The historical change similarity between the positive samples is calculated according to the following formula: Among them, s2 represents the historical change similarity between each positive sample, J(Hist(m i ),Hist(m k )) represents the Jaccard similarity coefficient between the historical change set of the i-th positive sample and the historical change set of the k-th positive sample, Hist(m i ) represents the historical change set of the i-th positive sample, Hist(m k ) represents the historical change set of the kth positive sample.
5. The class imbalance processing method based on code entity attribute similarity according to claim 2 is characterized in that: The S403 specifically includes: S4031: Convert the code text of each positive sample into multiple embedding vectors through the CodeBERT model; S4032: Calculate the similarity index between each embedding vector: Among them, d(X,Y) represents the similarity index between the embedding vector X and the embedding vector Y, and X n Represents the component value of the embedding vector X in the nth dimension, Y n represents the component value of the embedding vector Y in the nth dimension, n = 1, 2, ..., N, N represents the total number of dimensions; S4033: normalizing the similarity index to obtain the code text similarity between each of the positive samples: Among them, s3 represents the code text similarity between each positive sample.
6. The class imbalance processing method based on code entity attribute similarity according to claim 1, characterized in that: The S5 is specifically: According to the following formula, the dependency similarity, the historical change similarity and the code text similarity are weighted and fused to determine the comprehensive similarity between the positive samples: Among them, S(m i ,m k ) represents the comprehensive similarity between the i-th positive sample and the k-th positive sample, j = 1, 2, 3, α1 represents the weight coefficient of dependency similarity, α2 represents the weight coefficient of historical change similarity, and α3 represents the weight coefficient of code text similarity.
7. The class imbalance processing method based on code entity attribute similarity according to claim 6 is characterized in that: After S5 and before S6, the method further includes: S8: Optimizing the weight coefficients of the dependency similarity, the historical change similarity, and the code text similarity respectively through a differential evolution algorithm.
8. The class imbalance processing method based on code entity attribute similarity according to claim 1, characterized in that: The S6 is specifically: According to the following formula, similar samples of each positive sample are determined: Among them, m similarity Represents the i-th positive sample m i similar samples, argmax means taking the maximum value, and M represents the positive sample set.
9. The class imbalance processing method based on code entity attribute similarity according to claim 1, characterized in that: The S7 is specifically: Generate positive samples according to the following formula: m new =m i +λ·(m similarity -m i ) Among them, m new represents the generated positive sample, λ represents the synthesis ratio parameter, and r represents a random number in the range of (0,1) and not equal to 0.
5.
10. A class imbalance processing system based on code entity attribute similarity, characterized in that: include: memory and one or more processors; One or more applications are stored in the memory, and the one or more applications are suitable for being executed by the one or more processors to implement the class imbalance processing method based on code entity attribute similarity according to any one of claims 1 to 9.