A Dimensionality Reduction Method for Source Code Security Detection Large Model Based on Homomorphic Merging
Through homomorphic merging and dynamic similarity threshold adjustment methods, the source code security detection of large-scale artificial intelligence models is optimized, which solves the problems of high computing requirements and low detection efficiency, and achieves efficient and accurate source code security detection.
Patent Information
- Application Number
- CN202510300860.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2045-03-14
AI Technical Summary
When large-scale artificial intelligence models process multimodal source code data, the computing demand has increased sharply, resulting in high hardware costs and limited detection efficiency and accuracy. The existing static analysis methods are prone to false positives or missed reports.
Homomorphic merging technology is used to perform redundant detection and merging of source code data, combined with dynamic similarity threshold adjustment, optimized models through feature extraction and dimensional reduction, and generate representative feature data and train them.
It significantly reduces the data dimension, improves the efficiency, accuracy and flexibility of source code security detection, reduces hardware costs, and improves the generalization ability and detection accuracy of the model.
Smart Images

Figure CN119807721B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of large model dimensionality reduction, and more specifically, to a method for reducing the dimensionality of a large model for source code security detection based on homomorphic merging. Background Art
[0002] Currently, artificial intelligence models are developing towards large-scale and ultra-large-scale directions, such as GPT-4, which has hundreds of billions or even trillions of parameters. More tokens mean more data needs to be processed and learned by the model, and the training of large-scale models itself requires powerful computing power support to complete complex matrix operations and backpropagation and other operations. The development of multi-modal artificial intelligence makes the model need to process various types of data such as text, code, functions, processes, audio, and source code at the same time. The number of tokens of source code and data such as high-resolution code, functions, and processes usually far exceeds that of text. For example, a few minutes of source code may contain millions of tokens. When processing a large amount of multi-modal data, the model needs to consume more computing power to encode, decode, and fuse different modal data to extract valuable information, which makes the computing power demand increase sharply;
[0003] To meet the growing computing power demand, enterprises and research institutions need to purchase a large number of high-performance computing chips. To maintain competitiveness, they need to continuously update and upgrade the chips, which makes the hardware purchase cost continue to increase. In addition to chips, a large number of servers and storage devices are also needed to support the operation of computing power and data storage. As the amount of data and computing tasks increase, the scale of servers and storage devices also needs to be continuously expanded, which further increases the hardware cost investment. At the same time, as the project scale increases, the amount of code grows exponentially. The currently used static analysis and dynamic analysis methods may face performance bottlenecks. Especially for static analysis methods, as the amount of code and complexity increase, it may lead to a large number of false positives or false negatives, thus affecting the detection efficiency and accuracy. Therefore, a method for reducing the dimensionality of a large model for source code security detection based on homomorphic merging is proposed. Summary of the Invention
[0004] The purpose of the present invention is to provide a method for reducing the dimensionality of a large model for source code security detection based on homomorphic merging to solve the problems raised in the above background art.
[0005] To achieve the above object, a method for reducing the dimensionality of a large model for source code security detection based on homomorphic merging is provided, including the following steps:
[0006] S1. Collect source code data from the development environment, and at the same time obtain the performance data of the computer;
[0007] S2. Extract and analyze the data features of the source code data, perform token series conversion on the source code data based on the analysis results using the data features, and simultaneously assign position encodings to each token;
[0008] S3. Perform redundancy detection on all tokens with each other, thereby extracting identical tokens for homomorphic merging, replacing multiple tokens after homomorphic merging with a new token for representation, and then establishing a token sequence library to collect the tokens;
[0009] S4. Calculate the similarity values of the data features of all tokens in the token sequence library with each other, and simultaneously perform token quantity adaptation analysis based on the performance data and the complexity between the source code data, and set a dynamic similarity threshold according to the analysis results;
[0010] S5. Filter the similarity values between tokens in combination with the dynamic similarity threshold, then perform homomorphic merging between tokens that meet the dynamic similarity threshold according to the filtering results, and simultaneously perform data dimension reduction analysis on the tokens after homomorphic merging, thereby generating representative feature data;
[0011] S6. Feed the tokens processed by S5 back to the token sequence library for compression and replacement, then input the token sequence library into the source code model for training, and adjust the dynamic similarity threshold according to the training results.
[0012] As a further improvement of this technical solution, S1 collects source code data from different development environments and simultaneously performs formatting adjustment processing on the source code data, so as to unify the format of the source code data.
[0013] As a further improvement of this technical solution, S1 accesses the management port of the computer, thereby extracting the performance data of the computer in real time in the management port, so as to monitor the load status of the computer.
[0014] As a further improvement of this technical solution, the steps of S2 are as follows:
[0015] S2.1. Extract the features of the source code data, thereby obtaining the static features and semantic features of the source code data, then summarize the static features and semantic features into data features, and simultaneously perform security detection on the source code data according to the data features, and only retain the source code data whose data features pass the security detection;
[0016] S2.2. Convert the source code data retained in S2.1 into a series of tokens through data feature transformation. The token series consists of multiple tokens, and each token represents a basic unit, function name, function block, and process in the source code;
[0017] S2.3. Assign position encodings to each token so that each token corresponds to a separate position encoding. The application of the token is represented by the position encoding, thereby describing the source code data as a series of token reference encodings.
[0018] As a further improvement of this technical solution, the steps of S3 are as follows:
[0019] S3.1. Use a redundancy detection algorithm to perform redundancy detection on all tokens with each other, thereby extracting the same tokens. Then, perform homomorphic merging on multiple continuously combined and repeated tokens so that multiple repeated tokens form a new token;
[0020] S3.2. Establish a token sequence library, and at the same time use the token sequence library to collect the tokens reorganized in S3.1 and the unmerged tokens.
[0021] As a further improvement of this technical solution, the steps of S4 are as follows:
[0022] S4.1. Extract the data features of all tokens in the token sequence library, and then perform similarity value calculations on the data features of all tokens with each other to obtain the similarity values between each token;
[0023] S4.2. Combine the source code data corresponding to all tokens in the token sequence library for complexity value analysis. According to the analysis results, obtain the complexity value of the token sequence library. Then, combine the performance data with the complexity value for token quantity adaptation analysis. According to the analysis results, obtain the token quantity adaptable to the performance data, and at the same time set the analyzed token quantity as the dynamic similarity threshold.
[0024] As a further improvement of this technical solution, in S4.2, when the latest performance data shows that the load status of the computer is overloaded, substitute the latest performance data into S4.2 again to reset the dynamic similarity threshold. Conversely, when the latest performance data shows that the load status of the computer is running normally, continue to monitor.
[0025] As a further improvement of this technical solution, the steps of S5 are as follows:
[0026] S5.1. Combine the similar numerical values between tokens with a dynamic similarity threshold for screening, so as to form a token merge list that meets the dynamic similarity threshold according to the similar numerical values of the tokens. Then, calculate the average similarity value of the obtained token merge list, and preferentially retain the token merge list with the highest average similarity value.
[0027] S5.2. Perform homomorphic merging on all tokens according to the token merge list retained in S5.1. At the same time, extract the data features of the tokens after homomorphic merging, and then combine the extracted data features for feature selection and data dimension reduction analysis. Generate corresponding representative feature data for the tokens after homomorphic merging according to the analysis results.
[0028] As a further improvement of this technical solution, the steps of S6 are as follows:
[0029] S6.1. Feed the tokens after homomorphic merging in S5.2 back to the token sequence library for compression and replacement, so that the latest compressed and replaced tokens in the token sequence library are used as the training set.
[0030] S6.2. Input the training set into the source code model for training, and adjust the dynamic similarity threshold according to the training results. Then, feedback the training results to the developer. When the feedback result of the developer is to continue to adjust, increase the dynamic similarity threshold.
[0031] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0032] 1. In this method for reducing the dimension of a large-scale source code security detection model based on homomorphic merging, through the homomorphic merging technology, similar or duplicate tokens in the source code are merged into a unified representation. This process effectively reduces the redundant information in the source code and significantly reduces the data dimension, which is crucial for large-scale source code security detection models. Because the code library usually contains thousands of tokens, the merged data set is more compact, reducing the computational complexity of model training and inference, improving the efficiency of data processing, enabling the source code model to complete training in a shorter time, and being able to process large-scale code libraries more efficiently.
[0033] 2. In this method for reducing the dimension of the large-scale source code security detection model based on homomorphic merging, through the dynamic similarity threshold adjustment mechanism, it is allowed to dynamically optimize the token merging strategy according to the training results and developers' feedback. If developers feedback that the similarity requirements for certain tokens need to be stricter or looser, the adjustment of the threshold can help the system handle the situation of blurred boundaries more accurately when merging tokens. This dynamic adjustment ensures that the model is more flexible when dealing with different types of source code, and can adjust its merging and recognition strategies according to specific code features. Through this dynamic similarity threshold adjustment, the large-scale source code security detection model can continuously optimize itself, adapt to different code styles and security requirements, and improve the flexibility and adaptability of the detection process.
[0034] 3. In this method for reducing the dimension of the large-scale source code security detection model based on homomorphic merging, by combining technologies such as homomorphic merging, feature extraction, feature selection, dimension reduction, and dynamic similarity threshold adjustment, the efficiency, accuracy, and scalability of source code security detection are effectively improved. By reducing redundancy, improving the generalization ability of the model, flexibly adjusting the similarity threshold, reducing the model complexity, improving the detection accuracy, and expanding the coverage, this system can help developers discover potential security vulnerabilities in a shorter time, reduce the time and cost of vulnerability repair, and ultimately enhance the security of the system. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 is the overall flowchart of the present invention;
[0036] Figure 2 is the flowchart of the source code data of the present invention for retaining data features through security detection;
[0037] Figure 3 is the flowchart of the present invention for combining multiple duplicate tokens into a new token;
[0038] Figure 4 is the flowchart of the present invention for obtaining the similarity values between each token;
[0039] Figure 5 is the flowchart of the present invention for generating corresponding representative feature data for the homomorphically merged tokens according to the analysis results;
[0040] Figure 6 is the flowchart of the present invention for using the tokens replaced with the latest compression in the token sequence library as the training set. DETAILED DESCRIPTION OF THE INVENTION
[0041] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0042] Please refer to Figure 1 As shown, the purpose of this embodiment is to provide a method for dimension reduction of a large model for source code security detection based on homomorphic merging, including the following steps:
[0043] S1. Collect source code data from the development environment and obtain the performance data of the computer at the same time;
[0044] S1 collects source code data from different development environments and formats and adjusts the source code data at the same time, so as to unify the format of the source code data. The specific steps are as follows:
[0045] Source code data collection: Identify the development environment or version control system where the source code comes from, obtain the source code files from different environments, and in addition to the code files themselves, collect the relevant metadata of the source code;
[0046] Remove useless comments and white characters: Use regular expressions or specialized code parsing tools to remove the comment parts in the code, remove extra spaces, tab characters, blank lines, and line breaks to ensure the standardization of the code;
[0047] Unify the code and naming specifications: Ensure that all variables, functions, and class names conform to the unified naming rules, ensure that the definition positions of functions and classes in the code, the parameter order, etc. conform to the unified standards, and at the same time unify the comment format. Each function should have a clear docstring, and then use tools to automatically format the source code to ensure a unified code style.
[0048] S1 accesses the management port of the computer, and extracts the performance data of the computer in real time in the management port, so as to monitor the load status of the computer. The steps are as follows:
[0049] Obtain performance data through the management port: Use the management port to directly obtain the performance data of the computer at the hardware level. These data include CPU usage, memory usage, hard disk health status, power management status, etc.;
[0050] Load Status Monitoring and Real - time Extraction: The process of real - time extracting performance data is a continuous process. Therefore, it is necessary to set up automated scripts and collect data regularly to monitor the load status of the computer. Using Python scripts or other programming languages, combined with scheduled tasks (such as cron), data can be fetched regularly and saved to a log file or database. For example, performance data can be extracted every 5 minutes, and the load index can be calculated.
[0051] S2. Analyze and extract data features from the source code data, perform a token series conversion on the source code data according to the analysis results using the data features, and at the same time assign position encodings to each token;
[0052] As Figure 2 shown, the steps of S2 are as follows:
[0053] S2.1. Extract features from the source code data to obtain the static features and semantic features of the source code data, then summarize the static features and semantic features into data features, and at the same time perform security detection on the source code data according to the data features, and only retain the source code data whose data features pass the security detection. The steps are as follows:
[0054] Static feature extraction: For lexical features, obtain the basic structural information of the code through keywords, identifiers, constants, operators, etc. in the code. For syntactic structure features, through functions, variable definitions, conditional statements, loop structures, etc. in the code. For control - flow features, obtain the control - flow information of the program by analyzing the control - flow graph, such as judging whether there are potential risks such as infinite loops and callbacks. Then, through the parsing of the abstract syntax tree, extract structural features such as function calls, variable usage, and expressions. For dependency relationships, analyze the call relationships, variable dependencies, data flows, etc. between modules or functions in the code;
[0055] Semantic feature extraction: For data - flow analysis, extract the data flow between variables in the code, analyze the processes of variable passing, modification, and usage in functions, and identify potential vulnerabilities or risks. For control - flow analysis, analyze the execution paths of the program according to the control - flow graph to identify whether there are potential logical errors. For sensitive operation detection, extract specific system calls, network accesses, database operations, file reads and writes, etc. sensitive operations, and check whether there are possible security risks in the source code;
[0056] Feature summarization: Integrate the static features and semantic features together to form a unified feature set;
[0057] Security Detection: Security detection can be carried out through different machine learning algorithms or rule bases. Using an annotated code dataset (including secure and non-secure source code), a model is trained based on a feature set. The trained model is then evaluated for its performance. Methods such as cross-validation can be used to evaluate the accuracy, precision, recall, etc. of the model. Subsequently, the trained model is used to perform a security analysis on the source code to be detected, and the security of the source code is judged based on the prediction results. Finally, according to the security detection results, the source code that passes the security detection is retained, and the source code with security risks is eliminated.
[0058] S2.2. Convert the source code data retained in S2.1 into a series of tokens through data feature transformation. The series of tokens consists of multiple tokens, and each token represents a basic unit, function name, function block, and process in the source code.
[0059] Decompose the source code into basic syntax units. In the code, each token represents a basic unit in the source code, such as keywords, identifiers, operators, constants, etc., thus converting the source code into a sequence of tokens.
[0060] S2.3. Assign a position encoding to each token so that each token corresponds to a separate position encoding. The application of the token is represented by the position encoding, thereby describing the source code data as a series of token reference encodings. The steps are as follows:
[0061] Position Encoding: To enable the model to understand the order and structure of tokens, a position encoding needs to be assigned to each token. The position encoding is achieved by mapping the position of the token to a vector space. The commonly used position encoding method is based on the calculation of sine and cosine functions, and the most common one is the encoding method used in the Transformer model.
[0062] Fusion of Position Encoding and Token: After each token undergoes position encoding, it will be fused with its original representation (such as word embedding). Word embedding is the process of mapping a token to a vector space, usually using a pre-trained word vector model or a code-specific embedding method.
[0063] Create a Series of Tokens: The entire process finally generates a sequence represented by multiple tokens, and each token is accompanied by a position encoding. This sequence of tokens will be used as the input to the deep learning model.
[0064] S3. Redundancy detection is performed on all tokens to extract identical tokens for homomorphic merging. After homomorphic merging, multiple tokens are replaced with a new token for representation, and then a token sequence library is established to collect the tokens.
[0065] As Figure 3 shown, the steps of S3 are as follows:
[0066] S3.1. Use the redundancy detection algorithm to perform redundancy detection on all tokens to extract identical tokens, and then perform homomorphic merging on multiple continuously combined and repeated tokens so that multiple repeated tokens form a new token. The steps are as follows:
[0067] Statistical token frequency: Count the frequency of each token. Through the frequency table, it can be known which tokens are repeated, preparing for redundancy detection. The formula is as follows:
[0068] ;
[0069] where f(t) represents the frequency of the token, that is, the number of times the token appears in all texts, t is a single token, and I(t, x i ) is an indicator function indicating whether t appears in text x i . If it appears, the value is 1; if it does not appear, the value is 0, and n is the total number of texts;
[0070] Identify redundant tokens: Based on the frequency table, identify those tokens that appear frequently. These tokens are redundant tokens. During the redundancy detection process, tokens with a frequency exceeding a certain threshold will be marked as redundant. For example, the threshold can be set to 2. If a token appears more than the threshold (2 times), it is considered redundant;
[0071] Combine redundant tokens: For repeated tokens, merge them. The merge operation reduces the redundant representation by turning the repeated tokens into a new token. The formula is as follows:
[0072] ;
[0073] where t new is the new merged token, merge is the merge function that combines multiple repeated tokens into a new token, and t1, t 2, t n are the repeated tokens;
[0074] Update token representation: Update the original data, replace redundant tokens with the merged new tokens, remove redundant data, and make the data more compact.
[0075] S3.2. Establish a token sequence library, and at the same time use the token sequence library to collect the tokens reorganized in S3.1 and the unmerged tokens. The steps are as follows:
[0076] Establish a Token sequence library: Create an empty token sequence library. The representation of each token in the library can be its string form. Through the tokenization process, convert the input text data into a token sequence and store them in the token sequence library. At the same time, store each token as a key in the library and may attach frequency information or other metadata. If the token already exists, update its frequency or other information.
[0077] Collect the reorganized tokens:
[0078] Collect the tokens reorganized in S3.1, that is, the tokens after redundancy detection and merging, and add them to the token sequence library, and ensure that they can be effectively managed and queried with other unmerged tokens.
[0079] Collect the unmerged tokens: Store all unmerged tokens (i.e., original tokens) and the merged tokens together in the token sequence library to ensure data consistency and integrity, so that the token sequence library contains all merged tokens and original tokens at the same time, and their frequency information is updated and consistent.
[0080] S4. Calculate the similarity values of the data features of all tokens in the token sequence library with each other, and at the same time perform token quantity adaptation analysis according to the complexity between performance data and source code data, and set a dynamic similarity threshold according to the analysis results.
[0081] As Figure 4 shown, the steps of S4 are as follows:
[0082] S4.1. Extract the data features of all tokens in the token sequence library, and then calculate the similarity values of the data features of all tokens with each other to obtain the similarity values between each token. The steps are as follows:
[0083] Data feature extraction: Extract the feature vector of each token from the token sequence library.
[0084] Calculate the similarity between tokens: Calculate the similarity between all tokens in the vector space. The formula is as follows:
[0085] ;
[0086] where, are the feature vectors of two different tokens, is the dot product of the vectors, and are the lengths of the two token feature vectors, is the similarity value of two different tokens. The range of the similarity value is [-1, 1], where 1 means the two vectors are exactly the same, -1 means exactly the opposite, and 0 means no similarity at all.
[0087] S4.2. Combine the source code data corresponding to all tokens in the token sequence library for complexity value analysis. Obtain the complexity value of the token sequence library according to the analysis results. Then, combine the performance data with the complexity value for token quantity adaptability analysis. Obtain the token quantity adaptable to the performance data according to the analysis results. At the same time, set the analyzed token quantity as the dynamic similarity threshold. The steps are as follows:
[0088] Complexity value analysis of the token sequence library: It is necessary to extract the source code data of each token from the token sequence library. These source code data can be text data or related program execution or resource occupancy data. The main purpose of the complexity value analysis is to measure the complexity of each token sequence, which can be achieved by calculating some common complexity metrics, such as entropy, redundancy, information content, etc., to quantify the complexity of the sequence;
[0089] [[ID=...]]Performance data combined with complexity value for token quantity adaptability analysis: When performing performance analysis, it is necessary to evaluate the relationship between the token quantity and the system performance. Specifically, the token quantity has an impact on the system performance (such as running speed, resource usage, etc.). When analyzing complexity, performance metrics such as memory usage and execution time will change with the change of the token sequence. Therefore, it is necessary to combine the complexity value and the performance data for adaptability analysis;
[0090] Obtain the token quantity adaptable to the performance data: According to the performance data and complexity analysis, determine the maximum acceptable range of performance degradation, and combine the adaptability analysis results to obtain the corresponding token quantity;
[0091] Set a dynamic similarity threshold: Based on the above analysis, the final number of tokens needs to be set as the dynamic similarity threshold, and the optimal number of tokens obtained through the adaptability analysis is used as the initial threshold;
[0092] The key to this analysis method is to determine an adaptability parameter through the mathematical relationship between complexity and performance, so that the system can maintain good performance under different numbers of tokens.
[0093] In S4.2, when the latest performance data shows that the load status of the computer is overloaded, the latest performance data is re-substituted into S4.2 to reset the dynamic similarity threshold. Conversely, when the latest performance data shows that the load status of the computer is running normally, continuous monitoring is maintained.
[0094] By monitoring the latest performance data, first judge the current load status of the computer (overloaded or normal). If it is found that the system load is overloaded, the dynamic similarity threshold will be reset based on the current performance data so that the system can adapt to the changing load conditions. If the load status is normal, the system will continue to monitor and will not change the current threshold. In this way, through dynamic adjustment and continuous monitoring, the system can ensure efficient and stable operation.
[0095] S5. Combine the similarity values between tokens with the dynamic similarity threshold for screening, and then perform homomorphic merging between the tokens that meet the load dynamic similarity threshold according to the screening results. At the same time, perform data dimension reduction analysis on the tokens after homomorphic merging to generate representative feature data;
[0096] As Figure 5 shown, the steps of S5 are as follows:
[0097] S5.1. Combine the similarity values between tokens with the dynamic similarity threshold for screening, so as to form a token merge list that meets the dynamic similarity threshold according to the similarity values of the tokens. Then calculate the average similarity value of the obtained token merge list, and preferentially retain the token merge list with the highest average similarity value. The steps are as follows:
[0098] Token screening based on the dynamic similarity threshold: According to the calculated similarity values, set the dynamic similarity threshold to screen out those token pairs that meet the threshold conditions and form a new candidate list. This threshold is dynamically adjusted based on the current system or performance state to screen those tokens with similarity higher than this threshold;
[0099] Merge the tokens that meet the threshold conditions into a list: The token pairs selected through the above screening will be merged into a token merge list according to the similarity relationship. Here, the merge means combining the tokens with higher similarity together to form a new token set. The merged token set should contain all the eligible token pairs;
[0100] Calculate the average similarity value of each token merge list: For each token merge list, its average similarity will be calculated. The average similarity is the average of the similarities of all token pairs in the list, representing the overall similarity of the list;
[0101] Select the token merge list with the highest average similarity value: After obtaining multiple token merge lists and their average similarities, it is necessary to select the merge list with the highest average similarity as the final result. This is to retain the most similar token set to improve the efficiency and accuracy of subsequent operations;
[0102] The whole process is to calculate the similarity between tokens and screen them with a dynamic similarity threshold, finally forming multiple token merge lists that meet the threshold. Among all the merge lists, calculate the average similarity of each list, and preferentially retain the token merge list with the highest average similarity. Through this process, the final token set will be the most concise and highly similar set, thus improving the efficiency and performance of subsequent calculations.
[0103] S5.2. Homomorphically merge all tokens according to the token merge list retained in S5.1. At the same time, extract data features from the tokens after homomorphic merge, and then combine the extracted data features for feature selection and data dimension reduction analysis. Generate corresponding representative feature data for the tokens after homomorphic merge according to the analysis results. The steps are as follows:
[0104] Homomorphically merge the retained token list: Perform a homomorphic merge operation on the retained token merge list. Homomorphic merge means combining them into a unified representation while maintaining the information of each token. This usually includes unifying the representation of similar tokens or merging them according to a certain rule (such as vectorization);
[0105] Extract data features from the tokens after homomorphic merge: Once the tokens after homomorphic merge are obtained, data features need to be extracted from these tokens. These features can be various statistical characteristics of the original tokens, such as frequency, position, part of speech, context information, etc., or a vector table generated by a model (such as an embedding model);
[0106] Feature Selection and Data Dimension Reduction: Since a large number of features may be generated during the feature extraction process, the next step is to perform feature selection and data dimension reduction. The purpose of feature selection is to remove redundant or irrelevant features, while data dimension reduction reduces the dimension of the feature space through techniques such as Principal Component Analysis (PCA) and Linear Discriminant Analysis (LDA), making the data easier to process;
[0107] Generating Representative Feature Data: Based on feature selection and dimension reduction, representative feature data can be generated for the tokens after homomorphic merging. These representative feature data are the final "feature vector" representations of each token and are used for subsequent analysis, classification, clustering, and other tasks;
[0108] Integrate the tokens into a unified representation through homomorphic merging, and then extract relevant features from these merged tokens. These features will be optimized through feature selection and dimension reduction techniques, and finally, the representative feature data of each token will be generated. This series of operations can reduce redundant data and transform the tokens into a feature form that is easier to analyze and process, ensuring the efficiency and accuracy of the model and analysis.
[0109] S6. Feed the tokens processed in S5 back to the token sequence library for compression and replacement, then input the token sequence library into the source code model for training, and adjust the dynamic similarity threshold according to the training results.
[0110] As Figure 6 shown, the steps of S6 are as follows:
[0111] S6.1. Feed the tokens after homomorphic merging in S5.2 back to the token sequence library for compression and replacement, so that the latest compressed and replaced tokens in the token sequence library are used as the training set;
[0112] S6.2. Input the training set into the source code model for training, adjust the dynamic similarity threshold according to the training results, and then feedback the training results to the developers. When the feedback result from the developers is to continue adjusting, increase the dynamic similarity threshold. The steps are as follows:
[0113] Feedback to the Token Sequence Library: The merged tokens will be returned and stored in the token sequence library, which is used for subsequent token compression and replacement, enabling more effective model training;
[0114] Training Set Generation: Extract the compressed and replaced tokens from the token sequence library to construct a new training set. These compressed tokens are based on the merged results, thus improving the quality of the training set and removing redundant information;
[0115] Input the source code model for training: Input the generated training set into the source code model for training. The source code model learns patterns and structures based on these new, compressed tokens for semantic understanding and prediction;
[0116] Adjust the dynamic similarity threshold: After model training, evaluate the similarity between tokens based on the training results. The model calculates the similarity between tokens. If the model finds sufficient similarity, the dynamic similarity threshold will be adjusted. This threshold is used to control which tokens are considered "similar" and thus determine whether to merge into a new token;
[0117] Feedback to the developer: After training is completed, the model will feedback the training results to the developer based on the adjusted threshold. If the developer requests further adjustment (such as increasing or decreasing the tolerance of similarity), the dynamic similarity threshold will be appropriately increased or decreased;
[0118] Increase the dynamic similarity threshold: According to the developer's feedback, increase the dynamic similarity threshold (i.e., expand the range of accepted similarity) or decrease this threshold. This adjustment helps the system to be more lenient or strict when identifying token similarity, thereby further optimizing the merging and compression process.
[0119] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. The above embodiments and descriptions in the specification are only preferred examples of the present invention and are not used to limit the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of protection of the present invention is defined by the appended claims and their equivalents.
Claims
1. A method for reducing the dimension of a large model for source code security detection based on homomorphic merging, characterized in that: It includes the following steps: S1. Collect source code data from the development environment and obtain the performance data of the computer at the same time; S2. Extract and analyze the data features of the source code data, and perform token series conversion on the source code data according to the analysis results with the data features. At the same time, assign position encoding to each token. Each token represents a basic unit in the source code, and the basic unit includes keywords, identifiers, operators, and constants; S3. Perform redundancy detection on all tokens with each other, so as to extract the same tokens for homomorphic merging. Replace multiple tokens after homomorphic merging with a new token for representation, and then establish a token sequence library to collect the tokens; S4. Perform similar value calculation on the data features of all tokens in the token sequence library with each other. At the same time, perform token quantity adaptation analysis according to the complexity between the performance data and the source code data, and set a dynamic similarity threshold according to the analysis results; S5. Screen the similarity values between tokens in combination with the dynamic similarity threshold, and then perform homomorphic merging between the tokens that meet the dynamic similarity threshold according to the screening results. At the same time, perform data dimension reduction analysis on the tokens after homomorphic merging to generate representative feature data; S6. Feed the tokens processed in S5 back to the token sequence library for compression and replacement, and then input the token sequence library into the source code model for training, and adjust the dynamic similarity threshold according to the training results.
2. A method for dimension reduction of a large model for source code security detection based on homomorphic merging according to claim 1, characterized in that: In S1, by collecting source code data from different development environments and formatting and adjusting the source code data at the same time, the format of the source code data is unified.
3. A method for dimension reduction of a large model for source code security detection based on homomorphic merging according to claim 1, characterized in that: In S1, by accessing the management port of the computer, the performance data of the computer is extracted in real time in the management port, so as to monitor the load status of the computer.
4. A method for reducing the dimension of a large model for source code security detection based on homomorphic merging according to claim 1, characterized in that: The steps of S2 are as follows: S2.
1. Extract the features of the source code data to obtain the static features and semantic features of the source code data, and then summarize the static features and semantic features into data features. At the same time, perform security detection on the source code data according to the data features, and only retain the source code data whose data features pass the security detection; S2.
2. Convert the source code data retained in S2.1 into a token series with data features. The token series consists of multiple tokens; S2.
3. Assign position encoding to each token so that each token corresponds to a separate position encoding, and represent the application of the token through the position encoding, so as to describe the source code data as a series of token reference encodings.
5. A dimensionality reduction method for a large-scale source code security detection model based on homomorphic merging according to claim 1, characterized in that: The steps of S3 are as follows: S3.
1. Use the redundancy detection algorithm to perform redundancy detection on all tokens with each other, so as to extract the same tokens, and then perform homomorphic merging on multiple continuously combined and repeated tokens, so that multiple repeated tokens form a new token; S3.
2. Establish a token sequence library, and at the same time, use the token sequence library to collect the tokens reorganized in S3.1 and the tokens that have not been merged.
6. A dimensionality reduction method for a large model of source code security detection based on homomorphic merging according to claim 1, characterized in that: The steps of S4 are as follows: S4.
1. Extract the data features of all tokens in the token sequence library, and then perform similarity value calculations on the data features of all tokens to obtain the similarity values between each token. S4.
2. Conduct a complexity value analysis on the source code data corresponding to all tokens in the token sequence library. According to the analysis results, obtain the complexity value of the token sequence library. Then, combine the performance data with the complexity value to conduct a token quantity adaptation analysis. According to the analysis results, obtain the token quantity that the performance data can adapt to, and at the same time, set the analyzed token quantity as the dynamic similarity threshold.
7. A dimensionality reduction method for a source code security detection large model based on homomorphic merging according to claim 6, characterized in that: In S4.2, when the latest performance data shows that the load state of the computer is overloaded, substitute the latest performance data into S4.2 again to reset the dynamic similarity threshold. On the contrary, when the latest performance data shows that the load state of the computer is running normally, continue to monitor.
8. A dimensionality reduction method for a large-scale source code security detection model based on homomorphic merging according to claim 1, characterized in that: The steps of S5 are as follows: S5.
1. Screen the similarity values between tokens in combination with the dynamic similarity threshold, so as to form a token merge list that meets the dynamic similarity threshold according to the similarity values of the tokens. Then, calculate the average similarity value of the obtained token merge list, and preferentially retain the token merge list with the highest average similarity value. S5.
2. Perform homomorphic merging on all tokens according to the token merge list retained in S5.
1. At the same time, extract the data features of the tokens after homomorphic merging, and then conduct feature selection and data dimension reduction analysis on the extracted data features. According to the analysis results, generate corresponding representative feature data for the tokens after homomorphic merging.
9. A dimensionality reduction method for a large model of source code security detection based on homomorphic merging according to claim 1, characterized in that: The steps of S6 are as follows: S6.
1. Feed back the tokens after homomorphic merging in S5.2 to the token sequence library for compression and replacement, so that the latest compressed and replaced tokens in the token sequence library are used as the training set. S6.
2. Input the training set into the source code model for training, and adjust the dynamic similarity threshold according to the training results. Then, feed back the training results to the developer. When the feedback result of the developer is to continue adjusting, increase the dynamic similarity threshold.
Citation Information
Patent Citations
Software security vulnerability detection method based on text features and function dependency features
CN115186272A
Systems and methods for detecting code duplication in codebases
US20230185550A1