Code file packaging method and device, equipment and storage medium
Through the dependency graph and description information calculation semantic vectors, and combined with large language models to optimize code file packaging, the problems of inaccurate segmentation and relying on manual configuration in the existing technology are solved, and more efficient code file packaging and performance improvement are achieved.
Patent Information
- Application Number
- CN202510595633.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-09
- Publication Date
- 2025-08-22
AI Technical Summary
The existing code segmentation strategy relies on static analysis tools, lacks semantic understanding of the functions and module description information of the code module, resulting in inaccurate segmentation or excessive granularity, and relies on manual configuration by developers, affecting the code packaging efficiency and performance of large projects.
By obtaining the dependency diagram and description information of the code file, using large language models such as BERT and GNN to calculate semantic vectors, combined with the alignment matrix to optimize the similarity calculation, dynamically adjust the packaging granularity of the code file to ensure that the file is highly correlated and has a suitable size.
Improves the precision and accuracy of code file packaging, reduces page rendering time, improves user experience, and reduces the number of HTTP requests.
Smart Images

Figure CN120523508A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a code file packaging method, apparatus, device, and storage medium. Background Art
[0002] With the continuous development of front-end technology, the front-end code base is becoming larger and larger, which has led to the introduction of code splitting strategy. Code splitting strategy is a strategy for code packaging. Through this strategy, code compression can be achieved, code packaging efficiency can be improved, packaging size can be reduced, and application loading performance can be improved.
[0003] Existing code splitting strategies mainly rely on static analysis tools, such as webpack and splitChunks, which split the code into multiple blocks by analyzing the dependencies and module size of the code, and use dynamic imports to implement on-demand component loading.
[0004] Using static analysis tools to segment code improves its efficiency to a certain extent. However, static analysis tools lack semantic understanding of code module functionality and module descriptions, and lack dynamic adaptability during segmentation. This can lead to inaccurate segmentation or excessive granularity when segmenting code in large projects. Furthermore, the effectiveness of code segmentation using static analysis tools relies heavily on manual tool configuration by developers, posing challenges to their experience and capabilities. Summary of the Invention
[0005] The present application provides a code file packaging method, apparatus, device and storage medium to improve the precision of code file packaging.
[0006] In a first aspect, the present application provides a code file packaging method, the method comprising:
[0007] Obtain at least one code file to be packaged;
[0008] Determining a dependency graph and code description information corresponding to a first code file; the dependency graph represents a dependency relationship between the first code file and other code files in the at least one code file; the first code file is any one of the at least one code file;
[0009] Determine a semantic vector corresponding to the first code file according to the dependency graph and the code description information;
[0010] Similarity is calculated based on the semantic vector corresponding to the first code file and the semantic vector corresponding to the second code file. If the similarity satisfies a first condition and the file size after packaging the first code file and the second code file meets a second condition, the first code file and the second code file are packaged; the second code file is any code file in the dependency graph other than the first code file.
[0011] In this application, the semantic vector corresponding to the code file is determined based on the dependency graph and the code description information, which improves the accuracy of the semantic vector; further, the similarity calculated based on the more accurate semantic vector is also more accurate. In the process of packaging the code files, it is necessary not only to judge whether the similarity between the two files meets the first condition, but also to judge whether the size of the code file after the two code files are merged meets the second condition, so as to avoid the two packaged files being large in size and affecting the loading performance; it ensures a high degree of correlation between the packaged code files, and makes the code file packaging granularity more uniform.
[0012] In one possible design, determining the semantic vector corresponding to the first code file according to the dependency graph and the code description information includes:
[0013] A semantic vector corresponding to the first code file is determined according to the code information corresponding to the first code file, the dependency graph, and the code description information.
[0014] The semantic vector is determined not only by considering the dependency graph and code description information, but also by considering the code information. In the embodiment of the present application, the code information mainly refers to the coding or program content in the code file, which further improves the accuracy of the semantic vector.
[0015] In a possible design, before calculating similarity based on the semantic vector corresponding to the first code file and the semantic vector corresponding to the second code file, the method further includes:
[0016] Calculating an alignment matrix based on the semantic vector corresponding to the first code file and the semantic vector corresponding to the second code file; an element at any position in the alignment matrix represents a local difference between the semantic vector corresponding to the first code file and the semantic vector corresponding to the second code file at the position;
[0017] Determining an updated semantic vector corresponding to the first code file according to the semantic vector corresponding to the first code file and the alignment matrix;
[0018] An updated semantic vector corresponding to the second code file is determined according to the semantic vector corresponding to the second code file and the alignment matrix.
[0019] In one possible design, calculating an alignment matrix according to the semantic vector corresponding to the first code file and the semantic vector corresponding to the second code file includes:
[0020] For an element at any position in the alignment matrix, the element value at the position is calculated based on the element value corresponding to the semantic vector corresponding to the first code file at the position, the element value corresponding to the semantic vector corresponding to the second code file at the position, a preset number of element values of the semantic vector corresponding to the first code file around the position, and a preset number of element values of the semantic vector corresponding to the second code file around the position.
[0021] The core of the alignment matrix is to obtain the degree of differentiation between the two vectors by calculating the nonlinear distance between the two vectors in the same spatial coordinate system. In the process of calculating the alignment matrix, the present application not only considers the difference in elements between the semantic vector corresponding to the first code file and the semantic vector corresponding to the second code file at the current calculation position, but also considers the difference in elements in the area around the current calculation position, that is, the above-mentioned local difference; in this way, the calculated alignment matrix can capture element alignment errors and suppress outliers. In addition, the updated semantic vector corresponding to the first code file determined according to the semantic vector corresponding to the first code file and the alignment matrix, and the updated semantic vector corresponding to the second code file determined according to the semantic vector corresponding to the second code file and the alignment matrix, can, to a certain extent, reduce the similarity misjudgment caused by the semantic ambiguity of the annotations and improve the accuracy of subsequent similarity calculations.
[0022] In one possible design, the method further includes:
[0023] Analyzing the dependency relationship between the first code file and other code files in the at least one code file to obtain a dependency relationship file corresponding to the first code file;
[0024] Determine a code file path and / or a code file function description of the first code file according to the dependency relationship file;
[0025] The code file path and / or code file function description is used as code description information corresponding to the first code file.
[0026] During the coding process, developers generally store the code files corresponding to the same page or the same content area in the same local directory. This way, files with similar code file paths or similar code file function descriptions are more likely to belong to the same content area. By packaging them in the same package, when loading the current page, only one request is needed to obtain the data required for page loading from each file in the packaged package, reducing page rendering time and improving user experience. Therefore, the code description information determined based on the code file path and / or code file function description is used as one of the bases for the semantic vector, and similarity is calculated based on the semantic vector. Using similarity as one of the bases for determining whether to package the code files has a high reliability.
[0027] In one possible design, determining the semantic vector corresponding to the first code file according to the code information corresponding to the first code file, the dependency graph, and the code description information includes:
[0028] The code information corresponding to the first code file, the dependency graph and the code description information are input into a trained large language model to obtain a semantic vector corresponding to the first code file; the large language model includes a bidirectional transformer model BERT and a graph neural network GNN; wherein the BERT is used to process the code information and the code description information, and the GNN is used to process the dependency graph.
[0029] In one possible design, the code information corresponding to the first code file, the dependency graph, and the code description information are input into a trained large language model to obtain a semantic vector corresponding to the first code file, including:
[0030] Input the code information corresponding to the first code file into the BERT for forward propagation to obtain a first word segmentation vector set; perform maximum pooling on each word segmentation vector in the first word segmentation vector set to obtain a second word segmentation vector set; input the second word segmentation vector into the BERT for backward propagation to obtain a text embedding vector;
[0031] Obtain a feature matrix corresponding to the dependency graph through a preset plug-in in the GNN, input the feature matrix into the heterogeneous graph neural network in the GNN, and obtain a dependency embedding vector;
[0032] Input the code description information into the BERT to obtain a description information embedding vector;
[0033] The semantic vector is determined according to the text embedding vector, the dependency embedding vector, and the description information embedding vector.
[0034] In a second aspect, the present application further provides a code file packaging device, the device comprising: a processing unit and a transceiver unit;
[0035] The transceiver unit is used to obtain at least one code file to be packaged;
[0036] The processing unit is configured to determine a dependency graph and code description information corresponding to a first code file; the dependency graph represents a dependency relationship between the first code file and other code files in the at least one code file; the first code file is any one of the at least one code file;
[0037] The transceiver unit is further configured to determine a semantic vector corresponding to the first code file based on the dependency graph and the code description information;
[0038] The processing unit is further configured to calculate similarity based on a semantic vector corresponding to the first code file and a semantic vector corresponding to the second code file, and package the first code file and the second code file if the similarity satisfies a first condition and a file size after packaging the first code file and the second code file satisfies a second condition; the second code file is any code file other than the first code file in the dependency graph.
[0039] In one possible design, the transceiver unit, when used to determine the semantic vector corresponding to the first code file according to the dependency graph and the code description information, is specifically used to:
[0040] A semantic vector corresponding to the first code file is determined according to the code information corresponding to the first code file, the dependency graph, and the code description information.
[0041] In one possible design, before calculating the similarity based on the semantic vector corresponding to the first code file and the semantic vector corresponding to the second code file, the processing unit is further used to calculate an alignment matrix based on the semantic vector corresponding to the first code file and the semantic vector corresponding to the second code file; the element at any position in the alignment matrix represents the local difference between the semantic vector corresponding to the first code file and the semantic vector corresponding to the second code file at the position; the updated semantic vector corresponding to the first code file is determined based on the semantic vector corresponding to the first code file and the alignment matrix; and the updated semantic vector corresponding to the second code file is determined based on the semantic vector corresponding to the second code file and the alignment matrix.
[0042] In one possible design, the processing unit, when used to calculate the alignment matrix according to the semantic vector corresponding to the first code file and the semantic vector corresponding to the second code file, is specifically used to:
[0043] For an element at any position in the alignment matrix, the element value at the position is calculated based on the element value corresponding to the semantic vector corresponding to the first code file at the position, the element value corresponding to the semantic vector corresponding to the second code file at the position, a preset number of element values of the semantic vector corresponding to the first code file around the position, and a preset number of element values of the semantic vector corresponding to the second code file around the position.
[0044] In one possible design, the processing unit is further used to analyze the dependency relationship between the first code file and other code files in the at least one code file to obtain a dependency file corresponding to the first code file; determine the code file path and / or code file function description of the first code file based on the dependency file; and use the code file path and / or code file function description as code description information corresponding to the first code file.
[0045] In one possible design, the transceiver unit is used to determine the semantic vector corresponding to the first code file based on the code information corresponding to the first code file, the dependency graph and the code description information, and is specifically used to: input the code information corresponding to the first code file, the dependency graph and the code description information into a trained large language model to obtain the semantic vector corresponding to the first code file; the large language model includes a bidirectional transformer model BERT and a graph neural network GNN; wherein the BERT is used to process the code information and the code description information, and the GNN is used to process the dependency graph.
[0046] In one possible design, the transceiver is configured to input the code information corresponding to the first code file, the dependency graph, and the code description information into a trained large language model to obtain a semantic vector corresponding to the first code file, specifically for:
[0047] Input the code information corresponding to the first code file into the BERT for forward propagation to obtain a first word segmentation vector set; perform maximum pooling on each word segmentation vector in the first word segmentation vector set to obtain a second word segmentation vector set; input the second word segmentation vector into the BERT for backward propagation to obtain a text embedding vector;
[0048] Obtain a feature matrix corresponding to the dependency graph through a preset plug-in in the GNN, input the feature matrix into the heterogeneous graph neural network in the GNN, and obtain a dependency embedding vector;
[0049] Input the code description information into the BERT to obtain a description information embedding vector;
[0050] The semantic vector is determined according to the text embedding vector, the dependency embedding vector, and the description information embedding vector.
[0051] In a third aspect, the present application further provides a code file packaging device, the device comprising: a processor, and a memory communicatively connected to the processor;
[0052] The memory stores computer-executable instructions;
[0053] The processor executes the computer-executable instructions stored in the memory to implement the method described in the first aspect above.
[0054] In a fourth aspect, the present application further provides a computer-readable storage medium, wherein the computer-readable storage medium includes a program, and when the program is executed on a device, the device executes the method as described in any one of the above-mentioned first aspects.
[0055] In a fifth aspect, the present application further provides a computer program product, which includes a computer program, and when the computer program is executed by a processor, it implements the method described in the first aspect above. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0057] Figure 1 A flowchart of a code file packaging method provided in an embodiment of the present application;
[0058] Figure 2 A schematic diagram of dependency files provided for an embodiment of the present application;
[0059] Figure 3 A diagram illustrating the dependency relationships provided in the embodiments of the present application;
[0060] Figure 4 Schematic diagram of the code file packaging device structure provided in the embodiment of this application Figure 1 ;
[0061] Figure 5 Schematic diagram of the code file packaging device structure provided in the embodiment of this application Figure 2 . DETAILED DESCRIPTION
[0062] To make the objectives, technical solutions, and advantages of this application more clear, this application will be further described in detail below with reference to the accompanying drawings. Obviously, the embodiments described are only some of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making any creative efforts are within the scope of protection of this application.
[0063] The application scenarios described in the embodiments of this application are intended to more clearly illustrate the technical solutions of the embodiments of this application and do not constitute a limitation on the technical solutions provided by the embodiments of this application. Persons skilled in the art will appreciate that, as new application scenarios emerge, the technical solutions provided by the embodiments of this application are equally applicable to similar technical problems. In the description of this application, unless otherwise specified, "multiple" means two or more.
[0064] As business functions evolve, the number of code files in code development repositories increases. In modern front-end development, code bundling and compression are key steps in improving website performance. Code bundling allows multiple code files to be combined into one or several smaller files, thereby reducing the number of Hypertext Transfer Protocol (HTTP) requests. Compression can further reduce file size, speed up page loading, and improve application loading performance.
[0065] As mentioned in the background technology section, the existing front-end code file packaging method lacks semantic understanding of code module description information. To this end, this application proposes the following Figure 1 The code file packaging method shown in the figure can be executed by a server, a chip in the server, or a functional module in the server. This application does not limit this. The following uses the server as the execution subject as an example to illustrate the method, which includes:
[0066] Step 101: Obtain at least one code file to be packaged.
[0067] For example, in the embodiments of this application, code modules exist independently in the form of code files. One code file is equivalent to one code module. For example, the homepage of a website corresponds to one code file, and the search function on the homepage corresponds to one code file. Each code file in the code library has a corresponding code file identifier, which can be the file name of the code file or the path of the code file in the code library. This application does not limit this. It is sufficient to uniquely identify the code file.
[0068] Step 102: Determine a dependency graph and code description information corresponding to a first code file; the dependency graph represents a dependency relationship between the first code file and other code files in at least one code file; the first code file is any one of the at least one code file.
[0069] For example, the dependency graph and code description information corresponding to each code file are pre-stored and can be stored in a database or in memory. Since the code description information is relatively small, it is generally stored in memory. Compared with determining the code description information from a database, determining the code description information from memory is more efficient.
[0070] First, the storage process of the dependency graph of each code file is explained:
[0071] Based on the preset tool, the dependency relationship of each code file is analyzed to obtain the dependency file corresponding to each code file. The preset tool can be a webpack tool; then, the dependency graph corresponding to each code file is determined based on the dependency file corresponding to each code file. The dependency graph includes nodes and edges. The edge connecting the nodes indicates that there is a dependency relationship between the code files represented by the two nodes; finally, the code file identifier of each code file is used as the primary key, and the dependency graph corresponding to each code file is stored as the key value in the database. The database can be called a structural database.
[0072] Taking the webpack tool as an example, the webpack tool starts from the entry file and determines the dependencies between the code files based on the import and require statements in the code files, and then obtains the dependency files corresponding to each code file. Figure 2 The following are the dependency files corresponding to the code file named module1. Figure 2 As can be seen from the "dependencies" field in the code file named module1, the code file named module2 and the code file named module3 are dependent on the code file named module1. Figure 2 The dependency graph determined by the dependency file is as follows Figure 3 As shown, there are 3 nodes and 2 edges. The 3 nodes correspond to module1 to module3 respectively. One of the 2 edges is used to indicate that there is a dependency relationship between the two code files module1 and module2, and the other edge is used to indicate that there is a dependency relationship between the two code files module1 and module3. Figure 3 After the dependency graph, module1 can be used as the primary key. Figure 3 The dependency graph shown is stored as key values in the structure database.
[0073] Next, the storage process of the code description information is explained:
[0074] First, the dependency relationship between the first code file and at least one other code file in the code file is analyzed to obtain a dependency file corresponding to the first code file; then, the code file path and / or code file function description of the first code file is determined based on the dependency file; finally, the code file path and / or code file function description is used as the code description information corresponding to the first code file. When storing, the code description information can be stored in the form of an object. For example, the code description information of each code file can be stored in memory in the format of {id:, path:, describe:}, where id is the code file identifier, path is the code file path, and describe is the code file function description.
[0075] Still Figure 2 Take the dependency file shown in the figure as an example. For module1, its path is determined by the "identifier" field, that is, the path of the code file named module1 is " / project / src / index.js". The function description of the code file is obtained according to the "describe" field in the dependency file. Figure 2 It can be seen that the content of the "describe" field of module1 is "implementing page rendering". Therefore, the code file function description of the code file named module1 is "implementing page rendering". Further, the code description information of the code file named module1 is stored in the memory according to {id: module1, path: / project / src / index.js, describe: implementing page rendering}.
[0076] Similarly, the path of the code file named module2 is ". / src / utils / validator.js", and the function description of the code file is "implementing permission verification"; therefore, the code description information of the code file named module2 is stored in the memory according to {id: module2, path: . / src / utils / validator.js, describe: implementing permission verification}. Figure 2For module3, the code file path cannot be obtained. The code file's function description is "Implement public functions." Note that for code files whose code file paths cannot be obtained from dependency files, the code file path defaults to public. Therefore, the code description information for the code file named module3 is stored in memory as {id: module3, path: public, describe: Implement public functions}.
[0077] In addition to the dependency graph and code description information, the embodiment of the present application also stores the code information of each code file. The process is as follows:
[0078] Each code file is stringified to obtain the character strings corresponding to each code file as the code information corresponding to each code file; the character strings corresponding to each code file mainly include the coding content or program content in the code file; then the code file identifier of each code file is used as the primary key, and the corresponding character strings are stored as key values in the database, which can be called a module code database.
[0079] On the basis that the dependency graph, code description information and code information have been stored, a query in the structure database according to the code file identifier of the first code file can obtain the dependency graph corresponding to the first code file, a query in the memory according to the code file identifier of the first code file can obtain the code description information corresponding to the first code file, and a query in the module code database according to the identifier of the first code file can obtain the code information corresponding to the first code file.
[0080] Step 103: Determine the semantic vector corresponding to the first code file according to the dependency graph and the code description information.
[0081] In fact, the semantic vector corresponding to the first code file is determined based on the code information, dependency graph and code description information corresponding to the first code file; specifically, the code information, dependency graph and code description information corresponding to the first code file are input into the trained large language model to obtain the semantic vector corresponding to the first code file; wherein, the large language model includes a bidirectional transformer model (Bidirectional Encoder Representations from Transformers, BERT) and a graph neural network (Graph Neural Network, GNN); wherein, BERT is used to process the code information and code description information, and GNN is used to process the dependency graph.
[0082] When BERT processes code information, it includes:
[0083] (1) Input the code information corresponding to the first code file into BERT for forward propagation to obtain the first word segmentation vector set:
[0084] BERT includes a preset word segmenter, such as a tokenizer; the preset word segmenter preprocesses the input code information to obtain a string array. Specifically, the preset word segmenter first clears the comment information and error codes in the code string to obtain a cleaned code string, and then adds special tags CLS and SEP to the cleaned code string. CLS is used to characterize the category of the string. If the CLS categories corresponding to two substrings in the string array obtained after segmentation are the same, the two substrings can be merged; SEP is used to indicate segmentation. A SEP tag is added every 512 tokens, and the cleaned string is segmented at the SEP tag to obtain a string array.
[0085] After obtaining the string array, the string array is fed into BERT for forward propagation to obtain the first set of word segmentation vectors. For example, suppose the preset word segmenter segments the string, obtaining set 1. The data in set 1 is similar to [value1, value2, value3...], and each value in set 1 is filled with 512 tokens after segmentation. Set 1 is fed into BERT for forward propagation, and the gradient descent update parameters are set to generate the first set of word segmentation vectors, set 2. The data in set 2 is similar to [vect1, vect2...], and each vector in set 2 can be represented as an i×j matrix, where j is the number of hidden layers in BERT and i is the number of tokens, which is 512.
[0086] In the process of segmenting the code string, the embodiment of the present application does not adopt the common word segmentation based on each token, but fills 512 tokens into one value. Through multiple comparative tests, the segmentation method adopted by the embodiment of the present application can not only ensure the word segmentation accuracy, but also significantly reduce the size of the set obtained after segmentation, thereby improving the efficiency of obtaining text embedding vectors.
[0087] (2) Take the maximum pooling of each word segmentation vector in the first word segmentation vector set to obtain the second word segmentation vector set:
[0088] Taking the maximum pooling for each word segmentation vector in the first word segmentation vector set means taking the maximum value for each dimension of each word segmentation vector, and the second word segmentation vector set obtained by taking the maximum pooling for each word segmentation vector in the first word segmentation vector set set2 is recorded as set3.
[0089] (3) Input the second word segmentation vector set into BERT for backward propagation to obtain the text embedding vector:
[0090] Finally, set3 is input into BERT for backward propagation, and the gradient descent parameters are dynamically updated to obtain the text embedding vector.
[0091] When BERT processes the code description information, the code description information is input into BERT to obtain the description information embedding vector.
[0092] When GNN processes the dependency graph, it obtains the feature matrix corresponding to the dependency graph through the preset plug-in in GNN, and then inputs the feature matrix into the heterogeneous graph neural network in GNN to obtain the dependency embedding vector. The specific process is as follows:
[0093] The preset plug-in in GNN can be an APOC plug-in. The preset plug-in analyzes the node data and edge data of the dependency graph to obtain at least one feature matrix corresponding to the dependency graph. The number of at least one feature matrix is equal to the number of nodes included in the dependency graph. Different feature matrices are obtained after the preset plug-in analyzes the node data and edge data with different nodes in the dependency graph as the entry.
[0094] by Figure 3 Take the dependency diagram shown as an example, Figure 3 There are 3 nodes and 2 edges in the dependency graph. Each node has corresponding node data. Similarly, each edge has corresponding edge data. Figure 3 By analyzing the node data and edge data of the dependency graph shown, three feature matrices can be obtained, namely, feature matrix ① obtained by analyzing the node data and edge data with module1 node as the entrance, feature matrix ② obtained by analyzing the node data and edge data with module2 node as the entrance, and feature matrix ③ obtained by analyzing the node data and edge data with module3 node as the entrance.
[0095] Exemplarily, after obtaining at least one feature matrix corresponding to the dependency graph through a preset plug-in, the heterogeneous graph neural network is used to perform mean aggregation on at least one feature matrix to obtain the overall feature matrix corresponding to the dependency graph. The process of determining the overall feature matrix is as follows: first, at least one feature matrix is merged to obtain a merged feature matrix, and then the average value of each column of the merged feature matrix is calculated to obtain the overall feature matrix, and finally, the dependency embedding vector is obtained based on the overall feature matrix.
[0096] For example, assuming that feature matrix ① is [0.2, 0.5, 0.1], feature matrix ② is [0.4, 0.3, 0.7], and feature matrix ③ is [0.1, 0.9, 0.2], first merge feature matrices ① to ③. The merged feature matrix is as follows:
[0097]
[0098] Then calculate the average value of each column of the merged feature matrix. The average value of the first column is 0.233, the average value of the second column is 0.567, and the average value of the third column is 0.333. Therefore, the overall feature matrix corresponding to the dependency graph is [0.233, 0.567, 0.333]. Then determine the dependency embedding vector based on [0.233, 0.567, 0.333].
[0099] At this point, the text embedding vector, description information embedding vector, and dependency embedding vector corresponding to the first code file are obtained. The semantic vector can be determined based on the text embedding vector, description information embedding vector, and dependency embedding vector. The semantic vector can be determined by the following formula:
[0100] Semantic vector = α × text embedding vector + γ × dependency embedding vector + θ × description information embedding vector
[0101] Among them, α and γ are adjustable weight coefficients. The initial value of α can be set to 0.6, the initial value of γ can be set to 0.4, and θ is the learning rate, and the initial value can be set to 0.1. During the training process of the model, the θ learning rate, α and γ weight coefficients can be adjusted according to the model training effect. Of course, the number of BERT and GNN layers can also be adjusted until the accuracy of the semantic vector output by the model meets the requirements.
[0102] The semantic vector determined in this application not only takes into account the dependency relationship between code files, but also takes into account the code description information, including the code file path and the code file function description, making the subsequent calculation of the similarity between semantic vectors more accurate, and further improving the accuracy of code file packaging based on the similarity of semantic vectors.
[0103] The semantic vector of the code file obtained in this step can be stored in the memory in the form of key:value, where the key is the code file identifier and the value is the semantic vector. In this way, it is possible to subsequently determine whether the semantic vector already exists in the memory based on the code file identifier, avoiding the repeated acquisition of the semantic vector corresponding to the same code file through a large language model, thus saving resources.
[0104] For example, Figure 3 For example, Figure 3 It includes module1 to module3. Assume that the semantic vector corresponding to module1 is E module1 , the semantic vector corresponding to module2 is E module2 , the semantic vector corresponding to module3 is E module3, you can follow {module1:E module1 , module2: E module2 , module3:E module3} format is stored in memory.
[0105] Step 104: Calculate similarity based on the semantic vector corresponding to the first code file and the semantic vector corresponding to the second code file. If the similarity satisfies a first condition and the file size of the first code file and the second code file after packaging satisfies a second condition, the first code file and the second code file are packaged. The second code file is any code file other than the first code file in the dependency graph.
[0106] Exemplarily, before calculating the similarity based on the semantic vector corresponding to the first code file and the semantic vector corresponding to the second code file, an alignment matrix is first calculated based on the semantic vector corresponding to the first code file and the semantic vector corresponding to the second code file. The element at any position in the alignment matrix represents the local difference in position between the semantic vector corresponding to the first code file and the semantic vector corresponding to the second code file.
[0107] For an element at any position in the alignment matrix, the element value at that position is calculated based on the element value corresponding to the semantic vector corresponding to the first code file at that position, the element value corresponding to the semantic vector corresponding to the second code file at that position, a preset number of element values of the semantic vector corresponding to the first code file around that position, and a preset number of element values of the semantic vector corresponding to the second code file around that position.
[0108] Assume that the semantic vector E1 corresponding to the first code file is [v1,v2,v3,...,v n ], v1~v1 are all matrices with n rows and 1 column, so E1 is a matrix with n rows and n columns; suppose the semantic vector E2 corresponding to the second code file is [w1,w2,w3,...,w n ],w1~w n Both are matrices with n rows and 1 column, so E2 is a matrix with n rows and n columns. Since both E1 and E2 are matrices with n rows and n columns, the alignment matrix calculated based on E1 and E2 is also a matrix with n rows and n columns. When calculating the element in the i-th row and j-th column of the alignment matrix, it can be calculated using the following formula:
[0109]
[0110] Among them, V i Represents the weight of the i-th dimension, which must satisfy The default value is V iThe size indicates the weight ratio of different dimensions in the alignment matrix, γ is the adjustment coefficient, usually set to 0.1, a ij is the average value of the preset number of element values of the semantic vector corresponding to the first code file around the position, b ij The value is the average value of a preset number of element values of the semantic vector corresponding to the second code file around the position.
[0111] With a ij Take the calculation of as an example, select the average value of the elements in the rectangular window of m×m around the position [i,j] in E1, where m is usually 5, then:
[0112]
[0113] Among them, u and v are the forward and backward coordinates calculated based on i and j, u∈[im / 2,i+m / 2], v∈[jm / 2,j+m / 2]. If im / 2 or jm / 2 is less than 0, it is taken as 0 by default. Similarly,
[0114] When designing the calculation formula of the alignment matrix, this application combines the local differences between the semantic vectors corresponding to the first code file and the semantic vectors corresponding to the second code file to generate an alignment matrix that can capture alignment errors and suppress outliers.
[0115] Furthermore, after obtaining the alignment matrix, an updated semantic vector corresponding to the first code file is determined based on the semantic vector corresponding to the first code file and the alignment matrix. For example, the semantic vector corresponding to the first code file can be multiplied by the alignment matrix to obtain the updated semantic vector corresponding to the first code file. Furthermore, an updated semantic vector corresponding to the second code file is determined based on the semantic vector corresponding to the second code file and the alignment matrix. Similarly, the semantic vector corresponding to the second code file can be multiplied by the alignment matrix to obtain the updated semantic vector corresponding to the second code file. Finally, similarity is calculated based on the updated semantic vectors corresponding to the first code file and the updated semantic vectors corresponding to the second code file.
[0116] The updated semantic vector determined by the present application based on the alignment matrix and the semantic vector corresponding to the first code file, as well as the updated semantic vector determined based on the alignment matrix and the semantic vector corresponding to the second code file, can reduce the similarity misjudgment caused by the semantic ambiguity of the annotations to a certain extent, and improve the accuracy of subsequent similarity calculations.
[0117] For example, the alignment matrix is recorded as W, and the updated semantic vector corresponding to the first code file is recorded as E1 ′=W×E1, the updated semantic vector corresponding to the second code file is recorded as E2 ′ =W×E2, the updated semantic vector E1 of the first code file ′ And the updated semantic vector E2 of the second code file ′ The similarity between them can be calculated by the following formula:
[0118]
[0119] sim(E1 ′ , E2 ′ ) can indicate the degree of association between the first code file and the second code file, sim(E1 ′ , E2 ′ ) value is larger, indicating that the correlation between the first code file and the second code file is higher; on the contrary, sim(E1 ′ , E2 ′ ), the smaller the value is, the lower the correlation between the first code file and the second code file is.
[0120] For example, if the similarity satisfies the first condition, the similarity is greater than a first preset threshold, and if the file size of the first and second code files after being packaged satisfies the second condition, the file size of the first and second code files after being packaged is less than a second preset threshold. If the similarity is greater than the first preset threshold and the file size of the first and second code files after being packaged is less than the second preset threshold, the first and second code files should be packaged into one package. If the similarity is greater than the first preset threshold, but the file size of the first and second code files after being packaged is greater than the second preset threshold, the first and second code files should be packaged separately.
[0121] For example, assuming that the first preset threshold is 0.7, the second preset threshold is 300 kb, and sim(E1 ′ , E2 ′ ) is greater than 0.7, and the size of the packaged code file of the first code file and the second code file is less than 300kb, the first code file and the second code file are packaged into one package; if sim(E1 ′ , E2 ′ ) is greater than 0.7, but the size of the first code file and the second code file after packaging is 410kb, which is greater than 300kb, then the first code file and the second code file should be packaged independently.
[0122] In addition, after calculating the similarity between the first code file and the second code file, the calculation result can be stored in the memory. When needed later, it can be directly retrieved from the memory to improve efficiency. When storing, it can be stored in the format of {first code file identifier: {second code file identifier: similarity between the first code file and the second code file}}. For example, assuming that the first code file is identified as module5 and the second code file is identified as module6, and the similarity between the first code file and the second code file is 0.5, then {module5: {module6: 0.5}} will be stored in the memory.
[0123] The reason for preferring to bundle code files with high similarity into one package is that when the front-end page loads, if the content of the current page is not in the same package, multiple requests will be generated, each used to obtain the data required for the current page load from different packages, increasing page rendering time and resulting in a poor user experience. For example, suppose the current page load requires the data required for page load from three code files, A1 to A3. If these three files are packaged into three different packages due to unreasonable packaging, with A1 packaged into package A, A2 packaged into package B, and A3 packaged into package C, then the current page load will generate three requests, one for each of the A1 file in package A, the A2 file in package B, and the A3 file in package C.
[0124] If, through steps 101 through 104, files A1 through A3 are packaged together, the current page load only requires one request to retrieve the content needed from files A1 through A3 within that package. This shows that placing highly similar files in the same package while controlling the size of the request package can reduce the number of requests, thereby shortening rendering time and improving the user experience.
[0125] The code file packaging method proposed in this application converts code strings, dependency graphs between code files, and code module description information into semantic vectors based on a large language model, and innovatively proposes a method for calculating an alignment matrix, which reduces the impact of semantic ambiguity in annotations on similarity calculations. During the code file packaging process, the code files are dynamically packaged based on the similarity between the code files and the file size, ensuring the rationality of the code file packaging and the uniformity of the packaging granularity, thereby improving the application loading performance.
[0126] Figure 4 and Figure 5The following is a schematic diagram of a possible code file packaging device provided in an embodiment of the present application. These code file packaging devices can be used to implement the functions of the server in the above method embodiment, and thus can also achieve the beneficial effects possessed by the above method embodiment.
[0127] like Figure 4 As shown, the code file packaging device 400 includes a processing unit 410 and a transceiver unit 420. The code file packaging device 400 is used to implement the above Figure 1 The functions of the server in the method embodiment are shown.
[0128] When the code file packaging device 400 is used to implement Figure 1 The functions of the server in the method embodiment shown are:
[0129] The transceiver unit 420 is used to obtain at least one code file to be packaged;
[0130] The processing unit 410 is configured to determine a dependency graph and code description information corresponding to a first code file; the dependency graph represents a dependency relationship between the first code file and other code files in the at least one code file; the first code file is any one of the at least one code file;
[0131] The transceiver unit 420 is further configured to determine a semantic vector corresponding to the first code file based on the dependency graph and the code description information;
[0132] The processing unit 410 is further configured to calculate a similarity based on a semantic vector corresponding to the first code file and a semantic vector corresponding to the second code file, and package the first code file and the second code file if the similarity satisfies a first condition and a file size after packaging the first code file and the second code file satisfies a second condition; the second code file is any code file in the dependency graph other than the first code file.
[0133] In one possible design, the transceiver unit 420, when used to determine the semantic vector corresponding to the first code file according to the dependency graph and the code description information, is specifically used to:
[0134] A semantic vector corresponding to the first code file is determined according to the code information corresponding to the first code file, the dependency graph, and the code description information.
[0135] In one possible design, before calculating the similarity based on the semantic vector corresponding to the first code file and the semantic vector corresponding to the second code file, the processing unit 410 is further used to calculate an alignment matrix based on the semantic vector corresponding to the first code file and the semantic vector corresponding to the second code file; the element at any position in the alignment matrix represents the local difference between the semantic vector corresponding to the first code file and the semantic vector corresponding to the second code file at the position; the updated semantic vector corresponding to the first code file is determined based on the semantic vector corresponding to the first code file and the alignment matrix; and the updated semantic vector corresponding to the second code file is determined based on the semantic vector corresponding to the second code file and the alignment matrix.
[0136] In one possible design, the processing unit 410, when calculating the alignment matrix according to the semantic vector corresponding to the first code file and the semantic vector corresponding to the second code file, is specifically configured to:
[0137] For an element at any position in the alignment matrix, the element value at the position is calculated based on the element value corresponding to the semantic vector corresponding to the first code file at the position, the element value corresponding to the semantic vector corresponding to the second code file at the position, a preset number of element values of the semantic vector corresponding to the first code file around the position, and a preset number of element values of the semantic vector corresponding to the second code file around the position.
[0138] In one possible design, the processing unit 410 is further used to analyze the dependency relationship between the first code file and other code files in the at least one code file to obtain a dependency file corresponding to the first code file; determine the code file path and / or code file function description of the first code file based on the dependency file; and use the code file path and / or code file function description as code description information corresponding to the first code file.
[0139] In one possible design, the transceiver unit 420 is used to determine the semantic vector corresponding to the first code file based on the code information corresponding to the first code file, the dependency graph and the code description information, and is specifically used to: input the code information corresponding to the first code file, the dependency graph and the code description information into a trained large language model to obtain the semantic vector corresponding to the first code file; the large language model includes a bidirectional transformer model BERT and a graph neural network GNN; wherein the BERT is used to process the code information and the code description information, and the GNN is used to process the dependency graph.
[0140] In one possible design, the transceiver unit 420 is configured to input the code information corresponding to the first code file, the dependency graph, and the code description information into a trained large language model to obtain a semantic vector corresponding to the first code file, specifically for:
[0141] Input the code information corresponding to the first code file into the BERT for forward propagation to obtain a first word segmentation vector set; perform maximum pooling on each word segmentation vector in the first word segmentation vector set to obtain a second word segmentation vector set; input the second word segmentation vector into the BERT for backward propagation to obtain a text embedding vector;
[0142] Obtain a feature matrix corresponding to the dependency graph through a preset plug-in in the GNN, input the feature matrix into the heterogeneous graph neural network in the GNN, and obtain a dependency embedding vector;
[0143] Input the code description information into the BERT to obtain a description information embedding vector;
[0144] The semantic vector is determined according to the text embedding vector, the dependency embedding vector, and the description information embedding vector.
[0145] For more detailed description of the processing unit 410 and the transceiver unit 420, please refer to Figure 1 The relevant description in the method embodiment shown is directly obtained and will not be repeated here.
[0146] like Figure 5 As shown, the code file packaging device 500 includes a processor 510 and an interface circuit 520. The processor 510 and the interface circuit 520 are coupled to each other. It is understood that the interface circuit 520 can be a transceiver or an input / output interface. Optionally, the code file packaging device 500 may further include a memory 530 for storing instructions executed by the processor 510, or storing input data required by the processor 510 to execute instructions, or storing data generated after the processor 510 executes instructions.
[0147] When the code file packaging device 500 is used to implement Figure 1 When the method is shown, the processor 510 is used to implement the functions of the processing unit 410, and the interface circuit 520 is used to implement the functions of the transceiver unit 420.
[0148] The division of units in the embodiments of the present application is illustrative and is merely a logical functional division. In actual implementation, other division methods may be used. Furthermore, the functional units in the various embodiments of the present application may be integrated into a single processor, or may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0149] Although the preferred embodiments of the present application have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present application.
[0150] Obviously, those skilled in the art may make various changes and modifications to the present application without departing from the scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalents, the present application is intended to include these modifications and variations.
Claims
1. A code file packaging method, characterized in that: The method includes: Obtain at least one code file to be packaged; Determining a dependency graph and code description information corresponding to a first code file; the dependency graph represents a dependency relationship between the first code file and other code files in the at least one code file; the first code file is any one of the at least one code file; Determine a semantic vector corresponding to the first code file according to the dependency graph and the code description information; Similarity is calculated based on the semantic vector corresponding to the first code file and the semantic vector corresponding to the second code file. If the similarity satisfies a first condition and the file size after packaging the first code file and the second code file meets a second condition, the first code file and the second code file are packaged; the second code file is any code file in the dependency graph other than the first code file.
2. The method according to claim 1, wherein The determining, according to the dependency graph and the code description information, a semantic vector corresponding to the first code file includes: A semantic vector corresponding to the first code file is determined according to the code information corresponding to the first code file, the dependency graph, and the code description information.
3. The method according to claim 1, wherein Before calculating similarity based on the semantic vector corresponding to the first code file and the semantic vector corresponding to the second code file, the method further includes: Calculating an alignment matrix based on the semantic vector corresponding to the first code file and the semantic vector corresponding to the second code file; an element at any position in the alignment matrix represents a local difference between the semantic vector corresponding to the first code file and the semantic vector corresponding to the second code file at the position; Determining an updated semantic vector corresponding to the first code file according to the semantic vector corresponding to the first code file and the alignment matrix; An updated semantic vector corresponding to the second code file is determined according to the semantic vector corresponding to the second code file and the alignment matrix.
4. The method according to claim 3, wherein Calculating an alignment matrix according to the semantic vector corresponding to the first code file and the semantic vector corresponding to the second code file includes: For an element at any position in the alignment matrix, the element value at the position is calculated based on the element value corresponding to the semantic vector corresponding to the first code file at the position, the element value corresponding to the semantic vector corresponding to the second code file at the position, a preset number of element values of the semantic vector corresponding to the first code file around the position, and a preset number of element values of the semantic vector corresponding to the second code file around the position.
5. The method according to claim 1, wherein The method further comprises: Analyzing the dependency relationship between the first code file and other code files in the at least one code file to obtain a dependency relationship file corresponding to the first code file; Determine a code file path and / or a code file function description of the first code file according to the dependency relationship file; The code file path and / or code file function description is used as code description information corresponding to the first code file.
6. The method according to claim 2, wherein The determining, according to the code information corresponding to the first code file, the dependency graph, and the code description information, a semantic vector corresponding to the first code file includes: The code information corresponding to the first code file, the dependency graph and the code description information are input into a trained large language model to obtain a semantic vector corresponding to the first code file; the large language model includes a bidirectional transformer model BERT and a graph neural network GNN; wherein the BERT is used to process the code information and the code description information, and the GNN is used to process the dependency graph.
7. The method according to claim 6, wherein Inputting the code information corresponding to the first code file, the dependency graph, and the code description information into a trained large language model to obtain a semantic vector corresponding to the first code file includes: Input the code information corresponding to the first code file into the BERT for forward propagation to obtain a first word segmentation vector set; perform maximum pooling on each word segmentation vector in the first word segmentation vector set to obtain a second word segmentation vector set; input the second word segmentation vector into the BERT for backward propagation to obtain a text embedding vector; Obtain a feature matrix corresponding to the dependency graph through a preset plug-in in the GNN, input the feature matrix into the heterogeneous graph neural network in the GNN, and obtain a dependency embedding vector; Input the code description information into the BERT to obtain a description information embedding vector; The semantic vector is determined according to the text embedding vector, the dependency embedding vector, and the description information embedding vector.
8. A code file packaging device, characterized in that: The device includes a processing unit and a transceiver unit; The transceiver unit is used to obtain at least one code file to be packaged; The processing unit is configured to determine a dependency graph and code description information corresponding to the first code file; The dependency graph represents the dependency relationship between the first code file and other code files in the at least one code file; The first code file is any one of the at least one code file; The transceiver unit is further configured to determine a semantic vector corresponding to the first code file based on the dependency graph and the code description information; The processing unit is further configured to calculate similarity based on a semantic vector corresponding to the first code file and a semantic vector corresponding to the second code file, and package the first code file and the second code file if the similarity satisfies a first condition and a file size after packaging the first code file and the second code file satisfies a second condition; the second code file is any code file other than the first code file in the dependency graph.
9. A code file packaging device, characterized in that: include: a processor, and a memory communicatively connected to the processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory to implement the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer-executable instructions, which are used to implement the method according to any one of claims 1 to 7 when executed by a processor.
11. A computer program product, characterized in that The invention comprises a computer program, which implements the method according to any one of claims 1 to 7 when executed by a processor.