A code file relevance analysis method based on deep learning

CN122653680APending Publication Date: 2026-08-28BEIJING KEHANG HUAGUANG TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610818934.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-08
Publication Date
2026-08-28

AI Technical Summary

Technical Problem

[0004]本发明的一个目的在于提出一种基于深度学习的代码文件关联性分析方法,本发明通过生成迁移记录、模块域、路径段通道和文件邻接通道,将输入样本输入改进GraphCodeBERT模型生成归属校准结果,并根据归属校准结果对代码文件邻接图执行图域更新,解决目录迁移导致基于路径的模块归属判断失效和代码文件关联结果失真的问题,提高目录迁移场景下代码文件关联性分析的准确性

Benefits of technology

本发明根据原文件路径和变更路径执行路径层级解析,提取原模块路径段和变更模块路径段,生成模块域,并结合路径段通道和文件邻接通道生成输入样本,使目录迁移前后的路径变化和工程邻接状态能够共同参与分析,提高目录迁移场景下模块域判断的准确性;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122653680A_ABST
    Figure CN122653680A_ABST
Patent Text Reader

Abstract

The application discloses a code file correlation analysis method based on deep learning and relates to the technical field of computer software analysis, and comprises the following steps: S1, a code file path before directory migration is taken as an original file path, a code file path after directory migration is taken as a changed path, and a migration record is generated; S2, a module domain is generated; S3, a path segment channel and a file adjacency channel are constructed, and input samples are combined and generated; S4, an improved GraphCodeBERT model is inputted, and attribution calibration results are generated; S5, graph domain writing instructions are generated; S6, an updated code file adjacency graph is constructed; and S7, graph domain adjacency analysis is performed, and correlation analysis results are generated. The application reduces wrong connections and broken links caused by direct judgment based on the changed path, and improves the accuracy and stability of the code file correlation analysis results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer software analysis technology, and in particular to a method for code file correlation analysis based on deep learning. Background Technology

[0002] With the widespread application of large-scale software engineering, microservice architecture, and continuous integration development, the size of code files is constantly increasing, and the calls, references, module ownership, and project adjacency states between code files are becoming increasingly complex. Therefore, automated analysis technologies for code file relationships have received widespread attention. Existing code analysis tools mainly rely on file paths, package paths, import declarations, call chains, and dependency graphs to determine code file relationships, and use static scanning or code representation models to assist in identifying the relationship states between files. However, in practical applications, the following problems commonly exist: During code engineering processes such as refactoring, module splitting, directory reorganization, and service migration, code file paths frequently change. Existing methods often determine module affiliation directly based on the changed code file paths, which can easily misinterpret directory hierarchy adjustments as changes in module affiliation. This can lead to incorrect classification of code files that originally belonged to the same business module or project module. While the path names of code files may change before and after directory migration, their adjacency, reference, and module inheritance states within the project may remain continuous. Existing analysis methods based on path strings or directory hierarchies struggle to distinguish between path changes and actual module changes. After code files are migrated to candidate module directories, traditional dependency graph update methods can easily write files directly into the new module scope, causing candidate module paths to influence the association results. This can result in misconnections, broken links, and distorted strong association results in the code file association graph, affecting the accuracy of code review, impact scope analysis, and refactoring quality assessment.

[0003] Therefore, how to provide a deep learning-based method for code file correlation analysis is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0004] One objective of this invention is to propose a deep learning-based code file association analysis method. This invention generates migration records, module domains, path segment channels, and file adjacency channels. Input samples are fed into an improved GraphCodeBERT model to generate attribution calibration results. Based on the attribution calibration results, the graph domain is updated on the code file adjacency graph. This solves the problems of path-based module attribution judgment failure and code file association result distortion caused by directory migration, and improves the accuracy of code file association analysis in directory migration scenarios.

[0005] A method for code file correlation analysis based on deep learning according to an embodiment of the present invention includes the following steps: S1. Use the code file path before the directory migration in the target code project as the original file path, and the code file path after the directory migration in the target code project as the change path, and generate a migration record. S2. Perform path hierarchy parsing based on the original file path and the changed path, extract the original module path segment and the changed module path segment, and generate the module domain; S3. Construct path segment channels based on migration records, construct file adjacency channels based on the project adjacency status of code files before and after directory migration, and combine path segment channels and file adjacency channels to generate input samples. S4. Input the input sample into the improved GraphCodeBERT model to generate the attribution calibration result; the improved GraphCodeBERT model introduces a path denaming and reconstruction mechanism in the association attention module and a counterfactual path verification mechanism in the attribution calibration module; S5. Determine the writing domain corresponding to the code file after the directory migration based on the attribution calibration results, and generate the domain writing instruction. S6. Construct a code file adjacency graph based on module domains and file adjacency channels, and perform a graph domain update on the code file adjacency graph according to the graph domain write instruction to generate the updated code file adjacency graph; S7. Perform graph domain adjacency resolution based on the updated code file adjacency graph to generate association analysis results.

[0006] Optionally, S1 specifically includes: Obtain the set of code file paths in the two code project versions before and after directory migration of the target code project; Match the executable files in the code file path set to generate path-file pairs; Perform a directory hierarchy comparison on the pre-migration and post-migration paths of the file pairs corresponding to the paths, and generate path difference results; Filter the path difference results where the directory hierarchy has changed, and generate migration path pairs; The original file path of the directory in the migration path pair is marked as the original file path, and the changed path of the directory in the migration path pair is marked as the changed path, thus generating a migration record.

[0007] Optionally, S2 specifically includes: Perform path segmentation on the original file path to generate an original path segment sequence; perform path segmentation on the changed path to generate a changed path segment sequence; Hierarchical position annotation is performed on the original path segment sequence and the changed path segment sequence respectively to generate the original path hierarchical sequence and the changed path hierarchical sequence; Extract the original module path segment based on the original path hierarchy sequence, and extract the changed module path segment based on the changed path hierarchy sequence; The original module domain is generated based on the original module path segment, and the candidate module domain is generated based on the changed module path segment. The original module domain and the candidate module domain are combined to generate the module domain.

[0008] Optionally, S3 specifically includes: Extract the original file path, changed path, original module path segment, and changed module path segment from the migration record; Perform path segment alignment on the original file path and the changed path to generate common path segments and different path segments; Construct path segment channels based on common path segments, different path segments, original module path segments, and changed module path segments; Get the project adjacency status of the code file before the directory migration and the project adjacency status of the code file after the directory migration; Construct file adjacency channels based on the project adjacency status before and after the directory migration. The path segment channel and the file adjacency channel are concatenated to generate the input sample.

[0009] Optionally, S4 specifically includes: Convert the path segment channels in the input sample into path segment tag sequences, and convert the file adjacency channels in the input sample into file adjacency tag sequences; Map the original module path segment, changed module path segment, common path segment, and different path segment in the path segment tagging sequence to path segment tag vectors respectively. Map the project adjacency state before directory migration and the project adjacency state after directory migration in the file adjacency tagging sequence to file adjacency tag vectors respectively. The path segment marker vector and the file adjacency marker vector are input into the migration representation module. The migration representation module configures segment embedding and position embedding for the path segment marker vector and the file adjacency marker vector, and arranges them according to the corresponding order of the code files to generate a sequence of migration input vectors. The transfer representation module inputs the transfer input vector sequence into the Transformer encoding layer, performs multi-head attention encoding and feedforward mapping on the vectors in the transfer input vector sequence, and generates a transfer encoding sequence. The migration coding sequence is input into the association attention module, and a path denaming reconstruction mechanism is introduced to convert the directory name expression in the migration coding sequence into a path role expression, generating a denaming migration coding sequence. An attention visibility matrix is ​​constructed based on the denaming migration coding sequence, and graph-guided attention encoding is performed on the path role expression and file adjacency expression in the denaming migration coding sequence to generate a migration association representation. The migration association representation is input to the attribution calibration module, and a counterfactual path verification mechanism is introduced. Based on the migration association representation, true branch input and counterfactual branch input are constructed. The true branch input and counterfactual branch input are respectively input into the attribution mapping layer with shared parameters to generate true branch attribution output and counterfactual branch attribution output. The true branch attribution output and counterfactual branch attribution output are compared to generate the attribution calibration result.

[0010] Optionally, the path denaming and reconstruction mechanism specifically includes: Convert the original module path segment markers, modified module path segment markers, common path segment markers, and different path segment markers in the path segment marker sequence into original module role markers, candidate module role markers, common path role markers, and different path role markers, respectively, to generate a path role sequence; Based on the coding position of the path role sequence in the migration coding sequence, extract the coding representation of the corresponding coding position to generate the path role coding sequence; The path role encoding sequence and the encoding representation corresponding to the file adjacency tag are concatenated according to the corresponding order of the code files to generate the path adjacency encoding sequence; An attention visibility matrix is ​​generated based on the path role marker, file adjacency marker, and module domain marker in the path adjacency coding sequence. The coding position corresponding to the directory name marker is the blocking position, and the coding positions corresponding to the path role marker, file adjacency marker, and module domain marker are the connected positions. Graph-guided attention encoding is performed on the path adjacency encoding sequence based on the attention visibility matrix to generate a migration association representation.

[0011] Optionally, the counterfactual path verification mechanism specifically includes: Extract the original module role representation, candidate module role representation, common path role representation, difference path role representation and file adjacency representation from the migration association representation, and concatenate them in the order of their encoding positions to generate the real branch input; Replace the candidate module role representations in the true branch input with non-candidate module role representations, while keeping the original module role representations, common path role representations, different path role representations, and file adjacency representations unchanged, to generate counterfactual branch inputs; The true branch input is fed into the shared parameter mapping layer to generate the true branch output; The counterfactual branch input is fed into the attribution mapping layer with shared parameters to generate the counterfactual branch attribution output. If the response value of the candidate module domain in the true branch attribution output is greater than the response value of the original module domain, and the response value of the non-candidate module domain in the counterfactual branch attribution output is greater than the response value of the candidate module domain, then a path traction identifier is generated. If the response value of the candidate module domain in the true branch attribution output is greater than the response value of the original module domain, and the response value of the non-candidate module domain in the counterfactual branch attribution output is not greater than the response value of the candidate module domain, then a candidate access identifier is generated. If the original module domain response value in the true branch attribution output is not less than the candidate module domain response value, and the original module domain response value in the counterfactual branch attribution output is not less than the non-candidate module domain response value, then an original domain continuation identifier is generated. The attribution calibration result is generated based on the path traction identifier, candidate access identifier, and original domain continuation identifier.

[0012] Optionally, S5 specifically includes: Determine the identifier type corresponding to the attribution calibration result; If the calibration result corresponds to the original domain continuation identifier, then the original module domain is determined as the writing domain, and the original domain writing instruction is generated based on the code file after directory migration and the original module domain. If the attribution calibration result corresponds to the candidate access identifier, then the candidate module domain is determined as the write graph domain, and a candidate write instruction is generated based on the code file after directory migration and the candidate module domain. If the path traction identifier corresponds to the calibration result, the temporary storage area will be determined as the writing area, and an isolation write instruction will be generated based on the code file after the directory migration and the temporary storage area. Generate graph domain write instructions based on the original domain write instructions, candidate write instructions, or isolated write instructions.

[0013] Optionally, S6 specifically includes: Module domain nodes are constructed based on module domains, and code file nodes and initial adjacency paths are generated based on file adjacency channels; Construct a code file adjacency graph from module domain nodes, code file nodes, and initial adjacency paths; If the graph domain write instruction is the original domain write instruction, then the code file node corresponding to the code file after the directory migration is attached to the module domain node corresponding to the original module domain, and the original module domain mark is written into the corresponding initial adjacency path. If the graph domain write instruction is a candidate write instruction, then the code file node corresponding to the code file after the directory migration is attached to the module domain node corresponding to the candidate module domain, and the candidate module domain mark is written into the corresponding initial adjacency path. If the graph domain write instruction is an isolated write instruction, then the code file node corresponding to the code file after the directory migration will be attached to the module domain node corresponding to the temporary storage domain, and the associated output path between the code file node corresponding to the code file after the directory migration and the module domain node corresponding to the candidate module domain will be blocked. Based on the attached code file nodes, the initial adjacency paths marked in the module field, and the associated output paths after blocking, an updated code file adjacency graph is generated.

[0014] Optionally, S7 specifically includes: In the updated code file adjacency graph, adjacency paths within the graph domain are extracted according to the module domain marker; Generate associated code files based on the code file nodes at both ends of adjacent paths within the graph domain; Stop generating associated code files for the associated output paths corresponding to the isolated write command; The association analysis results are generated based on the associated code file, the adjacency path within the graph domain, and the module domain markers.

[0015] The beneficial effects of this invention are: This invention performs path hierarchy parsing based on the original file path and the changed path, extracts the original module path segment and the changed module path segment, generates a module domain, and generates input samples by combining the path segment channel and the file adjacency channel, so that the path changes before and after the directory migration and the project adjacency status can participate in the analysis together, thereby improving the accuracy of module domain judgment in the directory migration scenario. By improving the GraphCodeBERT model, a path denaming and reconstruction mechanism is introduced in the association attention module to convert the directory name expression into a path role expression, so that the directory name in the changed path cannot directly control the attribution judgment. In addition, a counterfactual path verification mechanism is introduced in the attribution calibration module to generate attribution calibration results by comparing the true branch attribution output and the counterfactual branch attribution output, thereby reducing the risk of candidate module domains being led by path names. Based on the attribution calibration results, the write domain corresponding to the code file after directory migration is determined, and a domain write instruction is generated. A code file adjacency graph is constructed based on the module domain and file adjacency channel. The domain is then updated according to the domain write instruction. This ensures that the correlation analysis results obtained from the domain adjacency resolution can reflect the continuity status between the original file path, the changed path, the module domain, and the attribution calibration results, reducing misconnections and broken links in the code file adjacency graph and improving the reliability of code review, refactoring impact analysis, and module attribution verification. Attached Figure Description

[0016] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings: Figure 1 This is a flowchart of a deep learning-based code file correlation analysis method proposed in this invention; Figure 2 This is a schematic diagram of the improved GraphCodeBERT model proposed in this invention; Figure 3 This is a schematic diagram of the code file adjacency graph update proposed in this invention. Detailed Implementation

[0017] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.

[0018] refer to Figures 1-3 A deep learning-based method for analyzing code file correlation includes the following steps: S1. Use the code file path before the directory migration in the target code project as the original file path, and the code file path after the directory migration in the target code project as the change path, and generate a migration record. S2. Perform path hierarchy parsing based on the original file path and the changed path, extract the original module path segment and the changed module path segment, and generate the module domain; S3. Construct path segment channels based on migration records, construct file adjacency channels based on the project adjacency status of code files before and after directory migration, and combine path segment channels and file adjacency channels to generate input samples. S4. Input the input samples into the improved GraphCodeBERT model to generate attribution calibration results; the improved GraphCodeBERT model introduces a path denaming and reconstruction mechanism in the association attention module and a counterfactual path verification mechanism in the attribution calibration module; S5. Determine the writing domain corresponding to the code file after the directory migration based on the attribution calibration results, and generate the domain writing instruction. S6. Construct a code file adjacency graph based on module domains and file adjacency channels, and perform a graph domain update on the code file adjacency graph according to the graph domain write instruction to generate the updated code file adjacency graph; S7. Perform graph domain adjacency resolution based on the updated code file adjacency graph to generate association analysis results.

[0019] In this embodiment, S1 specifically refers to: Obtain the code project version before directory migration and the code project version after directory migration for the target code project; Iterate through the root directories of the two code project versions respectively, extract the relative paths of the code files, and generate a set of code file paths before the directory migration and a set of code file paths after the directory migration. Read the filenames, file extensions, and file contents of code files from the set of code file paths; normalize the newline characters, whitespace characters, and comment delimiters in the file contents to generate normalized code content; hash the normalized code content to generate a file content digest; The filename, file extension, and file content summary are used as the basis for matching code files. The code files in the set of code file paths before and after the directory migration are matched to generate path-to-file pairs. The path-to-file pairs include the path before and the path after the directory migration. The paths before and after directory migration in the path-corresponding file pairs are divided into path segment sequences according to the path separator. The path segment sequence corresponding to the path before directory migration and the path segment sequence corresponding to the path after directory migration are compared layer by layer to generate path difference results. The path difference results include common path segments, different path segments, the number of path levels before directory migration and the number of path levels after directory migration. Filter out path difference results where the difference path segment is not empty, and determine the path-file pairs corresponding to the path difference results where the difference path segment is not empty as migration path pairs. Mark the original file path of the directory in the migration path pair as the original file path; Write the original file path, the changed path, and the path difference results into the same record to generate a migration record.

[0020] In this embodiment, S2 specifically refers to: Delete the filename field in the original file path, but keep the directory field; delete the filename field in the changed path, but keep the directory field. The original file path is divided into directory segments according to the path separator, and an original path segment sequence is generated according to the order from the project root directory to the directory where the code file is located. The modified path is divided into directory segments according to the path separator, and a modified path segment sequence is generated according to the order from the project root directory to the directory where the code file is located. According to the order of the path segments in the original path segment sequence, a hierarchical number is assigned to each path segment in the original path segment sequence to generate an original path hierarchical sequence. According to the order of the path segments in the modified path segment sequence, a hierarchical number is assigned to each path segment in the modified path segment sequence to generate a modified path hierarchical sequence. Path segments with the same level number and name in both the original and changed path hierarchy sequences are retained to generate common path segments. Path segments with different names that follow the common path segments in both sequences are extracted to generate differing path segments. The first directory segment closest to the common path segment is extracted from the differing path segments in the original path hierarchy sequence to generate the original module path segment. The first directory segment closest to the common path segment is extracted from the differing path segments in the changed path hierarchy sequence to generate the changed module path segment. Use the original module path segment as the domain identifier of the original module domain, and write the code file paths containing the original module path segment into the original module domain of the code file path set corresponding to the directory before migration; use the changed module path segment as the domain identifier of the candidate module domain, and write the code file paths containing the changed module path segment into the candidate module domain of the code file path set corresponding to the directory after migration; associate the original module domain and the candidate module domain according to the same migration record to generate a module domain.

[0021] In this embodiment, S3 specifically refers to: The original file path and the modified path are split according to the path separator. The file name segments in the original file path and the modified path are deleted to obtain the original path segment sequence and the modified path segment sequence. According to the order from the project root directory to the directory where the code files are located, the original path segment sequence and the changed path segment sequence are compared layer by layer, and the path segments with the same level position and the same path segment name are identified as common path segments. The path segments in the original path segment sequence and the changed path segment sequence that are located after the common path segment and have different path segment names are identified as the differential path segments. The common path segments, the different path segments, the original module path segments, and the changed module path segments are labeled according to the path segment type to form the path segment labeling results; Based on the order of the path segment annotation results in the original file path and the changed path, the common path segment, the difference path segment, the original module path segment, and the changed module path segment are arranged into path segment channels; Retrieve the code file adjacency data in the code project version before directory migration. The code file adjacency data includes reference adjacency, call adjacency, and same-directory adjacency between code files. Retrieve the code file adjacency data in the code project version after directory migration. The code file adjacency data includes reference adjacency, call adjacency, and same-directory adjacency between code files. The adjacency data corresponding to the original file path in the code file adjacency data before the directory migration is determined as the project adjacency status before the directory migration; the adjacency data corresponding to the changed path in the code file adjacency data after the directory migration is determined as the project adjacency status after the directory migration. Based on the code file correspondence, the project adjacency status before the directory migration and the project adjacency status after the directory migration are arranged into file adjacency channels. Place the path segment channel at the path input position of the input sample, place the file adjacency channel at the adjacency input position of the input sample, and set a separator between the path segment channel and the file adjacency channel to generate the input sample.

[0022] In this embodiment, S4 specifically refers to: Convert the path segment channels in the input sample into path segment tag sequences, and convert the file adjacency channels in the input sample into file adjacency tag sequences; Map the original module path segment, changed module path segment, common path segment, and different path segment in the path segment tagging sequence to path segment tag vectors respectively. Map the project adjacency state before directory migration and the project adjacency state after directory migration in the file adjacency tagging sequence to file adjacency tag vectors respectively. The path segment marker vector and the file adjacency marker vector are input into the migration representation module. The migration representation module configures segment embedding and position embedding for the path segment marker vector and the file adjacency marker vector, and arranges them according to the corresponding order of the code files to generate a sequence of migration input vectors. The transfer representation module inputs the transfer input vector sequence into the Transformer encoding layer, performs multi-head attention encoding and feedforward mapping on the vectors in the transfer input vector sequence, and generates a transfer encoding sequence. The migration coding sequence is input into the association attention module, and a path denaming reconstruction mechanism is introduced to convert the directory name expression in the migration coding sequence into a path role expression, generating a denaming migration coding sequence. An attention visibility matrix is ​​constructed based on the denaming migration coding sequence, and graph-guided attention encoding is performed on the path role expression and file adjacency expression in the denaming migration coding sequence to generate a migration association representation. The migration association representation is input to the attribution calibration module, and a counterfactual path verification mechanism is introduced. Based on the migration association representation, true branch input and counterfactual branch input are constructed. The true branch input and counterfactual branch input are respectively input into the attribution mapping layer with shared parameters to generate true branch attribution output and counterfactual branch attribution output. The true branch attribution output and counterfactual branch attribution output are compared to generate the attribution calibration result.

[0023] In this embodiment, the GraphCodeBERT model is improved by inheriting the Transformer bidirectional encoding structure, graph guided mask attention structure, and code structure representation capability of the GraphCodeBERT model. It retains the basic framework of the GraphCodeBERT model for jointly encoding code sequences and structural information. The input expression mode of the original GraphCodeBERT model, which is oriented towards source code tags, variable nodes, and data flow edges, is adjusted to a joint input expression mode of path segment channels, file adjacency channels, and module domain tags, which is oriented towards directory migration scenarios. This expands the model input from ordinary code semantic representation to a joint representation of path migration semantics and module adjacency semantics. The improved GraphCodeBERT model includes a transfer representation module, an association attention module, and a home calibration module. The transfer representation module inherits the input embedding and Transformer encoding layer processing methods from the GraphCodeBERT model, converting path segment channels into path segment label sequences and file adjacency channels into file adjacency label sequences. It also configures segment embeddings and position embeddings for the path segment label sequences and file adjacency label sequences to generate a transfer input vector sequence. Unlike the original GraphCodeBERT model, which primarily encodes source code labels and data flow nodes, the transfer representation module incorporates original module path segments, changed module path segments, common path segments, differing path segments, and the project adjacency state before and after directory migration into the same encoding space, enabling a unified representation of path changes and project adjacency changes before and after directory migration. The Associative Attention module inherits the encoding idea of ​​graph-guided mask attention in the GraphCodeBERT model. However, the graph-guided mask attention in the original GraphCodeBERT model mainly restricts attention interaction based on the data flow edges between variables. The Associative Attention module adjusts the attention constraint objects to path role labels, file adjacency labels, and module domain labels. It also introduces a path denaming and reconstruction mechanism in the Associative Attention module, converting the directory name expression into the original module role label, candidate module role label, common path role label, and different path role label. This prevents the specific directory name in the changed path from directly entering the attribution determination path, reducing the direct influence of the directory name on the module attribution determination. The attribution calibration module is a task-oriented output modification based on the general representation output of the GraphCodeBERT model. The original GraphCodeBERT model typically outputs code semantic representations and hands them over to downstream tasks for processing. The attribution calibration module directly judges the attribution calibration results of the module domain after directory migration and introduces a counterfactual path verification mechanism in the attribution calibration module. By comparing the attribution outputs of the real branch input and the counterfactual branch input, it determines whether the attribution of the candidate module domain is generated by the change of the module path segment. If the counterfactual branch output moves to the non-candidate module domain along with the non-candidate module role representation, a path traction identifier is generated. If the real branch output stably points to the candidate module domain and the counterfactual branch output does not move to the non-candidate module domain, a candidate access identifier is generated. If both the real branch output and the counterfactual branch output maintain the original module domain response, an original domain continuation identifier is generated. Through the above improvements, the GraphCodeBERT model, while inheriting the deep representation of code structure information, shifts its processing objective from code semantic understanding to module domain affiliation calibration in directory migration scenarios. The migration representation module enhances the joint expression capability of path migration information and project adjacency information, the association attention module reduces the direct influence of changing path directory names on affiliation judgment, and the affiliation calibration module identifies path traction results through counterfactual path verification, enabling code files after directory migration to be determined as continuing the original domain, candidate access, or path traction isolation states, thereby reducing the distortion of association analysis caused by directly judging module affiliation based on changing paths.

[0024] In this embodiment, the path denaming and reconstruction mechanism is specifically as follows: Replace the original module path segment markers in the path segment marker sequence with the original module role markers, replace the changed module path segment markers in the path segment marker sequence with the candidate module role markers, replace the common path segment markers in the path segment marker sequence with the common path role markers, and replace the different path segment markers in the path segment marker sequence with the different path role markers to generate a path role sequence. The encoding position of each path role marker in the path role sequence is retained in the migration encoding sequence. The corresponding encoding representation is extracted from the migration encoding sequence according to the encoding position to generate the path role encoding sequence. The encoding representations in the path role encoding sequence are saved in the order of original module role marker, candidate module role marker, common path role marker and different path role marker, and the path level position corresponding to each encoding representation is retained. The encoded representations corresponding to the file adjacency tags are arranged in the order of the code files after the path role encoded sequence to generate the path adjacency encoded sequence. Each encoded representation in the path adjacency encoded sequence is configured with a position number and a tag type. The tag types include path role tags, file adjacency tags, module domain tags, and directory name tags. An attention visibility matrix is ​​constructed based on the position number of the path adjacency coding sequence. The row and column positions of the attention visibility matrix correspond to the coding representation positions in the path adjacency coding sequence, respectively. Set the row and column positions corresponding to the directory name marker as blocking positions, and set the matrix elements corresponding to the blocking positions to 0; set the row and column positions corresponding to the path role marker, file adjacency marker, and module domain marker as connected positions, and set the matrix elements corresponding to the connected positions to 1. For any target encoding representation in the path adjacency encoding sequence, retain the connected position encoding representation corresponding to the target encoding representation in the attention visibility matrix, and mask the blocking position encoding representation corresponding to the target encoding representation in the attention visibility matrix; The values ​​of the target encoding representation and the connected position encoding representation in the same dimension are multiplied one dimension at a time, and the values ​​after multiplication are accumulated to obtain the correlation value between the target encoding representation and the connected position encoding representation. The correlation values ​​between the target encoding representation and the encoding representations of all connected positions are normalized to obtain the normalized correlation values ​​corresponding to the target encoding representation. The normalized correlation values ​​are multiplied dimension-by-dimensionally by the corresponding connected position encoding representations, and the resulting encoding representations are accumulated dimension-by-dimensionally to obtain the updated encoding representation corresponding to the target encoding representation. The updated encoding representations corresponding to all encoding representations in the path adjacency encoding sequence are rearranged in their original order to generate a migration association representation.

[0025] In this embodiment, the counterfactual path verification mechanism is specifically as follows: Extract the original module role representation, candidate module role representation, common path role representation, difference path role representation, and file adjacency representation from the migration association representation; Arrange the original module role representation, candidate module role representation, common path role representation, difference path role representation and file adjacency representation in the order of their encoding positions in the migration association representation, connect the first and last of the adjacent representations to generate the real branch input; Select a module domain that is different from the candidate module domain from the module domain as a non-candidate module domain, and convert the module path segment corresponding to the non-candidate module domain into a non-candidate module role representation; Replace the candidate module role representations in the real branch input with non-candidate module role representations, while retaining the original module role representations, common path role representations, different path role representations, and file adjacency representations in the real branch input to generate counterfactual branch inputs; The attribution mapping layer is configured with the same set of mapping parameters, which are applied to the true branch input and the counterfactual branch input respectively. The true branch input is assigned to the mapping layer, the representation vector and mapping parameters in the true branch input are multiplied one dimension at a time, and the results of the one-dimensional multiplication are accumulated to generate the true branch assigned output. The counterfactual branch input is fed into the mapping layer. The representation vector in the counterfactual branch input is multiplied dimension by dimension with the same set of mapping parameters. The results of the dimension-by-dimensional multiplication are accumulated to generate the counterfactual branch output. The true branch attribution output includes the original module domain response value and the candidate module domain response value. The original module domain response value represents the output value of the original module domain pointed to by the true branch input, and the candidate module domain response value represents the output value of the candidate module domain pointed to by the true branch input. The counterfactual branch attribution output includes the original module domain response value, the candidate module domain response value, and the non-candidate module domain response value. The non-candidate module domain response value represents the output value of the counterfactual branch input pointing to the non-candidate module domain. If the response value of the candidate module domain in the true branch attribution output is greater than the response value of the original module domain, and the response value of the non-candidate module domain in the counterfactual branch attribution output is greater than the response value of the candidate module domain, then the module domain output redirection caused by the directory name replacement will be marked as a path traction identifier. If the response value of the candidate module domain in the true branch attribution output is greater than the response value of the original module domain, and the response value of the non-candidate module domain in the counterfactual branch attribution output is not greater than the response value of the candidate module domain, then the stable state of the candidate module domain output is marked as the candidate access identifier. If the response value of the original module domain in the true branch attribution output is not less than the response value of the candidate module domain, and the response value of the original module domain in the counterfactual branch attribution output is not less than the response value of the non-candidate module domain, then the original module domain output is marked as the original domain continuation identifier. The generated results from the path traction identifier, candidate access identifier, and original domain continuation identifier are written into the attribution calibration results.

[0026] In this embodiment, S5 specifically refers to: Obtain the identifier type from the attribution calibration results. The identifier types include original domain continuation identifier, candidate access identifier, and path traction identifier. Establish a corresponding item between the code file after directory migration and the identifier type in the attribution calibration results, and generate a code file write item. If the code file write item corresponds to the original domain continuation identifier, then the original module domain is marked as the write graph domain, and the code file after directory migration, the original module domain, and the original domain continuation identifier are written into the same instruction field to generate the original domain write instruction; If the code file write item corresponds to the candidate access identifier, then the candidate module field is marked as the write graph field, and the code file after directory migration, the candidate module field, and the candidate access identifier are written into the same instruction field to generate a candidate write instruction; If the code file write item corresponds to the path traction identifier, then the temporary field is marked as a write field, and the code file after directory migration, the temporary field, and the path traction identifier are written to the same instruction field to generate an isolated write instruction. The original domain write instruction includes the code file identifier after directory migration, the original module domain identifier, and the original domain continuation identifier; the candidate write instruction includes the code file identifier after directory migration, the candidate module domain identifier, and the candidate access identifier; the isolated write instruction includes the code file identifier after directory migration, the temporary storage domain identifier, and the path traction identifier; the generated original domain write instruction, candidate write instruction, or isolated write instruction is used as the graph domain write instruction.

[0027] In this embodiment, S6 specifically refers to: Module domain nodes are generated based on the original module domain and candidate module domain in the module domain, and temporary domain nodes are generated based on the temporary domain. Based on the project adjacency status before and after directory migration in the file adjacency channel, extract the code file identifier, adjacent code file identifier, and adjacency type identifier. Convert the code file identifier and adjacent code file identifier into code file nodes respectively; Based on the correspondence between code file identifiers and adjacent code file identifiers, connect code file nodes and adjacent code file nodes to generate initial adjacency paths; Write the module domain node, temporary storage node, code file node, and initial adjacency path into the same graph structure to generate the code file adjacency graph; The graph domain to be written is determined according to the graph domain write instruction. The graph domain to be written is the original module domain, the candidate module domain, or the temporary domain. If the graph domain write instruction is the original domain write instruction, then connect the code file node corresponding to the code file after the directory migration to the module domain node corresponding to the original module domain, and write the original module domain mark into the initial adjacency path corresponding to the code file after the directory migration. If the graph domain write instruction is a candidate write instruction, then the code file node corresponding to the code file after the directory migration is connected to the module domain node corresponding to the candidate module domain, and the candidate module domain mark is written into the initial adjacency path corresponding to the code file after the directory migration. If the graph domain write instruction is an isolated write instruction, then the code file node corresponding to the code file after the directory migration is connected to the temporary domain node, and the temporary domain mark is written into the initial adjacency path corresponding to the code file after the directory migration. If the graph domain write instruction is an isolated write instruction, then delete the associated output path between the code file node corresponding to the code file after directory migration and the module domain node corresponding to the candidate module domain. Update the initial adjacency paths written to the module domain tag to graph domain adjacency paths; The connected module domain nodes, temporary domain nodes, code file nodes, graph domain adjacency paths, and the graph structure after deleting associated output paths are used to generate an updated code file adjacency graph.

[0028] In this embodiment, S7 specifically refers to: Traverse the updated code file adjacency graph, including code file nodes, module domain nodes, graph domain adjacency paths, and module domain markers; based on the module domain markers, divide the graph domain adjacency paths into adjacency paths within the original module domain, adjacency paths within the candidate module domain, and adjacency paths within the temporary domain. Extract the code file nodes connecting both ends of the adjacent paths within the original module domain to generate the associated code file corresponding to the original module domain; extract the code file nodes connecting both ends of the adjacent paths within the candidate module domain to generate the associated code file corresponding to the candidate module domain. Identify the associated output paths corresponding to the isolation write command in the adjacent paths within the temporary storage domain, and mark the associated output paths corresponding to the isolation write command as blocked; remove the code file nodes connected to both ends of the associated output paths corresponding to the blocked state, and retain the graph domain adjacent paths that are not in the blocked state; generate associated code files to participate in the correlation analysis based on the code file nodes connected to both ends of the retained graph domain adjacent paths. The association analysis results are generated by combining the associated code files, graph domain adjacency paths, module domain tags, original file paths, and changed paths.

[0029] Example 1: To verify the feasibility of this invention in implementation, it was applied to a code refactoring scenario of an enterprise-level business platform. This platform adopts a microservice architecture and includes six business modules: transaction, settlement, account, risk control, reporting, and public capabilities. The code repository contains 2360 Java, XML, and configuration class code files. This implementation selected a directory specification adjustment as the test object. The development team migrated settlement-related code files from the " / trade / order / settle / " directory to the " / finance / settlement / application / " directory, while retaining some file calls and references between the transaction module and other modules. Traditional module attribution judgment based on path changes would directly classify the migrated files into the financial settlement module, leading to the omission of previously associated files within the transaction module and the incorrect inclusion of some files that only underwent directory reorganization into the candidate module domain.

[0030] In implementation, the system first obtains the set of code file paths in the two code project versions before and after directory migration. It then matches the executable files of the code files to obtain migration path pairs, marking the paths before directory migration as original file paths and the paths after migration as changed paths, generating migration records. Based on the original file paths and changed paths, the system extracts original module path segments and changed module path segments, generating original module domains and candidate module domains. Subsequently, the system constructs path segment channels based on the migration records and file adjacency channels based on the project adjacency states before and after directory migration. These two channels are combined as input samples and fed into the improved GraphCodeBERT model. The model uses a path denaming and reconstruction mechanism to convert directory name representations into path role representations and a counterfactual path verification mechanism to determine whether candidate module domains exhibit path traction. Based on the attribution calibration results, the system generates graph domain write instructions, updates the code file adjacency graph, and generates correlation analysis results through graph domain adjacency resolution.

[0031] Table 1 Comparison of Module Domain Attribution Determination Results

[0032] As shown in Table 1, the accuracy of traditional path determination in all four test objects did not reach 71%. The main reason is that changing the candidate module name in the path directly affects the module ownership determination, causing directory reorganization migrations to be misjudged as genuine candidate accesses. This invention inputs both the path segment channel and the file adjacency channel into the improved GraphCodeBERT model, and weakens the direct influence of directory names on ownership determination through a path denaming and reconstruction mechanism. It also identifies path traction in candidate module domain ownership through a counterfactual path verification mechanism. Test results show that the accuracy of this invention exceeds 90% in all four test objects, with a reduction in misjudgments of 31, 27, 32, and 21, respectively, indicating that this invention can more stably distinguish between original domain continuation, candidate access, and path traction states.

[0033] Table 2 Comparison of Correlation Analysis Results

[0034] As shown in Table 2, traditional methods exhibit significant issues with incorrect connections and broken links in code file association analysis. Taking the settlement service file group as an example, with 286 manually confirmed associations, the traditional method identified 49 incorrect connections and 31 broken links, indicating that some candidate module files were over-connected after path migration, while some real adjacency paths within the original module domain were severed. This invention generates graph domain writing instructions based on the attribution calibration results and performs graph domain updates on the code file adjacency graph. This ensures that files continuing from the original domain remain in the original module domain, candidate access files enter the candidate module domain, and path-driven files enter the isolation processing flow. After graph domain adjacency resolution, the number of incorrect connections and broken links in the four analyzed objects is significantly reduced, and the review time is reduced by 36.9% to 41.2%, improving the efficiency of code review and refactoring impact scope analysis.

[0035] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A deep learning-based code file correlation analysis method, characterized in that, Includes the following steps: S1. Use the code file path before the directory migration in the target code project as the original file path, and the code file path after the directory migration in the target code project as the change path, and generate a migration record. S2. Perform path hierarchy parsing based on the original file path and the changed path, extract the original module path segment and the changed module path segment, and generate the module domain; S3. Construct path segment channels based on migration records, construct file adjacency channels based on the project adjacency status of code files before and after directory migration, and combine path segment channels and file adjacency channels to generate input samples. S4. Input the input sample into the improved GraphCodeBERT model to generate the attribution calibration result; the improved GraphCodeBERT model introduces a path denaming and reconstruction mechanism in the association attention module and a counterfactual path verification mechanism in the attribution calibration module; S5. Determine the writing domain corresponding to the code file after the directory migration based on the attribution calibration results, and generate the domain writing instruction. S6. Construct a code file adjacency graph based on module domains and file adjacency channels, and perform a graph domain update on the code file adjacency graph according to the graph domain write instruction to generate the updated code file adjacency graph; S7. Perform graph domain adjacency resolution based on the updated code file adjacency graph to generate association analysis results.

2. The code file correlation analysis method based on deep learning according to claim 1, characterized in that, Specifically, S1 is: Obtain the set of code file paths in the two code project versions before and after directory migration of the target code project; Match the executable files in the code file path set to generate path-file pairs; Perform a directory hierarchy comparison on the pre-migration and post-migration paths of the file pairs corresponding to the paths, and generate path difference results; Filter the path difference results where the directory hierarchy has changed, and generate migration path pairs; The original file path of the directory in the migration path pair is marked as the original file path, and the changed path of the directory in the migration path pair is marked as the changed path, thus generating a migration record.

3. The method for code file correlation analysis based on deep learning according to claim 1, characterized in that, Specifically, S2 is: Perform path segmentation on the original file path to generate an original path segment sequence; perform path segmentation on the changed path to generate a changed path segment sequence; Hierarchical position annotation is performed on the original path segment sequence and the changed path segment sequence respectively to generate the original path hierarchical sequence and the changed path hierarchical sequence; Extract the original module path segment based on the original path hierarchy sequence, and extract the changed module path segment based on the changed path hierarchy sequence; The original module domain is generated based on the original module path segment, and the candidate module domain is generated based on the changed module path segment. The original module domain and the candidate module domain are combined to generate the module domain.

4. The code file correlation analysis method based on deep learning according to claim 1, characterized in that, Specifically, S3 is: Extract the original file path, changed path, original module path segment, and changed module path segment from the migration record; Perform path segment alignment on the original file path and the changed path to generate common path segments and different path segments; Construct path segment channels based on common path segments, different path segments, original module path segments, and changed module path segments; Get the project adjacency status of the code file before the directory migration and the project adjacency status of the code file after the directory migration; Construct file adjacency channels based on the project adjacency status before and after the directory migration. The path segment channel and the file adjacency channel are concatenated to generate the input sample.

5. The method for code file correlation analysis based on deep learning according to claim 1, characterized in that, Specifically, S4 is: Convert the path segment channels in the input sample into path segment tag sequences, and convert the file adjacency channels in the input sample into file adjacency tag sequences; Map the original module path segment, changed module path segment, common path segment, and different path segment in the path segment tagging sequence to path segment tag vectors respectively. Map the project adjacency state before directory migration and the project adjacency state after directory migration in the file adjacency tagging sequence to file adjacency tag vectors respectively. The path segment marker vector and the file adjacency marker vector are input into the migration representation module. The migration representation module configures segment embedding and position embedding for the path segment marker vector and the file adjacency marker vector, and arranges them according to the corresponding order of the code files to generate a sequence of migration input vectors. The transfer representation module inputs the transfer input vector sequence into the Transformer encoding layer, performs multi-head attention encoding and feedforward mapping on the vectors in the transfer input vector sequence, and generates a transfer encoding sequence. Input the migration coding sequence into the association attention module, introduce the path denaming reconstruction mechanism, convert the directory name expression in the migration coding sequence into the path role expression, and generate the denaming migration coding sequence; An attention visibility matrix is ​​constructed based on the nameless transfer coding sequence. Graph-guided attention coding is then performed on the path role representation and file adjacency representation in the nameless transfer coding sequence to generate a transfer association representation. The migration association representation input attribution calibration module introduces a counterfactual path verification mechanism, constructs true branch inputs and counterfactual branch inputs based on the migration association representation, and inputs the true branch inputs and counterfactual branch inputs into the attribution mapping layer with shared parameters to generate true branch attribution outputs and counterfactual branch attribution outputs. The true branch attribution output and the counterfactual branch attribution output are compared to generate attribution calibration results.

6. The code file correlation analysis method based on deep learning according to claim 5, characterized in that, The path denaming and reconstruction mechanism is as follows: Convert the original module path segment markers, modified module path segment markers, common path segment markers, and different path segment markers in the path segment marker sequence into original module role markers, candidate module role markers, common path role markers, and different path role markers, respectively, to generate a path role sequence; Based on the coding position of the path role sequence in the migration coding sequence, extract the coding representation of the corresponding coding position to generate the path role coding sequence; The path role encoding sequence and the encoding representation corresponding to the file adjacency tag are concatenated according to the corresponding order of the code files to generate the path adjacency encoding sequence; An attention visibility matrix is ​​generated based on the path role marker, file adjacency marker, and module domain marker in the path adjacency coding sequence. The coding position corresponding to the directory name marker is the blocking position, and the coding positions corresponding to the path role marker, file adjacency marker, and module domain marker are the connected positions. Graph-guided attention encoding is performed on the path adjacency encoding sequence based on the attention visibility matrix to generate a migration association representation.

7. The method for code file correlation analysis based on deep learning according to claim 5, characterized in that, The counterfactual path verification mechanism is as follows: Extract the original module role representation, candidate module role representation, common path role representation, difference path role representation and file adjacency representation from the migration association representation, and concatenate them in the order of their encoding positions to generate the real branch input; Replace the candidate module role representations in the true branch input with non-candidate module role representations, while keeping the original module role representations, common path role representations, different path role representations, and file adjacency representations unchanged, to generate counterfactual branch inputs; The true branch input is fed into the shared parameter mapping layer to generate the true branch output; The counterfactual branch input is fed into the attribution mapping layer with shared parameters to generate the counterfactual branch attribution output. If the response value of the candidate module domain in the true branch attribution output is greater than the response value of the original module domain, and the response value of the non-candidate module domain in the counterfactual branch attribution output is greater than the response value of the candidate module domain, then a path traction identifier is generated. If the response value of the candidate module domain in the true branch attribution output is greater than the response value of the original module domain, and the response value of the non-candidate module domain in the counterfactual branch attribution output is not greater than the response value of the candidate module domain, then a candidate access identifier is generated. If the original module domain response value in the true branch attribution output is not less than the candidate module domain response value, and the original module domain response value in the counterfactual branch attribution output is not less than the non-candidate module domain response value, then an original domain continuation identifier is generated. The attribution calibration result is generated based on the path traction identifier, candidate access identifier, and original domain continuation identifier.

8. The method for code file correlation analysis based on deep learning according to claim 1, characterized in that, Specifically, S5 is: Determine the identifier type corresponding to the attribution calibration result; If the calibration result corresponds to the original domain continuation identifier, then the original module domain is determined as the writing domain, and the original domain writing instruction is generated based on the code file after directory migration and the original module domain. If the attribution calibration result corresponds to the candidate access identifier, then the candidate module domain is determined as the write graph domain, and a candidate write instruction is generated based on the code file after directory migration and the candidate module domain. If the path traction identifier corresponds to the calibration result, the temporary storage area will be determined as the writing area, and an isolation write instruction will be generated based on the code file after the directory migration and the temporary storage area. Generate graph domain write instructions based on the original domain write instructions, candidate write instructions, or isolated write instructions.

9. The code file correlation analysis method based on deep learning according to claim 1, characterized in that, Specifically, S6 is: Module domain nodes are constructed based on module domains, and code file nodes and initial adjacency paths are generated based on file adjacency channels; Construct a code file adjacency graph from module domain nodes, code file nodes, and initial adjacency paths; If the graph domain write instruction is the original domain write instruction, then the code file node corresponding to the code file after the directory migration is attached to the module domain node corresponding to the original module domain, and the original module domain mark is written into the corresponding initial adjacency path. If the graph domain write instruction is a candidate write instruction, then the code file node corresponding to the code file after the directory migration is attached to the module domain node corresponding to the candidate module domain, and the candidate module domain mark is written into the corresponding initial adjacency path. If the graph domain write instruction is an isolated write instruction, then the code file node corresponding to the code file after the directory migration will be attached to the module domain node corresponding to the temporary storage domain, and the associated output path between the code file node corresponding to the code file after the directory migration and the module domain node corresponding to the candidate module domain will be blocked. Based on the attached code file nodes, the initial adjacency paths marked in the module field, and the associated output paths after blocking, an updated code file adjacency graph is generated.

10. The code file correlation analysis method based on deep learning according to claim 1, characterized in that, Specifically, S7 is: In the updated code file adjacency graph, adjacency paths within the graph domain are extracted according to the module domain marker; Generate associated code files based on the code file nodes at both ends of adjacent paths within the graph domain; Stop generating associated code files for the associated output paths corresponding to the isolated write command; The association analysis results are generated based on the associated code file, the adjacency path within the graph domain, and the module domain markers.