A programming review method and system based on multi-modal data
Patent Information
- Application Number
- CN202510877233.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-27
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2045-06-27
AI Technical Summary
虽然该方法获取了多模态的信息,但仅对多模态信息进行语义提取,得到目标语义向量,对代码本身元信息、语义关联、以及多模态信息的利用不充分,导致代码审查过程中易忽略潜在风险,审查结果仍旧过于依赖人工经验与主观判断
[0037](1)现有技术对代码元信息的处理较为简单,仅是局部的提取和查看,导致信息分散且关联性差。本发明通过采集源代码片段的多种元信息,如入口声明、变量列表和注释文本记录等,并详细记录版本编号、提交者标识和提交日期等关键信息,再进行字符格式匹配和归类,生成初始节点映射。这种全面且有组织的元信息处理方式,使得代码审查过程中能够从元信息层面建立更完整、细致的初始信息架构,为后续深入审查奠定良好基础,有效避免了以往因元信息分散而导致的审查漏洞。
Smart Images

Figure CN120804726B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of programming review technology, and in particular to a programming review method and system based on multimodal data. Background Technology
[0002] The field of code review technology primarily involves technical methods for systematically reviewing, analyzing, and evaluating software code. Its aim is to identify and resolve defects, security vulnerabilities, or non-compliance issues in program design and coding. Through a combination of automated tools and manual review, it helps developers promptly identify and correct errors in their code, improving software security, reliability, and maintainability.
[0003] CN117648931A discloses a code review method, apparatus, electronic device, and medium. The method includes: semantic extraction of multimodal data of the code to be reviewed to obtain a target semantic vector corresponding to the code to be reviewed; inputting the target semantic vector and the code to be reviewed into a code review model to obtain a target review result corresponding to the code to be reviewed. The code review model is trained on a preset language model using a set of defective code and optimized through a human feedback mechanism. This method directly obtains the target review result corresponding to the code to be reviewed by inputting the target semantic vector and the code to be reviewed into the code review model, realizing automatic review of defects in the code and improving the applicability of code review. Although this method obtains multimodal information, it only performs semantic extraction on the multimodal information to obtain the target semantic vector. It does not fully utilize the meta-information, semantic associations, and multimodal information of the code itself, which makes it easy to overlook potential risks during the code review process. The review result still relies too much on human experience and subjective judgment. In practice, the semantic relationships between reviewed content are not effectively constructed and dynamically updated, and there is a lack of correlation analysis between code metadata extraction and code logic, resulting in an information silo effect in the review process. Furthermore, the review model lacks analytical and fusion methods for unstructured information or cross-modal data, causing a large number of valuable review clues to go unidentified or be incorrectly judged. For example, when code comments conflict with actual program logic, a single-dimensional static review method cannot promptly identify such hidden risks, thereby reducing review accuracy, increasing program security vulnerabilities, and affecting software reliability and stability. Summary of the Invention
[0004] The purpose of this invention is to provide a programming review method and system based on multimodal data, which makes full use of the code's metadata, semantic associations, and multimodal information to improve the accuracy of code review.
[0005] The objective of this invention can be achieved through the following technical solutions:
[0006] A programming review method based on multimodal data includes the following steps:
[0007] Collect source code fragments, extract metadata including entry declarations, variable lists, and comment text records, record version number, committer ID, and commit date, perform character format matching, classify the matched metadata according to committer ID and commit date, output consistency entries, and generate initial node mapping;
[0008] Based on the initial node mapping, pointer references and call order are extracted, node association and reorganization are performed, and node levels are filtered by combining the preset incremental update parameter set and node relationship table. Semantic tag matching is performed on each selected node, connection paths are summarized, and the connectable relationships between nodes are output. At the same time, code security scanning indicators are checked in the connection paths, and dynamic semantic structure is generated based on the connectable relationships between nodes and the check results.
[0009] Based on the dynamic semantic structure, the semantic tag group is parsed and the corresponding text range is measured. Layered matching is performed to compare pixel distribution and tag similarity, and layered matching items are generated. Based on the layered matching items, image content tags and text description items are loaded, and edge coordinates are scanned. The corresponding positions of tags and the distribution of visible areas are checked to generate multimodal loading information. Based on the multimodal loading information, the same elements in image content tags and text description items are compared, feature co-occurrence is determined and corresponding positions are marked to generate multimodal association features.
[0010] Based on the aforementioned multimodal association features, program logic is checked, variable initialization order is parsed, a set of check elements is generated, code security scanning indicators are compared, memory management markers and permission call records are retrieved, security scanning information is generated, risk identification level is determined based on the security scanning information, and a list of review instructions is generated.
[0011] The specific steps for generating the initial node mapping are as follows:
[0012] Construct a feedforward neural network to extract node names, version ascending order, and the mapping relationship between submitter ID and submission date. The input data includes submitter ID, submission date, and version number. The output data is a two-dimensional vector that identifies the level label of a node and its corresponding mapping index. Entries with the same node position and the same submitter ID are grouped into the same level and their mapping indices are merged to generate the initial node mapping.
[0013] The specific steps of performing node association and reorganization are as follows:
[0014] Based on the initial node mapping, the pointer information for each node entry is read and compared with the pre-compiled pointer reference table. The node names are compared word by word to confirm their matching relationship in the pointer reference table. Based on the upper and lower limits of the number of nodes, it is determined whether the number of times the same name appears exceeds the preset number to confirm whether the name is reused or spelled similarly. A pointer offset threshold is defined. If the difference between the starting address of a pointer and the starting address of the corresponding called function is greater than the pointer offset threshold, the pointer information is recorded for subsequent matching to eliminate potential conflicts and count the frequency of occurrence.
[0015] After traversing all nodes and recording the corresponding pointer matching status, the call chain is sorted according to the order of the node name sequence number. If the order is disordered and crosses the threshold, the nodes in that part are marked as needing to be rearranged. The actual function call sequence is rechecked through the pointer reference reference table and a new sequence number mapping is generated. The adjusted node sequence numbers are assembled into a call order list to obtain the associated reorganization item.
[0016] The specific hierarchy of the filtering nodes is as follows:
[0017] Set a threshold for the starting address difference. Based on the node association and reorganization results, determine whether the pointer address difference of a certain node in two consecutive versions is less than the threshold for the starting address difference. If so, the node is determined to be a node prone to conflict and is marked as a conflict node in the node relationship table. For nodes marked as having conflict, compare them with the incremental update parameter set to see if there are duplicate or short-interval merging operations across versions. If the number of operations on a single node exceeds the preset number, the node is marked as a conflict node.
[0018] Perform the above steps on all nodes to filter out conflicting nodes, retain the available node levels, and summarize the conflicting nodes to form a filtered node set.
[0019] The specific method for generating dynamic semantic structures is as follows:
[0020] A pre-compiled semantic tag reference table is obtained, which includes the functional tags and key attributes of nodes. The index value of each selected node in the selected node set is matched with the index value in the semantic tag reference table, and the tag feature similarity is calculated for verification to complete the tag matching. Based on the calling order and dependency relationship of the nodes with matched semantic tags, a connection path is formed, and the connectable relationship between nodes is output. Code security scanning indicators are checked on the connection path. If the code or attribute in the connection path triggers the code security scanning indicator, the security scanning result is marked in the connection path. The nodes with matched semantic tags, the connectable relationship between nodes, and the security scanning result marking are summarized to obtain the dynamic semantic structure.
[0021] The specific method for generating the stacked matching items is as follows:
[0022] Based on the dynamic semantic structure, the text feature vector is obtained by reading the semantic tag group marked in it, and the corresponding text range is decomposed in the form of character or word boundary coordinates. The text range coordinates are further mapped to the image coordinate system to obtain the image, the pixel density value is counted, the image is input into the convolutional neural network, and the image feature vector is output.
[0023] Obtain a label reference table and extract the specific meaning and typical keywords of each label to form a label vector. Calculate the similarity between the text feature vector and the image feature vector and the label vector respectively. Perform a layered matching judgment on the two similarities. If the similarity exceeds the set benchmark value, it is considered a match and recorded in the subsequent process to generate a layered matching item.
[0024] The specific steps for generating multimodal loading information are as follows:
[0025] Based on the explicit image content labels and text description entries in the overlay matching terms, the corresponding key coordinate information is extracted. Using a basic reference table containing standard resolution edge coordinates, the four corner points of the key coordinate information are compared. Based on the regions determined by the corner points, it is determined whether the visible pixel rate of the region is less than the visible region threshold. If so, it is considered that there is occlusion or edge cropping, and the region is segmented. For each region, a preset neighborhood range is searched in the horizontal and vertical directions according to the label content. If a complete labeled boundary is matched in the neighborhood range, it is confirmed that the label and text description entry are consistent. The matching coordinate information is associated with the image feature vector and summarized into the same mapping table. This mapping table is combined into multimodal loading information.
[0026] The specific steps for generating multimodal association features are as follows:
[0027] Based on the multimodal loading information, the specific element names mentioned in the image content tags and text description entries are compared one by one. If the same key terms are found, it is determined that there may be feature co-occurrence. Then, the coordinate index recorded in the mapping table is used to confirm whether the pixel distribution of their occurrence in the same area exceeds the overlap threshold. If so, it is determined that there is feature co-occurrence, and the pixel position of the element is marked to obtain the multimodal association feature.
[0028] The method further includes:
[0029] Based on the review instruction list, the task scheduling is parsed, the queue length is calculated and the asynchronous loading trigger signal is called, the review priority is superimposed for comparison, and the merged queue is output according to the operation time interval and batch synchronization rules to generate a synchronization processing package.
[0030] Based on the aforementioned synchronization processing package, integration and comparison are performed, review matching reference relationships are extracted and the number of tags is counted, execution stage identifiers and security inspection parameters are associated, a comprehensive list is output, version retrieval records are linked, and comprehensive review data is generated.
[0031] A programming review system based on multimodal data, used to implement the method, includes:
[0032] Corpus extraction module: Collects source code fragments, extracts metadata including entry declarations, variable lists, and comment text records, records version number, committer ID, and commit date, performs character format matching, categorizes the matched metadata according to committer ID and commit date, outputs consistency entries, and generates initial node mapping;
[0033] Dynamic construction module: Based on the initial node mapping, extract pointer references and call order, perform node association and reorganization, combine the preset incremental update parameter set and node relationship table, filter node levels, perform semantic tag matching on each selected node, summarize connection paths, output the connectable relationships between nodes, and at the same time check code security scanning indicators in the connection paths, and generate dynamic semantic structure based on the connectable relationships between nodes and the check results.
[0034] Multimodal analysis module: Based on the dynamic semantic structure, it parses semantic tag groups and measures the corresponding text range, performs layered matching to compare pixel distribution and tag similarity, generates layered matching items, loads image content tags and text description items based on the layered matching items, scans edge coordinates, checks the corresponding tag positions and visible area distribution, generates multimodal loading information, compares the same elements in image content tags and text description items based on the multimodal loading information, determines feature co-occurrence and marks the corresponding positions, and generates multimodal association features;
[0035] Review generation module: Based on the multimodal association features, it performs program logic verification, parses the variable initialization order, generates a set of verification elements, compares code security scanning indicators, retrieves memory management markers and permission call records, generates security scanning information, determines the risk identification level based on the security scanning information, and generates a review instruction list.
[0036] Compared with the prior art, the present invention has the following beneficial effects:
[0037] (1) Existing technologies handle code metadata in a relatively simple way, involving only partial extraction and viewing, resulting in scattered information and poor correlation. This invention collects various metadata from source code fragments, such as entry point declarations, variable lists, and comment text records, and records key information such as version number, submitter identifier, and submission date in detail. Then, it performs character format matching and classification to generate an initial node mapping. This comprehensive and organized metadata processing method enables a more complete and detailed initial information architecture to be established at the metadata level during code review, laying a good foundation for subsequent in-depth review and effectively avoiding review loopholes caused by the previous scattered metadata.
[0038] (2) In existing technologies, semantic relationships are often difficult to maintain dynamically, easily leading to outdated or missing information. This invention, based on initial node mapping, extracts pointer references and call order, and combines incrementally updated parameter sets and node relationship tables to filter node levels, perform semantic tag matching, and summarize connection paths, constructing a dynamic semantic structure and achieving high-precision semantic relationships between nodes. This can reflect the evolution of code logic over time and with modifications in real time, promptly capture newly added or changed semantic relationships, ensuring the review process is always based on the latest and most accurate semantic relationship information, reducing review bias caused by outdated semantic relationships, and improving the timeliness and accuracy of the review.
[0039] (3) In the prior art, there is a lack of effective semantic association construction and dynamic updates between different review contents, resulting in information silos. This invention organically connects the various isolated review contents through the classification of meta-information and the dynamic construction of semantic associations, so that code review can be carried out from a more macro and systematic perspective. It can more comprehensively understand the logical dependencies between different parts of the code, thereby discovering potential cross-module and cross-level hidden dangers, avoiding the review blind spots caused by the information silo effect in the past, and improving the quality of review.
[0040] (4) Existing technologies have limited capabilities in processing unstructured information and cross-modal data. This invention introduces multimodal analysis methods, further employing a semantic layering matching method based on dynamic semantic structure. It combines pixel distribution and tag similarity, integrating image content tags and text descriptions to form more complete and fine-grained multimodal association features. This series of operations can fully mine and integrate multiple modal information involved in code review, avoiding situations where conflicts between code comments and actual program logic cannot be identified. It effectively utilizes multimodal clues to improve review accuracy and reduce security risks.
[0041] (5) This invention uses program logic verification and variable initialization sequence analysis, and simultaneously uses code security scanning indicators, memory management markers and permission call records to determine the risk identification level and output a review instruction list. This reduces the reliance on human experience and subjective judgment. Through systematic multimodal data processing, semantic association construction and dynamic updates, and in-depth mining of meta-information, it improves the information utilization rate and association analysis depth of the code review process, enabling code review to be based on more comprehensive, objective and accurate information. This enhances the accuracy of program defect and risk discovery, makes programming review more automated and intelligent, reduces review blind spots and omissions caused by human intervention, and improves review efficiency and program quality control level. Attached Figure Description
[0042] Figure 1 This is a flowchart of the method of the present invention. Detailed Implementation
[0043] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments. These embodiments are based on the technical solution of the present invention and provide detailed implementation methods and specific operating procedures. However, the scope of protection of the present invention is not limited to the following embodiments.
[0044] Example 1
[0045] This embodiment provides a programming review method based on multimodal data, such as Figure 1 As shown, it includes the following steps:
[0046] S1, Corpus Extraction
[0047] S11 parses the version number and number of commits, compares the timestamps, collects source code fragments, extracts meta-information including entry declarations, variable lists, and comment text records, records the version number, committer identifier, and commit date, and generates node collection data.
[0048] Specifically, based on the previously read list of version numbers and corresponding commit counts, the timestamp content associated with each record is compared. A record is selected as the reference standard and parsed into multiple fields, including the version number, committer identifier, and corresponding date tag. To determine if the commit frequency is within an acceptable range, a time difference threshold based on empirical statistics is set; in this embodiment, it is set to 12 hours. The difference in date tags between adjacent records is calculated and compared with this threshold. If the difference exceeds 12 hours, it is determined to be a cross-day commit and the record is marked for further investigation. If the difference is within 12 hours, it is considered an intra-day commit and is categorized accordingly. To further improve the version number parsing process, a simplified identifier function can be used. To obtain its integer part and round the decimal part, where This indicates a rounding up operation. If the version number increments by more than 1 in adjacent records, it indicates that the version may be a major iteration. Each record in the above operation will be processed sequentially. When splitting date tags, it is necessary to compare time zone information. For example, if the time zone parameter is set to UTC+8, the timestamp of the comparison item needs to be converted to a unified UTC benchmark before the difference is calculated to ensure effective alignment of time nodes. In this way, multiple key points that need to be statistically analyzed in all records are marked and relevant information is accumulated and merged to finally obtain node collection data.
[0049] S12, based on node-collected data, filters the submitter identifier and submission date in the records, performs character format comparison, distinguishes duplicate items from new items, outputs consistent entries, and generates matching entry information.
[0050] After acquiring the node-collected data, further checks are performed on the submitter identifier and submission date of each record. To distinguish between duplicates and new entries, a consistency check needs to be performed on the submitter identifier, which requires setting up a string comparison method. Where δ represents the character-by-character comparison function, m represents the number of characters, s1 and s2 are the strings to be compared, and s 1k ,s 2k This represents the k-th character in the string. A value of 0 is used if the corresponding character is the same, otherwise 1. If the sum exceeds 3, the two are considered inconsistent. This threshold of 3 is based on common naming differences and can be adjusted according to the company's internal coding standards. To determine if records with the same submission date contain duplicate entries, the date tag and submitter identifier can be combined into a key-value pair for comparison. If the same date and version number already exist under the same key-value pair, it is considered a duplicate; otherwise, it is included in the new entry statistics. A fixed date range is also considered, for example, setting the 1st to 30th of the current month as the valid range. If the date in a record is earlier than the 1st or later than the 30th, it is considered an invalid record and excluded. The above comparison process can be adjusted according to the actual project scale. If the daily submission volume is usually around 50, a day with more than 100 submissions can be considered an anomaly and marked. These consistent entries are accumulated one by one and then summarized to finally generate matching entry information.
[0051] S13: Based on the matching entry information, compare the node positions in the entries with the submitter identifier, summarize the node levels to which they can be attributed and merge the mapping indexes to generate the initial node mapping.
[0052] After obtaining the aforementioned matching entry information, in order to compare the node positions with the submitter identifiers and summarize the node levels to which they can be attributed, a feedforward neural network with three hidden layers is constructed. Each layer has 128 neurons to extract the node name, version ascending order, and the mapping relationship between submitter and date. The input data includes the previously obtained submitter identifier, submission date, and version number. The output data is a two-dimensional vector used to identify the node's level label and corresponding mapping index. The training process of this feedforward neural network adopts the standard backpropagation algorithm. First, the network weight matrix ω = {ω ij The initial values were set to be between -0.05 and 0.05. 2000 historical submission records were used as the training set, and labels were constructed for each record. The loss function L = ∑(y) was calculated iteratively. pred -y true ) 2 , where y pred Represents the predicted value, y true Represents the true value, and adjusts the weights ω of each layer in reverse. ij and bias term b i The iteration stops when the training epochs reach 500 and the loss is less than 0.001. Here, 0.001 is a convergence threshold determined based on prior testing experience. To illustrate the origin of this threshold, five random initialization training iterations are performed initially, and the average convergence speed is calculated. If the average loss only drops to around 0.005 after 300 iterations, the learning rate can be appropriately reduced and the number of iterations increased until it reaches around 0.001. The inference process involves feeding new input data layer by layer into the feedforward network to calculate activation values and obtain node-level labels at the output layer. Finally, based on the node-level labels, entries with the same node position and submitter identifier are grouped into the same level and their mapping indices are merged to generate the initial node mapping.
[0053] S2, Dynamic Construction
[0054] S21: Based on the initial node mapping, retrieve pointer references, match node names with node numbers, perform association reorganization, obtain the calling order, and generate association reorganization items.
[0055] Based on the initial node mapping already obtained, the pointer information for each node entry is first read and compared with a pre-compiled pointer reference table. This table consists of several typical address ranges and function call locations, obtained through statistical analysis of historical code in the project. Then, each node name is compared literally to confirm its matching relationship in the pointer reference table. Using the upper and lower limits of the number of nodes as a basis, it is determined whether the same name appears more than three times. If so, it is further checked whether the name is reused or has similar spelling. A pointer offset threshold is also defined in this process; in this embodiment, it is set to 128 bytes to determine whether cross-domain access is reasonable. This threshold is derived from the average value calculated after observing the memory usage patterns of several example projects. For example, if the difference between the starting address of a pointer and the starting address of the corresponding called function is greater than 128 bytes, it is considered to be outside the safe range, and the pointer information is temporarily recorded for subsequent matching to eliminate potential conflicts and count the frequency of occurrence. After traversing all nodes and recording the corresponding pointer matching status, the call chain is sorted according to the order of the node name sequence number. If the order is disordered and crosses the threshold, the nodes in that part are marked as needing to be rearranged. The actual function call sequence is rechecked through the pointer reference reference table and a new sequence number mapping is generated. Finally, the corrected node sequence numbers are assembled into a call order list, and the associated reorganization item is obtained.
[0056] S22, based on the associated reorganization item, combined with the preset incremental update parameter set and node relationship table, locate pointer conflict points, filter available node levels, and generate a filtered node set.
[0057] After obtaining the associated reorganization items, the incremental update parameter set and node relationship table are compared to extract the change range information of each node between different versions. The incremental update parameter set contains several numerical parameters, such as the number of lines added and deleted, and is calculated based on the average change trend of each branch version in multiple iterations of the actual project. To locate potential pointer conflict points, it is necessary to first confirm whether the pointer reference address of the same node is reused or the adjacent distance is too small between different versions. For example, a starting address difference threshold is set, which is set to 64 bytes in this embodiment. This threshold is obtained by statistically averaging the minimum span distance of the key functions within the project and supplementing it with a floating coefficient of 0.1. If the pointer address difference of a node in two consecutive versions is less than 64 bytes, it is determined that it is prone to conflict and is marked in the node relationship table for subsequent screening. For nodes with conflict markings, the incremental update parameter set is further compared to see if there are cross-version duplication or merge operations with too short intervals. If the number of such operations on a single node exceeds 2, the node is marked as a conflict node, added to the list of prone conflicts, and prepared for further investigation. After cross-checking all nodes in this way, the marked conflicting nodes can be filtered out from the node list, the usable node levels can be retained, and these nodes can be summarized to form a filtered node set, thus obtaining the final filtered node set.
[0058] S23, based on the filtered node set, matches semantic tags and associates node indexes, summarizes connection paths, outputs the connectable relationships between nodes, and verifies code security scanning indicators in the connection paths. Based on the connectable relationships between nodes and the verification results, a dynamic semantic structure is generated.
[0059] After determining the set of nodes to be filtered, the semantic tags of each node need to be parsed and checked to see if they match their corresponding calling order. A pre-compiled semantic tag reference table is obtained, which records the functional labels and several key attributes of the nodes and can be supplemented by reading the incremental update parameter set from the previous stage. To match these labels with the node indices, the node's sequence number in the set of nodes needs to be matched with the index value in the semantic tag reference table.
[0060] During the matching process, if a node's label and index are found to be missing or have multiple mappings, a rule needs to be defined for verification. In this embodiment, this is achieved through label feature similarity. Label feature similarity is calculated as follows: The similarity benchmark of 0.8 is selected based on multiple test experiences. If the similarity calculation result of the label features exceeds 0.8, it is considered a match. Otherwise, the labels need to be manually recalibrated.
[0061] Based on the call order and dependencies of nodes that match semantic tags, connection paths are formed, and the connectable relationships between nodes are output. Code security scanning metrics are compared item by item along these connection paths. These metrics include keyword checks and access permission verification, and are obtained through statistical analysis of common vulnerabilities in existing projects. If sensitive keywords or abnormal access permission levels are found, they are marked in the path for further investigation. Finally, the nodes that match semantic tags, the connectable relationships between nodes, and the security scan results are summarized to obtain a dynamic semantic structure.
[0062] S3, Multimodal Analysis
[0063] S31, based on dynamic semantic structure, parse semantic tag groups and measure the corresponding text range, perform layered matching to compare pixel distribution with tag similarity, and generate layered matching items.
[0064] Specifically, based on the dynamic semantic structure, the text feature vector is first obtained by reading the pre-annotated semantic tag groups, and the corresponding text range is decomposed into character or word boundary coordinates. A tag reference table collected from historical projects is obtained, and the specific meaning and typical keywords of each tag are extracted to form a tag vector. Then, a similarity matrix is constructed for the text portion, and a baseline value of 0.8 is set based on past testing experience. This baseline value is obtained by calculating the average similarity from over a hundred actual semantic tag and text comparison data, with an additional 0.05 added as a safety redundancy. A similarity score is obtained by weighted comparison of the text feature vector and the tag vector. When the score reaches or exceeds 0.8, it is considered that the semantics are basically consistent. To perform synchronous analysis of pixel distribution, the previously obtained text range coordinates can be further mapped to the image coordinate system, and the pixel density values within the corresponding regions can be statistically analyzed. To improve the precision of similarity matching, a three-layer convolutional neural network can be trained. The training samples consist of 500 labeled images, which are pre-converted to grayscale or RGB matrices as input. Each convolutional kernel has a fixed size of 3×3 and a stride of 1. The first layer contains 32 convolutional kernels, the second layer contains 64 convolutional kernels, and the third layer contains 128 convolutional kernels. The output layer generates image feature vectors in a fully connected manner. The cosine similarity between the image feature vector and the label vector is calculated and denoted as... Where p i With q i Let represent the components of the image feature vector and label vector, respectively, by minimizing ∑(SY). 2The loss function is used for backpropagation iteration, where Y is the ground truth. Training stops when the loss is below 0.01 after approximately 300 training iterations. This 0.01 is determined by comparing the risk of overfitting during the initial trials. When a new sample is input into the network, the similarity between the image and the label in the feature vector space is obtained. This similarity is then combined with the aforementioned text similarity for layered matching. If both results exceed a set benchmark value, the network is considered a match and recorded in subsequent processes, ultimately generating layered matching items.
[0065] S32, based on the layered matching terms, loads image content labels and text description entries, scans edge coordinates, checks the corresponding positions of labels and the distribution of visible areas, and generates multimodal loading information.
[0066] After obtaining the overlay matching items, the corresponding key coordinate information is extracted based on the explicit image content labels and text descriptions. Then, a basic reference table containing standard resolution edge coordinates is used to compare the recorded top, bottom, left, and right corner points to determine the actual range of the visible area. A visible area threshold of 0.9 is set based on pixel distribution statistics. This threshold is derived from the average proportion of the total pixels in the visible area of 50 images of different sizes in the rendered view. If the visible pixel rate of a certain area is lower than this value, it is considered that there is occlusion or edge clipping, and the area needs to be segmented. Otherwise, the area is directly regarded as the visible area. Subsequently, for each region, a preset neighborhood range is searched in the horizontal and vertical directions according to the label content. For example, a maximum deviation of 20 pixels is allowed in the width direction and a maximum deviation of 15 pixels is allowed in the height direction. These ranges are derived from the alignment deviation of images captured by multiple camera devices in real application scenarios. If a complete label boundary can be matched within the neighborhood range, it is confirmed that the label and the text description entry are in the same position. The matching coordinate information and pixel density feature values are associated and summarized into the same mapping table. Finally, the mapping table is combined into multimodal loading information.
[0067] S33. Based on the multimodal loading information, compare the same elements in the image content labels and text description entries, determine feature co-occurrence and mark the corresponding positions, and generate multimodal association features.
[0068] After obtaining the multimodal loading information, it is necessary to compare the specific element names mentioned in the image content labels and text descriptions one by one. If both are found to have the same key terms, it is determined that there may be feature co-occurrence. Then, the coordinate index recorded in the mapping table is used to confirm whether the pixel distribution of their occurrence in the same area exceeds an overlap threshold of 0.75 derived from experimental statistics. This threshold is calculated by averaging the overlap ratio of image labels and text elements in 30 test scenarios. If the overlap reaches or exceeds 0.75, it is considered that the element is clearly indicated in both the image and the text, and the pixel position of the element is marked in subsequent steps. To further improve the accuracy of feature co-occurrence judgment, a support vector machine can be introduced for auxiliary classification. The training samples of the support vector machine are obtained by comparing the existing multimodal loading information with manually segmented and labeled data. 200 positive samples with high similarity and 200 non-matching negative samples are selected. In the feature space, the kernel function κ(u,v)=exp(-γ||uv|| 2 The distance between samples is measured, where u and v are two samples whose similarity is to be calculated, γ is a parameter for adjusting the width of the kernel function, and the classification result is made closer to the actual scene by setting the penalty parameter C. If the classification confidence is higher than 80%, it is used as the basis for confirming co-occurrence and the final position is recorded in the mapping table. After the above process, multimodal association features are obtained.
[0069] S4, Review and Generation
[0070] S41, based on multimodal correlation features, scans the program logic structure and flow entry point, verifies the call order and abnormal branches, parses the variable initialization order, and generates a set of verification elements.
[0071] Specifically, based on multimodal association features, the system first reads the call order and connectable relationships between nodes contained in the dynamic semantic deconstruction, and locates the process entry point, comparing each logical branch with the function call order. A concurrency threshold determined based on practical engineering experience is set; in this embodiment, it is set to three call chains. If more than three call chains exist simultaneously and their start timestamps are within 50 milliseconds of each other, it is marked as a high-concurrency state. Simultaneously, the variable initialization order is checked separately, and the first assignment position of each variable within the code segment is recorded. A variable reference verification function is defined. The reference density of variable i is used to calculate the reference density, where n is the reference check range, and δ is the reference density when variable i is referenced by the k-th reference source or location. ikThe value is 1 if the variable is not explicitly defined, and 0 otherwise. It is then compared to a baseline value of 1.5 calculated from historical data. If the reference density of a variable is greater than 1.5, it indicates that the variable may be called repeatedly in multiple places, requiring extra attention to its initialization on abnormal branches. Subsequently, all abnormal branches at the process entry points are summarized, and it is confirmed whether they have skipped connections with the normal call flow. For example, if a piece of logic skips necessary initialization steps between two consecutive calls with an interval of less than 10 milliseconds, it is determined that there is an abnormal jump. Such abnormal information is included in a temporary list for further comparison. After completing the above steps, all integrated call relationships, variable initialization records, and abnormal branch identifiers are summarized to generate a verification element set.
[0072] S42, based on the verification element set, compares code security scanning indicators, retrieves memory management markers and permission call records, and generates security scanning information.
[0073] Based on the set of verification elements, the branch entries marked as abnormal jumps or high concurrency states are first read, and pre-compiled code security scanning indicators are activated. These indicators include specific entries such as lists of sensitive function calls and memory access permissions. Each entry is matched one by one using string hash comparison. If a branch references a function name in the list and also contains dynamic memory allocation behavior exceeding 8KB, it is marked as a suspicious memory reference and recorded separately. When comparing these suspicious records with permission call records, attention is paid to whether the access level information exceeds the upper limit defined in the table. For example, an access level threshold of 5 is set. This threshold is determined by a comprehensive evaluation after statistical analysis of the operating system kernel call range. If the level of a call is marked as greater than 5, it indicates an irregular request for unauthorized privileges. After scanning all abnormal jump branches, the aforementioned suspicious memory references and unauthorized privilege information are summarized into a pending list, which is then automatically categorized using a trained deep learning classifier. The classifier's structure includes a bidirectional long short-term memory network and an attention layer for identifying the association between code fragment context and invocation intent. First, 10,000 historical normal calls and 10,000 abnormal calls are collected as training samples. Each code fragment is segmented into several vector sequences, which are then fed into the bidirectional long short-term memory network to extract sequence features. Finally, a weighted score 'a' is calculated in the attention layer. i =exp(e i ) / ∑ j exp(e i To obtain feature components that focus more on high-risk segments, where e i For eigenvalues, a i Attention score is calculated. A custom cross-entropy loss function L = -∑[y] is used during training. true ln(y pred )+(1-y true)ln(1-y pred The training process stops when the loss drops below 0.02 after approximately 500 iterations. New suspicious records are input, and classification confidence scores are obtained. Entries with confidence scores above 85% are listed as key risk items. Finally, these key risk items are combined with suspicious memory references and overstepping privilege information to generate security scan information.
[0074] S43, based on security scan information, summarizes anomaly indicators and determines risk level, and generates a list of review instructions.
[0075] Based on security scanning information, records marked as key risk items are first retrieved and each record is assigned a unique anomaly identifier. Then, the severity of each anomaly identifier is compared against pre-established risk grading rules. These risk grading rules are obtained by weighting the severity and scope of a large number of historical vulnerability events, and a kernel density estimation method is used to determine the cutoff point. The specific calculation process can be defined by defining a kernel function. The scope of harm identified by the anomaly is defined as x. If the risk score obtained after kernel density estimation exceeds 0.7, it is classified as high-risk; otherwise, if the risk score is below 0.4, it is classified as low-risk. The intermediate range is the medium-risk level. After completing the classification and determination of all key risk items, the records under different risk levels are summarized and a review instruction list is finally generated.
[0076] S5, Data Synchronization
[0077] S51, based on the review instruction list, reads the number of tasks and associated execution time, splits the queue, calculates the queue length, calls the asynchronous loading trigger signal, and generates the scheduling parsing result.
[0078] Specifically, based on the review instruction list, the task entries are first read, and the identifier and corresponding estimated execution time of each task are extracted. This information is then broken down into several sub-queues, and the length of all sub-queues is calculated. To ensure smooth processing under high load, a queue length threshold of 100 is predefined. The specific value is obtained by selecting the peak load average of 90 queues from the statistical data of multiple batch processing operations, plus a margin of 10 queues. If the length of a sub-queue exceeds 100, an additional sub-queue splitting process is triggered to further break it down into smaller units for easier subsequent scheduling. Simultaneously, to accommodate multi-task synchronous processing, an asynchronous loading trigger signal is set. This signal is determined based on system resource usage after the sub-queue splitting is completed. For example, if the CPU utilization is within a pre-defined effective range (0% to 80%) and the memory utilization is between 0GB and 4GB, the resources are considered available. Otherwise, if the CPU or memory load exceeds this range, the loading of some queues is temporarily suspended, and this status is recorded. After completing the above operations, all the split sub-queues and trigger signal statuses are summarized to generate a scheduling parsing result.
[0079] S52, based on the scheduling analysis results, identify the review priority and retrieve high priority sequences, compare the remaining sequences to determine the merging order, and generate priority comparison items.
[0080] After obtaining the scheduling analysis results, the queue information is read one by one and the review priority of each queue is retrieved. The priority is defined in a classification table and represented by an integer value. The larger the value, the more urgent the review or the higher the dependence. This classification table is formulated with reference to the distribution of the urgency of the company's review tasks over the years and the median is set to 2. If the priority of a queue item is greater than 2, it is classified as a high priority; otherwise, it is considered a regular priority. Then, all items marked as high priority are aggregated to form a high-priority sequence, which is then compared sequentially with the other lower-priority sequences to distinguish between parts that need to be merged and parts that can be delayed. For example, if a high-priority item appears more than 5 times in adjacent time periods, it is considered that it should be executed immediately in the first time period. If a low-priority item appears more than 10 times in the same time period in multiple queues, it is moved to the back and attached to the idle time period. The values of the thresholds 5 and 10 are derived from the statistical calculation of the queue execution of multiple items in the past. Among them, 5 is the average upper limit of emergency events in a single time period plus 1 reserve, and 10 is the median frequency of ordinary events in multiple time periods plus 2 reserves. After completing this comparison process, priority comparison items can be generated.
[0081] S53, based on priority comparison items, cross-merges queues according to the operation time interval and batch synchronization rules, allocates summary information to a unified channel, and generates a synchronization processing package.
[0082] After generating priority comparison items, further cross-merging processing is required based on the computation time interval and batch synchronization rules. The computation time interval refers to reserving a certain number of time blocks within a fixed period to start task execution or insert idle segments. The size of this time block is usually set to 30 minutes and is determined by the project management team after statistical analysis of the daily task peak distribution. If the number of queues in a batch exceeds the pre-set baseline of 50, batch synchronization rules need to be activated for batch merging. These 50 queues are derived from the common deployment scale of a single batch of tasks with a small amount of redundancy. The specific merging method is to mix high-priority items with some low-priority items in the same time period at a ratio of 1:2 to avoid long waiting times or blockages. At the same time, all merged queue information is aggregated into a unified channel for direct retrieval in subsequent execution periods. Finally, each queue is appended with its corresponding merging sequence number and time period information record in this channel to form a complete synchronization processing package.
[0083] S6, Output Results
[0084] S61, based on the synchronous processing package, verifies the list of review instructions and matching references, counts the number of tags and compares the differences in the merged information, and generates integrated comparison entries.
[0085] Specifically, based on the previously obtained synchronization processing package, all entries related to the review instruction list are first extracted and compared item by item with a pre-organized reference document. This reference document lists the index numbers and key identification information corresponding to the review instructions, and includes matching rules to judge possible duplicate or missing references. When each instruction is read and matched with the reference document, it is necessary to first check whether the index numbers are consistent and whether the key identification information matches the items in the same row. If the numbers are the same but the key identification information is different, it is considered an abnormal reference and temporarily summarized into an abnormal pending review list. At the same time, to confirm whether the tags are missing during the merging process, an empirical threshold of 5 tags is set according to the number of tags recorded in the synchronization processing package. For example, if a review instruction should originally contain 5 tags but only 4 or fewer after merging, it is marked as missing tag information. If this situation occurs multiple times, the entry is further split and checked. After the check is completed, the information difference rate before and after merging is calculated. This difference rate can be defined as... Compared to a baseline value of 0.1 obtained from historical review tasks, if the difference rate exceeds 0.1, it indicates that there may be significant omissions or duplications in the merging process. Here, 0.1 is derived by the enterprise from an average of approximately 0.08 obtained through statistical analysis of differences in nearly one hundred instructions, plus a redundancy of 0.02. Items exceeding this value are grouped into key verification targets, and all information difference check results are finally summarized. Finally, normally matching items and abnormal pending verification lists are merged to form integrated comparison items.
[0086] S62, based on the integrated comparison entries, read the execution stage identifier and retrieve the security inspection parameter record, confirm the processed and pending reference stages, and generate associated reference information.
[0087] After obtaining the integrated comparison entries, the execution stage identifier of each entry is read first, and it is checked whether there is a corresponding data row in the previously recorded security inspection parameters. These security inspection parameters contain several pairs of numerical or identifier descriptions of potential risk characteristics, and are gradually summarized from multiple historical documents and actual code review scenarios. In order to clarify which reference stages have been processed and which have not, it is necessary to first search the mapping relationship between the execution stage identifier and the security inspection parameters. If both appear under the same identifier number and the risk characteristic value does not exceed the benchmark range defined by the current project, such as between 0 and 100, it means that the stage has been processed; otherwise, it is classified as pending and marked with a special mark. To make this process more accurate, an allowable difference threshold of 3 can be set. This threshold is obtained by calculating the difference of hundreds of stage identifier numbers from different time periods, obtaining an average value of 2.5, and rounding up. If the identifier of the same link differs from the number of the safety inspection parameter by no more than 3, it is considered to be the same link; otherwise, it is judged as a mismatch or a new link. Following this logic, all entries are traversed and the marked results of processed and pending processing are recorded. Then, the results are cross-linked with the integrated comparison entries to form associated reference information.
[0088] S63, based on associated reference information, outputs a comprehensive list and locates the version retrieval record position, synchronizes the number of tags and archives security inspection parameters, and generates comprehensive review data.
[0089] Based on the associated reference information, references with clearly marked processing status in each entry are reclassified and a comprehensive list is output. To ensure that the merged content corresponds correctly with the version retrieval records, all available version identifiers need to be found in the project data and compared one by one with the same number appearing in the comprehensive list. If a number match is found, the version position corresponding to the list entry is marked, and the tag quantity information is read and synchronized with the tag records in the overall process. If there are more than two missing tags, the number of missing tags is recorded in an anomaly tracking statistics item. The value of this threshold 2 is determined after investigating the tag loss rate in the actual use of the project. If the tag loss rate is within 3%, it is approximated that the probability of multiple missing tags should not exceed 2. After the verification is completed, the security inspection parameters are associated with all corresponding list entries and archived, and their final numbers are recorded. Finally, the processing status of the comprehensive reference is marked in the record, and the above information is combined into a comprehensive review data.
[0090] Example 2
[0091] This embodiment provides a programming review system based on multimodal data, used to implement the method of Embodiment 1, including:
[0092] Corpus extraction module: Collects source code fragments, extracts metadata including entry declarations, variable lists, and comment text records, records version number, committer ID, and commit date, performs character format matching, categorizes the matched metadata according to committer ID and commit date, outputs consistency entries, and generates initial node mapping;
[0093] Dynamic Module: Based on the initial node mapping, it extracts pointer references and call order, performs node association and reorganization, combines the preset incremental update parameter set and node relationship table, filters node levels, performs semantic tag matching on each filtered node, summarizes connection paths, outputs the connectable relationships between nodes, and verifies code security scanning indicators in the connection paths. Based on the connectable relationships between nodes and the verification results, it generates a dynamic semantic structure.
[0094] Multimodal analysis module: Based on dynamic semantic structure, it parses semantic tag groups and measures the corresponding text range, performs layered matching to compare pixel distribution and tag similarity, generates layered matching items, loads image content tags and text description items based on layered matching items, scans edge coordinates, checks the corresponding tag positions and visible area distribution, generates multimodal loading information, compares the same elements in image content tags and text description items based on multimodal loading information, determines feature co-occurrence and marks the corresponding positions, and generates multimodal association features;
[0095] Review generation module: Based on multimodal association features, it performs program logic verification, parses variable initialization order, generates a set of verification elements, compares code security scanning indicators, retrieves memory management markers and permission call records, generates security scanning information, determines risk identification level based on security scanning information, and generates a review instruction list;
[0096] Data synchronization module: Based on the review instruction list, it performs task scheduling parsing, calculates queue length and calls asynchronous loading trigger signals, compares the review priorities, outputs merged queues according to the operation time interval and batch synchronization rules, and generates synchronization processing packages.
[0097] Results output module: Based on the synchronous processing package, it integrates and compares data, extracts review matching reference relationships and counts the number of tags, associates execution stage identifiers and security inspection parameters, outputs a comprehensive list, connects version retrieval records, and generates comprehensive review data.
[0098] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the described module can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0099] The preferred embodiments of the present invention have been described in detail above. It should be understood that those skilled in the art can make numerous modifications and variations based on the concept of the present invention without creative effort. Therefore, all technical solutions that can be obtained by those skilled in the art based on the concept of the present invention through logical analysis, reasoning, or limited experimentation on the basis of existing technology should be within the scope of protection defined by the claims.
Claims
1. A programming review method based on multimodal data, characterized in that, Includes the following steps: Collect source code fragments, extract metadata including entry declarations, variable lists, and comment text records, record version number, committer ID, and commit date, perform character format matching, classify the matched metadata according to committer ID and commit date, output consistency entries, and generate initial node mapping; Based on the initial node mapping, pointer references and call order are extracted, node association and reorganization are performed, and node levels are filtered by combining the preset incremental update parameter set and node relationship table. Semantic tag matching is performed on each selected node, connection paths are summarized, and the connectable relationships between nodes are output. At the same time, code security scanning indicators are checked in the connection paths, and dynamic semantic structure is generated based on the connectable relationships between nodes and the check results. Based on the dynamic semantic structure, the semantic tag group is parsed and the corresponding text range is measured. Layered matching is performed to compare pixel distribution and tag similarity, and layered matching items are generated. Based on the layered matching items, image content tags and text description items are loaded, and edge coordinates are scanned. The corresponding positions of tags and the distribution of visible areas are checked to generate multimodal loading information. Based on the multimodal loading information, the same elements in image content tags and text description items are compared, feature co-occurrence is determined and corresponding positions are marked to generate multimodal association features. Based on the aforementioned multimodal association features, program logic is checked, variable initialization order is parsed, a set of check elements is generated, code security scanning indicators are compared, memory management markers and permission call records are retrieved, security scanning information is generated, risk identification level is determined based on the security scanning information, and a list of review instructions is generated.
2. The programming review method based on multimodal data according to claim 1, characterized in that, The specific steps for generating the initial node mapping are as follows: Construct a feedforward neural network to extract node names, version ascending order, and the mapping relationship between submitter ID and submission date. The input data includes submitter ID, submission date, and version number. The output data is a two-dimensional vector that identifies the level label of a node and its corresponding mapping index. Entries with the same node position and the same submitter ID are grouped into the same level and their mapping indices are merged to generate the initial node mapping.
3. The programming review method based on multimodal data according to claim 1, characterized in that, The specific steps of performing node association and reorganization are as follows: Based on the initial node mapping, the pointer information for each node entry is read and compared with the pre-compiled pointer reference table. The node names are compared word by word to confirm their matching relationship in the pointer reference table. Based on the upper and lower limits of the number of nodes, it is determined whether the number of times the same name appears exceeds the preset number to confirm whether the name is reused or spelled similarly. A pointer offset threshold is defined. If the difference between the starting address of a pointer and the starting address of the corresponding called function is greater than the pointer offset threshold, the pointer information is recorded for subsequent matching to eliminate potential conflicts and count the frequency of occurrence. After traversing all nodes and recording the corresponding pointer matching status, the call chain is sorted according to the order of the node name sequence number. If the order is disordered and crosses the threshold, the nodes in that part are marked as needing to be rearranged. The actual function call sequence is rechecked through the pointer reference reference table and a new sequence number mapping is generated. The adjusted node sequence numbers are assembled into a call order list to obtain the associated reorganization item.
4. The programming review method based on multimodal data according to claim 1, characterized in that, The specific hierarchy of the filtering nodes is as follows: Set a threshold for the starting address difference. Based on the node association and reorganization results, determine whether the pointer address difference of a certain node in two consecutive versions is less than the threshold for the starting address difference. If so, the node is determined to be a node prone to conflict and is marked as a conflict node in the node relationship table. For nodes marked as having conflict, compare them with the incremental update parameter set to see if there are duplicate or short-interval merging operations across versions. If the number of operations on a single node exceeds the preset number, the node is marked as a conflict node. Perform the above steps on all nodes to filter out conflicting nodes, retain the available node levels, and summarize the conflicting nodes to form a filtered node set.
5. The programming review method based on multimodal data according to claim 1, characterized in that, The specific method for generating dynamic semantic structures is as follows: A pre-compiled semantic tag reference table is obtained, which includes the functional tags and key attributes of nodes. The index value of each selected node in the selected node set is matched with the index value in the semantic tag reference table, and the tag feature similarity is calculated for verification to complete the tag matching. Based on the calling order and dependency relationship of the nodes with matched semantic tags, a connection path is formed, and the connectable relationship between nodes is output. Code security scanning indicators are checked on the connection path. If the code or attribute in the connection path triggers the code security scanning indicator, the security scanning result is marked in the connection path. The nodes with matched semantic tags, the connectable relationship between nodes, and the security scanning result marking are summarized to obtain the dynamic semantic structure.
6. The programming review method based on multimodal data according to claim 1, characterized in that, The specific method for generating the stacked matching items is as follows: Based on the dynamic semantic structure, the text feature vector is obtained by reading the semantic tag group marked in it, and the corresponding text range is decomposed in the form of character or word boundary coordinates. The text range coordinates are further mapped to the image coordinate system to obtain the image, the pixel density value is counted, the image is input into the convolutional neural network, and the image feature vector is output. Obtain a label reference table and extract the specific meaning and typical keywords of each label to form a label vector. Calculate the similarity between the text feature vector and the image feature vector and the label vector respectively. Perform a layered matching judgment on the two similarities. If the similarity exceeds the set benchmark value, it is considered a match and recorded in the subsequent process to generate a layered matching item.
7. The programming review method based on multimodal data according to claim 1, characterized in that, The specific steps for generating multimodal loading information are as follows: Based on the explicit image content labels and text description entries in the overlay matching terms, the corresponding key coordinate information is extracted. Using a basic reference table containing standard resolution edge coordinates, the four corner points of the key coordinate information are compared. Based on the regions determined by the corner points, it is determined whether the visible pixel rate of the region is less than the visible region threshold. If so, it is considered that there is occlusion or edge cropping, and the region is segmented. For each region, a preset neighborhood range is searched in the horizontal and vertical directions according to the label content. If a complete labeled boundary is matched in the neighborhood range, it is confirmed that the label and text description entry are consistent. The matching coordinate information is associated with the image feature vector and summarized into the same mapping table. This mapping table is combined into multimodal loading information.
8. The programming review method based on multimodal data according to claim 1, characterized in that, The specific steps for generating multimodal association features are as follows: Based on the multimodal loading information, the specific element names mentioned in the image content tags and text description entries are compared one by one. If the same key terms are found, it is determined that there may be feature co-occurrence. Then, the coordinate index recorded in the mapping table is used to confirm whether the pixel distribution of their occurrence in the same area exceeds the overlap threshold. If so, it is determined that there is feature co-occurrence, and the pixel position of the element is marked to obtain the multimodal association feature.
9. The programming review method based on multimodal data according to claim 1, characterized in that, The method further includes: Based on the review instruction list, the task scheduling is parsed, the queue length is calculated and the asynchronous loading trigger signal is called, the review priority is superimposed for comparison, and the merged queue is output according to the operation time interval and batch synchronization rules to generate a synchronization processing package. Based on the aforementioned synchronization processing package, integration and comparison are performed, review matching reference relationships are extracted and the number of tags is counted, execution stage identifiers and security inspection parameters are associated, a comprehensive list is output, version retrieval records are linked, and comprehensive review data is generated.
10. A programming review system based on multimodal data, characterized in that, For implementing the method as described in any one of claims 1-9, comprising: Corpus extraction module: Collects source code fragments, extracts metadata including entry declarations, variable lists, and comment text records, records version number, committer ID, and commit date, performs character format matching, categorizes the matched metadata according to committer ID and commit date, outputs consistency entries, and generates initial node mapping; Dynamic construction module: Based on the initial node mapping, extract pointer references and call order, perform node association and reorganization, combine the preset incremental update parameter set and node relationship table, filter node levels, perform semantic tag matching on each selected node, summarize connection paths, output the connectable relationships between nodes, and at the same time check code security scanning indicators in the connection paths, and generate dynamic semantic structure based on the connectable relationships between nodes and the check results. Multimodal analysis module: Based on the dynamic semantic structure, it parses semantic tag groups and measures the corresponding text range, performs layered matching to compare pixel distribution and tag similarity, generates layered matching items, loads image content tags and text description items based on the layered matching items, scans edge coordinates, checks the corresponding tag positions and visible area distribution, generates multimodal loading information, compares the same elements in image content tags and text description items based on the multimodal loading information, determines feature co-occurrence and marks the corresponding positions, and generates multimodal association features; Review generation module: Based on the multimodal association features, it performs program logic verification, parses the variable initialization order, generates a set of verification elements, compares code security scanning indicators, retrieves memory management markers and permission call records, generates security scanning information, determines the risk identification level based on the security scanning information, and generates a review instruction list.
Citation Information
Patent Citations
Software vulnerability detection model method based on feature fusion and code visualization technology
CN117574383A
Source code vulnerability classification detection method based on multi-feature fusion and self-attention encoder neural network
CN118761058A