Computer software analysis system
Through the in-depth analysis methods of computer software analysis systems, the problem of neglecting the underlying function call path and loop call phenomenon in the existing technology is solved, and the accuracy of fault diagnosis and the stability and performance of software operation are improved.
Patent Information
- Application Number
- CN202510363745.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-07-08
AI Technical Summary
The existing technology lacks in-depth analysis of the underlying function call path, interface association and loop call phenomenon in computer software analysis, resulting in potential failures not being discovered in time.
The computer software analysis system is adopted to perform source code recording and comparison, function interface association, version difference identification, cross call statistics and resource load monitoring through dependency building modules, cycle detection modules, node screening modules, exception monitoring modules, access inspection modules, and resource tracking modules, combined with debugging breakpoints and program component information, source code recording and comparison, function interface association, version difference identification, cross call statistics and resource load monitoring, and generate dependency path indexes, exception event tags and resource load marks.
It improves the accuracy and timeliness of software fault diagnosis, reduces the impact of circular calls on performance, enhances the stability and security of software operation, and improves system performance and user experience.
Smart Images

Figure CN120276951A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer systems, and in particular, to a computer software analysis system. Background Art
[0002] A computer software analysis system is a specialized system for analyzing and evaluating the running state, quality, and performance of computer software. Its main uses include detecting software defects, optimizing software running efficiency, enhancing software stability, and providing analysis support required for software maintenance and updates, which helps improve software reliability and user experience.
[0003] The existing technology only relies on the superficial evaluation of the overall state and performance of the software during actual operation, lacking effective analysis of details such as the underlying function call paths, interface association situations, and loop call phenomena. It is easy to overlook deep-level version conflicts and cross-call hidden dangers, resulting in some potential faults not being discovered in time. Therefore, improvements are needed. Summary of the Invention
[0004] The purpose of the present invention is to solve the deficiencies existing in the prior art and propose a computer software analysis system.
[0005] To achieve the above purpose, the present invention adopts the following technical solutions: A computer software analysis system includes:
[0006] A dependency construction module, based on debugging breakpoints and program component information, extracts keyword fields and checks function interface associations when comparing source code records, while reviewing version differences and identifying conflict references, screening out invalid calls through error marker comparison, integrating function call chains and version tags, constructing a code snippet reference index, and generating a dependency path index;
[0007] A loop detection module, based on the dependency path index and stack trace information, counts the number of cross-calls and checks the source of duplicate paths when checking function jump nodes, marks the intertwined positions of circular calls and records the offset points of error segments during the interactive screening process, and obtains loop interference parameters;
[0008] A node screening module, based on the loop interference parameters and debugging log content, reads the abnormal trigger frequency records, compares the fault level assignment scheme, while screening associated nodes and calculating the dependency density, and evaluates the influence range in combination with the fault distribution situation to generate node influence weights;
[0009] An exception monitoring module, based on the node influence weights and runtime information flow, tracks the cross-thread access direction when analyzing data fluctuations, identifies the stack inputs that trigger errors, while retrieving the potential hazard factors generated at the interruption points and counting the number of repeated exceptions to obtain exception event tags.
[0010] Preferably, it further includes: an access check module, which, based on the abnormal event label and access audit record parameters, checks for differences in access operations and locks unauthorized actions when comparing user session data, synchronously retrieves violation request tags to analyze the source of cross-domain calls, counts the number of abnormal behaviors, and generates a permission check instruction;
[0011] A resource tracking module, which, based on the permission check instruction and running monitoring data, checks the occupancy ratio and locates the peak memory consumption when allocating server threads, and at the same time compares the disk read and write frequencies to determine the bandwidth tension situation, and establishes a resource load tag.
[0012] Preferably, the dependency construction module includes:
[0013] A field verification sub-module, which, based on debug breakpoints and program component information, compares source code records, verifies the record structure and extracts field positions, then checks the function interface associations and filters out duplicate records to generate a field set;
[0014] A conflict troubleshooting sub-module, which, based on the field set, reviews version differences, reads associated version numbers and compares them with the field list, and at the same time identifies conflict references and records the positions of duplicate references to generate a conflict record;
[0015] A call integration sub-module, which, based on the conflict record, filters out invalid calls through error tag comparison, eliminates unmatched functions, and then integrates the function call chain and version tags and associates the code fragment positions to build a code fragment reference index and generate a dependency path index.
[0016] Preferably, the loop detection module includes:
[0017] A node check sub-module, which, based on the dependency path index and stack trace information, locates function jump nodes, calculates the call depth and compares it with the record table, and then counts the number of intersections to generate a node intersection number;
[0018] A path verification sub-module, which, based on the node intersection number, checks the source of duplicate paths, compares the jump sequences and marks the circular intersection points, and then records the intersection type to generate circular reference points;
[0019] An interaction marking sub-module, which, based on the circular reference points, records the offset points of error segments, confirms the function entry positions and classifies the circular ranges, and then summarizes all circular information and merges adjacent positions to obtain loop interference parameters.
[0020] Preferably, the node screening module includes:
[0021] A frequency reading sub-module, which, based on the loop interference parameters and debug log content, reads abnormal trigger records, compares the fault allocation items and confirms the number of errors, and then obtains the frequency statistical value to generate an abnormal comparison value;
[0022] The risk screening sub-module, based on the abnormal comparison value, screens associated nodes, checks the dependency density and determines the fault label, then sorts out the node list to generate a risk distribution set;
[0023] The weight generation sub-module, based on the risk distribution set, evaluates the fault scope, records the dependency relationship between nodes and confirms the impact area, then integrates the evaluation results to generate the node impact weight.
[0024] Preferably, the abnormal monitoring module includes:
[0025] The data fluctuation sub-module, based on the node impact weight and the information flow during operation, tracks the cross-thread access direction, checks the degree of data change and identifies abnormal peaks, then marks the suspected abnormal positions to generate an abnormal trend set;
[0026] The stack analysis sub-module, based on the abnormal trend set, verifies the stack input, reads the error reference and records the interruption situation, then compares the continuous abnormal paragraphs and marks the positions to generate stack hidden danger items;
[0027] The abnormal summary sub-module, based on the stack hidden danger items, counts the number of repeated abnormalities, summarizes the trigger time points and classifies the abnormal types, then establishes an abnormal list index to obtain the abnormal event label.
[0028] Preferably, the access check module includes:
[0029] The session comparison sub-module, based on the abnormal event label and the access audit record parameters, checks the user session data, screens the access operation differences and records the abnormal marks to generate a difference identification set;
[0030] The privilege violation retrieval sub-module, based on the difference identification set, locks the privilege violation actions, retrieves the illegal requests and confirms the cross-privilege references, then sorts out the abnormal action table to generate a privilege violation list;
[0031] The cross-domain record sub-module, based on the privilege violation list, analyzes the cross-domain call source, locates the access path and counts the number of abnormal behaviors, then summarizes the illegal data to generate a privilege check instruction.
[0032] Preferably, the resource tracking module includes:
[0033] The thread verification sub-module, based on the privilege check instruction and the operation monitoring data, checks the server thread allocation, records the number of thread occupations and finds the memory peak value to generate a thread load value;
[0034] The disk comparison sub-module, based on the thread load value, compares the disk read and write frequencies, calculates the data transfer volume and identifies the overloaded nodes to generate disk utilization;
[0035] The bandwidth statistics sub-module summarizes all load information based on the disk utilization rate, analyzes the network occupancy level, confirms the bandwidth saturation, and establishes a resource load mark.
[0036] Compared with the prior art, the advantages and positive effects of the present invention are as follows:
[0037] In the present invention, by introducing keyword extraction and function interface checking during the comparison of source code records, reviewing version differences and identifying conflicting references, and screening out invalid calls, the software analysis process is made more accurate, and the accuracy of fault diagnosis is improved. During the function call process, the cross-call times are counted and the sources of duplicate paths are checked, so as to quickly locate intertwined circular calls, clarify the offset position of the error segment, and avoid the repeated impact of circular calls on software performance. Based on the analysis of debug logs, the abnormal trigger frequency is read and the scope of fault impact is evaluated, so as to intuitively calculate the dependency density and impact weight of nodes, predict the fault diffusion trend and severity, and improve the timeliness and pertinence of abnormal event handling. In addition, by tracking the cross-thread access direction and analyzing the stack input, potential risk factors can be keenly identified, the propagation speed of exceptions can be controlled, and the stability and reliability of software operation can be enhanced. In the permission access audit, the differences in user session data are actively checked and unauthorized behaviors are locked, the number of abnormal behaviors is accurately counted, and the threat posed by unauthorized operations to system security is prevented. During the analysis of the operation monitoring data of server threads and memory resources, the resource occupancy ratio and bandwidth pressure are checked, the memory occupancy peak and bandwidth bottleneck are timely discovered and processed, the resource allocation efficiency and the overall operation efficiency of the system are improved, and the user experience and system performance of the software are enhanced. Brief Description of the Drawings
[0038] Figure 1 It is the system flow chart of the present invention. Detailed Embodiments
[0039] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention, and are not used to limit the present invention.
[0040] Please refer to Figure 1 , the present invention provides a technical solution: a computer software analysis system includes:
[0041] The dependency construction module, based on debug breakpoints and program component information, extracts keyword fields and checks the function interface associations when comparing source code records, reviews version differences and identifies conflicting references at the same time, screens out invalid calls through error mark comparison, integrates the function call chain and version tags, constructs a code segment reference index, and generates a dependency path index;
[0042] The loop detection module, based on the dependency path index and stack trace information, counts the number of cross calls and verifies the source of repeated paths when checking function jump nodes, marks the interleaving position of circular calls and records the offset point of the error segment during the interactive screening process, and obtains loop interference parameters;
[0043] The node screening module reads the abnormal trigger frequency records based on the cycle interference parameters and debug log content, compares the fault level allocation scheme, screens the related nodes and calculates the dependency density, evaluates the impact range based on the fault distribution, and generates the node impact weight;
[0044] The exception monitoring module tracks cross-thread access directions when analyzing data fluctuations based on node impact weights and runtime information flows, identifies stack inputs that trigger errors, retrieves hidden dangers at record interruptions, and counts the number of repeated exceptions to obtain exception event labels;
[0045] The access check module, based on the abnormal event tags and access audit record parameters, checks for access operation differences and locks unauthorized actions when comparing user session data, simultaneously retrieves illegal request tags to resolve the cross-domain call source, counts the number of abnormal behaviors, and generates permission check instructions;
[0046] The resource tracking module, based on permission check instructions and operation monitoring data, checks the occupancy ratio and locates the peak memory consumption when allocating server threads. It also compares the disk read and write frequencies to determine bandwidth constraints and establishes resource load tags.
[0047] Dependency building blocks include:
[0048] The field verification submodule compares source code records based on debugging breakpoints and program component information, verifies record structures and extracts field locations, then checks function interface associations and removes duplicate records to generate a field set.
[0049] The conflict troubleshooting submodule reviews version differences based on field sets, reads the associated version number and compares it with the field list, identifies conflicting references, records the locations of duplicate references, and generates conflict records;
[0050] The call integration submodule is used to filter invalid calls based on conflict records and error markers, and unmatched functions are removed. The function call chain and version labels are then integrated and associated with the code snippet location, and a code snippet reference index is constructed to generate a dependency path index.
[0051] Specifically, based on the debugging breakpoint and program component information, firstly, all context records related to the source code running are extracted from the debugging breakpoint, and these records are matched one by one with the function interface description and variable definition listed in the program component information, and the parameter type and position mark of each function interface are read, and the function call sequence captured at the debugging breakpoint is compared with the interface description. If it is detected that the difference between the formal parameter and the actual parameter of any function exceeds α=1, it is temporarily recorded as an inconsistent item, where α=1 is obtained by counting the minimum error critical value of the input parameters of each function based on multiple tests. When the cumulative inconsistent items reach β=3, the detailed verification process will be automatically triggered. The setting source of β=3 is determined comprehensively according to the scenario that common function interfaces contain multiple parameters on average. At the same time, the number of fields recorded when the variable is defined is compared. If the number of fields deviates from the initially set range of 1≤γ≤10, it indicates that there is a field parsing anomaly. γ represents the expected field number interval, which is an acceptable range obtained through statistics in the program components of daily maintenance. For the phenomenon of repeated field positions or naming conflicts, the repeated naming statistics δ is compared with the predetermined limit value δ max =2 for comparison, when δ>δ max When this is done, list this as an object of follow-up attention and add it to the list of inconsistent items. Finally, after completing the corresponding comparison of all field positions and function interfaces, record all verified field information and use it as a field set that can be directly referenced in subsequent steps.
[0052] Based on the field set, check the information related to the version difference one by one. First, read the version numbers and their corresponding field change records retained in the source code management, and compare each field name in the field set with the update list under each version number one by one. If it is found that the field name or type conflicts with the current field in any version, it will be marked. When comparing, the threshold of the severity of the conflict is set to θ=2. θ is obtained by empirical statistics. When the number of rewrites caused by the same field name is greater than 2, it is recorded as a high degree of conflict and needs to be recorded. If there is only a slight conflict in one version, it is classified as a low conflict. When it is detected that the field is repeatedly updated in different periods and the version number difference is less than ω=0.1, it will be regarded as a high-risk conflict and an additional record will be made, where ω=0.1 represents the judgment value of adjacent version numbers and small functional differences. Through this item-by-item analysis, all repeated reference locations and their corresponding conflict information can be summarized, and all conflict details are arranged in version number order and recorded in a list, which will be used as a conflict record for subsequent processing.
[0053] Based on the conflict records, the function calls marked as invalid or repeated are additionally screened, the number of invalid calls is counted and it is determined whether these invalid calls exceed μ=5. If it exceeds μ, the relevant function is considered incompatible with the current version. μ=5 is based on the tolerance of the frequency of function call errors in ordinary projects. Then all valid function calls are grouped by version labels, and their corresponding start and end lines in the source code are checked. Then the call chains of the same function are compared under each group. Once it is found that the call chain sequence of the same function between different version labels is inconsistent and the difference is greater than ρ=2, the call association degree of the function is recorded as a low match. ρ=2 is mainly obtained by multiple calculations on typical functions. The adaptation threshold is obtained, all low-matching functions are summarized, and overlapping or abandoned function entries are eliminated. Finally, an index information that integrates the function call chain and version label is obtained, which is associated according to the position of the code snippet to form a dependency path index used to represent all valid reference relationships.
[0054] The loop detection module includes:
[0055] The node check submodule locates the function jump node based on the dependency path index and stack trace information, calculates the call depth and compares it with the record table, and then counts the number of crosses to generate the node cross count;
[0056] The path verification submodule verifies the source of repeated paths based on the number of node intersections, compares the jump sequence and marks the circular interweaving points, then records the interweaving type and generates circular reference points;
[0057] The interactive marking submodule records the error segment offset point based on the circular reference point, confirms the function entry position and classifies the circular range, then summarizes all the circular information and merges the adjacent positions to obtain the loop interference parameters.
[0058] Specifically, based on the dependency path index and stack trace information, the function jumps that occur during the current code execution are analyzed line by line. First, all jump nodes are listed in the order in which the function jumps occur and the call depths between adjacent nodes are given. When the call depth is greater than η=3, an additional record will be made. η=3 comes from the average function call complexity measurement of the target code, which is obtained by counting the recursion and nesting of common functions. At the same time, the jump order registered in each record table is checked. If a function jump record deviates from the same function entry in the corresponding index information by more than κ=2 positions, the jump is judged as an abnormal crossover. κ=2 is set with reference to the general stack size and the typical number of branches. After all abnormal crossovers are summarized, their ratio to normal jumps is calculated and a crossover count statistic is obtained. This statistic can be used to measure the call dislocation of functions in different time periods during analysis. Finally, the crossover counts of all nodes are merged to obtain an overall node crossover number.
[0059] Based on the number of node crossings, first compare the function jump sources corresponding to each crossing node in turn. If the number of occurrences of the source of the repeated path in the same function is greater than ζ = 2, which is obtained from the statistical results of the reuse frequency of common functions, then list it in the high-repetition path list. Then, make a vertical comparison of all jump sequences. When the matching degree between a certain sequence and another sequence is too high and the interleaving range is greater than ξ = 3, where ξ = 3 is obtained from the function similarity test, indicating that the circular structure is relatively obvious, then mark this position as a circular intersection point and record the interleaving type. During this process, if it is found that the distance between adjacent intersection points is too small, then merge them into a larger circular interval for overall management. Finally, centrally organize these circular reference points in the form of a table to obtain the set of all node positions containing circular structures, and regard it as the final output of the circular reference points.
[0060] Based on the circular reference points, first count the error segments that appear near these points, and mainly record them by line number position or function entry position. Extend each discovered error segment forward and backward by λ = 5 lines to confirm its coverage range, where λ = 5 is the measured value of the conventional error impact span in this code environment. When the cumulative number of markers of a certain error segment exceeds ν = 3, where ν = 3 comes from the calculation of test cases of the same type of error, then add this offset point to the subsequent tracking list. At the same time, classify and identify these offset points according to the archived source code structure, merge the offset points that may be within the same function entry range into a unified circular area, and sum up all the information of adjacent areas into a new merged item. If there are repeated error references or variable records during this merge, they will also be added to the record corresponding to this offset point. After completing this series of classification and integration operations, all interference positions with circular call characteristics will be obtained, and these information will be summarized as the final output of the loop interference parameters.
[0061] The node screening module includes:
[0062] The frequency reading sub-module, based on the loop interference parameters and the content of the debug log, reads the abnormal trigger records, compares the fault allocation items and confirms the number of errors, then obtains the frequency statistical value and generates the abnormal comparison value;
[0063] The risk screening sub-module, based on the abnormal comparison value, screens the associated nodes, compares the dependence density and determines the fault label, then sorts out the node list and generates the risk distribution set;
[0064] The weight generation sub-module, based on the risk distribution set, evaluates the fault range, records the dependence relationship between nodes and confirms the impact area, then integrates the evaluation results and generates the node impact weight.
[0065] Specifically, the frequency reading sub-module, based on the loop interference parameters and the content of the debug log, reads the stored abnormal trigger records item by item according to all the node information and trigger time periods listed in the loop interference parameters when executing. These abnormal trigger records already contain the abnormal code identifier, timestamp, and function entry involved in the previous collection phase. The corresponding fault assignment items are obtained by compiling a list for common fault types in the deployment environment. Each fault type in the list contains a standardized error code and a reference level. At the same time, the threshold for the allowed number of error occurrences is set to 5. This threshold is determined by the operation and maintenance personnel through statistical analysis of historical data. For example, when counting the total number of errors occurred in the recent month, it is found that for most fault types, if the number of repeated occurrences reaches 5 in the short term, it indicates a relatively high degree of potential problems in the related functions. When the actual trigger count of any fault type exceeds 5 in the current cycle, it is regarded as an item to be concerned. After the comparison is completed, the number of errors occurred is added up. If the trigger intervals of some errors are particularly short, they can be regarded as potential repetitive faults. At this time, the trigger intervals of these faults are compared with the judgment standard that should not exceed 3 times within one hour. This 3 - time standard is pre - specified by the operation and maintenance team during routine maintenance to identify abnormal intensive trigger situations. If it indeed exceeds this value, it is marked as a high - density error segment. Then, the total number of all matching error counts is summarized and the trigger frequency is calculated by dividing the total number of error occurrences by the recording duration to obtain the frequency statistical value. According to the comparison between the frequency statistical value and the standard level in the fault assignment item, if the comparison result shows that the frequency of a certain fault item is relatively high, it will be additionally listed as a key observation object in the subsequent steps. Finally, the obtained comparison result is refined to generate an abnormal comparison value.
[0066] Risk screening sub-module: Based on the abnormal comparison value, compared with the dependency density sorted out in the previous stage. When the dependency density is derived from a detailed description document of the interaction relationships among system nodes, which is sorted out by the operation and maintenance personnel before the system goes online to list the call or association information between nodes, check the positions of the corresponding nodes and the number of associated nodes one by one according to the fault items with higher numerical values in the abnormal comparison value. If the reference structures between nodes overlap more than 2 times, they are classified as highly coupled nodes. The "more than 2 times" here means that when the operation and maintenance personnel calculate the system dependencies, they find that the references between nodes and other nodes in the common logic often have 1 to 2 overlaps. Once there are more than 3 times, it indicates that the dependencies are overly dense. During the checking process, also check whether the functions or business logics covered by each node are the same. If the node function scopes are the same and the number of fault items is greater than 5 at the same time, it means that there is a potential risk of synchronous spread of faults among these nodes. At this time, map the numbers of these nodes with the types of fault items, list the corresponding node list, compare the time series of faults and the node call depth among them, and count the repeated combinations as a set that may have linkage risks. At the same time, refer to the marked dependency density value. If the dependency density exceeds 2.0, it is classified as a high-risk area. The value of the dependency density is calculated by measuring the ratio of the number of node interactions to the total number of nodes and taking one decimal place, so as to further distinguish between mild and severe dependency relationships. After completing the above screening, package all risk nodes to form a risk distribution set.
[0067] Weight generation sub-module: Based on the risk distribution set, first collect the fault types and fault statistical values recorded for each node in the distribution set, and then compare these nodes with each other with reference to an auxiliary table of the dependency relationships between nodes. The acquisition method of this auxiliary table is to record the direct reference times from node to node when performing graph structure analysis on all function calls within the system. When the reference times are greater than 3, it is marked as a tightly coupled relationship in the auxiliary table. The value of 3 is determined by researching the general function repeated call frequency in historical projects and adding a small amount of redundancy. After completing the comparison, calculate a superposition degree according to the number of fault labels of the node and its associated nodes. If the superposition degree of a single node exceeds 5, it means that the node itself is involved in a wide range of functions. The threshold of 5 comes from the comprehensive statistics of multiple projects. In order to clearly quantify the influence degree of these high-superposition-degree nodes, it is necessary to further set a reference coefficient k to represent the proportion of each node in the overall. The calculation method of k is to divide the number of errors of the node by the total number of times the node is called by other modules. For example, when the number of errors of a certain node is 10 and the number of times it is called by other modules is 50, then k = 0.2. Finally, multiply the superposition degree by k according to a certain function. If the product result exceeds 2, assign a higher weight value. The weights generated for all nodes are between 0 and 5, with 5 being the highest, so that the node influence weights can be obtained.
[0068] The anomaly monitoring module includes:
[0069] The data fluctuation sub-module, based on the node influence weight and the runtime information flow, tracks the cross-thread access direction, checks the degree of data change and identifies abnormal peaks, then marks the suspected abnormal positions and generates a set of abnormal trends;
[0070] The stack analysis sub-module, based on the set of abnormal trends, verifies the stack input, reads the error references and records the interruption situations, then compares consecutive abnormal paragraphs and identifies the positions, generating a set of stack potential hazards;
[0071] The anomaly summary sub-module, based on the set of stack potential hazards, counts the number of repeated anomalies, summarizes the trigger time points and classifies the anomaly types, then establishes an index for the anomaly list to obtain the anomaly event labels.
[0072] Specifically, the data fluctuation sub-module, based on the node influence weight and the runtime information flow, first summarizes the node information with a node influence weight greater than 3. The criterion of greater than 3 is an empirical result formed by observing the gradually expanding influence area of node failures during testing. It finds out the corresponding access time sequences of these nodes in the log and compares them item by item with the cross-thread access direction. When the number of cross-thread operations differs too much from the pre-agreed threshold, it will additionally record the data change range. For example, the threshold can be set that the cross-thread operations should not exceed 10 times within one hour. The value of 10 times is specified according to the research on the actual execution frequency of the business logic. If the actual statistical result exceeds 10, it is determined that the thread dispersion degree is too high. At this time, it checks the traffic peak in the read-write metrics and compares whether the instantaneous traffic recorded in the log exceeds the upper limit of 100MB / s preset at the deployment design. The setting of 100MB / s is derived from the measured network bandwidth tolerance at the project startup. If there is an overstep, it is marked as an abnormal peak and stored in a temporary list. It compares all the abnormal peaks and also pays attention to whether the time periods when they occur are concentrated under the same operation or function. When the peaks appear in different time periods but are concentrated on the same node, they need to be specially marked in the list. Finally, it combines and statistically analyzes the marked peak positions to generate a set of abnormal trends.
[0073] The stack analysis submodule, based on the abnormal trend set, first checks the time period and function name of each abnormal peak record in the trend set, and maps this information one by one to the sequence data of the stack input. The sequence data is generated by tracing the stack pointer upward step by step when the system is deployed, and contains key fields such as function name, call depth and error indication mark. When it is detected that the function name corresponds to the abnormal peak position and the error indication mark coincides with the peak time period, it is determined that a stack exception has occurred there. When checking the number of these exceptions, a comparison is made using the interruption status table prepared in advance. This table is generated for each function that may be interrupted during the development process. If the cumulative number of stack exceptions is greater than 3, the standard of 3 is mainly determined by the statistics of the average failure rate of the function. Then these exceptions are recorded in blocks and regarded as continuous abnormal paragraphs. Then, it is compared whether the adjacent positions of each paragraph are no more than 5 lines apart. The setting of no more than 5 lines is to identify the joint failures in the same range. When the above conditions are met, they are merged into one paragraph and marked. If the abnormal parameter value is too different from the expected input of the specified function, for example, the difference exceeds 10%, the place is marked as a potential hidden danger position. These marked positions will be summarized to obtain the stack hidden danger items.
[0074] The exception summary submodule summarizes the timestamp and function entry of each potential risk item based on the stack potential risk items, and checks the number of repeated occurrences of the same type of exceptions one by one. If the number of repeated occurrences exceeds 4 in a short period of time, the upper limit of 4 is based on the comprehensive statistical value of the density of the same type of errors in the historical usage records. In this case, the exception is included in a priority processing list, and the corresponding trigger time points are compared. When it is found that multiple triggers are concentrated in the alternation between early morning and daytime or during the peak business access period, the trigger density is aligned with the preset time segmentation scheme. The segmentation scheme divides 24 hours into four segments, each with 6 hours, and makes an initial estimate of the maximum number of visits for each segment of 50,000 times. If the number of visits rises to more than 50,000 times, it is regarded as a peak access period. For repeated exceptions in these peak periods, they are further distinguished according to the fault classification table. The fault classification table is established based on the error types that have occurred in the comprehensive project. Each error type is equipped with a unique code. According to the actual code, it is marked as an event of the same type and grouped. After completion, an exception list index can be established, and finally the exception event label can be obtained.
[0075] The access checking modules include:
[0076] The session comparison submodule checks user session data based on abnormal event tags and access audit record parameters, screens access operation differences, records abnormal tags, and generates a difference identification set;
[0077] The over-authority retrieval sub-module locks the over-authority actions based on the difference identification set, retrieves the violation requests and confirms the cross-authority references, then sorts out the abnormal action table and generates an over-authority list;
[0078] The cross-domain record sub-module parses the cross-domain call sources based on the over-authority list, locates the access paths and counts the number of abnormal behaviors, then summarizes the violation data and generates a permission check instruction.
[0079] Specifically, the session comparison sub-module, based on the abnormal event tags and access audit record parameters, first reads the timestamp and event type corresponding to each abnormality from the abnormal event tags, then retrieves the user session data item by item from the access audit record parameters, and matches and compares the two. If it is found that the overlap degree between the operation time of a certain session data and the timestamp in the abnormal event tags is greater than 20%, it is marked as potentially associated. This 20% is set according to the common maximum time offset range of the system. For example, tests show that under a high-load state, there may be an operation record delay of about one-tenth to one-fifth of a second. When a session operation item is matched, the operation type is compared with the existing legal operation list in the access audit record parameters. This legal operation list is manually compiled by the operation and maintenance personnel before deployment according to the business requirements and common access patterns. If it is found that a certain operation does not match the operation types registered in the list, or the operation parameters exceed the specified range, for example, the parameter limit is specified to be between 0 and 100, but the result is 120, it can be regarded as an abnormal operation. During the above matching and screening process, if the same user has multiple operations corresponding to the abnormal event tags within the same session cycle, and these operation forms are quite different from the legal operation list, the number of abnormal marks is accumulated. When the number of abnormal marks exceeds 3, this user is included in the session risk list. The benchmark of 3 is selected by calculating the weighted average of the number of abnormal user behaviors in previous historical access records. After all the screening is completed, all the differential operations are recorded and form a difference identification set.
[0080] The unauthorized retrieval submodule, based on the difference identification set, first searches each record in the difference identification set to check whether there is any situation that exceeds the assigned authority. The assigned authority range cited here is formulated according to the department functions and role levels when the system is initialized. For example, in a certain operation scenario, the highest operation level of different roles ranges from 1 to 5. Exceeding the highest operation level of the corresponding role is considered to be unauthorized. If the operation level appearing in the difference identification set is greater than the maximum available level of the user on record, the operation is included in the category of unauthorized actions. Then, the details of the violation request are retrieved, including the identification code of the initiator of the request, the time of request issuance, and the actual operation path. When the identification code matches the corresponding record value of the risk level file referenced across permissions and is higher than the role level that should be possessed, it is confirmed that the unauthorized action is established here. In the process, it is also necessary to screen whether there are multiple consecutive cross-authority references. If it is detected that three cross-authority operations occur in a short period of time, such as within one hour, these three standards are comprehensively determined by referring to historical violation statistics, and they are uniformly included in the high-frequency violation list. Finally, all abnormal actions are integrated, combined with the level differences of different operation items, an abnormal action table is obtained and merged into an unauthorized list.
[0081] The cross-domain record submodule, based on the unauthorized list, first parses the cross-domain call information of each illegal request. These call information are usually marked from the request header or access path. If it is detected that the main domain name is inconsistent with the domain name belonging to the system, it is recorded as a cross-domain source, and verifies whether the queried cross-domain address is in the legal domain name whitelist established before the project goes online during the comparison process. The whitelist was originally established by the security maintenance team through a survey of commonly used business collaboration domain names. If a domain name is not in the whitelist or conflicts with the list of secondary domain names registered in the system, it is considered an abnormal cross-domain call. At the same time, it is necessary to check whether the same access path has multiple abnormal cross-domain operations in succession. If the number of cross-domain calls exceeds 5 within ten minutes, it is marked as a high-frequency abnormal behavior. Ten minutes and 5 times are both empirically determined based on the system traffic peak and the general business call interval. After counting the time period and IP address of these abnormal calls, they are summarized and associated with the data information of the illegal actions. After confirming all cross-data, a complete set of illegal data can be obtained, and finally a permission check instruction is generated.
[0082] The resource tracking module includes:
[0083] The thread verification submodule verifies server thread allocation based on permission check instructions and operation monitoring data, records the number of threads occupied, finds the memory peak, and generates thread load values;
[0084] The disk comparison submodule compares the disk read and write frequencies based on the thread load value, calculates the data transfer volume, identifies overloaded nodes, and generates disk utilization;
[0085] Bandwidth statistics sub-module, based on disk utilization, summarizes all load information, analyzes the network occupancy level and confirms the bandwidth saturation, and establishes a resource load mark.
[0086] Specifically, the thread verification sub-module, based on the permission check instruction and the running monitoring data, first reads all the high-risk operations and time periods included in the permission check instruction, and then retrieves the server thread allocation situation at that time from the real-time monitoring of the system. When recording the number of threads, an estimated thread occupancy baseline is calculated based on the number of CPU cores and the thread management mechanism. Usually, in an eight-core CPU environment, if the number of threads exceeds 200, resource contention may occur. The 200 is selected comprehensively according to the bandwidth and memory stress test results under the same hardware configuration. Once it is confirmed that the number of threads exceeds 200, the current time period is recorded as a thread-intensive interval. At the same time, the corresponding relationship between thread occupancy and high-risk operations is analyzed. If the thread occupancy climbs to more than 200 during the trigger periods of multiple high-risk operations, it can be regarded as the thread allocation reaching a dangerous level. During this period, the memory usage situation also needs to be tracked. If the instantaneous memory occupancy approaches 90% of the available upper limit, and this 90% is the limit value obtained by the operation and maintenance personnel after estimation, it is regarded as approaching the memory peak. The peak time period is compared one by one with the thread-intensive interval. If thread tension and memory peak occur simultaneously at different times multiple times, it is a sign of frequent high load. After accumulating this information, it is interacted with the occupancy level standard defined by the thread management, and finally the thread load value is obtained.
[0087] Disk comparison sub-module, based on the thread load value, after confirming the specific thread load value, compares the read and write frequencies of the disk. The read and write frequencies can be extracted from the system IO monitoring record. When observing the historical performance test results, when the read and write frequency exceeds 300 times per second, it is regarded as the high-frequency band. The value of 300 times is set as the general upper limit for mechanical hard disks. When it is found that the average read and write frequency of the disk exceeds 300 times per second for more than 10 minutes during the high-load period, and this 10-minute standard is an empirical value obtained from the disk heat growth curve in previous troubleshooting, it is judged that disk stacking problems may occur during this period. To further troubleshoot the overloaded nodes, the most frequently accessed path is located, and the actual load times of the access path are compared with the maximum tolerance value set in the system. For example, the maximum tolerance value is 400 times per second. If it exceeds this value, it is listed as an overloaded node. When recording the overloaded nodes, the frequency of their occurrence also needs to be checked. If the same node appears on the list three times in a row, it is recorded as a high-load node. Finally, the read and write amounts of all high-load nodes are summarized to form disk utilization data.
[0088] The bandwidth statistics submodule combines the disk load information in all time periods based on disk utilization and cross-references it with the network occupancy record in the system. The network occupancy record is output by the monitoring program deployed on the network card interface, which includes the amount of data transmitted per second and the number of connection requests. If the total amount of data transmitted exceeds the preset 400MB / s, which is calculated based on the network environment with a bandwidth upper limit of 1Gbps, it is considered that the bandwidth may be in a congested state. At this time, the actual bandwidth utilization rate is calculated as Calculations are performed to compare the initial setting of 80% as the saturation critical value. Once 80% is exceeded multiple times, it can be identified as a period of high saturation. After continuous detection of multiple periods, if it is found that more than one-third of the periods are in a state of saturation greater than 80%, it can be regarded as a significant network pressure at the current stage. These key data are summarized together with the disk utilization and thread load values to establish a resource load mark.
Claims
1. A computer software analysis system, characterized in that, The system includes: A dependency construction module, which, based on debugging breakpoints and program component information, extracts keyword fields and checks the association of function interfaces when comparing source code records, reviews version differences and identifies conflicting references, filters out invalid calls through error tag comparison, integrates the function call chain and version tags, constructs a code snippet reference index, and generates a dependency path index; A loop detection module, which, based on the dependency path index and stack trace information, counts the number of cross-calls and checks the source of duplicate paths when checking function jump nodes, marks the intertwined positions of circular calls and records the offset points of error segments during the interactive screening process, and obtains loop interference parameters; A node screening module, which, based on the loop interference parameters and debug log content, reads the abnormal trigger frequency records, compares the fault level assignment scheme, screens associated nodes and calculates the dependency density, and evaluates the impact range in combination with the fault distribution, and generates node impact weights; An exception monitoring module, which, based on the node impact weights and the runtime information flow, tracks the cross-thread access direction when analyzing data fluctuations, identifies the stack inputs that trigger errors, retrieves the potential factors generated at the interrupt location and counts the number of repeated exceptions, and obtains exception event tags.
2. The computer software analysis system according to claim 1, wherein It further includes: An access check module, which, based on the exception event tags and access audit record parameters, checks for differences in access operations and locks unauthorized actions when comparing user session data, synchronously retrieves violation request tags to resolve the source of cross-domain calls, counts the number of abnormal behaviors, and generates permission check instructions; A resource tracking module, which, based on the permission check instructions and runtime monitoring data, checks the occupancy ratio and locates the peak memory consumption when allocating server threads, and compares the disk read / write frequencies to determine the bandwidth tension, and establishes resource load tags.
3. The computer software analysis system according to claim 1, wherein The dependency construction module includes: A field verification sub-module, which, based on debugging breakpoints and program component information, compares source code records, verifies the record structure and extracts field positions, then checks the association of function interfaces and filters out duplicate records, and generates a field set; A conflict troubleshooting sub-module, which, based on the field set, reviews version differences, reads associated version numbers and checks against the field list, identifies conflicting references and records the positions of duplicate references, and generates conflict records; A call integration sub-module, which, based on the conflict records, filters out invalid calls through error tag comparison, eliminates mismatched functions, then integrates the function call chain and version tags and associates the code snippet positions, constructs a code snippet reference index, and generates a dependency path index.
4. The computer software analysis system according to claim 1, characterized in that The loop detection module includes: A node check sub-module, which, based on the dependency path index and stack trace information, locates function jump nodes, calculates the call depth and checks against the record table, then counts the number of crosses, and generates the number of node crosses; A path verification sub-module, which, based on the number of node crosses, checks the source of duplicate paths, compares the jump sequences and marks the circular intertwined points, then records the intertwined types, and generates circular reference points; An interactive marking sub-module, which, based on the circular reference points, records the offset points of error segments, confirms the function entry positions and classifies the circular ranges, then summarizes all circular information and merges adjacent positions, and obtains loop interference parameters.
5. The computer software analysis system according to claim 1, characterized in that, The node screening module includes: The frequency reading submodule reads the abnormal trigger record based on the cyclic interference parameter and the debugging log content, compares the fault allocation item and confirms the number of errors, then obtains the frequency statistics value and generates the abnormal comparison value; The risk screening submodule screens related nodes based on the abnormal comparison value, compares the dependency density and determines the fault label, and then sorts the node list to generate a risk distribution set; The weight generation submodule evaluates the fault scope based on the risk distribution set, records the dependencies between nodes and confirms the affected area, and then integrates the evaluation results to generate the node impact weight.
6. The computer software analysis system according to claim 1, characterized in that The abnormal monitoring module includes: The data fluctuation submodule tracks the cross-thread access direction based on the node impact weight and runtime information flow, verifies the degree of data change and identifies abnormal peaks, then marks suspected abnormal locations and generates an abnormal trend set; The stack analysis submodule verifies the stack input, reads the error reference and records the interruption based on the abnormal trend set, and then compares the continuous abnormal sections and marks the positions to generate stack hidden danger items; The exception summary submodule counts the number of repeated exceptions based on the stack risk items, summarizes the triggering time points and classifies the exception types, and then establishes an exception list index to obtain an exception event label.
7. The computer software analysis system according to claim 2, wherein The access check module comprises: The session comparison submodule checks the user session data based on the abnormal event label and the access audit record parameters, screens the access operation differences and records the abnormal marks, and generates a difference identification set; The unauthorized search submodule locks the unauthorized actions based on the difference identification set, retrieves the illegal requests and confirms the cross-authority references, and then sorts out the abnormal action table to generate an unauthorized list; The cross-domain recording submodule analyzes the cross-domain call source based on the unauthorized list, locates the access path and counts the number of abnormal behaviors, then summarizes the violation data and generates permission check instructions.
8. The computer software analysis system according to claim 2, characterized in that The resource tracking module includes: The thread verification submodule verifies the server thread allocation based on the permission check instruction and the operation monitoring data, records the number of thread occupancy and searches for the memory peak value, and generates a thread load value; A disk comparison submodule compares the disk read and write frequencies based on the thread load value, calculates the data transmission volume and identifies the overloaded nodes, and generates disk utilization; The bandwidth statistics submodule summarizes all load information based on the disk utilization, analyzes the network occupancy level and confirms the bandwidth saturation, and establishes a resource load mark.
Citation Information
Cited By
Project working hour statistical analysis system for architectural design application behavior monitoring and algorithm analysis
CN120805224A
Real-time monitoring and blocking method and system for abnormal behaviors of Internet of Things terminal
CN121356829A
A real-time monitoring and blocking method and system for abnormal behavior of an internet of things terminal
CN121356829B