Information leakage detection method, device, equipment and medium

By obtaining the source code from the banking software system and converting it into an intermediate code format, the source and destination sets are determined. Data flow analysis technology is then used to quickly identify the leakage paths of sensitive information, solving the problems of delayed detection of sensitive information leakage and waste of resources in the banking software system, and improving the security and compliance quality of the software system.

CN120930146APending Publication Date: 2025-11-11AGRICULTURAL BANK OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511057944.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-30
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

In the current technology, the detection of sensitive information leakage in banking software systems mainly relies on manual detection after the fact, which leads to detection delays, untimely discovery of problems, and a large consumption of human resources, posing compliance risks.

Method used

By acquiring the source code to be detected and setting a dataset, converting it into an intermediate code format, determining the source and destination sets, identifying the information leakage path based on data flow analysis, and using static code analysis and data flow technology to quickly identify sensitive information leakage within the software lifecycle.

Benefits of technology

It enables early and rapid identification of sensitive information leakage risks throughout the software lifecycle, improving the security and compliance quality of software systems. It has the advantages of timely detection and rapid response, solving the problems of detection lag and resource waste in existing technologies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120930146A_ABST
    Figure CN120930146A_ABST
Patent Text Reader

Abstract

The invention discloses an information leakage detection method and device, equipment and a medium. The method comprises the following steps: acquiring a to-be-detected source code and a set data set; converting the source code to be detected into an intermediate code format to obtain target code data; determining a source point set and a sink point set according to the target code data and a set data set; and performing data flow analysis based on the source point set and the sink point set, and determining an information leakage path of the source code. According to the technical scheme, the risk of sensitive information leakage in the code can be quickly identified as early as possible in the software life cycle, and the safety production quality and compliance quality of a software system are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of software security technology, and in particular to a method, apparatus, device, and medium for detecting information leakage. Background Technology

[0002] As times change, the current security landscape is evolving. From traditional communication security to antivirus, to perimeter security, and now to data and content security, data security has become a focal point in the smart era.

[0003] Currently, the main way to detect the leakage of sensitive customer information in banking software systems is through post-event manual detection. That is, the sensitive customer information is only detected after it has been leaked, and then the sensitive information is manually cleaned up.

[0004] However, this method of detecting sensitive information leaks has problems such as detection lag, untimely discovery of problems, and consumption of a lot of human resources. At the same time, the leaked sensitive information may pose certain compliance risks to banking software systems in terms of code auditing and security review. Summary of the Invention

[0005] This invention provides a method, apparatus, device, and medium for detecting information leakage, which can identify the risk of sensitive information leakage in code as early and quickly as possible during the software lifecycle, thereby improving the security and compliance quality of software systems.

[0006] According to one aspect of the present invention, an information leakage detection method is provided, comprising:

[0007] Obtain the source code to be detected and set the dataset;

[0008] The source code to be detected is converted into an intermediate code format to obtain the target code data;

[0009] The source point set and the destination point set are determined based on the target code data and the set dataset.

[0010] Data flow analysis is performed based on the source set and the destination set to determine the information leakage path of the source code.

[0011] According to another aspect of the present invention, an information leakage detection device is provided, comprising:

[0012] The data acquisition module is used to acquire the source code to be tested and to set the dataset;

[0013] The format conversion module is used to convert the source code to be detected into an intermediate code format to obtain target code data.

[0014] The set determination module is used to determine the source point set and the destination point set based on the target code data and the set dataset.

[0015] The information leakage path determination module is used to perform data flow analysis based on the source point set and the destination point set to determine the information leakage path of the source code.

[0016] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:

[0017] At least one processor; and

[0018] A memory communicatively connected to the at least one processor; wherein,

[0019] The memory stores a computer program that can be executed by the at least one processor, which enables the at least one processor to perform the information leakage detection method according to any embodiment of the present invention.

[0020] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions, the computer instructions being configured to cause a processor to execute and implement the information leakage detection method according to any embodiment of the present invention.

[0021] The technical solution of this invention involves acquiring the source code to be detected and a set dataset; converting the source code to be detected into an intermediate code format to obtain target code data; determining a source point set and a destination point set based on the target code data and the set dataset; and performing data flow analysis based on the source point set and the destination point set to determine the information leakage path of the source code. This technical solution can identify the risk of sensitive information leakage in the code as early and quickly as possible during the software lifecycle, thereby improving the security and compliance quality of the software system.

[0022] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0023] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0024] Figure 1This is a flowchart of an information leakage detection method provided in Embodiment 1 of the present invention;

[0025] Figure 2 This is a flowchart of an information leakage detection method provided in Embodiment 2 of the present invention;

[0026] Figure 3 This is a schematic diagram of the structure of an information leakage detection device according to Embodiment 3 of the present invention;

[0027] Figure 4 This is a schematic diagram of the structure of an electronic device provided according to Embodiment 4 of the present invention. Detailed Implementation

[0028] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0029] It should be noted that the terms "target," "original," etc., used in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0030] Example 1

[0031] Figure 1 This is a flowchart of an information leakage detection method according to Embodiment 1 of the present invention. This embodiment is applicable to detecting information leakage in system software code. The method can be executed by an information leakage detection device, which can be implemented in hardware and / or software. This information leakage detection device can be configured in an electronic device with data processing capabilities. Figure 1 As shown, the method includes:

[0032] S110, Obtain the source code to be detected and set the dataset.

[0033] The source code to be detected can refer to the Java source code of any software application system that needs to be detected for sensitive information leakage. The dataset can be pre-defined data content. In this embodiment, the dataset can include information such as defined sensitive words and a set of rules. The defined sensitive words can be a combination of general sensitive words and some sensitive words customized by the user for the source code to be detected. The set of rules can be a combination of general rules and specific rules customized by the user for the source code to be detected. It is understood that in this embodiment, different datasets can be used for different source codes to be detected.

[0034] This embodiment allows for static code analysis of the source code to be detected. It's a code analysis method that scans program code using techniques such as lexical analysis, syntax analysis, control flow analysis, and data flow analysis without running the code, verifying whether the code meets indicators such as standardization, security, reliability, and maintainability. The technical solution of this embodiment can be executed by an information leakage detection system, capable of performing general-scenario checks and analyses on large Java system software. It automatically detects the leakage of sensitive information (including authentication credentials, passwords, and tokens), and accurately locates the source and path of information leakage based on the source and destination using information flow analysis technology. Specifically, the information leakage detection system can consist of four units: a data input unit, a data preprocessing unit, an information leakage detection unit, and an information leakage path output unit.

[0035] In this embodiment, the data input unit can be used to receive the JAVA source code of the software application system to be detected, as well as user-defined sensitive words and custom rule sets, thereby obtaining the source code to be detected and the set dataset, and transmitting the obtained information to the data preprocessing unit.

[0036] S120. Convert the source code to be tested into intermediate code format to obtain target code data.

[0037] The intermediate code format refers to an intermediate code representation of the code. In this embodiment, the intermediate code format can be Intermediate Representation (IR) format. The intermediate representation in this embodiment is an internal representation during the program compilation and parsing process. It is formed after the compiler scans and understands the source program, reflecting the semantic structure and syntactic features of the source program. The target code data can be understood as code data obtained by converting the target code format to IR code format.

[0038] In this embodiment, the data preprocessing unit first compiles the input JAVA source code to be detected. After compilation, Soot is used for source code parsing and initialization. Then, the JAVA source code is converted into IR (an intermediate code representation) to obtain the target code data.

[0039] S130. Determine the source point set and the destination point set based on the target code data and the set dataset.

[0040] The source point set can be a collection of many source points. A source point can be understood as the starting point of sensitive information leakage. For example, in this embodiment, for a Java web application, a source point typically refers to the entry method of the HTTP request and the method for obtaining sensitive customer information. The destination point set can be a collection of many destination points. A destination point can refer to the endpoint of sensitive information leakage; specifically, it can be the final point in a series of paths from the appearance of sensitive information in the program, its assignment and transmission, to its output to a non-customer-authorized channel. For example, in this embodiment, destination points can include file writing methods, log printing methods, and console printing methods, etc.

[0041] In this embodiment, the data preprocessing unit can parse information such as sensitive words and rule sets in the set dataset, and use the parsed sensitive words and rule data in the target code data in IR format. The sensitive words mark the sensitive variables in the target code data and the list of entry functions that may contain sensitive variables, which are used as the source point set. The output function list of the target code data is identified according to the rule set, which is called the destination point set. The obtained source point set, destination point set and IR are transmitted to the information leakage detection unit.

[0042] S140. Perform data flow analysis based on the source and destination sets to determine the information leakage path of the source code.

[0043] Data flow analysis can refer to a software verification technique that collects information about variables referenced in the code to analyze their assignment, referencing, and propagation within the program. An information leakage path refers to a specific path in the source code where sensitive information is leaked. In this embodiment, the information leakage path refers to the path from the source point to the destination point. Specifically, the format of each node on the information leakage path in this embodiment can be: package name + class name + method name + execution stack statement carrying the leaked information variable (the polluted variable will be displayed as the parameter number when passed as a parameter).

[0044] In this embodiment, entry points are determined and a call graph for the entry points is constructed based on the source points in the source point set and the destination points in the destination point set. Then, the call graph is processed according to the entry point's defined processing rules to obtain the target entry point. Based on the target entry point, the corresponding destination points are determined from each destination point, thereby constructing their respective control flow graphs. Then, based on the call graph and control flow graph, data flow reachability analysis is performed between the source points and destination points to obtain the specific paths where sensitive information leakage occurs. The data flow reachability analysis can include forward data flow analysis and reverse data flow analysis.

[0045] In this embodiment, the information leakage detection unit can detect sensitive information leakage based on the output data of the preprocessing unit. This unit consists of a graph construction subunit, an reachability analysis subunit, a forward data flow analysis subunit, and a reverse data flow analysis subunit.

[0046] The technical solution of this invention involves acquiring the source code to be detected and a set dataset; converting the source code to be detected into an intermediate code format to obtain target code data; determining the source point set and the destination point set based on the target code data and the set dataset; and performing data flow analysis based on the source point set and the destination point set to determine the information leakage path of the source code. This technical solution can identify the risk of sensitive information leakage in the code as early and quickly as possible within the software lifecycle, improving the security and compliance quality of the software system. The technical solution of this embodiment can be used in software systems developed using JAVA to detect whether sensitive customer information is leaked to non-customer authorized channels such as log systems, identifying code security audit risks as early and quickly as possible within the software lifecycle, fully leveraging the agility and responsiveness of DevOps methods, and improving the security and compliance quality of banking software systems.

[0047] Example 2

[0048] Figure 2 This is a flowchart of an information leakage detection method according to Embodiment 2 of the present invention. This embodiment is an optimization based on the above embodiment. Specifically, the optimization includes: setting a dataset including setting sensitive data and setting rules; determining a source set and a destination set based on the target code data and the set dataset, including: identifying multiple entry functions from the target code data according to the set sensitive information, as the source set; and identifying multiple output functions from the target code data according to the set rules, as the destination set. Figure 2 As shown, the method includes:

[0049] S210. Obtain the source code to be detected and set the dataset.

[0050] S220. Convert the source code to be tested into an intermediate code format to obtain the target code data.

[0051] The dataset setting includes setting sensitive data and setting rules. Sensitive data can be a combination of general sensitive data and user-defined sensitive data specific to the source code being detected. Rules can be a combination of general rules and user-defined rules specific to the source code being detected.

[0052] S230. Based on the set sensitive information, identify multiple entry functions from the target code data and use them as a set of source points.

[0053] In this embodiment, the target code data includes multiple entry functions. This embodiment can parse and define each sensitive variable in the sensitive information, then identify each entry function contained in the target code data, thereby marking the sensitive variables in the target code data and multiple entry functions that may contain sensitive variables, forming a list of entry functions and determining it as the source point set.

[0054] S240. Based on the set rules, identify multiple output functions from the target code data and use them as a set of destination points.

[0055] In this embodiment, the target code data includes multiple output functions. This embodiment can be achieved by parsing the various rules in the set rules, identifying each output function contained in the target code data, determining multiple output functions that satisfy each rule from the target code data, forming an output function list, and identifying it as the destination set.

[0056] S250. Based on the source and destination sets, perform data flow analysis to determine the information leakage path of the source code.

[0057] In this embodiment, optionally, data flow analysis is performed based on the source set and the destination set to determine the information leakage path of the source code, including: determining the program call graph, control flow graph, and method pair set based on the source set and the destination set; determining the data flow relationship between each method pair according to the method pair set, control flow graph, and program call graph; and performing data flow analysis on the data flow relationship to determine the source code information leakage path.

[0058] In this context, a call graph (CG) is a directed graph used to represent the call relationships between methods (functions) in a computer program. Nodes in the call graph represent methods, and edges represent call relationships. A control flow graph (CFG) is a representation in computer science that uses mathematical graph representations to identify all paths traversed during the execution of a computer program. It represents the execution flow within a method. Nodes in the control flow graph represent statements, and edges represent the execution flow. A set of method pairs can be a collection of multiple method pairs. In this embodiment, each method pair can represent a complete call chain from a source node to a destination node, and each method pair can be constructed from a source node and a destination node. In this embodiment, during program analysis, the data flow relationship between method pairs describes how data is passed from one method (source or intermediate method) to another method (intermediate method or destination node), ultimately affecting the program's security or functional behavior.

[0059] In this embodiment, based on the source and destination sets, tools can be used to analyze the control flow of the target code program and generate a call graph between methods. After the CG (Call Graph) is constructed, the statements in each method body are further parsed, and a Soot (Soot algorithm) is used to construct the method body's own control flow graph (CFG). Furthermore, in this embodiment, method nodes belonging to the source and destination sets can be marked in the call graph; then, starting from each source node, the call graph is traversed using depth-first search (DFS) or breadth-first search (BFS); if a destination node is encountered, all methods on the path are recorded, thus forming method pairs and obtaining a set of method pairs.

[0060] In this embodiment, for each method pair in the method pair set, the corresponding method node can be found in the program call graph. The reachability between two nodes in each method pair can be analyzed in conjunction with the control flow graph. Based on the analysis results, all possible data flow relationships between each method pair can be determined. Then, based on the data flow relationships between each method pair, forward data flow analysis and reverse data flow analysis can be performed to determine the source code information leakage path.

[0061] In this embodiment, by using such a setup, program calls can be converted into reachable edges of a graph based on control flow and data flow analysis techniques. The Tabulation algorithm can then be used to track sensitive variables, thereby accurately locating the source and path of information leakage in the source program.

[0062] In this embodiment, optionally, determining the program call graph, control flow graph, and method pair set based on the source point set and the destination point set includes: constructing a program call graph based on each target entry function in the source point set as an entry point; processing the program call graph using entry point processing rules to obtain a target entry point set; determining the method body of each entry point in the target entry point set, and constructing a corresponding control flow graph based on each entry point method body and the destination point set; and constructing a method pair set based on the source points and destination points contained in the program call graph.

[0063] The entry point processing rules can be specific rules for processing each entry point. In this embodiment, the entry point processing rules can be pre-defined. The target entry point set can refer to the set of entry points obtained by filtering all entry points in the program call graph using the entry point processing rules.

[0064] In this embodiment, the graph construction subunit receives the source point information, destination point information, and IR data output by the data preprocessing unit. It then finds the entry function list corresponding to the source point information within the IR and uses this list as the entry point to construct a preliminary CG (call graph). Since static analysis methods cannot fully simulate all execution processes in dynamic execution, if reflection calls exist, the method and parameter list of the reflection call are parsed. Furthermore, for applications developed using the Spring framework, this unit marks all classes annotated with `@Controller` as entry points, all methods annotated with `@requestMapping` as entry methods, and classes implementing `HandleInterceptor` or inheriting from `HandleInterceptorAdapter` to implement custom Spring interceptors are marked as entry points, and mock objects are created for their method parameters and receiving objects.

[0065] The entry point processing rules set in this embodiment may include:

[0066] ① Given an entry point method m and its declared type C, create a receiver object for a simulated object of type C. This invention assumes that the receiver object variable (this) points to the simulated object.

[0067] ② Given an entry point method m and a parameter of type T with index i: If T has a concrete subtype in the application, create a mock object for each subtype of T in the application. If T has no subtype in the application, find all type conversions T related to the subtype S in the entry point method. Create a mock object for each S. Mark the constructor of type T as the entry point, thus generating mock objects for the constructor's parameters.

[0068] In this embodiment, the rule of one simulation object per type can be followed to ensure that the analysis remains scalable regardless of the number of entry points. In this embodiment, parameter i can be assumed to point to the simulation object of the aforementioned compatible type.

[0069] ③ For each simulated object type T, mark the constructor of T as the entry point to ensure that the simulated object of T has obtained the state required for analysis, enabling full analysis of the entry point method and its callers. The parameters of the constructor are analyzed recursively under the same simulation strategy, i.e., they are assigned simulated objects of compatible types.

[0070] In this embodiment, bean construction supports building bean objects using annotations and XML configuration files. For aspect-oriented programming (AOP) scenarios, this embodiment parses the corresponding configurations or related annotations in the solution and simulates the relevant dynamic execution logic by adding corresponding processing edges to the CG graph. The list of AOP annotations to be identified in this embodiment is shown in Table 1 below. Taking the AOP annotations shown in Table 1 as an example, the system first scans and obtains relevant classes with the @Aspect annotation. These classes are aspect classes, and when processing normal business logic, they will simulate the dynamic execution process. Based on the specific characteristics of these aspect class methods and the pointcut expressions of the corresponding annotations, the execution process of the relevant methods is woven into the normal business logic. Taking the @Before and @After annotations as examples, by obtaining methods with these two annotations, the pointcut expressions within the relevant annotations are parsed. Based on the pointcut expressions, the business methods that need to be woven are found, and then, based on the annotation type and the pointcut expression, the relevant aspect code is inserted before or after the corresponding business execution logic.

[0071] Table 1 List of AOP annotations

[0072]

[0073] Understandably, in this embodiment, the previous preprocessing step only identified possible entry point function points. However, in a complete program, all possible entry point function points will form a program call graph, but some entry point function points are useless. Therefore, it is necessary to filter the program call graph using entry point processing rules to further filter out the useful entry points and obtain the target entry point set.

[0074] In this embodiment, after the program call graph (CG) is constructed, the statements in each method body can be further parsed to determine each entry point method body in the target entry point set. Within each method body, the control flow graph (CFG) of the method body itself is constructed using Soot, in conjunction with the destination point set. In this embodiment, the sensitive variable data output from the data preprocessing unit and the generated CG and CFG can be output to the reachability analysis subunit.

[0075] In this embodiment, a method pair <source node name> can be constructed based on the set of method names of the source and destination nodes contained in the graph in the program call subunit. i roll call i >, where i can represent that the element can be any element in the corresponding set, thus obtaining the method pair set.

[0076] In this embodiment, the corresponding program call graph, control flow graph, and method pair set can be constructed based on the source set and the destination set, which facilitates the analysis of the data flow relationship between various methods.

[0077] In this embodiment, optionally, determining the data flow relationship between each method pair based on the method pair set, control flow graph, and program call graph includes: for each method pair in the method pair set, determining the transmission path between two nodes in the method pair based on the control flow graph and program call graph; and determining the data flow relationship between two nodes in the method pair based on the transmission path.

[0078] Here, the transmission path refers to the transmission path from the source node to the destination node in a method pair. The data flow relationship refers to whether a data flow relationship exists between two nodes or not. It is understood that in this embodiment, data flow analysis will only be performed on method pairs with a data flow relationship if one exists between the two nodes, thereby determining whether there is a risk of information leakage.

[0079] In this embodiment, for each method pair in the method pair set, the corresponding method node and the program execution flow within a method in the program call graph (CG) can be found. The depth-first search algorithm is used to determine whether there is a reachable transmission path between the two nodes in the two method pairs. If there is a reachable transmission path between the two nodes, all reachable transmission paths between the two nodes are determined. Then, based on the reachable transmission paths, it is further determined whether there is a cleaning function. Based on the determination result, it can be determined whether there is a data flow relationship between the two nodes in the method pair or not.

[0080] In this embodiment, the data flow relationship between methods can be determined by deep first-order algorithm analysis based on the method set, control flow graph, and program call graph. This improves the reliability of data flow determination and facilitates the subsequent identification of information leakage paths in the code information.

[0081] In this embodiment, optionally, determining the data flow relationship between two nodes in the method pair based on the transmission path includes: if there is a transmission path between the two nodes, determining whether the transmission path includes a cleaning function; if the transmission path does not include a cleaning function, then there is a data flow relationship between the two nodes in the method pair.

[0082] The cleaning function refers to a type of variable transformation function, such as `replace()`. After the cleaning function is applied, the variable changes and is no longer its original value, thus eliminating the need for further data flow analysis. The variable can refer to sensitive information or other data initially input by the user.

[0083] In this embodiment, the corresponding method node and the program execution flow within a method in the control flow graph are found in the program call graph (CG). A depth-first search algorithm is used to determine whether a reachable transmission path exists between two nodes in a method pair. If a reachable transmission path exists, all reachable transmission paths between the two nodes are identified, and it is determined whether the transmission path contains a cleaning function. If the reachable transmission path contains a cleaning function, there is no data flow relationship between the two nodes; if no reachable transmission path exists between the two nodes, there is no data flow relationship; if a reachable transmission path exists between the two nodes and no cleaning function is included, then a data flow relationship exists between the two nodes. In this embodiment, all reachable paths with possible data flow relationships can be output to the forward data flow analysis subunit.

[0084] In this embodiment, by setting up such a method, reachability analysis can be performed between two nodes to determine all possible reachable paths with data flow relationships, thereby improving the reliability of data flow relationship determination and facilitating further data flow analysis based on these relationships.

[0085] In this embodiment, optionally, data flow analysis is performed on the data flow relationship to determine the source code information leakage path, including: performing forward data flow analysis and reverse data flow analysis between the two nodes in the method pair according to the transmission path of the data flow relationship, to obtain forward analysis results and reverse analysis results; and determining the source code information leakage path based on the forward analysis results and reverse analysis results.

[0086] Forward data flow analysis can start from the program's entry point (such as the main method or function entry point) and deduce the program's state information at each point along the control flow (such as sequential execution, branching, and looping), ultimately covering all possible execution paths. Reverse data flow analysis can start from the program's exit point (such as function return or program termination) and deduce the program's state information at each point in reverse, simulating the process of "tracing back from the result to the cause." Forward analysis results are those obtained through forward data flow analysis. Reverse analysis results are those obtained through reverse data flow analysis.

[0087] In this embodiment, the forward data flow analysis subunit can perform forward data flow analysis between the two nodes in the method pair according to the transmission path. Specifically, the analysis method can be as follows: perform data flow analysis on each reachable path in the reachability analysis subunit that may have a data flow relationship. For each reachable path output by the reachability analysis subunit, a local hypergraph is constructed, and the data flow function is defined as:

[0088] f d (x)=gen d ∪(x-kill d );

[0089] Where d represents the constant value generated by the current statement. A decomposition hypergraph is constructed based on the representational relationship of the stream function, and the Tabulation algorithm is used to calculate the reachability of each variable in the decomposition hypergraph. In this embodiment, the Tabulation algorithm is a dynamic programming algorithm based on worklist. Its main tasks include bracket matching and path exploration, handling call edges, return edges, and summary edges, and directly connecting two indirectly reachable nodes. The algorithm uses four functions: <1> returnSite: Connects the calling node and the returning node; <2> procOf: associates a function node with its function body; <3> calledProc: Associates the function call node with the name of the function being called; <4> `callers` associates function names with the set of calls made to that function. The algorithm uses the `PathEdge` set to record subsets of reachable edges at the same level on the decomposition hypergraph paths. Reachability at the same level means that function calls and returns are consistent, without mixing different function calls and returns together. For each function, there is a "summary edge" connecting `call` and `return`, used to summarize the results of the called function, thus allowing us to know whether variables are reachable after the function runs without entering the called function. The "summary function" utilizes the distributive law: the `return` node equals the union of information flowing into each path, preventing information loss; therefore, the result of the `return` node is equivalent to the result of the complete function body.

[0090] In this embodiment, data stream computation can be represented by the following formula:

[0091] OUT[s] = f s (IN[s]);

[0092] Where 's' represents a statement, 'IN[s]' represents the data stream value before the statement, and 'OUT[s]' represents the data stream value after statement 's'. When the tainted variables being tracked are stored in collection types such as Map and List, this invention, in order to improve time and space efficiency and avoid marking the entire collection as tainted variables, adopts a solution of recording the key for pruning. In this case, the unit no longer marks the entire collection as a tainted variable, but instead records the key index of the variable. When the key or its corresponding value is needed again, the corresponding value is retrieved directly based on the index, thus avoiding operations on the entire collection and improving running efficiency.

[0093] In this embodiment, if the sensitive information variables in the final data flow graph are reachable from the input parameters of the destination method, it can be determined that there is a data flow leakage in the current path, and the data flow path is transmitted to the information leakage path output unit. Here, the input parameters can refer to the preceding sensitive variables; the sensitive variables, once they reach the input parameters of the destination method, have a data flow relationship.

[0094] In this embodiment, the reverse data flow analysis subunit can perform reverse data flow analysis between the two nodes in the method pair according to the transmission path. Specifically, the analysis method can be as follows: perform reverse data flow analysis on each reachable path in the reachability analysis subunit that may have a data flow relationship. For each reachable path output by the reachability analysis subunit, a local hypergraph is constructed, and the data flow function is defined as:

[0095] f d (x)=gen d ∪(x-kill d );

[0096] Here, d represents the constant value generated by the current statement. A decomposition hypergraph is constructed based on the representational relationship of the stream function, and the Tabulation algorithm is used to calculate the reachability of each variable in the decomposition hypergraph.

[0097] Reverse data flow calculation can be represented by the following formula:

[0098] IN[s]=f s (OUT[s]);

[0099] Where s represents a statement, IN[s] represents the data stream value before the statement, and OUT[s] represents the data stream value after the statement.

[0100] In this embodiment, if the input parameters and sensitive information variables of the final data flow graph destination method are reachable, it is determined that there is a data flow leakage in the current path, and the data flow path is transmitted to the information leakage path output unit.

[0101] In this embodiment, the leaked data stream obtained from the forward analysis and the leaked data stream obtained from the reverse analysis can be used to determine the final data stream relationship between the two nodes in the method pair that contains sensitive information leakage. Then, the corresponding data stream path is output as the source code information leakage path based on the data stream relationship that contains sensitive information leakage.

[0102] Furthermore, in this embodiment, the information leakage path output unit processes the detection result passed in by the information leakage detection unit and finds the specific location of the data stream in the source code. The output content is the path from the source point to the destination point. The format of each node on the path is: package name + class name + method name + execution stack statement carrying leakage information variables (when the pollution variable is passed as a parameter, it will be shown as the nth parameter).

[0103] This embodiment, through such a setting, can accurately determine the information leakage path in the code information based on data flow analysis.

[0104] The technical solution of this invention involves acquiring the source code to be detected and setting a dataset; converting the source code to be detected into an intermediate code format to obtain target code data; setting the dataset includes setting sensitive data and setting rules; identifying multiple entry functions from the target code data based on the set sensitive information, as a source point set; identifying multiple output functions from the target code data based on the set rules, as a destination point set; and performing data flow analysis based on the source point set and the destination point set to determine the information leakage path of the source code. This technical solution can identify the risk of sensitive information leakage in the code as early and quickly as possible during the software lifecycle, improving the security and compliance quality of the software system. This embodiment utilizes data flow analysis methods from static code analysis to detect sensitive information leakage issues in the system during the source code writing phase. It identifies and resolves problems early in the software lifecycle, offering advantages such as timely detection and rapid response. Furthermore, based on data flow analysis algorithms from static analysis, it parses Java source code by analyzing custom sensitive word sets and rule sets, constructs CG and CFG algorithms, and performs forward and reverse data flow analysis using the Tabulation algorithm to track sensitive variables. This allows for the discovery of sensitive information leakage issues, identification of corresponding leakage paths, and addresses the current lack of systematic sensitive information leakage detection tools. This improves software quality and avoids audit risks.

[0105] Example 3

[0106] Figure 3 This is a schematic diagram of the structure of an information leakage detection device according to Embodiment 3 of the present invention. Figure 3 As shown, the device includes:

[0107] Data acquisition module 310 is used to acquire the source code to be detected and set the dataset;

[0108] The format conversion module 320 is used to convert the source code to be detected into an intermediate code format to obtain the target code data.

[0109] The set determination module 330 is used to determine the source point set and the destination point set based on the target code data and the set dataset.

[0110] The information leakage path determination module 340 is used to perform data flow analysis based on the source point set and the destination point set to determine the information leakage path of the source code.

[0111] Optionally, setting up the dataset includes setting sensitive data and setting rules;

[0112] Set determination module 330 is specifically used for:

[0113] Based on the set sensitive information, multiple entry functions are identified from the target code data and used as a set of source points;

[0114] Based on the set rules, multiple output functions are identified from the target code data and used as a set of destination points.

[0115] Optionally, the information leakage path determination module 340 includes:

[0116] The method pair determination unit is used to determine the program call graph, control flow graph, and method pair set based on the source set and the destination set.

[0117] The data flow relationship determination unit is used to determine the data flow relationship between each method pair based on the method pair set, control flow graph, and program call graph.

[0118] The analysis unit is used to perform data flow analysis on data flow relationships to determine the leakage path of source code information.

[0119] Optionally, the method for determining the unit is specifically used for:

[0120] The program call graph is constructed based on the target entry functions in the source point set as entry points;

[0121] The program call graph is processed using entry point processing rules to obtain the target entry point set;

[0122] Determine the method body of each entry point in the target entry point set, and construct the corresponding control flow graph based on each entry point method body and the destination set;

[0123] A set of method pairs is constructed based on the source and destination nodes contained in the program call graph.

[0124] Optionally, the data flow relationship determination unit includes:

[0125] The transmission path determination subunit is used to determine the transmission path between two nodes in each method pair in the method pair set, based on the control flow graph and the program call graph.

[0126] The relationship determination subunit is used to determine the data flow relationship between two nodes in a method pair based on the transmission path.

[0127] Optionally, the relationship determines the sub-unit, specifically used for:

[0128] If a transmission path exists between two nodes, determine whether the transmission path includes a cleaning function;

[0129] If the transmission path does not include a cleaning function, then there is a data flow relationship between the two nodes in the method pair.

[0130] Optional, analysis unit, specifically used for:

[0131] Based on the transmission path of the data flow relationship, forward data flow analysis and reverse data flow analysis are performed between the two nodes in the method pair respectively to obtain the forward analysis results and the reverse analysis results;

[0132] The source code information leakage path was determined based on the results of forward and reverse analysis.

[0133] The information leakage detection device provided in this embodiment of the invention can execute an information leakage detection method provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects of the method.

[0134] Example 4

[0135] Figure 4 This is a schematic diagram of an electronic device according to Embodiment 4 of the present invention. The electronic device 10 is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (such as helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0136] like Figure 4As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 may also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0137] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0138] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as information leakage detection methods.

[0139] In some embodiments, the information leakage detection method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the information leakage detection method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the information leakage detection method by any other suitable means (e.g., by means of firmware).

[0140] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0141] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0142] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0143] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0144] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0145] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.

[0146] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.

[0147] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A method for detecting information leakage, characterized in that, include: Obtain the source code to be detected and set the dataset; The source code to be detected is converted into an intermediate code format to obtain the target code data; The source point set and the destination point set are determined based on the target code data and the set dataset. Data flow analysis is performed based on the source set and the destination set to determine the information leakage path of the source code.

2. The method according to claim 1, characterized in that, Setting up a dataset includes defining sensitive data and setting rules; Based on the target code data and the defined dataset, the source point set and the destination point set are determined, including: Based on the defined sensitive information, multiple entry functions are identified from the target code data, which serve as a set of source points; Based on the established rules, multiple output functions are identified from the target code data and used as a set of destinations.

3. The method according to claim 1, characterized in that, Based on the source set and the destination set, data flow analysis is performed to determine the information leakage path of the source code, including: Based on the source node set and the destination node set, a program call graph, a control flow graph, and a method pair set are determined. The data flow relationships between each method pair are determined based on the method pair set, the control flow graph, and the program call graph. Data flow analysis is performed on the data flow relationships to determine the source code information leakage path.

4. The method according to claim 3, characterized in that, Based on the source set and the destination set, a program call graph, a control flow graph, and a set of method pairs are determined, including: A program call graph is constructed based on each target entry function in the source point set as an entry point. The program call graph is processed using entry point processing rules to obtain a target entry point set; Determine the method body of each entry point in the target entry point set, and construct the corresponding control flow graph based on each entry point method body and the destination point set; A set of method pairs is constructed based on the source and destination nodes contained in the program call graph.

5. The method according to claim 3, characterized in that, Determining the data flow relationships between method pairs based on the method pair set, the control flow graph, and the program call graph includes: For each method pair in the set of method pairs, the transmission path between two nodes in the method pair is determined based on the control flow graph and the program call graph. The data flow relationship between two nodes is determined according to the transmission path determination method.

6. The method according to claim 5, characterized in that, The method for determining the data flow relationship between two nodes based on the transmission path includes: If a transmission path exists between the two nodes, determine whether the transmission path includes a cleaning function; If the transmission path does not include a cleaning function, then there is a data flow relationship between the two nodes in the method.

7. The method according to claim 6, characterized in that, Data flow analysis is performed on the aforementioned data flow relationships to determine the source code information leakage path, including: Based on the transmission path of the data flow relationship, forward data flow analysis and reverse data flow analysis are performed between the two nodes in the method pair, respectively, to obtain forward analysis results and reverse analysis results; The source code information leakage path is determined based on the results of the forward analysis and the reverse analysis.

8. An information leakage detection device, characterized in that, include: The data acquisition module is used to acquire the source code to be tested and to set the dataset; The format conversion module is used to convert the source code to be detected into an intermediate code format to obtain target code data. The set determination module is used to determine the source point set and the destination point set based on the target code data and the set dataset. The information leakage path determination module is used to perform data flow analysis based on the source point set and the destination point set to determine the information leakage path of the source code.

9. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the information leakage detection method according to any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that are used to cause a processor to execute the information leakage detection method according to any one of claims 1-7.