Mobile application security assessment and automatic detection method
By performing syntax and semantic analysis of mobile applications, generating code semantic structured encoding features, and semantic embedding encoding of security vulnerabilities, the false alarm and missed response problems of traditional methods when identifying complex logical vulnerabilities are solved, and the accuracy and reliability of detection results are improved.
Patent Information
- Application Number
- CN202510402930.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-07-18
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional mobile application security evaluation methods cannot effectively identify complex business logic vulnerabilities across function calls, and due to the lack of dynamic context reasoning capabilities, the false positives and missed response rates are high, affecting the accuracy and reliability of the detection results.
Using artificial intelligence-based data processing technology, we conduct syntax analysis and semantic understanding of the code of mobile applications, generate code semantic structured encoding features, and semantic embedding encoding of the security vulnerability rule base, generate detection results through dynamic query analysis, and generate security alarm prompts.
It improves the accuracy and reliability of security vulnerability detection, reduces false alarms and missed reports, and builds a complete security threat coverage network.
Smart Images

Figure CN120337210A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of application security detection, and more specifically, to a method for mobile application security assessment and automated detection. Background Art
[0002] With the explosive growth of mobile applications and their deep penetration into key fields such as financial transactions and personal privacy data storage, the mobile terminal has become the main target of cyberattacks. Malicious attackers frequently exploit code vulnerabilities to steal sensitive information, implant malicious code, or disrupt business logic. Mobile application security assessment has become a core link to ensure user trust and business continuity.
[0003] Traditional automated detection methods scan the code line by line through a built-in security rule library, and use pattern matching algorithms to quickly locate code fragments that match predefined vulnerability features (such as hard-coded passwords, sensitive data leakage patterns), which are efficient in known vulnerability detection scenarios. However, its core defects lie in the static nature of the rule library and the shallowness of the matching logic: on the one hand, the rule engine can only identify vulnerabilities with explicit feature matches (such as directly hitting attack code patterns), but cannot penetrate the code semantics to understand complex logical vulnerabilities (such as business flow defects across function calls); on the other hand, although pattern matching can capture potential risk signals (such as unvalidated user input), due to the lack of dynamic reasoning ability for context logic (such as judging whether the input actually triggers dangerous operations), it may mislabel harmless code as vulnerabilities (false positives) or ignore hidden attack chains that are not explicitly defined but pose actual threats (false negatives). This limitation makes it difficult for traditional methods to build a complete security threat coverage network in the face of scenarios that require deep semantic association such as zero-day vulnerabilities and multi-step logical vulnerabilities, ultimately affecting the accuracy and reliability of the detection results.
[0004] Therefore, an optimized method for mobile application security assessment and automated detection is expected. Summary of the Invention
[0005] The present application aims at the disadvantages in the prior art and provides a method for mobile application security assessment and automated detection.
[0006] According to one aspect of the present application, there is provided a method for mobile application security assessment and automated detection, which includes:
[0007] Obtain the code of the mobile application program;
[0008] Perform code semantic analysis on the code of the mobile application program to obtain a code semantic structured encoding vector;
[0009] Extract a set of security vulnerability rules from the security vulnerability rule library;
[0010] Semantically embed and encode the set of security vulnerability rules to obtain a set of security vulnerability rule semantic embedding encoding vectors;
[0011] Perform a global semantic scan of the security vulnerability rules on the set of the code semantic structured encoding vectors and the set of the security vulnerability rule semantic embedding encoding vectors to obtain a detection result, including: performing a dynamic query analysis guided by the semantic features of the security vulnerability rules on the set of the code semantic structured encoding vectors and the set of the security vulnerability rule semantic embedding encoding vectors to obtain a code-vulnerability rule global semantic query response encoding vector; based on the code-vulnerability rule global semantic query response encoding vector, obtain the detection result, where the detection result is used to indicate whether the confidence level of the existence of security vulnerabilities in the mobile application exceeds a predetermined threshold;
[0012] Based on the detection result, determine whether to generate a security alert prompt.
[0013] Due to the adoption of the above technical solutions, the present application has remarkable technical effects:
[0014] The mobile application security assessment and automated detection method provided by the present application uses an artificial intelligence-based data processing technology to perform syntax analysis and semantic understanding on the code of the obtained mobile application to obtain code semantic structured encoding features, and at the same time semantically embed and encode the set of security vulnerability rules extracted from the security vulnerability rule library to obtain a set of security vulnerability rule semantic embedding encoding features. Subsequently, based on the dynamic query analysis representation of the code semantic structured encoding features and the set of security vulnerability rule semantic embedding encoding features, the detection result is automatically obtained, and a security alert prompt is generated in response to the detection result that the confidence level of the existence of security vulnerabilities in the mobile application exceeds a predetermined threshold. In this way, the accuracy and reliability of the security vulnerability detection result can be effectively improved. Description of the Drawings
[0015] By describing the embodiments of the present application in more detail in conjunction with the drawings, the above and other objects, features, and advantages of the present application will become more obvious. The drawings are used to provide a further understanding of the embodiments of the present application, and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the present application, and do not constitute a limitation to the present application. In the drawings, the same reference numerals generally represent the same components or steps.
[0016] Figure 1 It is a flowchart of the mobile application security assessment and automated detection method according to the embodiment of the present application.
[0017] Figure 2 It is a flowchart of step S2 in the mobile application security assessment and automated detection method according to the embodiment of the present application.
[0018] Figure 3 It is a flowchart of step S5 in the mobile application security assessment and automated detection method according to an embodiment of the present application.
[0019] Figure 4 It is a flowchart of step S51 in the mobile application security assessment and automated detection method according to an embodiment of the present application. Detailed implementation manners
[0020] Next, exemplary embodiments according to the present application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. It should be understood that the present application is not limited by the exemplary embodiments described herein.
[0021] The wide application of mobile applications in key fields such as financial transactions and personal privacy management makes them a major target of cyberattacks. Traditional automated detection methods are based on a preset rule library and can quickly identify explicit vulnerability features in code (such as hardcoded passwords, sensitive data leakage) through pattern matching algorithms, which have an efficiency advantage in detecting known vulnerabilities. However, its core limitations lie in the static rule library and shallow matching logic: on the one hand, it is difficult to identify complex business logic vulnerabilities across function calls through semantic analysis; on the other hand, due to the lack of dynamic context reasoning ability, it may both misreport code that does not trigger actual risks (such as isolated unverified inputs) and miss detecting multi-step hidden attack chains that are not clearly defined. This technical defect makes it difficult for traditional methods to build a complete security threat coverage system when facing zero-day vulnerabilities and deeply correlated logic vulnerabilities, ultimately affecting the accuracy and reliability of the detection results.
[0022] To address the above technical problems, the technical concept of the present application is to first obtain the code of the mobile application program and extract a set of security vulnerability rules from the security vulnerability rule library, then use artificial intelligence-based data processing technology to perform syntax analysis and semantic understanding on the code of the mobile application program to obtain code semantic structured encoding features, and at the same time perform semantic embedding encoding on the set of security vulnerability rules to obtain a set of security vulnerability rule semantic embedding encoding features. Subsequently, based on the dynamic query analysis representation of the code semantic structured encoding features and the set of security vulnerability rule semantic embedding encoding features, the detection result is automatically obtained, and a security warning prompt is generated in response to the confidence that there is a security vulnerability in the mobile application program exceeding a predetermined threshold. In this way, by deeply mining the semantic information of the code, the defect that traditional methods cannot penetrate the code semantics to understand complex logic vulnerabilities is made up, and through the dynamic semantic association between the code and the rules, it helps to accurately infer whether the code meets the trigger conditions of the vulnerability rules, thereby reducing false positives and false negatives, and ultimately a complete security threat network can be constructed to improve the accuracy and reliability of the detection results.
[0023] Figure 1 This is a flowchart of a mobile application security assessment and automated detection method according to an embodiment of the present application. As Figure 1 shown, the mobile application security assessment and automated detection method according to an embodiment of the present application includes: S1, obtaining the code of the mobile application; S2, performing code semantic analysis on the code of the mobile application to obtain a code semantic structured encoding vector; S3, extracting a set of security vulnerability rules from a security vulnerability rule library; S4, performing semantic embedding encoding on the set of security vulnerability rules to obtain a set of security vulnerability rule semantic embedding encoding vectors; S5, performing a security vulnerability rule global semantic scan on the code semantic structured encoding vector and the set of security vulnerability rule semantic embedding encoding vectors to obtain a detection result; S6, based on the detection result, determining whether to generate a security alert prompt.
[0024] In step S1, the code of the mobile application is obtained. It should be understood that the code of the mobile application contains rich information. From the syntactic level, it contains the syntactic structures of various programming languages, such as function definitions, class declarations, variable declarations and uses, control statements (such as conditional statements if-else, loop statements for, while, etc.) and expressions. These syntactic structures construct the basic framework and execution logic of the program. From the semantic level, the code contains business logic information, such as the data processing flow, the interaction method between different modules, the response logic to user input, etc. At the same time, it also contains data-related information, such as the data storage method (such as using a database, file system, etc.), the data transmission format (such as JSON, XML, etc.) and the operations on the data (such as adding, deleting, modifying, querying, etc.). In addition, the code may also contain information related to system resource interaction, such as the invocation of device hardware (such as cameras, microphones, sensors, etc.), the management of network connections, etc. Specifically, many security vulnerabilities are hidden in the business process. For example, injection attacks caused by unvalidated user input, etc. Understanding the business logic at the semantic level helps to discover these potential security risks. The data-related information can help detect whether there is a risk of data leakage, such as improper storage or transmission of sensitive data. The information related to system resource interaction can determine whether the application has unauthorized access to device hardware or network resources, ensuring that the use of system resources by the application complies with security specifications. In short, various types of information in the mobile application code are the basis for comprehensive security assessment and automated detection. By analyzing and processing these information, security vulnerabilities existing in the application can be identified, thereby ensuring the security of the mobile application.
[0025] In step S2, the code of the mobile application is subjected to code semantic analysis to obtain a code semantic structured encoding vector. Specifically, Figure 2Flowchart of step S2 in the mobile application security assessment and automated detection method according to an embodiment of the present application. As Figure 2 shown, step S2 includes: S21, performing syntax analysis on the code of the mobile application to obtain an abstract syntax tree of the code; S22, performing code semantic understanding based on a tree-shaped long short-term memory network on the abstract syntax tree of the code to obtain a structured encoding vector of the code semantics.
[0026] In step S21, syntax analysis is performed on the code of the mobile application to obtain an abstract syntax tree of the code. It should be understood that converting the code of the mobile application from the original text form to a highly structured abstract syntax tree can strip redundant information irrelevant to syntax (such as spaces, comments, formatting symbols), and only retain the syntax units reflecting the essence of the program logic and their nested relationships. This process can eliminate the syntactic ambiguity of the programming language itself (such as operator precedence, scope nesting ambiguity) by precisely parsing elements such as expressions, statements, and function declarations in the code and their hierarchical structures, and can provide a standardized and machine-readable logical framework for subsequent semantic analysis. At the same time, the abstract syntax tree (AST) of the code enables semantic understanding and vulnerability detection of cross-language code by unifying the intermediate representation forms of different programming languages, and lays a foundation for subsequent code semantic modeling - the tree structure naturally adapts to the modeling requirements of neural networks for hierarchical dependency relationships, enabling the model to penetrate the surface form of the code and capture the deep semantic associations of complex logics such as cross-function calls and multi-module interactions, thus effectively supporting the inference and detection of implicit vulnerabilities.
[0027] The following is a detailed elaboration of a specific implementation process of "performing syntax analysis on the code of the mobile application to obtain an abstract syntax tree":
[0028] First of all, code reading is the starting point of the entire syntax analysis process. The code file of the mobile application needs to be loaded into the syntax analysis module, and this process needs to ensure the integrity and accuracy of the code to avoid deviations in subsequent analysis caused by reading errors. When reading the code, different encoding formats need to be considered to ensure that the characters in the code can be correctly parsed. For example, when processing code containing Chinese characters, an appropriate encoding method, such as UTF-8, needs to be selected to ensure the correct recognition of characters.
[0029] After the code reading is completed, the lexical analysis stage begins. The lexical analyzer scans the code in the order of the character stream and, based on the lexical rules of the programming language, splits the code into meaningful lexical units. These lexical units include keywords, identifiers, operators, constants, and delimiters. Keywords are reserved words with specific meanings in the programming language, such as "public" and "class" in Java, "def" and "if" in Python. Identifiers are variable names, function names, etc. defined by the programmer. Operators are used to perform various operations, such as addition, subtraction, multiplication, and division. Constants are fixed values, such as integers and strings. Delimiters are used to separate different parts of the code, such as parentheses and semicolons. By matching and identifying the characters in the code, the lexical analyzer converts the code into a stream of lexical units, providing a basis for subsequent syntactic analysis. For example, for the code "int num = 10;", the lexical analyzer will identify the lexical units "int" (keyword), "num" (identifier), "=" (operator), "10" (constant), and ";" (delimiter).
[0030] After lexical analysis is completed, the resulting stream of lexical units serves as the input to the syntactic analyzer. The syntactic analyzer processes the lexical units according to the predefined syntactic rules of the programming language, usually described by context-free grammars, and constructs an abstract syntax tree for the code. The syntactic analyzer generally uses top-down or bottom-up parsing algorithms. Top-down parsing algorithms, such as recursive descent parsing, start from the start symbol of the syntactic rules and gradually derive the syntactic structure of the code. Bottom-up parsing algorithms, such as operator precedence parsing and LR parsing, start from the lexical units of the code and gradually reduce to the syntactic rules. Taking the code "if(a > 10){b = 20;}" as an example, the syntactic analyzer will identify the structure of the "if" statement, "(a > 10)" as the conditional expression, and "{b = 20;}" as the statement block. It will organize these elements into a tree structure according to the syntactic rules, with "if" as the root node and the conditional expression and statement block as the child nodes. During the construction process, if there are syntax errors in the code, such as mismatched parentheses or missing semicolons, the syntactic analyzer will detect and report an error.
[0031] However, during the syntax analysis process, syntax ambiguity problems may be encountered. Syntax ambiguity refers to the situation where the same code snippet can have multiple syntactic interpretations. For example, in some programming languages, for an expression like "a + b * c", due to operator precedence issues, there may be different ways of understanding it (whether to perform addition first or multiplication first). To eliminate this ambiguity, the syntax analyzer will follow pre-set operator precedence rules (such as multiplication having a higher precedence than addition), or clarify the syntactic structure of the code through special designs of syntax rules. In addition, for scope nesting ambiguities, such as the definition and use of variables in different scopes, the syntax analyzer will determine the correct scope of variables according to the scope rules of the programming language to ensure that the code abstract syntax tree can accurately reflect the logical structure of the code.
[0032] After obtaining the preliminary code abstract syntax tree, optimization and standardization processes are still required. The purpose of this step is to remove redundant information unrelated to syntax, making the code abstract syntax tree more concise and clear, facilitating subsequent semantic analysis. Redundant information includes spaces, comments, and formatting symbols in the code. These information have no practical effect on the logical execution of the code but will increase the complexity of syntax analysis. Removing them can improve the analysis efficiency. At the same time, to make the code abstract syntax tree better applicable to unified processing of different programming languages and subsequent vulnerability detection, standardization operations will also be performed. The standardization operations ensure that after the code of different programming languages is converted into a code abstract syntax tree, they have similar structures and representations, facilitating cross-language semantic understanding and vulnerability detection.
[0033] In step S22, code semantic understanding based on a tree-shaped long short-term memory network is performed on the code abstract syntax tree to obtain the code semantic structured encoding vector. It should be understood that although the code abstract syntax tree completely retains the syntax hierarchy of the code, it lacks explicit expression of semantic information such as logical association, data flow, and cross-node dependencies. In order to convert the syntax structure of the code into a semantic vector representation and break through the dependence of traditional methods on the surface pattern of the code, in this application, code semantic understanding based on a tree-shaped long short-term memory network needs to be performed on the code abstract syntax tree to obtain the code semantic structured encoding vector. Those of ordinary skill in the art should understand that the tree-shaped long short-term memory network (Tree-LSTM) is an extended variant of the traditional sequential LSTM and is designed specifically for processing tree-structured data. In this model, each tree node corresponds to an independent LSTM unit, and its core mechanism is to dynamically adjust its own state by integrating information from child nodes. Specifically, the information starts from the leaf nodes and is passed layer by layer from bottom to top along the branch structure of the tree to the root node. During this process, each internal node actively aggregates the state information of all its child nodes as input. As the information flow progresses, the LSTM unit of each node coordinates the features transmitted by the child nodes through a gating mechanism (including an input gate, a forget gate, and an output gate), and synchronously updates its own cell state and hidden state. This design enables the node to explicitly model the structural dependencies between parent and child nodes in the tree topology while retaining the characteristics of the data sequence. Finally, the encoding of the semantics of the overall tree structure is achieved through the state output of the root node. That is, Tree-LSTM can encode discrete syntax units into a code semantic structured encoding vector by modeling the topological relationships (parent-child and sibling node connections) and context dependencies (such as function call chains and variable scope nesting) between the nodes of the code abstract syntax tree, enabling the deep semantics of the code (such as business logic intentions and potential vulnerability patterns) to be directly perceived and processed by the model, thereby providing semantic-level feature support for subsequent vulnerability rule matching.
[0034] The following is a detailed elaboration of a specific implementation process of "performing code semantic understanding based on a tree-shaped long short-term memory network on the code abstract syntax tree to obtain the code semantic structured encoding vector":
[0035] Before formally applying the Tree-LSTM (Long Short-Term Memory) network to process the code abstract syntax tree, data preprocessing is essential. The primary task in this stage is to represent the features of each node in the code abstract syntax tree. The type of the node, such as variable declaration, function call, etc., is important feature information; the name of the node, if it is an identifier node, can also reflect the semantics of the code; the attributes of the node, like the data type, cannot be ignored either. To convert this information into a vector form suitable for network processing, methods such as one-hot encoding or word embedding can be adopted. For keyword nodes, one-hot encoding can clearly represent them as a high-dimensional vector, where each dimension corresponds to a specific keyword, and only the dimension corresponding to that keyword is 1, and the rest are 0. For identifier nodes, a pre-trained word embedding model can map them to a low-dimensional vector space, which can better capture the semantic associations between identifiers. In addition to node feature representation, the representation of the tree structure is also crucial. The Tree-LSTM network needs to clarify the parent-child relationships between nodes in the tree, which can be achieved through an adjacency matrix or an edge list. The adjacency matrix can intuitively show the connection relationships between nodes, while the edge list more concisely records the edge information between nodes, laying the foundation for the network to correctly process the dependencies between nodes.
[0036] The Tree-LSTM network cleverly combines the characteristics of the LSTM network and the tree structure, and has a powerful ability to capture the long-term dependencies between nodes in the tree. Its core consists of an input gate, a forget gate, an output gate, and a cell state. The input gate is responsible for determining which parts of the new input information will be added to the cell state. It calculates based on the current input and the previous hidden state, and filters out the information useful for the current moment. The forget gate controls which information in the cell state needs to be forgotten. It can judge which information is outdated based on the input and the hidden state, thus preventing the cell state from being occupied by too much useless information. The output gate determines which parts of the cell state will be used as the output of the current node. It comprehensively considers the cell state, the input, and the hidden state, and outputs information valuable for subsequent processing. The cell state is like an information container that stores the historical information from the root node to the current node, and continuously updates and maintains this information through the control of the input gate and the forget gate. When constructing the Tree-LSTM network, it is necessary to initialize the parameters of the network, including the weight matrix and the bias vector. Random initialization methods, such as Gaussian distribution initialization, can be used to assign initial values to these parameters. At the same time, determine the number of layers of the network and the dimension of the hidden layer. These hyperparameters will affect the performance and expressive power of the network.
[0037] After completing data preprocessing and network construction, the forward propagation process begins. This process starts from the leaf nodes of the code abstract syntax tree because leaf nodes are the basic units of the tree, and their input is the preprocessed feature vectors. For each leaf node, its feature vector is input into the units of the tree-shaped long short-term memory network. According to the calculation formula of the network, the states of the input gate, forget gate, and output gate, as well as the update of the cell state, are calculated. These gating mechanisms work together to enable the network to dynamically adjust the cell state based on the features and context information of the nodes. Then, according to the structural information of the tree, the output of the child nodes is passed to the parent nodes. After receiving the output of the child nodes, the parent nodes combine their own inputs and perform the calculation and update of the gating state and cell state again. This process is carried out recursively, propagating gradually from the leaf nodes upwards until reaching the root node. The output obtained at the root node is the encoding vector of the entire code abstract syntax tree.
[0038] However, the obtained encoding vector may need to be further optimized and adjusted to improve its ability to represent code semantics. This requires using a loss function to measure the difference between the encoding vector and the expected target. Common loss functions include the cross-entropy loss function and the mean squared error loss function, etc. The cross-entropy loss function is suitable for classification problems and can measure the difference between the probability distribution of the encoding vector over different classes and the true labels; the mean squared error loss function is commonly used in regression problems and calculates the squared error between the encoding vector and the target vector. Through the backpropagation algorithm, the gradient of the loss function with respect to the network parameters is calculated, that is, the influence degree of each parameter on the loss function. Then, optimization algorithms such as stochastic gradient descent, Adam, etc. are used to update the network parameters according to the gradient information, so that the value of the loss function continuously decreases. During the training process, to improve the generalization ability of the network, a large number of code samples are needed for training. At the same time, to prevent overfitting, regularization techniques such as L1 or L2 regularization can be adopted to constrain the network parameters and avoid the model over-relying on the noise in the training data.
[0039] After multiple iterations of training, the parameters of the tree-shaped long short-term memory network will gradually converge. At this time, the trained network can be used to encode the code abstract syntax tree to be processed. The code abstract syntax tree to be processed is first preprocessed and then input into the network. After the forward propagation process, the output finally obtained at the root node is the code semantic structured encoding vector. This vector contains the deep semantic information of the code and can reflect the features of the code in terms of function, logic, and structure, etc.
[0040] In step S3, a set of security vulnerability rules is extracted from the security vulnerability rule library. It should be understood that security vulnerability rules describe the specific manifestations and characteristics of security vulnerabilities. For example, the specific format of hard-coded passwords, the use of relevant functions or protocols without encryption during sensitive data transmission, code patterns with SQL injection risks, etc. These characteristics are the key basis for identifying vulnerabilities. A single security vulnerability rule can only detect specific vulnerabilities. By extracting a set of security vulnerability rules, multiple types of vulnerabilities can be integrated to form a complete threat coverage network, and the set of security vulnerability rules can expand the detection ability through continuous updates (such as adding new vulnerability patterns).
[0041] In step S4, semantic embedding encoding is performed on the set of security vulnerability rules to obtain a set of security vulnerability rule semantic embedding encoding vectors. Specifically, in the embodiment of the present application, step S4 includes: using a semantic encoder based on the BERT-TextCNN model to perform semantic embedding encoding on each security vulnerability rule in the set of security vulnerability rules to obtain the set of security vulnerability rule semantic embedding encoding vectors. It should be understood that traditional methods perform rule matching through predefined keywords or fixed code patterns. Although they can quickly identify explicit vulnerability features (such as string constants of hardcoded passwords), it is difficult to penetrate the code semantics to capture complex logical vulnerabilities (such as business flow defects across function calls). For example, the description of "sensitive data transmitted without encryption" in the rule not only involves explicit matching of keywords "encryption" and "sensitive data", but also requires combining the context to judge specific scenarios of data flowing to network interfaces, actual calls of encryption algorithms, and other implicit logics. Due to the lack of semantic abstraction ability for natural language descriptions and code patterns, the static rule library can neither dynamically associate vulnerability trigger conditions (such as constituting a risk only when data is persisted) nor generalize to new types of vulnerability variants that are not predefined (such as new encryption misuse patterns). By performing semantic embedding encoding on the set of security vulnerability rules, redundant information in the rule expression can be stripped, and its core logical features (such as the dynamic path of "performing a dangerous operation without verifying user input") can be refined, enabling the model to capture potential associations between rules in the vector space (such as "hardcoded key" and "plaintext storage of sensitive data" both pointing to data protection defects). That is, through semantic embedding, the rules are no longer limited to literal matching but are transformed into quantifiable semantic entities, thus supporting dynamic association with code semantic vectors. In particular, in a specific embodiment of the present application, a semantic encoder based on the BERT-TextCNN model is used to perform semantic embedding encoding on each security vulnerability rule in the set of security vulnerability rules to obtain the set of security vulnerability rule semantic embedding encoding vectors. Specifically, the core feature of the BERT-TextCNN model is to use the multi-layer self-attention mechanism of BERT to capture the deep context associations of rule texts. For example, in a rule like "calling a dangerous API without verifying user input", BERT can dynamically associate the long-distance logical dependencies among "user input", "lack of verification", and "dangerous API", breaking through the shallow parsing of isolated words by traditional bag-of-words models. At the same time, TextCNN extracts key phrase patterns in the rule description through multi-scale convolutional kernels, enhancing the sensitivity to the core features of vulnerabilities and avoiding dilution of semantic encoding by redundant information.This joint modeling mechanism enables the model to not only understand the general semantic framework of the rule text through the pre-trained knowledge transfer of BERT (based on massive vulnerability reports, code comments, and other corpora), but also accurately locate the vulnerability-specific patterns by leveraging the local features of TextCNN. Ultimately, a set of security vulnerability rule semantic embedding encoding vectors that integrate multi-granularity semantics of global logic and local details is generated. It should also be noted that this structural design not only enhances the parsing ability of the rule dynamic context but also quickly adapts to new rules through the fine-tuning mechanism of the pre-trained model, significantly reducing the model reconstruction cost caused by the update of the traditional rule library.
[0042] In step S5, a global semantic scan of the security vulnerability rules is performed on the set of the code semantic structured encoding vectors and the security vulnerability rule semantic embedding encoding vectors to obtain a detection result. Specifically, Figure 3 FIG. is a flowchart of step S5 in the mobile application security assessment and automated detection method according to an embodiment of the present application. As Figure 3 shown, the step S5 includes: S51, performing a dynamic query analysis guided by the security vulnerability rule semantic features on the set of the code semantic structured encoding vectors and the security vulnerability rule semantic embedding encoding vectors to obtain a code-vulnerability rule global semantic query response encoding vector; S52, based on the code-vulnerability rule global semantic query response encoding vector, obtaining the detection result, where the detection result is used to indicate whether the confidence level that there is a security vulnerability in the mobile application exceeds a predetermined threshold.
[0043] In step S51, a dynamic query analysis guided by the security vulnerability rule semantic features is performed on the set of the code semantic structured encoding vectors and the security vulnerability rule semantic embedding encoding vectors to obtain a code-vulnerability rule global semantic query response encoding vector. Specifically, Figure 4 FIG. is a flowchart of step S51 in the mobile application security assessment and automated detection method according to an embodiment of the present application. As Figure 4 shown, the step S51 includes: S511, calculating a code-vulnerability rule semantic query response Laplacian matrix based on the set of the code semantic structured encoding vectors and the security vulnerability rule semantic embedding encoding vectors; S512, performing spectral decomposition on the code-vulnerability rule semantic query response Laplacian matrix to obtain a set of code-vulnerability rule semantic query response node core encoding vectors;
[0044] S513, performing complementary information adaptive integration on the set of the code-vulnerability rule semantic query response node core encoding vectors to obtain the code-vulnerability rule global semantic query response encoding vector.
[0045] It should be understood that traditional query analysis methods based on shallow neural networks rely on simple non-linear functions to interact with features, and this mechanism is difficult to characterize the complex interaction patterns between code semantics and vulnerability rules. For example, a function call in the code may indirectly trigger the sensitive data leakage condition in the vulnerability rule through multiple levels of nested logic (such as conditional branches and loop structures). Such cross-level and non-linear correlation relationships cannot be accurately modeled by the limited non-linear transformation of the shallow network, resulting in the omission of key vulnerability signals. At the same time, query analysis in the high-dimensional original feature space (such as directly matching the semantic vectors of code semantic nodes and rule texts) not only faces low computational efficiency due to the curse of dimensionality, but also makes it difficult for the model to capture the intrinsic manifold structure of the data due to high-dimensional sparsity. In addition, traditional query analysis methods tend to ignore the graph structure associations between code and rules (such as the correspondence between code data flows and rule trigger paths), resulting in the inability to effectively identify complex vulnerability scenarios that require multi-hop reasoning (for example, a hidden vulnerability chain where user input is encrypted and stored, but the encryption fails due to hard-coded keys). Based on this, in this application, dynamic query analysis guided by security vulnerability rule semantic features is performed on the set of the code semantic structured encoding vectors and the security vulnerability rule semantic embedding encoding vectors to obtain code-vulnerability rule global semantic query response encoding vectors.
[0046] Specifically, in the embodiment of this application, the step S511 includes: respectively performing non-linear activation processing on each security vulnerability rule semantic embedding encoding vector in the set of the code semantic structured encoding vectors and the security vulnerability rule semantic embedding encoding vectors to obtain a set of code-vulnerability rule semantic query response node implicit encoding vectors, and this process can be represented by the formula:
[0047] V2 = {v 21 , v 22 ,..., v 2i ,..., v 2n}
[0048]
[0049] V 1,2 = {v 1,21 , v 1,22 ,..., v 1,2i ,..., v 1,2n}
[0050] Among them, V1 represents the code semantic structured encoding vector, V2 represents the set of security vulnerability rule semantic embedding encoding vectors, v 21 , v 22 , v 2i , v 2nrespectively represent the 1st, 2nd, ith, and nth security vulnerability rule semantic embedding coding vectors in the set of security vulnerability rule semantic embedding coding vectors, sigmoid represents the sigmoid function, represents matrix multiplication, W 1i and b1 i respectively represent the response weight matrix and the response bias vector, v 1,2i represents the implicit coding vector of the code-vulnerability rule semantic query response node between V1 and v 2i V represents the set of implicit coding vectors of the code-vulnerability rule semantic query response node, v 1,2 represents the set of implicit coding vectors of the code-vulnerability rule semantic query response node, v 1,21 , v 1,22 v 1,2i , v 1,2j , v 1,2n represent each implicit coding vector of the code-vulnerability rule semantic query response node in the set of implicit coding vectors of the code-vulnerability rule semantic query response node;
[0051] Based on the set of implicit coding vectors of the code-vulnerability rule semantic query response node, calculate the adjacency matrix of the code-vulnerability rule semantic query response node. This process can be expressed by the formula:
[0052]
[0053] where, ‖·‖ 2 is the square of the Euclidean norm of the vector, arccosh is the inverse hyperbolic cosine function, A ij represents the explicit association metric factor of the code-vulnerability rule semantic query response between v 1,2i and v 1,2j A represents the adjacency matrix of the code-vulnerability rule semantic query response node, A 11 , A n1 , A 1n , A nn represent the explicit association metric factors at each position in the adjacency matrix of the code-vulnerability rule semantic query response node;
[0054] Based on the set of implicit coding vectors of the code-vulnerability rule semantic query response node, calculate the degree distribution matrix of the code-vulnerability rule semantic query response node. This process can be expressed by the formula:
[0055]
[0056] where, represents the square of the one-norm of the vector, n represents the number of vectors in V minus one, D 1,2 in minus one, D i represents v1,2i and v 1,2j The implicit association metric factor between the code-vulnerability rule semantic query response, D represents the node degree distribution matrix of the code-vulnerability rule semantic query response, and D1 and D n represent the first and nth implicit association metric factors of the code-vulnerability rule semantic query response on the diagonal of the node degree distribution matrix of the code-vulnerability rule semantic query response;
[0057] Based on the node adjacency matrix of the code-vulnerability rule semantic query response and the node degree distribution matrix of the code-vulnerability rule semantic query response, calculate the Laplacian matrix of the code-vulnerability rule semantic query response. This process can be expressed by the formula:
[0058] L = D - A
[0059] where L represents the Laplacian matrix of the code-vulnerability rule semantic query response.
[0060] It should be understood that traditional shallow models are difficult to model the high-dimensional non-linear interactions between code and rules (such as vulnerability paths triggered by cross-function nested logic). Direct matching in high-dimensional space is vulnerable to the curse of dimensionality, while non-linear activation processing can construct an implicit space of the response state between code and vulnerability rules through non-linear mapping and interaction between features, enabling features that were originally difficult to distinguish to be effectively separated in this implicit space. Specifically, by dividing the structured encoding vector of code semantics by the embedded encoding vector of security vulnerability rule semantics to eliminate redundant feature differences (such as variable naming interference), then multiplying by the response weight matrix to introduce non-linear mapping (such as cross-data feature interaction), adding a bias vector to adjust the feature distribution, and using the Sigmoid activation function to further constrain the modulation result to the [0,1] interval to enhance the sensitivity to key response signals (such as highlighting the association between "unencrypted transmission" and network API calls). That is, the generated implicit encoding vector of the code-vulnerability rule semantic query response node can initially separate semantic association features and provide a basis for subsequent graph structure modeling.
[0061] Accordingly, traditional methods ignore the graph structure associations between code and rules (such as the correspondence between data flows and trigger paths), resulting in the failure of multi-hop reasoning (such as the encryption failure chain caused by hard-coded keys). By calculating the adjacency matrix of the code-vulnerability rule semantic query response nodes, the local structural relationships between the implicit coding vectors of the modeled code-vulnerability rule semantic query response nodes can be displayed. Specifically, the element values in the adjacency matrix reflect the state semantic association degree between node pairs (such as the similarity between the code function call chain and the rule sensitive conditions in the latent space), which is essentially a discrete approximation of the graph structure of the code-rule interaction pattern. That is, the generated adjacency matrix of the code-vulnerability rule semantic query response nodes transforms the complex non-linear relationships between high-dimensional features (such as cross-level conditional branches) into graph connection weights, which can provide a structural prior for subsequent spectral analysis and solve the blind spot of vulnerability chain identification caused by the lack of structure in traditional methods.
[0062] It should be understood that the sparsity of high-dimensional features easily causes the model to ignore key nodes (such as core functions that participate in data flow transmission multiple times). By recording the node connection strength with the diagonal elements of the code-vulnerability rule semantic query response node degree distribution matrix, its "structural importance" is quantified (such as the higher weight of code nodes that frequently participate in multi-hop vulnerability triggers). This matrix can achieve two functions: First, by weighting the connection quantity, the sensitivity to key nodes (such as encryption function call points) is enhanced, and the interference of marginal noise nodes is suppressed; Second, it provides a basis for adaptive adjustment of the Laplacian matrix, enabling spectral analysis to dynamically correct the structural representation according to the differences in node centrality (such as strengthening the dominant role of high-connection nodes in the frequency domain analysis of graph signals), so as to more accurately reflect the internal manifold characteristics of the data.
[0063] Accordingly, the code-vulnerability rule semantic query response Laplacian matrix, as the core of spectral graph theory, can transform the graph structure information of the code-vulnerability rule semantic query response node adjacency matrix and the code-vulnerability rule semantic query response node degree distribution matrix into an algebraic representation. The essence of its construction process is the preprocessing of the frequency domain analysis of the graph structure, which combines the local connection relationship (adjacency matrix) with the global node distribution (degree distribution matrix) to generate an operator that reflects the smoothness of the graph signal. Specifically, the eigenvalues of this matrix correspond to the frequency components of the graph, and the eigenvector corresponding to the smallest eigenvalue describes the low-dimensional embedding direction of the data manifold (such as the topological structure of the encryption-storage-key call chain), which can provide a basis for subsequent spectral decomposition and solve the problem that it is difficult to directly model complex structures in high-dimensional spaces.
[0064] Specifically, in the embodiment of this application, the step S512 includes: performing pseudo-inverse-driven topological closure optimization on the code-vulnerability rule semantic query response Laplacian matrix to obtain an optimized code-vulnerability rule semantic query response Laplacian matrix, and this process can be expressed by the formula:
[0065]
[0066]
[0067]
[0068] Among them, M1 represents the feature matrix of the adjacent connected components of the code-vulnerability rule semantic query response nodes, and M2 represents the feature matrix of the degree distribution connected components of the code-vulnerability rule semantic query response nodes. denotes addition by position points, exp represents the exponential function value with the natural constant e as the base, and L' represents the Laplacian matrix of the optimized code-vulnerability rule semantic query response nodes;
[0069] Performing spectral decomposition on the Laplacian matrix of the optimized code-vulnerability rule semantic query response to obtain the set of core coding vectors of the code-vulnerability rule semantic query response nodes. This process can be represented by the formula:
[0070]
[0071] Among them, Spectral Decomposition represents the spectral decomposition operation, U represents the set of core coding vectors of the code-vulnerability rule semantic query response nodes, and x1, x2, …, x k represent each core coding vector of the code-vulnerability rule semantic query response nodes in the set of core coding vectors of the code-vulnerability rule semantic query response nodes, and diag(λ1, λ2, …, λ k ) represents the code-vulnerability rule semantic query response core diagonal matrix with elements λ1, λ2, …, λ on the diagonal, Λ represents the code-vulnerability rule semantic query response core diagonal matrix, and λ1, λ2, …, λ k are the weight values of each core coding vector of the code-vulnerability rule semantic query response nodes respectively. k
[0072] Specifically, during the construction of the Laplacian matrix for code-vulnerability rule semantic query responses, since the limit point representation of nodes in the graph structure may lead to the discretization of topological information due to high-dimensional sparsity or noise interference (such as the disconnection of some subgraph connections or the interference of redundant loops), it is necessary to enhance the completeness and robustness of its structural representation through topological closure operations. Specifically, based on the pseudo-inverse operation, the connected component features of the code-vulnerability rule semantic query response node adjacency matrix and the code-vulnerability rule semantic query response node degree distribution matrix are extracted. The separated subspace that describes the direct association between nodes (such as the independent connected path of the encryption function call chain) can be disjunctively extracted from the adjacency matrix, and the loop subspace that reflects the node centrality dependence (such as the multi-hop data flow loop triggered by key hardcoding) can be disjunctively extracted from the degree distribution matrix. Subsequently, the local connectivity is strengthened through the separated subspace closure operation (such as the missing edge repair based on matrix completion) to ensure the integrity of the core vulnerability trigger path; at the same time, the interference of non-critical paths is suppressed by means of the loop subspace closure operation (such as loop weight normalization). Finally, the optimized code-vulnerability rule semantic query response Laplacian matrix extracts the global topological invariants in the graph structure (such as the cross-module propagation invariance of the encryption failure chain) by fusing the principal component features of the separated and loop spaces, so that the spectral features not only retain the connection strength of key nodes (such as the high-degree weight of the unencrypted transmission interface), but also eliminate the influence of noise subgraphs on low-frequency signals (such as the pseudo-association of temporary variable storage). This pseudo-inverse-driven topological closure mechanism essentially transforms the discretized topological information of the original matrix into a compact representation in the continuous manifold space, thereby significantly enhancing the structural characterization ability of the Laplacian matrix for complex vulnerability scenarios (such as multi-step logical defects), and laying a robust and interpretable mathematical foundation for subsequent spectral decomposition and dynamic query response encoding generation.
[0073] It should be understood that spectral decomposition is a process of performing eigenvalue decomposition on the Laplacian matrix of the optimized code-vulnerability rule semantic query response, aiming to extract the low-dimensional structure information contained in the graph spectrum. Specifically, through orthogonal decomposition of the Laplacian matrix of the optimized code-vulnerability rule semantic query response, the model can filter out the eigenvectors corresponding to the smallest eigenvalues (such as the first k eigenvectors). These eigenvectors constitute the low-frequency components of the graph signal, characterizing the main structure direction of the data in the manifold space (such as the core path of the interaction between vulnerability rules and code). For example, when analyzing the "sensitive data transmitted without encryption" vulnerability, the first eigenvector may correspond to the global defect of "the data transmission path is not encrypted", while the subsequent eigenvectors may characterize local features such as specific network interface calls or missing encryption algorithms. By retaining the first k eigenvectors, the model realizes non-linear dimensionality reduction, mapping high-dimensional nodes to the low-dimensional spectral domain space (such as from thousands of dimensions to dozens of dimensions), thereby concentrating the key semantic association features. This step not only significantly reduces the computational complexity but also suppresses high-dimensional noise (such as redundant variable naming interference), enabling the core coding vector to accurately represent the essential logic of vulnerability triggering.
[0074] Specifically, in the embodiment of the present application, the step S513 is used to: perform complementary information adaptive integration on the set of core coding vectors of the code-vulnerability rule semantic query response nodes to obtain the code-vulnerability rule global semantic query response coding vector, which can be expressed by the following formula:
[0075]
[0076] where AF(·) is the complementary information adaptive integration operation, W 2i and b 2i are the transformation scoring vector, linear transformation matrix, and linear bias vector corresponding to x i respectively, a i represents the code-vulnerability rule semantic query response feature compensation factor corresponding to x i , softmax represents the softmax function, τ represents the preset threshold, mask represents the gating function, w si represents the code-vulnerability rule semantic query response feature weight factor corresponding to x i , and v f represents the code-vulnerability rule global semantic query response coding vector.
[0077] It should be understood that although the set of core encoding vectors of the code-vulnerability rule semantic query response captures the principal component features of the data, a single vector may not be able to cover multi-granularity semantic information (such as the complementarity of global logic and local code patterns). Therefore, a multi-vector information is dynamically fused through an adaptive integration mechanism. Specifically, the model assigns dynamic weights to each core encoding vector, and the weight values are jointly determined by the current query context (such as the control flow structure of the target code snippet) and the semantic contribution degree of the vector itself. For example, when detecting the vulnerability of "user input bypassing verification to trigger dangerous operations", the model may assign a high weight to the vector representing "lack of input verification" and a low weight to the vector representing "temporary variable storage". This process normalizes the weights through a learnable gating function (such as Softmax) to ensure that the contribution degrees of different vectors are interpretable within the range of [0,1]. The finally generated code-vulnerability rule global semantic query response encoding vector not only aggregates multi-dimensional vulnerability features (such as data flow defects, lack of permission verification), but also enhances the robustness to complex scenarios (such as distinguishing the risks of temporary caching and persistent storage) through the adaptive adjustment of weights. This fusion mechanism enables the model to accurately identify multi-step attack chains (such as "input → unencrypted storage → key leakage"), while avoiding misjudgments caused by single feature deviations.
[0078] In step S52, based on the code-vulnerability rule global semantic query response encoding vector, the detection result is obtained, and the detection result is used to indicate whether the confidence level of the existence of a security vulnerability in the mobile application exceeds a predetermined threshold. Specifically, in the embodiment of the present application, step S52 includes: inputting the code-vulnerability rule global semantic query response encoding vector into a security assessment detection module based on a classifier to obtain the detection result. It should be understood that the code-vulnerability rule global semantic query response encoding vector is the result obtained after a series of semantic analyses and query matches on the mobile application code and security vulnerability rules. However, this vector itself is only a numerical representation and cannot directly determine whether there is a security vulnerability in the mobile application and the likelihood of the existence of a vulnerability. The security assessment detection module based on a classifier has such a judgment ability. Specifically, the classifier can learn the feature patterns in the encoding vector through training with a large amount of training data, and based on these features, judge the matching situation between the mobile application code represented by the input encoding vector and the security vulnerability rules, and then determine the confidence level of the existence of a security vulnerability in the mobile application. That is, by inputting the code-vulnerability rule global semantic query response encoding vector into the classifier, the complex semantic matching result can be converted into a specific and quantifiable detection result, thereby providing a basis for subsequent security decisions. Particularly, in a specific embodiment of the present application, inputting the code-vulnerability rule global semantic query response encoding vector into a security assessment detection module based on a classifier to obtain the detection result includes: using the fully connected layer of the classifier to perform fully connected encoding on the code-vulnerability rule global semantic query response encoding vector to obtain a code-vulnerability rule global semantic query response fully connected encoding feature vector; inputting the code-vulnerability rule global semantic query response fully connected encoding feature vector into the Softmax classification function of the classifier to obtain the probability values of the code-vulnerability rule global semantic query response encoding vector belonging to each classification label, where the classification labels include those indicating that the confidence level of the existence of a security vulnerability in the mobile application exceeds a predetermined threshold and those indicating that the confidence level of the existence of a security vulnerability in the mobile application does not exceed a predetermined threshold; and determining the classification label corresponding to the largest of the probability values as the detection result.
[0079] In step S6, based on the detection result, it is determined whether to generate a security alert prompt. Specifically, in the embodiment of the present application, step S6 includes: in response to the confidence level of the existence of a security vulnerability in the mobile application exceeding a predetermined threshold, generating the security alert prompt. It should be understood that when the confidence level of detecting a security vulnerability in the mobile application exceeds the predetermined threshold, the system automatically triggers a security alert prompt (such as "High-risk vulnerability: Hard-coded password (confidence level 92%)") and accurately locates it to the specific code location (such as file path, line number), while associating the vulnerability type and repair suggestions to provide a targeted solution for the development team to shorten the troubleshooting cycle. And by dynamically adjusting the confidence level threshold (such as setting a strict threshold of 95% for financial applications and a loose threshold of 80% for internal tools), the system can balance the risks of false positives and false negatives, optimize the resource allocation priority (such as giving priority to high-risk vulnerabilities), and generate structured logs (including timestamps, vulnerability details) to meet compliance audits. This can significantly improve the timeliness of the security threat response. For example, after a social application detects a "vulnerability of storing user privacy data in plain text", the security alert prompt is pushed to the security team in real time, driving the completion of the encryption logic rectification on the same day to avoid the continuous exposure of sensitive information; at the same time, the historical alert data can analyze the distribution law of vulnerabilities (such as frequent hard-coding problems), drive the iteration of the rule library (add dynamic key management detection) and the update of the development specification (mandate input verification in the mandatory code review link), and achieve cross-team collaboration by integrating into the DevOps tool chain (such as automatically creating Jira work orders), eliminating information silos, and finally forming a governance loop of "detection - alert - repair - verification" to promote the continuous evolution of the security protection ability.
[0080] In summary, the mobile application security assessment and automated detection method based on the embodiment of the present application is elucidated. It uses artificial intelligence-based data processing technology to perform syntax analysis and semantic understanding on the obtained code of the mobile application to obtain the code semantic structured coding features, and at the same time performs semantic embedding coding on the set of security vulnerability rules extracted from the security vulnerability rule library to obtain the set of security vulnerability rule semantic embedding coding features. Subsequently, based on the dynamic query analysis representation of the code semantic structured coding features and the set of security vulnerability rule semantic embedding coding features, the detection result is automatically obtained, and a security alert prompt is generated in response to the detection result that the confidence level of the existence of a security vulnerability in the mobile application exceeds a predetermined threshold. In this way, the accuracy and reliability of the security vulnerability detection result can be effectively improved.
Claims
1. A mobile application security assessment and automated detection method, characterized in that, Including: Obtaining the code of the mobile application; Performing code semantic analysis on the code of the mobile application to obtain a code semantic structured encoding vector; Extracting a set of security vulnerability rules from a security vulnerability rule library; Performing semantic embedding encoding on the set of security vulnerability rules to obtain a set of security vulnerability rule semantic embedding encoding vectors; Performing a security vulnerability rule global semantic scan on the code semantic structured encoding vector and the set of security vulnerability rule semantic embedding encoding vectors to obtain a detection result, including: performing a dynamic query analysis guided by security vulnerability rule semantic features on the code semantic structured encoding vector and the set of security vulnerability rule semantic embedding encoding vectors to obtain a code-vulnerability rule global semantic query response encoding vector; based on the code-vulnerability rule global semantic query response encoding vector, obtaining the detection result, and the detection result is used to represent whether the confidence level of the existence of security vulnerabilities in the mobile application exceeds a predetermined threshold; Based on the detection result, determining whether to generate a security alert prompt.
2. The mobile application security assessment and automated detection method according to claim 1, wherein Performing code semantic analysis on the code of the mobile application to obtain a code semantic structured encoding vector, including: Performing syntax analysis on the code of the mobile application to obtain a code abstract syntax tree; Performing code semantic understanding based on a tree-shaped long short-term memory network on the code abstract syntax tree to obtain the code semantic structured encoding vector.
3. The mobile application security assessment and automated detection method according to claim 1, wherein Performing semantic embedding encoding on the set of security vulnerability rules to obtain a set of security vulnerability rule semantic embedding encoding vectors, including: using a semantic encoder based on the BERT-TextCNN model to perform semantic embedding encoding on each security vulnerability rule in the set of security vulnerability rules to obtain the set of security vulnerability rule semantic embedding encoding vectors.
4. The mobile application security assessment and automated detection method according to claim 1, characterized in that Performing a dynamic query analysis guided by security vulnerability rule semantic features on the code semantic structured encoding vector and the set of security vulnerability rule semantic embedding encoding vectors to obtain a code-vulnerability rule global semantic query response encoding vector, including: Calculating a code-vulnerability rule semantic query response Laplacian matrix based on the code semantic structured encoding vector and the set of security vulnerability rule semantic embedding encoding vectors; Performing spectral decomposition on the code-vulnerability rule semantic query response Laplacian matrix to obtain a set of code-vulnerability rule semantic query response node core encoding vectors; Performing complementary information adaptive integration on the set of code-vulnerability rule semantic query response node core encoding vectors to obtain the code-vulnerability rule global semantic query response encoding vector.
5. The mobile application security assessment and automated detection method according to claim 4, characterized in that Calculating a code-vulnerability rule semantic query response Laplacian matrix based on the code semantic structured encoding vector and the set of security vulnerability rule semantic embedding encoding vectors, including: Performing non-linear activation processing on each security vulnerability rule semantic embedding encoding vector in the set of code semantic structured encoding vectors and the set of security vulnerability rule semantic embedding encoding vectors to obtain a set of code-vulnerability rule semantic query response node implicit encoding vectors; Based on the set of implicit encoding vectors of the code-vulnerability rule semantic query response nodes, calculate the adjacency matrix of the code-vulnerability rule semantic query response nodes; Based on the set of implicit encoding vectors of the code-vulnerability rule semantic query response nodes, calculate the degree distribution matrix of the code-vulnerability rule semantic query response nodes; Based on the adjacency matrix of the code-vulnerability rule semantic query response nodes and the degree distribution matrix of the code-vulnerability rule semantic query response nodes, calculate the Laplacian matrix of the code-vulnerability rule semantic query response; 6. The mobile application security evaluation and automated detection method according to claim 5, wherein Perform spectral decomposition on the Laplacian matrix of the code-vulnerability rule semantic query response to obtain a set of core encoding vectors of the code-vulnerability rule semantic query response nodes, including: Perform topology-closure optimization based on pseudo-inverse driving on the Laplacian matrix of the code-vulnerability rule semantic query response to obtain an optimized Laplacian matrix of the code-vulnerability rule semantic query response; Perform spectral decomposition on the optimized Laplacian matrix of the code-vulnerability rule semantic query response to obtain the set of core encoding vectors of the code-vulnerability rule semantic query response nodes.
7. The mobile application security evaluation and automated detection method according to claim 1, wherein Based on the code-vulnerability rule global semantic query response encoding vector, obtain the detection result, including: inputting the code-vulnerability rule global semantic query response encoding vector into a security assessment detection module based on a classifier to obtain the detection result.
8. The mobile application security assessment and automated detection method according to claim 7, wherein Based on the detection result, determine whether to generate a security alert prompt, including: in response to the confidence that there is a security vulnerability in the mobile application program in the detection result exceeding a predetermined threshold, generate the security alert prompt.
Citation Information
Cited By
Code-level security vulnerability detection method, electronic equipment and storage medium
CN120744938A
Method, system and device for realizing security detection of mobile application based on multi-engine cooperation, processor and readable storage medium thereof
CN122490511A