Software risk assessment method, device, storage medium, program product and equipment
By analyzing and feature extraction of software code, generating word vector features and combining static features to calculate the overall risk score, the problems of low efficiency and low accuracy in large-scale code evaluation are solved, efficient and accurate risk identification is achieved, and software quality and security are improved.
Patent Information
- Application Number
- CN202510131246.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-06
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-02-06
AI Technical Summary
Traditional software risk assessment methods are inefficient and have low accuracy when processing large-scale code, making it difficult to efficiently and accurately identify and analyze risks in software.
By parsing the code file based on the software to be evaluated, code information is extracted, and the feature extraction of generative word vector features is calculated. Combined with static feature risk scores, the overall risk score is calculated to achieve efficient and accurate risk identification.
It realizes efficient and accurate identification of risks in software, and improves software quality and software security.
Smart Images

Figure CN119577790B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of software engineering technology, and in particular to a software risk assessment method, device, storage medium, program product and equipment. Background Art
[0002] As the complexity and scale of software systems continue to increase, software component analysis has become an important task in software engineering. Traditional analysis methods usually require manual risk assessment of software source code according to set risk indicators. However, this has the problems of low efficiency and low accuracy when dealing with large-scale codes.
[0003] With the popularity of open source software, there are a large number of source codes on the Internet for developers to call. If the source code with poor code quality is accidentally called, it will pose a certain threat to the developed software project (such as causing the developed software to crash). Therefore, how to efficiently and accurately identify and analyze the risk level in the software has become an important issue in improving software quality and software security. Summary of the invention
[0004] In order to solve the above technical problems, the embodiments of the present application propose a software risk assessment method, device, storage medium, program product and equipment, which can efficiently and accurately identify the risks existing in the software to improve software quality and software security.
[0005] In a first aspect, an embodiment of the present application provides a software risk assessment method, comprising:
[0006] Parsing the code file of the software to be evaluated to obtain code information, wherein the code information is used to indicate the components and library functions corresponding to the code file;
[0007] Perform feature extraction based on the code information to obtain word vector features, and determine a word vector risk score based on the word vector features;
[0008] Based on the code file, performing risk assessment on the component to obtain a first static feature risk score, and performing risk assessment on the library function to obtain a second static feature risk score;
[0009] Determining a static feature risk score based on the first static feature risk score and the second static feature risk score;
[0010] An overall risk score is determined based on the word vector risk score and the static feature risk score.
[0011] Optionally, the code information includes an abstract syntax tree, and the abstract syntax tree is used to indicate the component and the library function. The code information is obtained by parsing the code file based on the software to be evaluated, including:
[0012] Preprocessing the code file to obtain target code;
[0013] Determine a lexical analyzer and a syntax analyzer based on grammatical rules matching the code file;
[0014] Calling the lexical analyzer to decompose the target code into at least one lexical unit;
[0015] The syntax analyzer is called to construct the abstract syntax tree based on the at least one token unit.
[0016] Optionally, the extracting features based on the code information to obtain word vector features, and determining the word vector risk score according to the word vector features, includes:
[0017] Using a neural network model, extracting features from the code information to generate at least one word vector;
[0018] Based on the at least one word vector, construct the word vector feature;
[0019] The absolute value of each dimension in the word vector feature is calculated, and the sum of the absolute values of each dimension is used as the word vector risk score.
[0020] Optionally, the using of a neural network model to extract features from the code information to generate at least one word vector includes:
[0021] Determine hyperparameters including window size and minimum word frequency based on the grammatical structure and contextual relationship indicated by the code information;
[0022] Configuring the determined hyperparameters for the neural network model;
[0023] The configured neural network model is used to extract features from the code information to generate at least one word vector.
[0024] Optionally, constructing the word vector feature based on the at least one word vector includes:
[0025] Based on the code information, performing TF-IDF statistics on each word vector in the at least one word vector to obtain a weight corresponding to each word vector;
[0026] Normalizing each word vector in the at least one word vector respectively;
[0027] Based on at least one normalized word vector and its corresponding weight, a weighted calculation is performed to obtain the word vector feature.
[0028] Optionally, the first static feature risk score includes at least one of the following: a first vulnerability severity score, a first usage frequency score, a first version score, and a first license compliance score;
[0029] The second static feature risk score includes at least one of the following: a second vulnerability severity score, a second usage frequency score, a second version score, and a second license compliance score.
[0030] Optionally, the method further comprises:
[0031] Using a pre-built LDA-based component analysis and risk assessment model, identifying each component in the code file, and performing risk assessment on each component in the code file to obtain an LDA component risk score;
[0032] The determining of the overall risk score based on the word vector risk score and the static feature risk score includes:
[0033] The overall risk score is obtained by performing a weighted sum based on the LDA component risk score, the word vector risk score and the static feature risk score.
[0034] Optionally, the pre-built LDA-based component analysis and risk assessment model is used to identify the components in the code file, including:
[0035] Decomposing the code file into a plurality of sub-documents, wherein the plurality of sub-documents include class fragments, method fragments and comment fragments;
[0036] The LDA-based component analysis and risk assessment model is used to identify the multiple sub-documents to obtain the components in the code file.
[0037] In a second aspect, an embodiment of the present application provides a software risk assessment device, including:
[0038] A parsing module, used to parse the code file of the software to be evaluated to obtain code information, wherein the code information is used to indicate the components and library functions corresponding to the code file;
[0039] A first risk assessment module is used to perform feature extraction based on the code information to obtain word vector features, and determine a word vector risk score based on the word vector features;
[0040] A second risk assessment module is used to perform risk assessment on the component based on the code file to obtain a first static feature risk score, and to perform risk assessment on the library function to obtain a second static feature risk score;
[0041] a third risk assessment module, configured to determine a static feature risk score based on the first static feature risk score and the second static feature risk score;
[0042] The overall risk scoring module is used to determine the overall risk score based on the word vector risk score and the static feature risk score.
[0043] Optionally, the code information includes an abstract syntax tree, and the abstract syntax tree is used to indicate the component and the library function. The code information is obtained by parsing the code file based on the software to be evaluated, including:
[0044] Preprocessing the code file to obtain target code;
[0045] Determine a lexical analyzer and a syntax analyzer based on grammatical rules matching the code file;
[0046] Calling the lexical analyzer to decompose the target code into at least one lexical unit;
[0047] The syntax analyzer is called to construct the abstract syntax tree based on the at least one token unit.
[0048] Optionally, the extracting features based on the code information to obtain word vector features, and determining the word vector risk score according to the word vector features, includes:
[0049] Using a neural network model, extracting features from the code information to generate at least one word vector;
[0050] Based on the at least one word vector, construct the word vector feature;
[0051] The absolute value of each dimension in the word vector feature is calculated, and the sum of the absolute values of each dimension is used as the word vector risk score.
[0052] Optionally, the using of a neural network model to extract features from the code information to generate at least one word vector includes:
[0053] Determine hyperparameters including window size and minimum word frequency based on the grammatical structure and contextual relationship indicated by the code information;
[0054] Configuring the determined hyperparameters for the neural network model;
[0055] The configured neural network model is used to extract features from the code information to generate at least one word vector.
[0056] Optionally, constructing the word vector feature based on the at least one word vector includes:
[0057] Based on the code information, performing TF-IDF statistics on each word vector in the at least one word vector to obtain a weight corresponding to each word vector;
[0058] Normalizing each word vector in the at least one word vector respectively;
[0059] Based on at least one normalized word vector and its corresponding weight, a weighted calculation is performed to obtain the word vector feature.
[0060] Optionally, the first static feature risk score includes at least one of the following: a first vulnerability severity score, a first usage frequency score, a first version score, and a first license compliance score;
[0061] The second static feature risk score includes at least one of the following: a second vulnerability severity score, a second usage frequency score, a second version score, and a second license compliance score.
[0062] Optionally, the device further comprises:
[0063] A fourth risk assessment module, for identifying each component in the code file using a pre-built LDA-based component analysis and risk assessment model, and performing risk assessment on each component in the code file to obtain an LDA component risk score;
[0064] The determining of the overall risk score based on the word vector risk score and the static feature risk score includes:
[0065] The overall risk score is obtained by performing a weighted sum based on the LDA component risk score, the word vector risk score and the static feature risk score.
[0066] Optionally, the pre-built LDA-based component analysis and risk assessment model is used to identify the components in the code file, including:
[0067] Decomposing the code file into a plurality of sub-documents, wherein the plurality of sub-documents include class fragments, method fragments and comment fragments;
[0068] The LDA-based component analysis and risk assessment model is used to identify the multiple sub-documents to obtain the components in the code file.
[0069] In a third aspect, an embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the steps of any of the software risk assessment methods described above are implemented.
[0070] In a fourth aspect, an embodiment of the present application provides a computer program product, comprising computer instructions, which, when executed by a processor, implement the steps of any of the software risk assessment methods described above.
[0071] In a fifth aspect, an embodiment of the present application provides a computer device, comprising a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein the processor implements the steps of any of the above-mentioned software risk assessment methods when executing the computer program.
[0072] In summary, the embodiments of the present application have at least the following beneficial effects:
[0073] According to an embodiment of the present application, code information is obtained by parsing the code file based on the software to be evaluated, wherein the code information is used to indicate the components and library functions corresponding to the code file; feature extraction is performed based on the code information to obtain word vector features, and a word vector risk score is determined based on the word vector features; based on the code file, a risk assessment is performed on the component to obtain a first static feature risk score, and a risk assessment is performed on the library function to obtain a second static feature risk score; based on the first static feature risk score and the second static feature risk score, a static feature risk score is determined; based on the word vector risk score and the static feature risk score, an overall risk score is determined, thereby being able to efficiently and accurately identify risks in the software to improve software quality and software security. BRIEF DESCRIPTION OF THE DRAWINGS
[0074] Figure 1 It is a flowchart of a software risk assessment method provided in an embodiment of the present application;
[0075] Figure 2 is a schematic diagram of the preprocessing provided in the embodiment of the present application;
[0076] Figure 3 It is a schematic diagram of lexical analysis and grammatical analysis provided by an embodiment of the present application;
[0077] Figure 4 It is a schematic diagram of constructing word vector features provided in an embodiment of the present application;
[0078] Figure 5 is a schematic diagram of risk assessment provided by an embodiment of the present application;
[0079] Figure 6is a schematic diagram of overall risk score calculation and risk rating provided in an embodiment of the present application;
[0080] Figure 7 is a structural diagram of a software risk assessment device provided in an embodiment of the present application;
[0081] Figure 8 It is a schematic diagram of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0082] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0083] In the description of the present application, the terms "first", "second", "third", etc. are used for descriptive purposes only and are not to be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Thus, features defined as "first", "second", "third", etc. may explicitly or implicitly include one or more of the features. In the description of the present application, unless otherwise specified, the meaning of "multiple" is two or more. In the description of the present application, the term "including" and its variations are open inclusions, i.e., "including but not limited to". The term "based on" means "at least partially based on". The term "according to" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one other embodiment"; the term "some embodiments" means "at least some embodiments".
[0084] In the description of this application, it should be noted that, unless otherwise clearly specified and limited, the terms "installed", "connected", and "connected" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or it can be indirectly connected through an intermediate medium, or it can be the internal communication of two components. For ordinary technicians in this field, the specific meanings of the above terms in this application can be understood according to specific circumstances.
[0085] In the description of this application, it should be noted that, unless otherwise defined, all technical and scientific terms used in this application have the same meaning as those commonly understood by those skilled in the art. The terms used in the specification of this application are only for the purpose of describing specific embodiments and are not intended to limit this application. For those of ordinary skill in the art, the specific meanings of the above terms in this application can be understood by specific circumstances.
[0086] The following is an explanation of some terminology concepts involved in the embodiments of the present application:
[0087] TF-IDF (Term Frequency-Inverse Document Frequency) is a statistical method used to evaluate the importance of a word in a document or corpus.
[0088] LDA (Latent Dirichlet Allocation) is a topic model used to discover abstract topics from a document collection.
[0089] First, see Figure 1 , shows a schematic flow chart of a software risk assessment method provided in an embodiment of the present application, the method comprising steps S101-S105, which are specifically as follows:
[0090] S101, parsing a code file of the software to be evaluated to obtain code information, wherein the code information is used to indicate components and library functions corresponding to the code file;
[0091] In an example, the language type corresponding to the code file may be Java (an object-oriented programming language), and thus the code file is a corresponding Java source code file.
[0092] S102, performing feature extraction based on the code information to obtain word vector features, and determining a word vector risk score based on the word vector features;
[0093] S103, based on the code file, performing risk assessment on the component to obtain a first static feature risk score, and performing risk assessment on the library function to obtain a second static feature risk score;
[0094] S104, determining a static feature risk score based on the first static feature risk score and the second static feature risk score; in an example, the static feature risk score can be obtained by summing / weighted summing the first static feature risk score and the second static feature risk score.
[0095] S105: Determine an overall risk score based on the word vector risk score and the static feature risk score.
[0096] In an optional implementation, the code information includes an abstract syntax tree, the abstract syntax tree is used to indicate the component and the library function, and the code information is obtained by parsing the code file based on the software to be evaluated, including:
[0097] Preprocessing the code file to obtain target code;
[0098] In one example, see Figure 2 , preprocessing the code file to obtain the target code may include: extracting the source code file and the binary file generated after the software is compiled from the code file, and ensuring the integrity of the source code file and the binary file during the extraction process to avoid missing or losing the file during the extraction process; then, cleaning the source code file, wherein the cleaning process may include at least one of the following: accurately removing irrelevant characters, removing comments, and maintaining the consistency of the source code file encoding through regular expression characteristics; in addition, the source code file and the binary file may be renamed according to the naming convention of the language type (such as Java) corresponding to the code file, to ensure that the file name is clear, intuitive and does not contain special characters; finally, in the case where the language type corresponding to the code file is Java, the source code file and the binary file may be classified and organized according to the package name and functional module according to the modular design principle of the Java project for subsequent analysis. In this embodiment, the cleanliness and consistency of the data can be ensured, the data quality can be improved, and a standardized data set is provided for subsequent component analysis and risk detection, laying a solid foundation.
[0099] Determine a lexical analyzer and a syntax analyzer based on grammatical rules matching the code file;
[0100] In one example, see Figure 3 If the language used in the code file is Java, you can use the ANTLR (ANotherTool for Language Recognition, a syntax analyzer) syntax definition language according to the Java language specification to write Java syntax rules including class definition, method definition, variable declaration, etc., so as to use the ANTLR tool to automatically generate efficient lexical analyzers and syntax analyzers according to the defined Java syntax rules.
[0101] Calling the lexical analyzer to decompose the target code into at least one lexical unit;
[0102] The syntax analyzer is called to construct the abstract syntax tree based on the at least one token unit.
[0103] In one example, see Figure 3, when the language used in the code file is Java, the lexical analyzer can be used to decompose the target code into a series of lexical units (tokens), where lexical units are basic elements that constitute Java programs, such as keywords, identifiers, literals, etc. The syntax analyzer can cleverly use these lexical units to construct an AST (Abstract Syntax Tree) based on the defined grammatical rules. The tree structure of the abstract syntax tree can clearly show the grammatical hierarchy and logical relationship of the Java source code, providing a solid foundation for subsequent in-depth analysis. During the parsing process, this embodiment pays special attention to the characteristics of the Java language, such as the use of rich standard libraries and third-party dependency packages, as well as complex grammatical structures (such as inner classes, anonymous classes, interfaces and implementations, etc.). Therefore, the lexical analyzer and syntax analyzer of this embodiment have been carefully designed and optimized to ensure that these characteristics can be accurately processed, thereby generating a complete and accurate AST.
[0104] In an optional implementation, the extracting features based on the code information to obtain word vector features, and determining the word vector risk score according to the word vector features, includes:
[0105] Using a neural network model, extracting features from the code information to generate at least one word vector;
[0106] In one example, see Figure 4 , the neural network model can be a Word2Vec model. In the case where the language used in the code file is Java, since Java code has a strict grammatical structure and rich contextual relationships, it is possible to use the Word2Vec model to capture the semantic relationship between functions, variables and classes. In this way, the vocabulary (including keywords, variable names, function names, class names, etc.) in the code can be extracted by performing vocabulary extraction preprocessing on the Java code, and then the vector representation of the extracted vocabulary is learned using the Word2Vec model to generate at least one word vector (i.e., generate word embedding). In this embodiment, special attention is paid to some unique characteristics of Java code, such as package import, class inheritance, interface implementation, etc. These characteristics usually appear in the code with specific grammatical structures. Therefore, when constructing the vocabulary library, this embodiment retains these grammatical structure information (i.e., the grammatical structure contained in the code information is suitable for representing the above characteristics) so that the Word2Vec model can better understand the contextual relationship in the code.
[0107] Based on the at least one word vector, construct the word vector feature;
[0108] In some cases, when the neural network model is a Word2Vec model and the language used in the code file is Java, since the Word2Vec model is divided based on vocabulary frequency, this may cause some important Java packages (such as core libraries) to be underestimated by importing only once, while unimportant words may be overestimated due to frequent appearance. In order to solve this problem, this embodiment can introduce a weight allocation mechanism when constructing word vector features, and assign weights according to the importance and frequency of use of each word in the code file. For example, a higher weight is given to the word vectors corresponding to important packages and classes (such as java.util, java.io, etc.); and a lower weight is given to the word vectors corresponding to some frequently appearing but less important words (such as temporary variable names, loop variables, etc.). In this way, when constructing word vector features, the word vectors corresponding to important words will have a greater impact on the construction results, thereby improving the quality of the constructed word vector features (i.e., word vector representations).
[0109] The absolute value of each dimension in the word vector feature is calculated, and the sum of the absolute values of each dimension is used as the word vector risk score.
[0110] In an optional implementation, the use of a neural network model to extract features from the code information to generate at least one word vector includes:
[0111] Based on the grammatical structure and contextual relationship indicated by the code information, determine hyperparameters including window size and minimum word frequency, wherein the window size is suitable for determining the scope of the context to ensure that the model can capture sufficient contextual information, and the minimum word frequency is suitable for filtering out words with too low frequency of occurrence to reduce the impact of noise on the quality of the word vector, wherein the window size and / or the minimum word frequency can be determined according to the grammatical structure and contextual relationship;
[0112] Configuring the determined hyperparameters for the neural network model;
[0113] The configured neural network model is used to extract features from the code information to generate at least one word vector.
[0114] In an optional implementation, constructing the word vector feature based on the at least one word vector includes:
[0115] Based on the code information, performing TF-IDF statistics on each word vector in the at least one word vector to obtain a weight (TF-IDF weight) corresponding to each word vector, wherein the TF-IDF weight is suitable for representing the frequency and importance of the vocabulary corresponding to the word vector in the code snippet;
[0116] Normalizing each word vector in the at least one word vector, respectively, wherein the normalization refers to normalizing each word vector to ensure comparability between different word vectors;
[0117] Based on at least one normalized word vector and its corresponding weight, a weighted calculation is performed to obtain the word vector feature.
[0118] In this embodiment, when the neural network model is a Word2Vec model, the advantages of TF-IDF and Word2Vec word vectors are combined to more comprehensively reflect the semantic features of the code snippet.
[0119] In an optional embodiment, the first static feature risk score includes at least one of the following: a first vulnerability severity score, a first usage frequency score, a first version score, and a first license compliance score;
[0120] The second static feature risk score includes at least one of the following: a second vulnerability severity score, a second usage frequency score, a second version score, and a second license compliance score.
[0121] In one example, see Figure 5 , the scoring criteria for risk assessment of static features may include: VSS (Vulnerability Severity Scoring) based on CVSS (Common Vulnerability Scoring System), UFS (Usage Frequency Score), VS (Version Scoring), and LCS (License Compliance Score). Among them, UFS means that if a component that is used frequently has a problem, the impact will be greater; VS means that the latest version scores higher and the outdated version scores lower; LCS means that the license type that complies with project requirements scores higher. Therefore, SF (Static Feature-based Risk Assessment) can be calculated by combining one or more of the above scoring criteria.
[0122] In one example, the first static feature risk score or the second static feature risk score It can be determined by the following formula:
[0123]
[0124] in, , , , These are the weights corresponding to VSS, UFS, VS, and LCS respectively.
[0125] In an optional embodiment, the method further includes:
[0126] Using a pre-built LDA-based component analysis and risk assessment model, identifying each component in the code file, and performing risk assessment on each component in the code file to obtain an LDA component risk score;
[0127] In one example, the LDA component risk score It can be determined by the following formula:
[0128]
[0129] in, It is a function used to calculate LCRS (LDA-based risk assessment), which is used to convert CI (Component Interdependencies) and RC (Risk Characteristics) into risk scores. It can be calculated based on the distribution of the LDA model output, indicating the similarity or degree of association between components. It may include factors such as known vulnerability patterns, code complexity, external library dependencies, etc. Wherein f may represent addition, multiplication, weighted sum, etc., which are not specifically limited here.
[0130] The determining of the overall risk score based on the word vector risk score and the static feature risk score includes:
[0131] The overall risk score is obtained by performing a weighted sum based on the LDA component risk score, the word vector risk score and the static feature risk score.
[0132] In one example, the overall risk score It can be determined by the following formula:
[0133]
[0134] in, , , They are respectively word vector risk score , LDA component risk score , Static feature risk score The corresponding weight.
[0135] In an optional implementation, the pre-built LDA-based component analysis and risk assessment model is used to identify the components in the code file, including:
[0136] Decomposing the code file into a plurality of sub-documents, wherein the plurality of sub-documents include class fragments, method fragments and comment fragments;
[0137] The LDA-based component analysis and risk assessment model is used to identify the multiple sub-documents to obtain the components in the code file.
[0138] In one example, see Figure 5 In order to enable LDA to better identify different components in the code file, the code file can be first parsed into multiple fragments such as classes, methods, comments, etc., and each fragment can be separated into a small "sub-document". Then build the LDA model and set the hyperparameters of LDA, such as the number of topics (tc), Dirichlet prior (α and β), and related parameters of Gibbs sampling (such as the number of sampling times n and the sampling interval si). The choice of these parameters will directly affect the performance and results of the LDA model. Through the LDA model, different topics or components can be extracted from code files (such as Java documents). These components may correspond to different modules, classes, or functions in the source code. By analyzing these components, you can have a deeper understanding of the structure and composition of the source code.
[0139] In one example, see Figure 6 , the risk rating can be divided into AE levels according to the overall risk score, where:
[0140] Level A (Secure): Risk score 0-1, no known vulnerabilities, and all components are the latest versions.
[0141] Level B (Low Risk): Risk score 1-3, a small number of low-risk vulnerabilities, and external libraries are all the latest versions.
[0142] Level C (Medium Risk): Risk score 3-5, multiple known vulnerabilities, and some external libraries are out of date.
[0143] Level D (High Risk): Risk score 5-7, severe vulnerabilities, and multiple key components and libraries are not updated.
[0144] Level E (unacceptable risk): Risk score is above 7, the system has major security risks and needs to be immediately deactivated and fully reviewed.
[0145] As described in the above related embodiments, this application can effectively identify potential security vulnerabilities and risks in software by combining word vector technology with machine learning classifiers. The use of neural network models (such as the Word2Vec model) to capture the semantic relationship of the code improves the depth of understanding of components and library functions, thereby improving the accuracy of risk assessment, so that defects can be analyzed automatically, manual intervention can be reduced, and more comprehensive risk detection can be ensured, thereby providing a solid foundation for software maintenance and security management.
[0146] In addition, a risk assessment model based on multi-dimensional scoring (i.e., a risk assessment system) has been established to conduct detailed risk analysis for different components and library functions. By combining factors such as vulnerability severity, usage frequency, version updates, and license compliance, an overall risk score is constructed to make the risk assessment results more scientific and reasonable. This flexible rating standard can help development teams quickly identify high-risk components and develop priority repair plans, significantly improving the security and reliability of the software.
[0147] Moreover, the constructed risk assessment model has strong adaptability and extensibility. When coping with the ever-changing software environment and technical requirements, the scoring criteria can be flexibly adjusted or new indicators can be added. For example, if new types of security vulnerabilities or component usage need to be considered, only the corresponding data inputs need to be updated without rebuilding the entire model. This modular design makes the present invention widely applicable and of high practical value in practical applications.
[0148] On the second aspect, accordingly, the embodiments of the present application also provide a software risk assessment device, which can implement all the processes of the software risk assessment method provided in the above embodiments.
[0149] See also Figure 7 , shows a schematic diagram of the structure of a software risk assessment device provided in an embodiment of the present application, the software risk assessment device comprising:
[0150] The parsing module 701 is used to parse the code file of the software to be evaluated to obtain code information, wherein the code information is used to indicate the components and library functions corresponding to the code file;
[0151] A first risk assessment module 702 is used to perform feature extraction based on the code information to obtain word vector features, and determine a word vector risk score based on the word vector features;
[0152] A second risk assessment module 703 is used to perform risk assessment on the component based on the code file to obtain a first static feature risk score, and to perform risk assessment on the library function to obtain a second static feature risk score;
[0153] A third risk assessment module 704, configured to determine a static feature risk score based on the first static feature risk score and the second static feature risk score;
[0154] The overall risk score module 705 is used to determine an overall risk score based on the word vector risk score and the static feature risk score.
[0155] In an optional implementation, the code information includes an abstract syntax tree, the abstract syntax tree is used to indicate the component and the library function, and the code information is obtained by parsing the code file based on the software to be evaluated, including:
[0156] Preprocessing the code file to obtain target code;
[0157] Determine a lexical analyzer and a syntax analyzer based on grammatical rules matching the code file;
[0158] Calling the lexical analyzer to decompose the target code into at least one lexical unit;
[0159] The syntax analyzer is called to construct the abstract syntax tree based on the at least one token unit.
[0160] In an optional implementation, the extracting features based on the code information to obtain word vector features, and determining the word vector risk score according to the word vector features, includes:
[0161] Using a neural network model, extracting features from the code information to generate at least one word vector;
[0162] Based on the at least one word vector, construct the word vector feature;
[0163] The absolute value of each dimension in the word vector feature is calculated, and the sum of the absolute values of each dimension is used as the word vector risk score.
[0164] In an optional implementation, the use of a neural network model to extract features from the code information to generate at least one word vector includes:
[0165] Determine hyperparameters including window size and minimum word frequency based on the grammatical structure and contextual relationship indicated by the code information;
[0166] Configuring the determined hyperparameters for the neural network model;
[0167] The configured neural network model is used to extract features from the code information to generate at least one word vector.
[0168] In an optional implementation, constructing the word vector feature based on the at least one word vector includes:
[0169] Based on the code information, performing TF-IDF statistics on each word vector in the at least one word vector to obtain a weight corresponding to each word vector;
[0170] Normalizing each word vector in the at least one word vector respectively;
[0171] Based on at least one normalized word vector and its corresponding weight, a weighted calculation is performed to obtain the word vector feature.
[0172] In an optional embodiment, the first static feature risk score includes at least one of the following: a first vulnerability severity score, a first usage frequency score, a first version score, and a first license compliance score;
[0173] The second static feature risk score includes at least one of the following: a second vulnerability severity score, a second usage frequency score, a second version score, and a second license compliance score.
[0174] In an optional embodiment, the device further comprises:
[0175] A fourth risk assessment module, for identifying each component in the code file using a pre-built LDA-based component analysis and risk assessment model, and performing risk assessment on each component in the code file to obtain an LDA component risk score;
[0176] The determining of the overall risk score based on the word vector risk score and the static feature risk score includes:
[0177] The overall risk score is obtained by performing a weighted sum based on the LDA component risk score, the word vector risk score and the static feature risk score.
[0178] In an optional implementation, the pre-built LDA-based component analysis and risk assessment model is used to identify the components in the code file, including:
[0179] Decomposing the code file into a plurality of sub-documents, wherein the plurality of sub-documents include class fragments, method fragments and comment fragments;
[0180] The LDA-based component analysis and risk assessment model is used to identify the multiple sub-documents to obtain the components in the code file.
[0181] In a third aspect, an embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the steps of any of the software risk assessment methods described above are implemented.
[0182] In a fourth aspect, an embodiment of the present application provides a computer program product, comprising computer instructions, which, when executed by a processor, implement the steps of any of the software risk assessment methods described above.
[0183] In a fifth aspect, an embodiment of the present application provides a computer device, comprising a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein the processor implements the steps of any of the above-mentioned software risk assessment methods when executing the computer program.
[0184] See also Figure 8 The computer device of this embodiment includes: a processor 801, a memory 802, and a computer program stored in the memory 802 and executable on the processor 801, such as a software risk assessment program. When the processor 801 executes the computer program, the steps in the above-mentioned software risk assessment method embodiments are implemented, such as Figure 1 Steps S101-S105 are shown.
[0185] Exemplarily, the computer program may be divided into one or more modules / units, which are stored in the memory 802 and executed by the processor 801 to complete the present application. The one or more modules / units may be a series of computer program instruction segments capable of completing specific functions, which are used to describe the execution process of the computer program in the computer device.
[0186] The computer device may be a computing device such as a desktop computer, a notebook, a PDA, a cloud server, etc. The computer device may include, but not limited to, a processor 801 and a memory 802. Those skilled in the art may understand that the schematic diagram is only an example of a computer device and does not constitute a limitation on the computer device. The computer device may include more or fewer components than shown in the figure, or may combine certain components, or different components. For example, the computer device may also include input and output devices, network access devices, buses, etc.
[0187] The processor 801 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or the processor 801 may also be any conventional processor, etc. The processor 801 is the control center of the computer device, and uses various interfaces and lines to connect various parts of the entire computer device.
[0188] The memory 802 can be used to store the computer program and / or module. The processor 801 implements various functions of the computer device by running or executing the computer program and / or module stored in the memory 802 and calling the data stored in the memory 802. The memory 802 can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc.; the data storage area can store data created according to the use of the mobile phone (such as audio data, a phone book, etc.), etc. In addition, the memory 802 can include a high-speed random access memory, and can also include a non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (SecureDigital, SD) card, a flash card (Flash Card), at least one disk storage device, a flash memory device, or other volatile solid-state storage devices.
[0189] Wherein, if the module / unit integrated in the computer device is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application implements all or part of the processes in the above-mentioned embodiment method, and can also be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium, and the computer program can implement the steps of the above-mentioned various method embodiments when executed by the processor 801. Wherein, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, U disk, mobile hard disk, disk, optical disk, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal and software distribution medium, etc.
[0190] In summary, the embodiments of the present application have at least the following beneficial effects:
[0191] According to an embodiment of the present application, code information is obtained by parsing the code file based on the software to be evaluated, wherein the code information is used to indicate the components and library functions corresponding to the code file; feature extraction is performed based on the code information to obtain word vector features, and a word vector risk score is determined based on the word vector features; based on the code file, a risk assessment is performed on the component to obtain a first static feature risk score, and a risk assessment is performed on the library function to obtain a second static feature risk score; based on the first static feature risk score and the second static feature risk score, a static feature risk score is determined; based on the word vector risk score and the static feature risk score, an overall risk score is determined, thereby being able to efficiently and accurately identify risks in the software to improve software quality and software security.
[0192] Through the description of the above implementation methods, those skilled in the art can clearly understand that the present application can be implemented by means of software plus the necessary hardware platform, and of course it can also be implemented entirely by hardware. Based on such an understanding, all or part of the contribution of the technical solution of the present application to the background technology can be embodied in the form of a software product, and the computer software product can be stored in a storage medium, such as ROM (Read-Only Memory) / RAM (Random Access Memory), a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment of the present application or some parts of the embodiments.
[0193] The above is a preferred embodiment of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications are also considered to be within the scope of protection of the present application.
Claims
1. A software risk assessment method, characterized in that: include: Parsing the code file of the software to be evaluated to obtain code information, wherein the code information is used to indicate the components and library functions corresponding to the code file; Perform feature extraction based on the code information to obtain word vector features, and determine a word vector risk score based on the word vector features; Based on the code file, performing risk assessment on the component to obtain a first static feature risk score, and performing risk assessment on the library function to obtain a second static feature risk score; Determining a static feature risk score based on the first static feature risk score and the second static feature risk score; Determining an overall risk score based on the word vector risk score and the static feature risk score; The extracting features based on the code information to obtain word vector features, and determining the word vector risk score according to the word vector features, includes: Using a neural network model, extracting features from the code information to generate at least one word vector; Based on the at least one word vector, construct the word vector feature; Calculate the absolute value of each dimension in the word vector feature, and take the sum of the absolute values of each dimension as the word vector risk score; The constructing the word vector feature based on the at least one word vector includes: Based on the code information, performing TF-IDF statistics on each word vector in the at least one word vector to obtain a weight corresponding to each word vector; Normalizing each word vector in the at least one word vector respectively; Based on at least one normalized word vector and its corresponding weight, a weighted calculation is performed to obtain the word vector feature.
2. The software risk assessment method according to claim 1, characterized in that: The code information includes an abstract syntax tree, and the abstract syntax tree is used to indicate the component and the library function. The code information is obtained by parsing the code file based on the software to be evaluated, including: Preprocessing the code file to obtain target code; Determine a lexical analyzer and a syntax analyzer based on grammatical rules matching the code file; Calling the lexical analyzer to decompose the target code into at least one lexical unit; The syntax analyzer is called to construct the abstract syntax tree based on the at least one token unit.
3. The software risk assessment method according to claim 1, characterized in that: The method of using a neural network model to extract features from the code information and generate at least one word vector includes: Determine hyperparameters including window size and minimum word frequency based on the grammatical structure and contextual relationship indicated by the code information; Configuring the determined hyperparameters for the neural network model; The configured neural network model is used to extract features from the code information to generate at least one word vector.
4. The software risk assessment method according to claim 1, characterized in that: The first static feature risk score includes at least one of the following: a first vulnerability severity score, a first usage frequency score, a first version score, and a first license compliance score; The second static feature risk score includes at least one of the following: a second vulnerability severity score, a second usage frequency score, a second version score, and a second license compliance score.
5. The software risk assessment method according to any one of claims 1 to 4, characterized in that: The method further comprises: Using a pre-built LDA-based component analysis and risk assessment model, identifying each component in the code file, and performing risk assessment on each component in the code file to obtain an LDA component risk score; The determining of the overall risk score based on the word vector risk score and the static feature risk score includes: The overall risk score is obtained by performing a weighted sum based on the LDA component risk score, the word vector risk score and the static feature risk score.
6. The software risk assessment method according to claim 5, characterized in that: The components in the code file are identified by using the pre-built LDA-based component analysis and risk assessment model, including: Decomposing the code file into a plurality of sub-documents, wherein the plurality of sub-documents include class fragments, method fragments and comment fragments; The LDA-based component analysis and risk assessment model is used to identify the multiple sub-documents to obtain the components in the code file.
7. A software risk assessment device, characterized in that: include: A parsing module, used to parse the code file of the software to be evaluated to obtain code information, wherein the code information is used to indicate the components and library functions corresponding to the code file; A first risk assessment module is used to perform feature extraction based on the code information to obtain word vector features, and determine a word vector risk score based on the word vector features; A second risk assessment module is used to perform risk assessment on the component based on the code file to obtain a first static feature risk score, and to perform risk assessment on the library function to obtain a second static feature risk score; a third risk assessment module, configured to determine a static feature risk score based on the first static feature risk score and the second static feature risk score; An overall risk scoring module, configured to determine an overall risk score based on the word vector risk score and the static feature risk score; The extracting features based on the code information to obtain word vector features, and determining the word vector risk score according to the word vector features, includes: Using a neural network model, extracting features from the code information to generate at least one word vector; Based on the at least one word vector, construct the word vector feature; Calculate the absolute value of each dimension in the word vector feature, and take the sum of the absolute values of each dimension as the word vector risk score; The constructing the word vector feature based on the at least one word vector includes: Based on the code information, performing TF-IDF statistics on each word vector in the at least one word vector to obtain a weight corresponding to each word vector; Normalizing each word vector in the at least one word vector respectively; Based on at least one normalized word vector and its corresponding weight, a weighted calculation is performed to obtain the word vector feature.
8. The software risk assessment device according to claim 7, characterized in that: The code information includes an abstract syntax tree, and the abstract syntax tree is used to indicate the component and the library function. The code information is obtained by parsing the code file based on the software to be evaluated, including: Preprocessing the code file to obtain target code; Determine a lexical analyzer and a syntax analyzer based on grammatical rules matching the code file; Calling the lexical analyzer to decompose the target code into at least one lexical unit; The syntax analyzer is called to construct the abstract syntax tree based on the at least one token unit.
9. The software risk assessment device according to claim 7, characterized in that: The method of using a neural network model to extract features from the code information and generate at least one word vector includes: Determine hyperparameters including window size and minimum word frequency based on the grammatical structure and contextual relationship indicated by the code information; Configuring the determined hyperparameters for the neural network model; The configured neural network model is used to extract features from the code information to generate at least one word vector.
10. The software risk assessment device according to claim 7, characterized in that: The first static feature risk score includes at least one of the following: a first vulnerability severity score, a first usage frequency score, a first version score, and a first license compliance score; The second static feature risk score includes at least one of the following: a second vulnerability severity score, a second usage frequency score, a second version score, and a second license compliance score.
11. The software risk assessment device according to any one of claims 7 to 10, characterized in that: The device also includes: A fourth risk assessment module, for identifying each component in the code file using a pre-built LDA-based component analysis and risk assessment model, and performing risk assessment on each component in the code file to obtain an LDA component risk score; The determining of the overall risk score based on the word vector risk score and the static feature risk score includes: The overall risk score is obtained by performing a weighted sum based on the LDA component risk score, the word vector risk score and the static feature risk score.
12. The software risk assessment device according to claim 11, characterized in that: The components in the code file are identified by using the pre-built LDA-based component analysis and risk assessment model, including: Decomposing the code file into a plurality of sub-documents, wherein the plurality of sub-documents include class fragments, method fragments and comment fragments; The LDA-based component analysis and risk assessment model is used to identify the multiple sub-documents to obtain the components in the code file.
13. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the software risk assessment method according to any one of claims 1 to 6 is implemented.
14. A computer program product comprising computer instructions, characterized in that When the computer instructions are executed by a processor, the software risk assessment method according to any one of claims 1 to 6 is implemented.
15. A computer device, characterized in that: The invention comprises a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein the processor implements the software risk assessment method according to any one of claims 1 to 6 when executing the computer program.
Citation Information
Patent Citations
Automatic vulnerability classification method and system based on weighted Word2vec
CN114462049A
Source code tracing analysis and evaluation system
CN118133278A
Cited By
Product information risk assessment method and system based on large model
CN121615142A
A product information risk assessment method and system based on a large model
CN121615142B