Multi-modal learning and deep learning technology-based operating system security evaluation system and method

Through the security assessment system of multimodal learning and deep learning technology, multi-source data is integrated and static dynamic analysis is combined to solve the accuracy and adaptability of operating system security assessment in the existing technology, and efficient and accurate risk assessment is achieved.

CN120337225APending Publication Date: 2025-07-18ZHEJIANG GONGSHANG UNIVERSITY
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510333259.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing operating system security assessment methods are insufficient in the face of complex network attacks, difficult to integrate multi-source data, and traditional methods are difficult to adapt to system updates and new threats, resulting in a decrease in the timeliness and accuracy of the assessment results.

Method used

A security assessment system based on multimodal learning and deep learning technology is adopted, through multi-source data integration and intelligent preprocessing, combined with static analysis and dynamic monitoring, and using multi-dimensional path analysis and machine learning models to generate accurate risk assessment reports.

Benefits of technology

Real-time reflection and adaptability assessment of operating system security status are achieved, comprehensive risk insights and priority guidance are provided, and the accuracy and reliability of assessments are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120337225A_ABST
    Figure CN120337225A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of multi-modal learning, and discloses an operating system security evaluation system and method based on multi-modal learning and deep learning technology, the system uses the multi-modal learning technology, integrates vulnerability description, code structure and runtime data, extracts cross-modal features through a deep learning model, and constructs a unified evaluation framework. The system firstly collects vulnerability data from databases such as CVE, NVD and the like, and forms multi-dimensional input in combination with a calling path generated by static analysis and a running log of dynamic test. Secondly, predicting the availability and risk degree of the vulnerabilities through a multi-modal fusion network, and generating a comprehensive risk score; and finally, outputting a visual report by the system, and clearly displaying risk distribution and priority. Potential and actual security threats of the operating system are comprehensively evaluated, and the accuracy and reliability of evaluation are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of multimodal learning, and particularly relates to an operating system security assessment system and method based on multimodal learning and deep learning technologies. Background Art

[0002] With the rapid development of information technology, the operating system (such as Linux, Windows), as the core component of a computer system, its security is directly related to the stability of user data and system operation. Modern operating systems (such as Linux, Windows) are widely used in servers, personal devices, and Internet of Things terminals, facing increasingly complex network attack threats, such as vulnerability exploitation, malicious code injection, and privilege escalation attacks. However, the existing operating system security assessment methods have significant deficiencies in dealing with these threats, and there is an urgent need for a more intelligent and efficient solution.

[0003] Current security assessment technologies mainly rely on static analysis and dynamic testing. Static analysis identifies potential vulnerabilities, such as buffer overflows or insecure function calls, by scanning source code or binary files. However, this method is difficult to accurately capture runtime behavior, resulting in a high false positive rate. For example, static analysis may mark a function that is not actually called as a risk point, wasting subsequent verification resources. Dynamic testing verifies the exploitability of vulnerabilities by simulating attacks, but its coverage is limited, and the testing time cost for complex systems is high. In addition, these two methods usually run independently, lacking data fusion, resulting in fragmented assessment results and being unable to comprehensively reflect system risks.

[0004] Another key issue is the insufficient utilization of multi-source data by existing methods. Vulnerability information (such as the CVE database), code structure, and runtime logs provide security insights from different dimensions, but traditional technologies are difficult to integrate these heterogeneous data. For example, the CVE description may indicate the attack vector of a vulnerability, while code analysis reveals the call path, and runtime logs record abnormal behavior, but there is no unified model to associate this information. In addition, with the iteration of operating system versions and the diversification of configurations, the traditional method of manually adjusting assessment parameters can no longer adapt to the dynamically changing environment, resulting in a decline in the timeliness and accuracy of assessment results. The improvement of security requirements also exposes the limitations of existing methods. As attackers use artificial intelligence technologies (such as generative adversarial networks) to generate new attacks, traditional rule-driven assessment systems are difficult to cope with. Summary of the Invention

[0005] The purpose of the present invention is to provide an operating system security assessment system and method based on multimodal learning and deep learning technologies to solve the above technical problems.

[0006] To solve the above technical problems, the specific technical solutions of the operating system security assessment system and method based on multi-modal learning and deep learning technologies of the present invention are as follows:

[0007] An operating system security assessment system based on multi-modal learning and deep learning technologies includes 4 stages:

[0008] Vulnerability data collection and preprocessing stage: In this stage, through multi-source data integration and intelligent preprocessing, a vulnerability data set is provided for the system. Data including CVE-ID, description, and version is collected from the database using standardized APIs. A pre-trained NLP model is used to screen vulnerabilities related to the target operating system, including extracting CVSS metrics AV, AC, PR, and UI; cleaning missing data and normalizing the format;

[0009] The final data set is stored in a distributed database in JSON format;

[0010] Vulnerability path analysis and call chain identification stage: In this stage, the preprocessed data is mapped to the target operating system code path, features are extracted, and call chain similarity is identified to provide core data support for risk assessment; The system loads CVE-ID and component names from the database, generates a mapping table between vulnerabilities and code, and uses regular matching to locate the positions of vulnerable functions in the target operating system source code; Through static analysis tools, the call chain from the system entry to the vulnerable function is traced, a structured path graph is generated, and path features are marked; The system compares the call chains of the target operating system and the reference operating system, calculates the similarity, and analyzes the security differences to deepen the understanding of vulnerability behavior; The similarity is quantified by the edit distance algorithm, and the security differences focus on the change in the number of checkpoints;

[0011] Risk difference quantification and assessment stage: In this stage, through multi-dimensional path comparison and machine learning technologies, the vulnerability risk differences between the target system and the reference system are quantified. The process includes loading feature data, analyzing path differences, applying the XGBoost model to predict changes in CVSS metrics, and using Bayesian optimization to adjust weights to calculate the comprehensive risk score. The risk formula considers prediction uncertainty. The assessment results are stored in a Parquet file to support big data analysis and fast query;

[0012] Risk assessment detection report generation stage: In this stage, the risk assessment results are converted into a professional report to provide users with comprehensive risk insights and priority guidance, including integrating risk data, calculating the weighted average risk, using K-means clustering to classify risks, and generating a structured report containing CVE details and confidence intervals. The report is output in JSON and PDF formats and uploaded to the cloud to support downloading and interactive views.

[0013] A security assessment method for an operating system security assessment system based on multi-modal learning and deep learning technologies, comprising the following steps:

[0014] Step 1: Vulnerability data collection and preprocessing;

[0015] Step 2: Vulnerability path analysis and call chain identification;

[0016] Step 3: Risk difference quantification and assessment;

[0017] Step 4: Generation of a risk assessment detection report.

[0018] Furthermore, the said Step 1 includes the following specific steps:

[0019] Step 1.1: Collect vulnerability data from CVE and NVD, including CVE-ID, description, and version information, to generate an initial dataset, supporting large-scale parallel pulling;

[0020] Step 1.2: Use a pre-trained NLP model to analyze the description, filter relevant vulnerabilities according to the target system version and context keywords, and generate a refined list; The NLP model uses the filter_vulnerabilities function to evaluate the relevance. This function takes a list of vulnerability entries and the NLP model as inputs, classifies the description text of each entry by calling the classify method of the model, outputs a relevance score between 0 and 1, and uses list comprehension to only retain the entries with a score greater than or equal to 0.9, indicating high relevance. The function returns the refined vulnerability list to ensure that subsequent analysis focuses on data related to the target system;

[0021] Step 1.3: Parse the vulnerability description and extract CVSS metrics: attack vector, attack complexity, privilege requirements, and user interaction to form a structured set;

[0022] Step 1.4: Data cleaning and normalization, removing entries with missing metrics, merging duplicate CVEs, and normalizing the format to generate a highly consistent and non-redundant dataset;

[0023] Step 1.5: Distributed data storage, storing the preprocessed data: CVE-ID, component name, and metrics in a distributed database, generating a JSON serialized file to support efficient retrieval.

[0024] Furthermore, the said Step 2 includes the following specific steps:

[0025] Step 2.1: Read the preprocessed data, load CVE-ID, component name, and metrics from the database, and generate a mapping table of the target operating system code;

[0026] Step 2.2: Locate vulnerable functions, parse the descriptions, extract key functions using regular expressions, and locate them in the source code of the target operating system;

[0027] Step 2.3: Generate the call chain of the target operating system. Use static analysis to trace the call chain from the entry of the target operating system to the vulnerable functions and generate a path graph;

[0028] Step 2.4: Label path features, analyze path logic, and label AV, AC, PR, and UI;

[0029] Step 2.5: Compare the call chains of the target operating system and the reference operating system, calculate the similarity and analyze the security differences; Call the call chain similarity comparison function, which receives the call chain C_ref of the reference operating system and the call chain C_target of the target operating system, calculates the similarity similarity through the edit distance, counts the number of security checkpoints using count_security_checks, and the difference diff reflects the security changes of the target system. The function returns the similarity and the difference value, which are used to evaluate the call chain behavior differences. Further, the said Step 3 includes the following specific steps:

[0030] Step 3.1: Load and align the feature dataset. Extract the path features of the target and reference systems from the database, and align and pair them through path hashing to ensure the comparison accuracy;

[0031] Step 3.2: Perform multi-dimensional path difference analysis. Compare the function implementations, parameters, and configurations in the paths, extract the differences that affect exploitability, generate a list, and analyze the differences;

[0032] Step 3.3: Predict the change in CVSS metrics. Apply the XGBoost model to predict the change ranges of AV, AC, PR, and UI based on historical data and output the confidence intervals; The function for predicting metric changes receives the path features and the pre-trained XGBoost model as inputs, calls the predict method to generate the predicted change values of the CVSS metrics AV, AC, PR, and UI; Then, calculate the 95% confidence interval of the prediction through the confidence_interval method and return the predicted values and the confidence intervals;

[0033] Step 3.4: Use Bayesian optimization to adjust the weight w i , and calculate the risk by combining the predicted changes:

[0034]

[0035] where is the metric change and uncertainty i is the prediction uncertainty;

[0036] At this stage, highly specialized risk assessment results are generated through multi-dimensional analysis, XGBoost prediction, and Bayesian optimization.

[0037] Furthermore, step 4 includes the following specific steps:

[0038] Step 4.1: Extract CVE-ID, risk score, and confidence interval from the Parquet file and integrate them into a multi-dimensional risk analysis view;

[0039] Step 4.2: Aggregate and visualize the risk scores and calculate the weighted risk:

[0040]

[0041] Step 4.3: Apply the K-means clustering algorithm to classify the risks as high, medium, and low according to R CVE The risk is classified into high, medium, and low levels to adapt to the system characteristics. The K-means clustering algorithm receives the risk score list risk_scores and the number of clusters k. First, initialize_centroids randomly selects k initial cluster centers. Then enter the loop: assign_to_clusters assigns each score to the nearest center according to the Euclidean distance; update_centroids calculates the new center means according to the assignment results. The loop continues until the centers converge, and the final clusters, that is, the risk levels of each score, are returned.

[0042] Step 4.4: Generate a structured assessment report containing CVE details, risk scores, priorities, and confidence intervals;

[0043] Step 4.5: Serialize the report into JSON and PDF and provide download links and interactive views.

[0044] The operating system security assessment system and method based on multi-modal learning and deep learning technologies of the present invention have the following advantages:

[0045] 1. The system can reflect the changes in the security state of the operating system in real time based on multi-source data collection and analysis, adapt to system updates, configuration changes, or new threats, and ensure that the assessment results reflect the latest risk situation.

[0046] 2. Quantitatively analyze the risks through multiple dimensions such as vulnerability features, code paths, and running behaviors to generate accurate risk scores. This multi-dimensional quantification method provides a more comprehensive risk insight.

[0047] 3. The system comprehensively evaluates the potential and actual security threats of the operating system through the combination of static analysis and dynamic monitoring. This hybrid method ensures that both hidden defects can be discovered and their exploitability in the real environment can be verified, improving the accuracy and reliability of the assessment. Brief Description of the Drawings

[0048] Figure 1 It is a schematic diagram of the system model structure of the present invention;

[0049] Figure 2 It is a flowchart of the system operation of the present invention;

[0050] Figure 3 It is a schematic diagram of the function for evaluating relevance using the NLP model of the present invention;

[0051] Figure 4 It is a schematic diagram of the pseudo-code for comparing the similarity of call chains of the present invention;

[0052] Figure 5 It is a schematic diagram of the pseudo-code for predicting measure changes of the present invention;

[0053] Figure 6 It is a schematic diagram of the pseudo-code for K-means classification of the present invention. Detailed Description of the Invention

[0054] In order to better understand the purpose, structure and function of the present invention, the operating system security evaluation system and method based on multi-modal learning and deep learning technology of the present invention will be further described in detail below with reference to the accompanying drawings.

[0055] As Figure 1 shown, the operating system security evaluation system and method based on multi-modal learning and deep learning technology of the present invention includes 4 stages:

[0056] Vulnerability data collection and preprocessing stage: In this stage, high-quality vulnerability data sets are provided for the system through multi-source data integration and intelligent preprocessing. Vulnerability data including CVE-ID, description and version are collected from databases such as CVE and NVD using standardized APIs, and a pre-trained NLP model is used to screen vulnerabilities related to the target operating system. The key steps include extracting CVSS measures (AV, AC, PR, UI), cleaning missing data and normalizing the format to ensure data consistency. The final data set is stored in a distributed database in JSON format, supporting efficient retrieval and expansion. The professional design of this stage significantly improves the accuracy and pertinence of the input data, laying a solid foundation for subsequent path analysis.

[0057] Vulnerability Path Analysis and Call Chain Identification Phase: In this phase, the preprocessed data is mapped to the target operating system code paths, features are extracted, and call chain similarities are identified to provide core data support for risk assessment. The system loads CVE-ID and component names from the database to generate a mapping table between vulnerabilities and code, and uses regular expression matching to locate the positions of vulnerable functions in the target operating system source code. The call chain from the system entry to the vulnerable function is traced through static analysis tools to generate a structured path graph, and path features (such as AV, AC, PR, UI) are marked. In addition, the system compares the call chains of the target operating system and the reference operating system (such as Windows or Unix), calculates the similarity, and analyzes the security differences to deepen the understanding of vulnerability behavior. The similarity is quantified by the edit distance algorithm, and the security differences focus on the changes in the number of checkpoints (such as permission verification). Through path generation, feature extraction, and call chain analysis in this phase, a comprehensive analysis foundation is constructed to ensure the accuracy and coherence of the data, providing high-quality input for subsequent risk quantification.

[0058] Risk Difference Quantification and Assessment Phase: In this phase, through multi-dimensional path comparison and machine learning techniques, the vulnerability risk differences between the target system and the reference system are quantified. The process includes loading feature data, analyzing path differences, applying the XGBoost model to predict the changes in CVSS metrics, and using Bayesian optimization to adjust weights to calculate the comprehensive risk score. The risk formula takes into account the prediction uncertainty to ensure the scientific nature of the results. The assessment results are stored in Parquet files to support big data analysis and fast query. The specialized quantification method in this phase, combined with statistical support, provides users with accurate risk insights and is suitable for in-depth analysis of complex security scenarios.

[0059] Risk Assessment Detection Report Generation Phase: In this phase, the risk assessment results are transformed into a specialized report to provide users with comprehensive risk insights and priority guidance. The core steps include integrating risk data, calculating the weighted average risk, using K-means clustering to classify risks (high, medium, low), and generating a structured report containing CVE details and confidence intervals. The report is output in JSON and PDF formats and uploaded to the cloud to support downloading and interactive views. The specialized process in this phase improves the readability and usability of the data, provides users with an efficient risk management tool, and reflects the professionalism and user-friendliness of the system output.

[0060] The General Operating System Structure Security Risk Automated Analysis System achieves comprehensive risk assessment through four stages. From multi-source data collection and preprocessing to ensure input quality, to path analysis and feature extraction to reveal vulnerability exploitability, then to risk difference quantification to evaluate system differences, and finally to generate professional reports to provide insights. The system is designed efficiently and precisely, leveraging NLP, machine learning, and big data technologies, and is applicable to various operating systems. It significantly enhances the depth and breadth of security analysis, provides scientific risk management support for users, and has significant technical advantages and application value.

[0061] As Figure 2 shown, the operating system security assessment method based on multi-modal learning and deep learning technologies of the present invention includes the following steps:

[0062] Step 1: Vulnerability data collection and preprocessing: This stage is the foundation of system analysis. By collecting and preprocessing vulnerability data from multiple sources, it ensures high-quality and consistent input, providing reliable support for subsequent analysis. The specific process is as follows:

[0063] Step 1.1: Collect vulnerability data from CVE, NVD, etc., including CVE-ID, description, and version information, to generate an initial dataset, supporting large-scale parallel pulling.

[0064] Step 1.2: Use a pre-trained NLP model to analyze the description, filter relevant vulnerabilities according to the target system version and context keywords, and generate a refined list. As Figure 3 shown, the NLP model uses the filter_vulnerabilities function to evaluate relevance. This function takes a list of vulnerability entries and the NLP model as inputs, classifies the description text of each entry by calling the classify method of the model, and outputs a relevance score between 0 and 1. Using list comprehensions, only the entries with a score greater than or equal to 0.9 are retained, indicating high relevance. The function returns the refined vulnerability list, ensuring that subsequent analysis focuses on data related to the target system.

[0065] Step 1.3: Parse the vulnerability description and extract CVSS metrics: Attack Vector (AV), Attack Complexity (AC), Privileges Required (PR), User Interaction (UI), to form a structured set.

[0066] Step 1.4: Data cleaning and normalization, removing entries with missing metrics, merging duplicate CVEs, and normalizing the format, to generate a highly consistent and non-redundant dataset.

[0067] Step 1.5: Distributed data storage. The preprocessed data (CVE-ID, component name, metrics) is stored in a distributed database, and a JSON serialization file is generated to support efficient retrieval.

[0068] Through the data integration, filtering, and cleaning processes, a high-quality preprocessed dataset is generated. The optimized steps enhance the professionalism and reliability of the data, ensuring the consistency and accuracy of the input for subsequent analysis and providing strong support for the overall system performance.

[0069] Step 2: Vulnerability path analysis and call chain identification. This stage focuses on mapping vulnerabilities to the operating system code paths and extracting features, providing the core basis for risk assessment. It reveals the exploitable features of vulnerabilities through static analysis and feature annotation, preparing data for subsequent differential quantification. The specific steps are as follows:

[0070] Step 2.1: Read the preprocessed data, load the CVE-ID, component name, and metrics from the database, and generate a mapping table for the target operating system code.

[0071] Step 2.2: Locate vulnerable functions, parse the descriptions, and use regular expressions to extract key functions and locate them in the target operating system source code.

[0072] Step 2.3: Generate the call chain of the target operating system. Use static analysis to trace the call chain from the entry of the target operating system to the vulnerable functions and generate a path diagram.

[0073] Step 2.4: Annotate path features, analyze path logic, and annotate AV (call type), AC (complexity), PR (permission), and UI (interaction).

[0074] Step 2.5: Compare the call chains of the target operating system and the reference operating system, calculate the similarity, and analyze the security differences. As Figure 4 shown, the following is the pseudocode for call chain similarity comparison. This function receives the call chain of the reference operating system C_ref and the call chain of the target operating system C_target, and calculates the similarity similarity (range [0,1]) through the edit distance. count_security_checks counts the number of security checkpoints, and the difference diff reflects the security changes in the target system. The function returns the similarity and the difference value, which are used to evaluate the call chain behavior differences.

[0075] This stage constructs an accurate analysis basis through path generation and call chain analysis.

[0076] Step 3: Risk Difference Quantification and Assessment: In this stage, based on path features and call chain analysis, the risk differences between the target system and the reference system are quantified to provide a scientific assessment. It reveals the impact of vulnerabilities through difference calculation and severity analysis, laying the foundation for report generation. The specific steps are as follows:

[0077] Step 3.1: Load and align the feature dataset, extract the path features of the target and reference systems from the database, and align and pair them through path hashing to ensure the comparison accuracy.

[0078] Step 3.2: Multi-dimensional path difference analysis, compare the function implementations, parameters, and configurations in the paths, extract the differences affecting exploitability, generate a list, and analyze the differences.

[0079] Step 3.3: Predict the change in CVSS metrics, apply the XGBoost model, predict the change amplitudes of AV, AC, PR, and UI based on historical data, and output the confidence interval. As Figure 5 shown, it is the pseudocode for predicting the metric change. This function receives the path features and the pre-trained XGBoost model as inputs, calls the predict method to generate the predicted values of the CVSS metrics (AV, AC, PR, UI) changes. Then, calculate the 95% confidence interval of the prediction (α = 0.05) through the confidence_interval method and return the predicted values and the confidence interval. This process utilizes the high prediction ability of gradient boosting trees to provide statistical support for risk quantification, ensuring the reliability and transparency of the evaluation results.

[0080] Step 3.4: Use Bayesian optimization to adjust the weight w i , and calculate the risk by combining the predicted changes:

[0081]

[0082] where is the metric change and uncertainty i is the prediction uncertainty.

[0083] In this stage, through multi-dimensional analysis, XGBoost prediction, and Bayesian optimization, highly specialized risk assessment results are generated. The optimized process improves the scientific nature and accuracy of the assessment, provides reliable insights for users, and ensures the applicability and technical depth of the system in a complex environment.

[0084] Step 4: Generation of Risk Assessment Detection Report: In this stage, the risk assessment results are transformed into a specialized report, providing users with comprehensive risk insights and priority guidance. It helps users efficiently manage vulnerabilities through the integration of risk data, score aggregation, and visual output. The specific steps are as follows:

[0085] Step 4.1: Extract CVE-ID, risk score, and confidence interval from the Parquet file and integrate them into a multi-dimensional risk analysis view.

[0086] Step 4.2: Aggregate and visualize the risk scores, and calculate the weighted risk:

[0087]

[0088] Step 4.3: Apply the K-means clustering algorithm to classify the risks as high, medium, and low according to R CVE to adapt to the system characteristics. As Figure 6 shown in the following is the pseudocode for classification by the K-means clustering algorithm. This function implements the core logic of K-means clustering. The function receives a list of risk scores risk_scores and the number of clusters k (default 3, representing high, medium, and low). First, initialize_centroids randomly selects k initial cluster centers. Then enter the loop: assign_to_clusters assigns each score to the nearest center according to the Euclidean distance; update_centroids calculates the new center means based on the assignment results. The loop continues until the centers converge (no longer change), and finally returns the clusters, that is, the risk levels of each score. This algorithm automatically discovers the internal structure of the score distribution through unsupervised learning, adapts to the risk characteristics of different systems, and improves the professionalism and accuracy of classification.

[0089] Step 4.4: Generate a structured assessment report, including CVE details, risk scores, priorities, and confidence intervals.

[0090] Step 4.5: Serialize the report into JSON and PDF, and provide download links and interactive views.

[0091] In this stage, through risk integration and classification, a highly professional assessment report is generated. The optimized process focuses on risk insight and visualization, improves the scientificity and readability of the report, and provides efficient risk management support for users.

[0092] It can be understood that the present invention is described through some embodiments. Those skilled in the art know that without departing from the spirit and scope of the present invention, various changes or equivalent replacements can be made to these features and embodiments. Additionally, under the teaching of the present invention, these features and embodiments can be modified to adapt to specific situations and materials without departing from the spirit and scope of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed herein, and all embodiments falling within the scope of the claims of this application belong to the scope protected by the present invention.

Claims

1. An operating system security assessment system based on multimodal learning and deep learning technologies, characterized in that, It includes 4 stages: Vulnerability data collection and preprocessing stage: In this stage, through multi-source data integration and intelligent preprocessing, a vulnerability data set is provided for the system. Vulnerability data including CVE-ID, description, and version is collected from the database using a standardized API. A pre-trained NLP model is used to filter vulnerabilities related to the target operating system, including extracting CVSS metrics AV, AC, PR, and UI; cleaning missing data and normalizing the format; The final data set is stored in a distributed database in JSON format; Vulnerability path analysis and call chain identification stage: In this stage, the preprocessed data is mapped to the code paths of the target operating system, features are extracted, and call chain similarity is identified, providing core data support for risk assessment; The system loads CVE-ID and component names from the database, generates a mapping table between vulnerabilities and code, and uses regular expressions to locate the positions of vulnerable functions in the source code of the target operating system; the call chain from the system entry to the vulnerable function is traced through a static analysis tool, a structured path graph is generated, and path features are marked; The system compares the call chains of the target operating system and the reference operating system, calculates the similarity, and analyzes the security differences to deepen the understanding of vulnerability behavior; the similarity is quantified by the edit distance algorithm, and the security differences focus on the change in the number of checkpoints; Risk difference quantification and assessment stage: In this stage, through multi-dimensional path comparison and machine learning techniques, the vulnerability risk differences between the target system and the reference system are quantified. The process includes loading feature data, analyzing path differences, applying the XGBoost model to predict changes in CVSS metrics, and using Bayesian optimization to adjust weights to calculate the comprehensive risk score. The risk formula considers prediction uncertainty, and the assessment results are stored in a Parquet file to support big data analysis and fast query; Risk assessment detection report generation stage: In this stage, the risk assessment results are converted into a professional report, providing users with comprehensive risk insights and priority guidance, including integrating risk data, calculating the weighted average risk, using K-means clustering to classify risks, and generating a structured report containing CVE details and confidence intervals. The report is output in JSON and PDF formats and uploaded to the cloud to support downloading and interactive views.

2. A security assessment method for an operating system security assessment system based on multi-modal learning and deep learning technologies as described in claim 1, characterized in that, It includes the following steps: Step 1: Vulnerability data collection and preprocessing; Step 2: Vulnerability path analysis and call chain identification; Step 3: Risk difference quantification and assessment; Step 4: Risk assessment detection report generation.

3. The security assessment method according to claim 2, wherein The specific steps included in the above Step 1 are as follows: Step 1.1: Collect vulnerability data from CVE and NVD, including CVE-ID, description, and version information, to generate an initial data set, supporting large-scale parallel pulling; Step 1.2: Use the pre-trained NLP model to analyze the description, filter relevant vulnerabilities by target system version and context keywords, and generate a refined list; the NLP model uses the filter_vulnerabilities function to evaluate relevance. This function receives a list of vulnerability entries and an NLP model as input, and classifies the description text of each entry by calling the model's classify method. It outputs a relevance score between 0 and 1, and uses list derivation to only retain entries with a score greater than or equal to 0.9, indicating high relevance. The function returns a refined vulnerability list to ensure that subsequent analysis focuses on data related to the target system; Step 1.3: Parse the vulnerability description and extract CVSS metrics: attack vector, attack complexity, permission requirements, user interaction, to form a structured set; Step 1.4: Data cleaning and normalization: remove missing measurement items, merge duplicate CVEs and normalize the format to generate a highly consistent and non-redundant data set; Step 1.5: Distributed data storage, store the pre-processed data: CVE-ID, component name, and measurement into a distributed database, generate a JSON serialized file, and support efficient retrieval.

4. The security assessment method according to claim 2, wherein The step 2 comprises the following specific steps: Step 2.1: Read the preprocessed data, load CVE-ID, component name and measurement from the database, and generate a mapping table of the target operating system code; Step 2.2: Locate the vulnerable function, parse the description, extract the key function using regular matching, and locate it in the target operating system source code; Step 2.3: Generate the target operating system call chain, use static analysis to trace the call chain from the target operating system entry to the vulnerable function, and generate a path graph; Step 2.4: Mark the path features, analyze the path logic, and mark AV, AC, PR, and UI; Step 2.5: Compare the call chains of the target operating system and the reference operating system, calculate the similarity and analyze the security differences; the call chain similarity comparison function receives the reference operating system call chain C_ref and the target operating system call chain C_target, calculates the similarity similarity through the edit distance, count_security_checks counts the number of security checkpoints, and the difference diff reflects the security changes of the target system. The function returns the similarity and difference values for evaluating the differences in call chain behavior.

5. The security assessment method according to claim 2, wherein The step 3 comprises the following specific steps: Step 3.1: Load and align the feature dataset, extract the path features of the target and reference systems from the database, and align and pair them through path hashing to ensure comparison accuracy; Step 3.2: Multi-dimensional path difference analysis, compare the function implementation, parameters and configuration in the path, extract the differences that affect the availability, generate a list, and analyze the differences; Step 3.3: Predict the change in CVSS metrics. Apply the XGBoost model to predict the change amplitude of AV, AC, PR, and UI based on historical data and output the confidence interval. The predicted metric change function takes the path features and the pre-trained XGBoost model as inputs, calls the predict method to generate the predicted values of the CVSS metric changes for AV, AC, PR, and UI. Then, calculate the 95% confidence interval of the prediction through the confidence_interval method and return the predicted values and the confidence interval. Step 3.4: Use Bayesian optimization to adjust the weight w i , and calculate the risk by combining the predicted changes: Among them, is the measurement change, uncertainty i is the prediction uncertainty; At this stage, highly specialized risk assessment results are generated through multi-dimensional analysis, XGBoost prediction, and Bayesian optimization.

6. The security assessment method according to claim 2, wherein The said Step 4 includes the following specific steps: Step 4.1: Extract the CVE-ID, risk score, and confidence interval from the Parquet file and integrate them into a multi-dimensional risk analysis view. Step 4.2: Aggregate and visualize the risk scores and calculate the weighted risk. Step 4.3: Apply the K-means clustering algorithm. According to R CVE classify the risks into high, medium, and low levels to adapt to the system characteristics. The K-means clustering algorithm takes the risk score list risk_scores and the number of clusters k as inputs. First, initialize_centroids randomly selects k initial cluster centers. Then enter the loop: assign_to_clusters assigns each score to the nearest center according to the Euclidean distance; update_centroids calculates the new center means based on the assignment results. The loop continues until the centers converge, and finally returns the clusters, which are the risk levels of each score. Step 4.4: Generate a structured assessment report, including CVE details, risk scores, priorities, and confidence intervals. Step 4.5: Serialize the report into JSON and PDF and provide download links and interactive views.

Citation Information

Cited By

  • Server firmware testing method and electronic equipment

    CN120743787A

  • A server firmware testing method and electronic device

    CN120743787B

  • Code review method, system and equipment and storage medium

    CN120832675A