Intelligent network connection automobile information security threat analysis method based on large language model
Through the combination of large language models and knowledge graphs, the security threats of intelligent connected vehicles are automatically identified and evaluated, and the lack of analysis caused by manual entry in the existing technology is solved, and efficient and accurate threat analysis is achieved.
Patent Information
- Application Number
- CN202510356748.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-11
AI Technical Summary
The existing methods of information security threat analysis of intelligent connected vehicles rely on manual entry, lack comprehensive and automated processing capabilities, and cannot effectively deal with massive and complex threat analysis tasks.
The large language model fine-tuning technology is adopted, combined with the knowledge graph, and the security threats in intelligent connected vehicles are automatically identified and evaluated. By building a security threat analysis model, automated threat scenario description and evaluation are achieved.
It improves the accuracy and efficiency of threat analysis of intelligent connected vehicles information security, can accurately extract keywords and generate threat scene description text, and supports efficient processing of complex threat analysis tasks.
Smart Images

Figure CN120297271A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of intelligent connected vehicles. More specifically, it relates to a method for analyzing information security threats of intelligent connected vehicles based on large language models. Background Art
[0002] In recent years, with the rapid development of Intelligent Connected Vehicles (ICVs), the global automotive industry is undergoing profound changes. Intelligent connected vehicles integrate advanced technologies such as autonomous driving, vehicle networking, intelligent sensing, and data analysis, greatly improving driving convenience. However, the widespread application of intelligent connected vehicles has also brought unprecedented security challenges. With the increasing complexity of vehicle network systems and electronic and electrical architectures, information security incidents in intelligent connected vehicles present problems such as an enlarged attack surface and increased attack paths. Automotive information security will become the fourth major security issue in the automotive field after active safety, passive safety, and functional safety. Currently, the analysis of information security threats in intelligent connected vehicles has not been fully applied to the design, research and development, manufacturing, and information security management processes of vehicles. Existing automotive information security threat analysis models, although able to identify security risks to a certain extent, have limited evaluation scope and rely on manual input, lacking comprehensiveness and automated processing capabilities. With the continuous development of intelligent connected vehicles, the amount of data of information security threat analysis objects involved inside vehicles has increased explosively. The data required for the evaluation of a single vehicle reaches tens of thousands of levels, and there are often complex one-to-many or many-to-many relationships between different evaluation objects. This makes the traditional manual evaluation method unable to effectively meet the needs of automotive information security threat analysis. Therefore, there is an urgent need for an automated threat analysis method that can efficiently and accurately handle these massive and complex threat analysis tasks. Summary of the Invention
[0003] The purpose of the present invention is to overcome the deficiencies of the prior art and provide a method for analyzing information security threats of intelligent connected vehicles based on large language models, which can automatically identify and evaluate security threats in intelligent connected vehicles by fine-tuning large language models, improving the accuracy and efficiency of threat analysis.
[0004] To achieve the above-mentioned invention purpose, the method for analyzing information security threats of intelligent connected vehicles based on large language models of the present invention includes the following steps:
[0005] S1: Collect several pieces of TARA example data of intelligent connected vehicles. For each piece of TARA example data, extract the natural language descriptions of the damage scenario and the attack path to form a threat scenario description text, and then extract the keyword set from it. Use the keyword set corresponding to each piece of TARA example data as the input and the corresponding threat scenario description text as the output to form the training samples for fine-tuning the large language model;
[0006] Label the threat level for each piece of TARA example data to obtain an N-dimensional threat level vector, where the nth element represents the evaluation score of this piece of TARA example data on the nth evaluation index; Use the description text of each piece of TARA example data as the input and the threat level vector as the output to form the training samples for the security threat analysis model;
[0007] S2: Select a large language model according to actual needs, and fine-tune the large language model using the large language model training samples obtained in step S1;
[0008] S3: Construct a security threat analysis model according to actual needs, with its input being the threat scenario description text and the output being an N-dimensional threat level vector, and train it using the training samples of the security threat analysis model obtained in step S1; The security threat analysis model includes a text feature extraction module and a classifier, where:
[0009] The text feature extraction module is used to extract text features from the description text and send them to the classifier;
[0010] The classifier is used to predict the N-dimensional threat level vector based on the text features;
[0011] S4: When it is necessary to conduct a security threat analysis on a certain typical business of an intelligent connected vehicle, obtain the original description text of the typical business to be analyzed, and extract the keyword set from the original description text using the same method as in step S1;
[0012] S5: Input the keyword set of the typical business to be analyzed into the fine-tuned large language model to obtain the threat scenario description text of this typical business, and then input the threat scenario description text into the trained security threat analysis model to obtain an N-dimensional threat level vector.
[0013] The information security threat analysis method for intelligent connected vehicles based on large language models in the present invention extracts training samples for large language models and security threat analysis models from the TARA example data of intelligent connected vehicles, fine-tunes the large language models, constructs a security threat analysis model according to actual needs and trains it with corresponding training samples. When it is necessary to conduct a security threat analysis on a certain typical business of an intelligent connected vehicle, a keyword set is extracted from the original description text of the typical business to be analyzed, and then it is input into the fine-tuned large model to obtain the threat scenario description text of this typical business. Then, the description text is input into the trained security threat analysis model to obtain the threat analysis result.
[0014] The present invention has the following beneficial effects:
[0015] 1) Through the large language model fine-tuning technology, the present invention learns the natural language descriptions of damage scenarios and attack paths in TARA example data, and can automatically generate damage scenarios and attack paths related to intelligent connected vehicles.
[0016] 2) The present invention can be combined with knowledge graph technology to more accurately extract keywords in the description text, improve the accuracy of the generated threat scenario description text, and thus improve the performance of security threat analysis. Description of the Drawings
[0017] Figure 1 is the flowchart of the specific implementation manner of the information security threat analysis method for intelligent connected vehicles based on large language models in the present invention;
[0018] Figure 2 is the flowchart of the keyword set extraction method based on the knowledge graph in this embodiment
[0019] Figure 3 is the sub-graph example diagram of the knowledge graph in this embodiment. Specific Implementation Manner
[0020] The following describes the specific implementation manner of the present invention with reference to the drawings, so that those skilled in the art can better understand the present invention. It should be particularly noted that in the following description, when the detailed descriptions of known functions and designs may dilute the main content of the present invention, these descriptions will be omitted here.
[0021] Embodiment
[0022] Figure 1 is the flowchart of the specific implementation manner of the information security threat analysis method for intelligent connected vehicles based on large language models in the present invention. As Figure 1 shown, the specific steps of the information security threat analysis method for intelligent connected vehicles based on large language models in the present invention include:
[0023] S101: Obtain training samples:
[0024] In the present invention, in order to perform information security threat analysis on intelligent connected vehicles, it is necessary to use a large language model and a security threat analysis model. In order to enable the large language model to adapt to the application scenario of the present invention, it is necessary to fine-tune the large language model. Therefore, it is necessary to construct training samples for the large language model, and at the same time, it is also necessary to construct training samples for the security threat analysis model. TARA (Threat Analysis and Risk Assessment) is the core method for evaluating the cybersecurity of intelligent vehicles. The TARA example data includes assets, attack paths, impact levels, threat scenarios, etc. Therefore, in the present invention, training samples are extracted from the TARA example data. The specific method is as follows:
[0025] Collect several pieces of TARA example data of intelligent connected vehicles. For each piece of TARA example data, extract the natural language descriptions of the damage scenario and the attack path to form a threat scenario description text, and then extract a keyword set from it. Use the keyword set corresponding to each piece of TARA example data as the input and the corresponding threat scenario description text as the output to form the training samples for fine-tuning the large language model. Through the above processing, the TARA example data can be converted into a training data set suitable for fine-tuning the large language model.
[0026] Label the threat level for each piece of TARA example data to obtain an N-dimensional threat level vector, where the nth element represents the evaluation score of the TARA example data on the nth evaluation index. Use the description text of each piece of TARA example data as the input and the threat level vector as the output to form the training samples for the security threat analysis model. The evaluation index can be set according to actual needs. Generally speaking, the evaluation indexes for the damage scenario include ratings in 4 dimensions: functional safety, finance, operation, and privacy. Each dimension is divided into the following four levels: None or Minor (Negligibel), Moderate, Major, and Severe. Table 1 is the rating table for the damage scenario evaluation indexes in this embodiment.
[0027]
[0028] Table 1
[0029] The evaluation indexes for the attack path include ratings in 5 dimensions: attack duration, expertise, product publicity, opportunity window, and attack equipment. The specific descriptions are as follows:
[0030] Attack time consumption: Evaluate the time required for the attack path to start and successfully complete. The longer the time, the lower the feasibility of the attack path. It is divided into the following five levels: less than one week, less than one month, less than six months, less than three years, and more than three years.
[0031] Expertise: Evaluate the level of technology and expertise required to execute the attack. The annotation of this dimension considers whether the attack path requires highly specialized knowledge, skills, or experience. It is divided into the following four levels: layman, proficient, expert, and multiple experts.
[0032] Product publicity: Evaluate whether the attack path depends on publicly available product information. If the attacker can easily obtain relevant information about the target product, the feasibility of the attack path will be higher. It is divided into the following four levels: public, restricted, confidential information, and highly confidential information.
[0033] Opportunity window: Evaluate whether the attacker can easily achieve the attack at a specific moment or under certain conditions. The annotation of this dimension usually considers whether the attacker can take advantage of system vulnerabilities or misconfigurations to launch an attack at a specific time. It is divided into the following four levels: unrestricted, easy, medium, and difficult.
[0034] Attack equipment: Evaluate the physical devices or software tools required to execute the attack. The annotation of this dimension analyzes whether the attack path depends on specific attack tools, malware, or physical devices. It is divided into the following four levels: ordinary equipment, special equipment, customized equipment, and multiple customized equipment.
[0035] It can be seen that in this embodiment, the dimension N of the threat level vector is 9.
[0036] Regarding the keyword set, in this embodiment, the keyword set includes asset names, vulnerability keywords, and attack path keywords. In practical applications, the keyword set can be obtained through manual analysis. In this embodiment, in order to more accurately extract the keyword set from the original description text, a keyword set extraction method based on a knowledge graph is proposed. Figure 2 is the flowchart of the keyword set extraction method based on the knowledge graph in this embodiment. As Figure 2 shown, the specific steps of the keyword set extraction based on the knowledge graph in this embodiment include:
[0037] S201: Obtain information security threat intelligence data:
[0038] First, obtain information security threat intelligence data from publicly available information security knowledge bases, including asset data, vulnerability data, weakness data, and attack pattern data, and associate assets, vulnerabilities, weaknesses, and attack patterns in different data.
[0039] In this embodiment, the asset data comes from CPE (Common Platform Enumeration), which is a standardized method for describing information technology systems, software, and devices. It provides a unified format to identify different products and versions. For example, cpe:2.3:o:microsoft:windows_10::-:x64:build:19044 represents the Microsoft Windows 10 operating system with a 64-bit version and an internal version number of 19044. In vulnerability management, CPE can help accurately determine which systems and software are affected by specific vulnerabilities, so as to take corresponding repair and protection measures. At the same time, it also facilitates information exchange and sharing between security tools and systems.
[0040] The vulnerability data comes from CVE (Common Vulnerabilities and Exposures), which is an open vulnerability list maintained by organizations such as the National Vulnerability Database (NVD) of the United States. It assigns a unique number (CVE number) to each identified vulnerability and exposure, such as CVE-2021-44228 (Log4j remote code execution vulnerability).
[0041] The weakness data comes from CWE (Common Weakness Enumeration), which is a classification list of software and system security weaknesses maintained by MITRE Corporation. It classifies and numbers various common security weaknesses, such as CWE-79 (Cross-Site Scripting, XSS), CWE-89 (SQL injection), etc.
[0042] The attack pattern data comes from CAPEC (Common Attack Pattern Enumeration and Classification), which is a framework for classifying and describing various attack patterns. It details the attack methods, techniques, and strategies that attackers may adopt, such as CAPEC-100 (brute force attack), CAPEC-250 (man-in-the-middle attack).
[0043] To make the constructed knowledge graph more accurate, the obtained information security threat intelligence data can be cleaned and preprocessed, including removing duplicate data, filling in missing values, standardizing data formats, etc., to ensure the compatibility and consistency between various data sources. The external structured descriptions of CVE entries are effectively associated with CWE and CPE information through the "cpe_match" and "problemtype" fields, and the hierarchical relationship between CAPEC and CWE is reconstructed. The "configurations" field in NVD-CVE is used to describe the enumeration relationship between vulnerabilities and assets; CWE entries establish a hierarchical relationship with other weaknesses through the "RelatedWeaknesses" field and are associated with attack patterns through the "Related Attack Patterns" field; CAPEC describes the hierarchical structure of its attack patterns through "Related Attack Patterns".
[0044] S202: Construct a knowledge graph:
[0045] Construct a knowledge graph based on the information security threat intelligence data obtained in step S201, where the nodes include assets, vulnerabilities, weaknesses, and attack models, and the relationships include the association between vulnerabilities and assets (i.e., the asset has a vulnerability), the association between vulnerabilities and weaknesses (the vulnerability has a corresponding weakness), and the association between weaknesses and attack patterns.
[0046] Figure 3 is an example subgraph of the knowledge graph in this embodiment. As Figure 3 shown, this subgraph contains 5 assets, 1 vulnerability, 1 weakness, and 1 attack pattern. The basic information of various nodes and their attributes in this subgraph includes:
[0047] CVE:
[0048] cpe:['cpe:2.3:a:gossamer_threads:dbman:2.0.4:*:*:*:*:*:*:*']
[0049] cve_id:CVE-2000-0381
[0050] cwe:['NVD-CWE-Other']
[0051] description:The Gossamer Threads DBMan db.cgi CGI script allowsremote attackers to view environmental variables and setup information byreferencing a non-existing database in the db parameter.
[0052] last_modified_date:2024-11-20T23:32:22.500
[0053] publish_date:2000-05-05T04:00:00.000
[0054] threat_score:6.4
[0055] CPE:
[0056] cpe_id:cpe:2.3:a:1password:1password:8.7.3:*:*:*:*:windows:*:*
[0057] cpe_name:1password 8.7.3for Windows
[0058] CWE:
[0059] common_consequences:::SCOPE:Access Control:IMPACT:Gain Privileges orAssume Identity:NOTE:An attacker could gain unauthorized access to the systemby retrieving legitimate user's authentication credentials.::SCOPE:Availability:IMPACT:DoS: Resource Consumption(Other):NOTE:An attacker coulddeny service to legitimate system users by launching a brute force attack onthe password recovery mechanism using user ids of legitimate users.::SCOPE:Integrity:SCOPE:Other:IMPACT:Other:NOTE:The system's security functionalityis turned against the system by the attacker.::
[0060] cwe_id:CWE-640
[0061] description:The product contains a mechanism for users to recover orchange their passwords without knowing the original password,but themechanism is weak.
[0062] extended_description:The product contains a mechanism for users torecover or change their passwords without knowing the original password,butthe mechanism is weak.
[0063] name:Weak Password Recovery Mechanism for Forgotten Password
[0064] related_attack_patterns:::50::(For example, this field describes the association relationship with CAPEC)
[0065] related_weaknesses:::NATURE:ChildOf:CWE ID:1390:VIEWID:1000:ORDINAL:Primary::NATURE:ChildOf:CWE ID:287:VIEWID:1003:ORDINAL:Primary::
[0066] CAPEC:
[0067] capec_id:CAPEC-61
[0068] consequences:::SCOPE:Confidentiality:SCOPE:Access Control:SCOPE:Authorization:TECHNICAL IMPACT:Gain Privileges::
[0069] description:The attacker induces a client to establish a session with the target software using a session identifier provided by the attacker.Once the user successfully authenticates to the target software, the attacker uses the (now privileged) session identifier in their own transactions. This attack leverages the fact that the target software either relies on client-generated session identifiers or maintains the same session identifiers after privilege elevation.
[0070] execution_flow:::STEP:1:PHASE:Explore:DESCRIPTION:[Setup the Attack]Setup a session:The attacker has to setup a trap session that provides a valid session identifier, or select an arbitrary identifier, depending on the mechanism employed by the application. A trap session is a dummy session established with the application by the attacker and is used solely for the purpose of obtaining valid session identifiers. The attacker may also be required to periodically refresh the trap session in order to obtain valid session identifiers.:TECHNIQUE:The attacker chooses a predefined identifier that they know.:TECHNIQUE:The attacker creates a trap session for the victim.::STEP:2:PHASE:Experiment:DESCRIPTION:[Attract a Victim]Fixate the session:The attacker now needs to transfer the session identifier from the trap session to the victim by introducing the session identifier into the victim's browser. This is known as fixating the session.The session identifiercan be introduced into the victim's browser by leveraging cross sitescripting vulnerability,using META tags or setting HTTP response headers in avariety of ways.:TECHNIQUE:Attackers can put links on web sites(such asforums,blogs,or comment forms).:TECHNIQUE:Attackers can establish rogue proxyservers for network protocols that give out the session ID and then redirectthe connection to the legitimate service.:TECHNIQUE:Attackers can emailattack URLs to potential victims through spam and phishing techniques.::STEP:3:PHASE:Exploit:DESCRIPTION:[Abuse the Victim's Session]Takeover the fixatedsession:Once the victim has achieved a higher level of privilege,possibly bylogging into the application,the attacker can now take over the session usingthe fixated session identifier.:TECHNIQUE:The attacker loads the predefinedsession ID into their browser and browses to protected data orfunctionality.:TECHNIQUE:The attacker loads the predefined session ID into their software and utilizes functionality with the rights of the victim.::.
[0071] name:Session Fixation
[0072] prerequisites:::Session identifiers that remain unchanged when the privilege levels change.::Permissive session management mechanism that accepts random user-generated session identifiers::Predictable session identifiers::
[0073] related_attack_patterns:::NATURE:ChildOf:CAPEC ID:593::
[0074] related_weaknesses:::384::664::732::
[0075] resources_required:::None:No specialized resources are required to execute this type of attack.::
[0076] skills_required:::SKILL:Only basic skills are required to determine and fixate session identifiers in a user's browser. Subsequent attacks may require greater skill levels depending on the attackers' motives.:LEVEL:Low::
[0077] typical_severity:High
[0078] S202: Determine alternative entities:
[0079] Determine alternative entities based on the original description text, including assets, potential vulnerabilities, and attack paths.
[0080] An example of the original description text in this embodiment is as follows:
[0081] In modern intelligent connected vehicles, manufacturers use the OTA (Over-the-Air) function to remotely push controller firmware updates to fix known vulnerabilities, optimize performance, or introduce new features. When the vehicle detects a new OTA update package, the system will notify the owner to select the update. The vehicle establishes communication with the OTA update server through the 4G / 5G network and downloads the update package. After the download is complete, the vehicle will verify the integrity of the update package and initiate the firmware installation. If the OTA update package lacks digital signature verification, an attacker can use a man-in-the-middle attack (MitM) to hijack and tamper with the update package during the download process, implant malicious firmware, resulting in abnormal vehicle systems, function failures, or privacy data leaks.
[0082] It can be analyzed that its assets are wireless communication modules (4G / 5G / Wi-Fi), the vulnerability is the lack of encryption protection for the wireless transmission link, and the attack path is to steal the OTA update package through the weak encryption link and forge version information.
[0083] S203: Match assets:
[0084] Match the assets in the alternative entities with each asset node in the knowledge graph, select the asset node with the highest matching degree as the corresponding matching asset, and then obtain the set of associated vulnerabilities and the set of attack paths of the matching asset from the knowledge graph.
[0085] In this embodiment, asset matching can be directly performed according to the asset name, that is, first generate the word embeddings of the asset names in the alternative entities and the names of each asset node in the knowledge graph, and then use the similarity of the word embeddings as the matching degree to achieve asset matching. The word embeddings are generated using the BERT model. When obtaining the set of associated vulnerabilities and the set of attack paths of the matching asset, screening can be performed based on vulnerabilities, that is, select the vulnerabilities directly associated with the matching asset to form the set of vulnerabilities, and the attack paths directly associated with these vulnerabilities form the set of attack paths.
[0086] S204: Match vulnerabilities and attack paths:
[0087] Match the vulnerabilities in the alternative entities with each vulnerability in the set of vulnerabilities, and select the vulnerability with the highest matching degree as the corresponding matching vulnerability. At the same time, match the attack paths in the alternative entities with each attack path in the set of attack paths, and select the attack path with the highest matching degree as the corresponding matching attack path.
[0088] The matching method is as follows:
[0089] Extract keywords for the weaknesses / attack paths in the alternative entities, extract keywords for the weaknesses / attack paths in the weakness set / attack path set, then generate word embeddings for each keyword, and then calculate the keyword word embedding similarity between the weaknesses / attack paths in the alternative entities and the weaknesses / attack paths in the weakness set / attack path set as the matching degree, so as to achieve the matching of weaknesses / attack paths.
[0090] In this embodiment, the keyword extraction algorithm adopts the KeyBERT algorithm. Examples of matching weaknesses, attack paths, and extracted keywords are as follows:
[0091] Matching weakness: The OTA update package lacks an appropriate digital signature verification mechanism, which may enable an attacker to intercept data transmission during the update process and inject malicious firmware into the vehicle's control unit. This may lead to system instability, unauthorized control, or data leakage.
[0092] Extracted weakness keywords: OTA update package, digital signature verification mechanism, interception, data leakage.
[0093] Matching attack path: An attacker can take advantage of the vulnerability of the lack of digital signature verification during the OTA update process to initiate a man-in-the-middle (MitM) attack. The attacker intercepts the data transmission between the OTA server and the vehicle, modifies the firmware package and injects malicious code. Then, the tampered firmware is installed on the vehicle's control unit, which may cause system failures, unauthorized remote control, or data theft.
[0094] Extracted attack path keywords: OTA update, lack of digital signature verification, man-in-the-middle (MitM) attack, interception, injection, tampering, unauthorized remote control.
[0095] S206: Construct a keyword set:
[0096] Construct a keyword set from the names of the matching assets obtained in step S203, the keywords of the matching weaknesses obtained in step S205, and the keywords of the matching attack paths.
[0097] S102: Fine-tune the large language model:
[0098] Select a large language model according to actual needs and fine-tune the large language model using the large language model training samples obtained in step S101.
[0099] In this embodiment, approximately 70%-80% of the data is selected from all large language model training samples as the training set for training the large language model, approximately 10%-15% of the data is selected as the validation set for adjusting hyperparameters during model training, and approximately 10%-15% of the data is reserved as the test set for evaluating the final performance of the model. The fine-tuning uses the P-Tuning algorithm, which is a way of parameter-efficient fine-tuning. By introducing "soft prompts", that is, a set of learnable embedding vectors, it replaces the traditional way of directly adjusting the model weights.
[0100] S103: Construct and train a security threat analysis model:
[0101] Construct a security threat analysis model according to actual needs. Its input is the description text, and the output is an N-dimensional threat level vector, and it is trained using the security threat analysis training samples obtained in step S101. The security threat analysis model in the present invention includes a text feature extraction module and a classifier, where:
[0102] The text feature extraction module is used to extract text features from the description text and send them to the classifier. In this embodiment, the text feature extraction module uses the BERT (Bidirectional Encoder Representations from Transformers) model.
[0103] The classifier is used to predict the N-dimensional threat level vector based on the text features. In this embodiment, the classifier uses a multilayer perceptron (MLP).
[0104] In this embodiment, in order to improve the training effect of the security threat analysis model, each evaluation index is used as a training task, and the loss function L of the multi-task is calculated using the following formula:
[0105]
[0106] where loss n represents the loss function of the nth evaluation index, and α n represents the weight corresponding to the n evaluation indexes, and n = 1, 2,..., N.
[0107] S104: Extract the keyword set of the typical business to be analyzed:
[0108] When it is necessary to perform security threat analysis on a certain typical business of an intelligent connected vehicle, obtain the original description text of the business scenario, and extract the keyword set from the original description text using the same method in step S101.
[0109] S105: Security threat analysis:
[0110] Input the keyword set of the typical business to be analyzed into the fine-tuned large model to obtain the text description of the threat scenario of the typical business, and then input the text description of the threat scenario into the trained security threat analysis model to obtain the N-dimensional threat level vector.
[0111] Although the above describes the illustrative specific embodiments of the present invention for the understanding of those skilled in the art of the present technology, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those of ordinary skill in the art of the present technology, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions and creations using the concept of the present invention are within the scope of protection.
Claims
1. An intelligent connected vehicle security threat analysis method based on a large language model, characterized in that, Including the following steps: S1: Collect a number of TARA example data of intelligent connected vehicles. For each piece of TARA example data, extract the natural language descriptions of the damage scenario and the attack path to form a threat scenario description text, and then extract a keyword set from it. Use the keyword set corresponding to each piece of TARA example data as the input and the corresponding threat scenario description text as the output to form the training samples for fine-tuning the large language model; Perform threat level annotation on each piece of TARA example data to obtain an N-dimensional threat level vector, where the nth element represents the evaluation score of the TARA example data on the nth evaluation index; Use the description text of each piece of TARA example data as the input and the threat level vector as the output to form the training samples for the security threat analysis model; S2: Select a large language model according to actual needs, and use the large language model training samples obtained in step S1 to fine-tune the large language model; S3: Build a security threat analysis model according to actual needs, whose input is the threat scenario description text and the output is an N-dimensional threat level vector, and use the training samples of the security threat analysis model obtained in step S1 for training; The security threat analysis model includes a text feature extraction module and a classifier, where: The text feature extraction module is used to extract text features from the description text and send them to the classifier; The classifier is used to predict an N-dimensional threat level vector based on the text features; S4: When it is necessary to perform security threat analysis on a certain typical business of an intelligent connected vehicle, obtain the original description text of the typical business to be analyzed, and extract a keyword set from the original description text using the same method as in step S1; S5: Input the keyword set of the typical business to be analyzed into the fine-tuned large language model to obtain the threat scenario description text of the typical business, and then input the threat scenario description text into the trained security threat analysis model to obtain an N-dimensional threat level vector.
2. The intelligent networked vehicle security threat analysis method according to claim 1, wherein The keyword set includes the asset name, vulnerability keywords, and attack path keywords.
3. The intelligent networked vehicle security threat analysis method according to claim 2, wherein The keyword sets in step S1 and step S4 are extracted based on the knowledge graph. The specific method is: 1) Obtain information security threat intelligence data from a public information security knowledge base, including asset data, vulnerability data, weakness data, and attack pattern data, and associate the assets, vulnerabilities, weaknesses, and attack patterns in different data; 2) Build a knowledge graph based on the information security threat intelligence data obtained in step 1), where the nodes include assets, vulnerabilities, weaknesses, and attack models, and the relationships include the associations between vulnerabilities and assets, between vulnerabilities and weaknesses, and between weaknesses and attack patterns; 3) Determine alternative entities from the original description text, including assets, potential weaknesses, and attack paths; 4) Match the assets in the alternative entities with each asset node in the knowledge graph, select the asset node with the highest matching degree as the corresponding matching asset, and then obtain the set of weaknesses and the set of attack paths associated with the matching asset from the knowledge graph; 5) Match the weaknesses in the alternative entities with each weakness in the weakness set, and select the weakness with the highest matching degree as the corresponding matching weakness; at the same time, match the attack paths in the alternative entities with each attack path in the attack path set, and select the attack path with the highest matching degree as the corresponding matching attack path; the matching method is as follows: Extract keywords from the weaknesses / attack paths in the alternative entities, extract keywords from the weaknesses / attack paths in the weakness set / attack path set, then generate word embeddings for each keyword, and then calculate the keyword word embedding similarity between the weaknesses / attack paths in the alternative entities and the weaknesses / attack paths in the weakness set / attack path set as the matching degree, so as to achieve weakness / attack path matching; 6) Construct a keyword set from the name of the matching asset obtained in step 3), the keywords of the matching weakness obtained in step 5), and the keywords of the matching attack path.
4. The intelligent networked vehicle security threat analysis method according to claim 1, wherein In step S1, the dimension N of the threat level vector is 9. The evaluation indicators of the damage scenario include ratings in 4 dimensions: functional safety, finance, operation, and privacy. The evaluation indicators of the attack path include ratings in 5 dimensions: attack time consumption, professional knowledge, product publicity, opportunity window, and attack equipment.
5. The intelligent networked vehicle security threat analysis method according to claim 1, characterized in that In step S3, the text feature extraction module uses the BERT model.
6. The intelligent networked vehicle security threat analysis method according to claim 1, wherein In step S3, the classifier uses a multi-layer perceptron.
7. The intelligent networked vehicle security threat analysis method according to claim 1, wherein, In the training process of the security threat analysis model in step S3, the calculation formula of the loss function L is as follows: Among them, loss n represents the loss function of the nth evaluation index, and α n represents the weight corresponding to the nth evaluation index, where n = 1, 2, …, N.
Citation Information
Cited By
Generation of TARA-based IDPS rules utilizing generative artificial intelligence
US12500915B1