A Supply Chain Security Risk Identification Method and System for Wind Power Stations
Through the risk identification method combined with NLP technology and random forest model, the speed and efficiency of the safety risk identification of wind farm supply chain is solved, and automated data analysis and real-time response to the wind farm supply chain are realized, which improves the flexibility and accuracy of safety management.
Patent Information
- Application Number
- CN202410568271.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-09
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2044-05-09
AI Technical Summary
Traditional methods are slow and inefficient in identifying safety risks in wind farm supply chains, unable to respond to emerging security threats in real time, and lack of automation tools lead to information silos and limited analysis capabilities.
NLP technology is used to extract security-related discussion topics, combine LDA and TF-IDF algorithms to evaluate community discussion activity, use random forest models to predict risk levels, and trigger corresponding risk response processes, including component isolation, vulnerability scanning and patch deployment.
It has achieved rapid and accurate identification and response to safety risks in the supply chain of wind farms, improved safety management efficiency, timely identification and response to potential threats, and ensured the safety and reliability of energy supply.
Smart Images

Figure CN118428731B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer security technology, and particularly to a method and system for identifying supply chain security risks for wind farm stations. Background Art
[0002] In a globalized economic environment, the complexity of the supply chain is increasing continuously, especially in the key energy industry of wind farm stations. Wind farm stations involve a large number of high-tech software components, precision hardware devices and diverse service providers. These factors together increase the security risks of the supply chain and pose a major challenge to enterprises. Identifying supply chain security risks for wind farm stations requires extracting relevant information from numerous data sources, evaluating potential security threats, and taking corresponding response measures.
[0003] Most traditional methods rely on manual operations. From data collection, analysis to risk assessment, there is a lack of effective automated tools, resulting in slow processing speed and inability to respond promptly to newly emerging security threats.
[0004] In the supply chain, information sharing among different organizations is not smooth, forming information silos. This leads to a lack of comprehensiveness in security risk assessment and makes it difficult to accurately identify cross-organization or cross-level security threats.
[0005] Traditional methods have limited analytical capabilities when dealing with large-scale data and complex security scenarios, and it is difficult to extract valuable security-related information from massive supply chain data to identify potential security risks.
[0006] Due to the lack of real-time analysis and automated response mechanisms, the response measures after risk identification by traditional methods are often not flexible and timely enough to effectively prevent or mitigate the impact of security incidents.
[0007] With the continuous emergence of new technologies and the continuous evolution of new security threats, traditional security management methods are difficult to adapt and update quickly, resulting in security protection measures lagging behind the development of security threats. Summary of the Invention
[0008] In view of the above existing problems, the present invention is proposed.
[0009] Therefore, the technical problem solved by the present invention is: to solve the problems of slow speed, low efficiency and inability to respond to newly emerging security threats in real time in traditional supply chain security risk identification for wind farm stations through highly automated data analysis and intelligent risk identification methods.
[0010] To solve the above technical problems, the present invention provides the following technical solution: A method for identifying supply chain security risks for wind farm stations, comprising:
[0011] Collecting wind farm station supply chain data and performing data preprocessing;
[0012] Extract the topics of security-related discussions based on NLP technology;
[0013] Train a random forest model to predict the risk level, and trigger a risk response process according to the risk level.
[0014] As a preferred solution of the supply chain security risk identification method for wind farm stations according to the present invention, wherein: the wind farm station supply chain data includes code submission records, security vulnerability reports, developer community discussion content, and component documentation and maintenance guides;
[0015] The data preprocessing includes removing duplicates, filling in missing values, and unifying the date format.
[0016] As a preferred solution of the supply chain security risk identification method for wind farm stations according to the present invention, wherein: the extracting the topics of security-related discussions based on NLP technology includes
[0017] Analyze the developer community discussion content, extract the text for analysis, use the LDA algorithm to analyze the community discussion text, identify the discussion topics, assign topics to each document, and each topic is defined by a set of keywords;
[0018] For each identified discussion topic, use the TF-IDF algorithm to extract keywords, evaluate the importance of each keyword in the document, assign weights to the keywords under each topic, construct a topic-keyword matrix, and obtain the keyword frequency;
[0019] For the vulnerabilities related to each topic, calculate the repair time and conduct statistical analysis;
[0020] Based on the extracted topics, keyword frequencies, and vulnerability repair times, evaluate the community discussion activity of each topic and convert the information into numerical features.
[0021] As a preferred solution of the supply chain security risk identification method for wind farm stations according to the present invention, wherein: the evaluating the community discussion activity of each topic is expressed as,
[0022]
[0023] wherein, represents the distribution weight of document d with respect to topic t i RT i represents the average vulnerability repair time associated with topic t i to evaluate the repair urgency, represents the TF weight of keyword k in the document j IDF kj represents keyword k jInverse document frequency, ∈ represents a constant, i represents traversing topics, and j represents traversing keywords.
[0024] As a preferred solution of the supply chain security risk identification method for wind farm stations according to the present invention, wherein: the conversion of information into numerical features is expressed as
[0025]
[0026]
[0027] wherein, S represents the activity score, C represents the comprehensive consideration factor, and P i represents the participation index of the i-th discussion thread, which is determined according to the number of replies, views, and interaction of the thread, and V i represents the visual influence of keywords in the i-th discussion thread, which is determined by analyzing the frequency of keyword appearance in community discussions, M represents the number of discussion threads, ΔT represents the time distribution of community discussions within the considered time period, which is the difference between the start and end times of the discussion, and ω represents the weight coefficient.
[0028] As a preferred solution of the supply chain security risk identification method for wind farm stations according to the present invention, wherein: the prediction of risk levels includes extracting the comprehensive consideration factor C, the distribution weight of the document regarding the topic and the TF-IDF weight of keywords in the document Perform and normalization processing, perform feature encoding using one-hot encoding, and convert the and of each document into a fixed-length vector;
[0029] Combine the comprehensive consideration factor C with the standardized topic distribution vector and the keyword TF-IDF vector in sequence to form a single comprehensive feature vector, determine C as a scalar, and are each converted into a vector of length N, generate a feature vector of length 1 + 2N, match the feature vector with the risk levels in the preset risk library to form a training data set, and use the constructed feature vector as the input of the machine learning model;
[0030] Select the random forest model as the prediction tool, use the training set data to train the random forest model, use the test set data to evaluate the performance of the model, and use the evaluated random forest model to predict the risk level;
[0031] The risk levels include low risk level, medium risk level, and high risk level.
[0032] As a preferred solution of the supply chain security risk identification method for wind farm stations according to the present invention, wherein: the trigger risk response process includes, when the high-risk level is judged, isolating the detected high-risk components, spreading potential security threats, and after isolation, transferring the components to a sandbox environment. While analyzing in the sandbox environment, perform a vulnerability scan on the high-risk components to identify and locate security vulnerabilities. When the vulnerabilities are determined, automatically deploy repair patches or update the components to a secure version to repair the vulnerabilities;
[0033] When the medium-risk level is judged, perform a security assessment on the medium-risk components. The security assessment includes vulnerability scanning and dependency checking, identify the dependency relationships of the medium-risk components with other components, formulate targeted protection measures, and adjust the security protection strategies of other components directly or indirectly connected to the medium-risk components to reduce the influence range of potential threats;
[0034] When the low-risk level is judged, perform an incremental update on the components identified as low-risk, perform automatic backup, regard the monitoring and recording of low-risk events as a continuous process, use the collected data for risk log analysis, and identify potential security threat patterns from the log data through machine learning and pattern recognition technologies. Based on the analysis results, adjust the sensitivity of the intrusion detection system.
[0035] Another object of the present invention is to provide a supply chain security risk identification system for wind farm stations. In the existing supply chain security risk management, there are challenges in quickly and accurately identifying and responding to security vulnerabilities. For the supply chain system of wind farm stations, the diverse software components and services included make it more difficult to identify security risks. Traditional security risk identification methods often rely on manual analysis or simple automated tools, which are inefficient in dealing with large-scale and complex data and cannot be updated in real time to cope with newly emerging threats.
[0036] To solve the above technical problems, the present invention provides the following technical solution: a supply chain security risk identification system for wind farm stations, including: a data acquisition module, a derivation generation module, and a vulnerability detection module; the data acquisition module is used to obtain the expression function of the software source code to be tested in the supply chain; the derivation generation module is used to generate constraint derivations by backward slicing of the expression function; the vulnerability detection module is used to obtain the software vulnerability detection result based on the similarity comparison of the constraint derivations.
[0037] A computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the steps of the supply chain security risk identification method for wind farm stations as described above.
[0038] A computer-readable storage medium stores a computer program thereon, and when the computer program is executed by a processor, the steps of the above-mentioned supply chain security risk identification method for a wind farm are implemented.
[0039] Advantages of the present invention: The supply chain security risk identification method for a wind farm provided by the present invention realizes the automated collection, processing, and analysis of wind farm supply chain data, as well as a dynamic response strategy based on risk levels, by integrating natural language processing (NLP) technology and machine learning models. This method significantly improves the efficiency and effectiveness of wind farm supply chain security management. Through automated data analysis and real-time risk assessment, it can timely identify and respond to security threats in the supply chain, thus providing a more flexible supply chain security solution.
[0040] Especially in wind farms, where the hardware and software components often need to operate under extreme climate conditions, security management is particularly important. The present invention can effectively identify potential security hazards that may be caused by extreme climate, equipment aging, or improper maintenance. In addition, by analyzing historical and real-time data, this method can predict possible component failures and supply interruptions, and take measures in advance to ensure the continuous operation of the wind farm and the stability of energy output.
[0041] The implementation of the present invention not only improves the accuracy of risk identification, but also shortens the response time, enabling wind farms to better cope with rapidly changing supply chain conditions and potential security threats, and ultimately ensuring the security and reliability of energy supply. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings based on these drawings without creative efforts.
[0043] Figure 1 It is the overall flowchart of a supply chain security risk identification method for a wind farm provided by an embodiment of the present invention.
[0044] Figure 2 It is the overall structure diagram of a supply chain security risk identification system for a wind farm provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0045] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following will provide a detailed description of the specific embodiments of the present invention in conjunction with the accompanying drawings of the specification. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative work shall fall within the scope of protection of the present invention.
[0046] Many specific details are set forth in the following description to facilitate a thorough understanding of the present invention, but the present invention may also be implemented in other ways different from those described herein. Those skilled in the art can make similar generalizations without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.
[0047] Embodiment 1
[0048] Referring to Figure 1 , an embodiment of the present invention provides a method for identifying supply chain security risks for a wind farm station, including:
[0049] Collecting wind farm station supply chain data and performing data preprocessing; extracting the themes of security-related discussions based on NLP technology; training a random forest model to predict the risk level, and triggering a risk response process according to the risk level.
[0050] Security-related discussions often contain the views and responses of the developer community on specific vulnerabilities, security practices, or emerging threats. These information cannot be directly reflected by traditional data features. By extracting the themes of these discussions through NLP technology, we can more comprehensively understand the community's concerns about security issues and potential risks that may exist. Quick response and repair of security vulnerabilities are an important indicator for evaluating component security. By statistically analyzing the vulnerability repair time, we can quantify the response capabilities of each part of the component or supply chain to security events, which is an important dimension for measuring the security situation. In the risk response process, different types of risks may require different handling measures. By introducing text analysis based on NLP and statistical analysis of vulnerability repair time, the model can more finely divide the risk levels and support the implementation of more precise and targeted security measures.
[0051] The wind farm station supply chain data includes code submission records, security vulnerability reports, developer community discussion content, and component documentation and maintenance guides; performing data preprocessing includes removing duplicates, filling in missing values, and unifying the date format.
[0052] Extracting the themes of security-related discussions based on NLP technology includes analyzing the developer community discussion content, extracting text for analysis, using the LDA algorithm to analyze the community discussion text, identifying the discussion themes, and assigning themes to each document. Each theme is defined by a set of keywords;
[0053] For each identified discussion topic, use the TF-IDF algorithm to extract keywords, evaluate the importance of each keyword in the document, assign weights to the keywords under each topic, construct a topic-keyword matrix, and obtain the keyword frequencies;
[0054] For the vulnerabilities related to each topic, calculate the repair time and conduct statistical analysis;
[0055] Based on the extracted topics, keyword frequencies, and vulnerability repair times, evaluate the community discussion activity of each topic and convert the information into numerical features.
[0056] Evaluating the community discussion activity of each topic is expressed as,
[0057]
[0058] where, represents the distribution weight of document d with respect to topic t i RT i represents the average vulnerability repair time associated with topic t i to evaluate the repair urgency, represents the TF weight of keyword k j in the document, IDF kj represents the inverse document frequency of keyword k j and ∈ represents a constant, i represents traversing topics, and j represents traversing keywords.
[0059] Traditional methods for evaluating the community discussion activity of each topic rely on keyword matching or simple text queries to identify security-related discussions, lack an understanding of the deep semantics and context of the text, and are unable to accurately capture complex discussion topics. They often evaluate the community discussion activity based on the number of discussion posts or reply frequencies, without considering the quality of the discussion content, the relevance to specific security issues, or the sense of urgency expressed in the discussion. Generally, the vulnerability repair time is used as the only indicator to measure the urgency of security issues, ignoring the degree of community attention to the vulnerability discussion and the sentiment tendency of the discussion.
[0060] Our invention method automatically identifies topics from community discussions through the LDA algorithm, which can understand the latent semantic structure of the text and identify multiple topics of specific discussions. Each topic is defined by a set of highly relevant keywords. This method can more accurately reflect the true content and focus of community discussions. Using the topic-keyword matrix and keyword frequency (obtained from TF-IDF analysis), combined with the vulnerability repair time, the community discussion activity of each topic is evaluated. This method not only considers the quantity of discussions but also comprehensively takes into account the depth of the discussion topics, the importance of keywords, and the urgency of related vulnerabilities, providing a multi-dimensional and more accurate measure for community discussion activity. The vulnerability repair time is closely combined with the content and sentiment of community discussions to dynamically evaluate the urgency of vulnerabilities related to each topic. This not only reflects the repair speed of vulnerabilities but also considers the community's attention to the vulnerability and the urgency of the discussion. Then, the results of this comprehensive evaluation are converted into numerical features to update the feature vector, providing more comprehensive and real-time data support for subsequent security risk assessment and decision-making.
[0061] Our invention method can more deeply understand and analyze security-related discussions in the community and more accurately evaluate the activity and urgency of discussions compared with traditional methods by introducing advanced NLP technologies and statistical analysis.
[0062] First, analyze the content of developer community discussions through the LDA algorithm to identify the topics of the discussions and assign topics to each document. Provide the data basis, that is, the distribution weight of each document d regarding topic t i
[0063] Next, use the TF-IDF algorithm to extract the keywords under each topic and evaluate their importance in the document to construct a topic-keyword matrix. This step provides the data basis for and in the formula, that is, the TF weight and inverse document frequency of keyword k j in the document.
[0064] Then, for the vulnerabilities related to each topic, calculate their repair time and conduct statistical analysis. Provide the data for RT i that is, the average vulnerability repair time associated with topic t i
[0065] Finally, based on the extracted topics, keyword frequencies, and vulnerability repair times, evaluate the community discussion activity of each topic. By quantifying the discussion activity of each topic and the urgency of vulnerabilities, a numerical feature is generated for each topic.
[0066] Convert the information into numerical features and represent it as
[0067]
[0068]
[0069] Among them, S represents the activity score, C represents the comprehensive consideration factor, P i represents the participation index of the i-th discussion thread, which is determined according to the number of replies, views, and interaction of the thread, V i represents the visual influence of keywords in the i-th discussion thread, which is determined by analyzing the frequency of appearance of keywords in community discussions, M represents the number of discussion threads, ΔT represents the time distribution of community discussions within the considered time period, which is the difference between the start and end times of the discussion, and ω represents the weight coefficient.
[0070] Traditional activity scoring methods focus on a single dimension and only pay attention to the number of posts or simple participation (such as the number of replies or likes). This model provides a more comprehensive scoring mechanism by combining the participation index P i , visual influence V i and time sensitivity ΔT.
[0071] Fusing multi-dimensional data enables the activity score to more comprehensively reflect the actual situation of community discussions, including the quality of discussions, the activity level of participants, and the urgency of discussions, thus improving the accuracy and practicality of the score.
[0072] The logarithmic sum and sigmoid functions are used to process UI d and C. This processing method non-linearly adjusts the original data, making the scoring result smoother and easier for subsequent processing. It enhances the model's ability to handle extreme values while retaining the subtle changes in the data, improving the sensitivity and robustness of the scoring system.
[0073] By introducing ΔT, the model can quantify the time distribution of discussions and capture the urgency and timeliness of discussions. Time sensitivity is a key factor in evaluating the urgency of security discussions. It ensures that the activity score not only reflects the quantity and quality of discussions but also reflects the timeliness of discussions, providing data support for the timely response to security issues.
[0074] P i = α·log(1 + R i ) + β·log(1 + L i + S i )
[0075] Among them, R i represents the number of replies to the i-th discussion thread, L i is the number of likes the post received, S i represents the number of shares of the post, and α and β are coefficients used to adjust the influence of different interaction types, and log represents the natural logarithm.
[0076]
[0077] Among them, F k is the occurrence frequency of the k-th keyword in the post, and I k represents the degree of interaction caused by the k-th keyword, which is the number of relevant comments in the present invention. γ is a coefficient for balancing various factors, and K is the total number of keywords identified in the post.
[0078] ΔT = T end - T start
[0079] Among them, T end is the time point when the discussion ends, and T start is the time point when the discussion starts.
[0080] The predicted risk level includes extracting comprehensive consideration factors C, the distribution weight of the document regarding the theme and the TF-IDF weight of the keywords in the document Perform normalization processing on and and perform feature encoding using one-hot encoding. Convert and of each document into a vector with a fixed length. Since the document may be associated with different numbers of themes and keywords, it is necessary to ensure that the lengths of the converted vectors are the same, which is achieved by selecting TOPN themes and keywords or using padding / truncation strategies;
[0081] Concatenate the comprehensive consideration factor C with the standardized theme distribution vector and the keyword TF-IDF vector in sequence to form a single comprehensive feature vector. Determine that C is a scalar, and are each converted into a vector with a length of N, generating a feature vector with a length of 1 + 2N. Match the feature vector with the risk levels in the preset risk library to form a training data set, and use the constructed feature vector as the input of the machine learning model;
[0082] Select the random forest model as the prediction tool, train the random forest model using the training set data, evaluate the performance of the model using the test set data, and use the evaluated random forest model to predict the risk level;
[0083] The risk levels include low risk level, medium risk level, and high risk level.
[0084] The trigger risk response process includes isolating the detected high-risk components when the high-risk level is judged, spreading potential security threats. After isolation, the components are transferred to the sandbox environment. While analyzing in the sandbox environment, vulnerability scanning is performed on the high-risk components to identify and locate security vulnerabilities. When the vulnerabilities are determined, repair patches are automatically deployed or the components are updated to a secure version to fix the vulnerabilities;
[0085] When the medium-risk level is judged, a security assessment is performed on the medium-risk components. The security assessment includes vulnerability scanning and dependency check, identifying the dependency relationships between the medium-risk components and other components, formulating targeted protection measures, and adjusting the security protection strategies of other components directly or indirectly connected to the medium-risk components to reduce the influence scope of potential threats;
[0086] When the low-risk level is judged, incremental updates are performed on the components identified as low-risk, automatic backups are made, and the monitoring and recording of low-risk events are carried out as a continuous process. Risk log analysis is performed using the collected data. Through machine learning and pattern recognition techniques, potential security threat patterns are identified from the log data. Based on the analysis results, the sensitivity of the intrusion detection system is adjusted.
[0087] The method of the present invention integrates the LDA algorithm, TF-IDF algorithm, and random forest model based on the specific requirements of supply chain security risk identification for wind farm stations. First, the LDA algorithm is used to automatically extract security-related discussion topics from the text data related to the supply chain. This step provides the necessary context information for the subsequent keyword importance evaluation. Immediately afterwards, the TF-IDF algorithm is used to accurately evaluate the importance of each keyword in these discussion topics, ensuring that the extracted feature vectors can accurately reflect the security relevance of the text content.
[0088] By combining the topic distribution weights extracted by LDA, the keyword importance weights obtained by the TF-IDF algorithm, and the comprehensive consideration factor C, a comprehensive feature vector for machine learning model training is formed. The construction of this feature vector takes into account the multi-dimensional characteristics of the supply chain text data, ensuring that the risk identification model can comprehensively understand and evaluate security risks.
[0089] The random forest model is trained based on these comprehensive feature vectors, which can not only accurately predict the risk level but also automatically trigger the corresponding risk response process according to the prediction results. This design reflects a complete closed-loop from data processing to risk prediction and then to response measures, and each step is carefully designed to solve the actual problems in supply chain security risk management.
[0090] The data processing and risk management processes are optimized, overcoming the limitations that may exist in the application of a single technology: When using the LDA algorithm for topic extraction, the specific context of supply chain security is considered, and through the optimization of the customized topic model, the relevance and accuracy of the topics are improved. This step solves the problem of context mismatch that may be encountered in the application of traditional topic extraction in specific fields.
[0091] By directly correlating the results of the TF-IDF algorithm with the topics extracted by LDA, your method can more accurately identify the keywords highly relevant to supply chain security. This precise matching is achieved through the optimized collaborative work between algorithms, solving the problems of information redundancy or missed detection that may be caused by simple algorithm stacking.
[0092] Based on the risk level predicted by the random forest model, the risk response strategy is automatically adjusted in the most suitable way for the current security state.
[0093] Example 2
[0094] Refer to Figure 2 , an embodiment of the present invention provides a supply chain security risk identification system for a wind farm, including:
[0095] A data collection module, a feature extraction module, and a risk identification module.
[0096] The data collection module is used to collect the supply chain data of the wind farm and perform data preprocessing.
[0097] The feature extraction module is used to extract the topics of security-related discussions based on NLP technology.
[0098] The risk identification module is used to train the random forest model to predict the risk level and trigger the risk response process according to the risk level.
[0099] Example 3
[0100] An embodiment of the present invention, which is different from the previous two embodiments, is:
[0101] If the above-mentioned functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs.
[0102] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a definite sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or in combination with these instruction execution systems, apparatus, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device.
[0103] More specific examples (non-exhaustive list) of computer-readable media include the following: electrical connection parts with one or more wirings (electronic devices), portable computer disk cartridges (magnetic devices), random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memories), optical fiber devices, and portable compact disc read-only memories (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, then editing, interpreting, or otherwise processing it as necessary, and then storing it in a computer memory.
[0104] It should be understood that each part of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one of the following techniques known in the art or a combination thereof can be used: discrete logic circuits with logic gate circuits for implementing logical functions on data signals, application specific integrated circuits with suitable combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.
[0105] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered by the scope of the claims of the present invention.
Claims
1. A supply chain security risk identification method for wind farms, characterized in that, Including: Collecting the supply chain data of the wind farm station and performing data preprocessing; Extracting the topics of security-related discussions based on NLP technology; Training a random forest model to predict the risk level, and triggering a risk response process according to the risk level; The supply chain data of the wind farm station includes code submission records, security vulnerability reports, developer community discussion content, and component documentation and maintenance guides; The data preprocessing includes removing duplicates, filling in missing values, and unifying the date format; The extracting of the security-related discussion topics based on NLP technology includes analyzing the developer community discussion content, extracting text for analysis, using the LDA algorithm to analyze the community discussion text, identifying the discussion topics, and assigning topics to each document, with each topic defined by a set of keywords; For each identified discussion topic, using the TF-IDF algorithm to extract keywords, evaluating the importance of each keyword in the document, assigning weights to the keywords under each topic, constructing a topic-keyword matrix, and obtaining the keyword frequency; For the vulnerabilities related to each topic, calculating the repair time and performing statistical analysis; Based on the extracted topics, keyword frequencies, and vulnerability repair times, evaluating the community discussion activity of each topic and converting the information into numerical features; The evaluating of the community discussion activity of each topic is expressed as Among them, represents the distribution weight of document d with respect to topic t i RT i represents the average vulnerability repair time associated with topic t i to evaluate the urgency of repair, represents the TF weight of keyword k in the document j IDF kj represents the inverse document frequency of keyword k j ∈ represents a constant, i represents traversing topics, and j represents traversing keywords; The converting of the information into numerical features is expressed as Among them, S represents the activity score, C represents the comprehensive consideration factor, and P i represents the participation index of the i-th discussion thread, which is determined according to the number of replies, views, and interaction of the thread, and V i represents the visual influence of keywords in the i-th discussion thread, which is determined by analyzing the frequency of keyword appearance in community discussions. M represents the number of discussion threads, ΔT represents the time distribution of community discussions within the considered time period, which is the difference between the start and end times of the discussion, and ω represents the weight coefficient; The predicted risk level includes extracting comprehensive consideration factors C, the distribution weight of the document regarding the theme and the TF-IDF weight of the keywords in the document Perform normalization processing on and Use one-hot encoding for feature encoding, and convert the and of each document into a vector with a fixed length; The comprehensive consideration factor C and the standardized topic distribution vector and the keyword TF-IDF vector are concatenated in sequence into a single comprehensive feature vector, determining that C is a scalar, and are each converted into a vector of length N, generating a feature vector of length 1 + 2N. The feature vector is matched with the risk levels in the preset risk database to form a training dataset, and the constructed feature vector is used as the input to the machine learning model; Selecting a random forest model as the prediction tool, training the random forest model using the training set data, evaluating the performance of the model using the test set data, and using the evaluated random forest model to predict the risk level; The risk levels include low risk level, medium risk level, and high risk level.
2. The supply chain security risk identification method for a wind farm as described in claim 1, characterized in that: The triggering of the risk response process includes when it is judged as the high risk level, isolating the detected high-risk components, diffusing potential security threats, and after isolation, transferring the components to the sandbox environment. While analyzing in the sandbox environment, performing a vulnerability scan on the high-risk components to identify and locate security vulnerabilities. When the vulnerabilities are determined, automatically deploying repair patches or updating the components to a secure version to repair the vulnerabilities; When it is judged as the medium risk level, performing a security assessment on the medium-risk components. The security assessment includes vulnerability scanning and dependency checking, identifying the dependency relationships between medium-risk components and other components, formulating targeted protection measures, and adjusting the security protection strategies of other components directly or indirectly connected to the medium-risk components to reduce the impact range of potential threats; When it is judged as the low risk level, performing an incremental update on the components identified as low risk, performing automatic backups, taking the monitoring and recording of low-risk events as a continuous process, using the collected data for risk log analysis, and identifying potential security threat patterns from the log data through machine learning and pattern recognition technologies. Based on the analysis results, adjusting the sensitivity of the intrusion detection system.
3. A system adopting the supply chain security risk identification method for a wind farm station as described in claim 1 or 2, characterized in that, Including: A data collection module, a feature extraction module, and a risk identification module; The data collection module is used to collect the supply chain data of the wind farm station and perform data preprocessing; The feature extraction module is used to extract the security-related discussion topics based on NLP technology; The risk identification module is used to train a random forest model to predict the risk level, and trigger a risk response process according to the risk level.
4. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the steps of the supply chain security risk identification method for a wind farm station described in claim 1 or 2 are implemented.
5. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the steps of the supply chain security risk identification method for a wind farm station described in claim 1 or 2 are implemented.
Citation Information
Patent Citations
Pipeline risk grade evaluation method and device based on support vector machine
CN113191599A
Data security assessment method and system
CN116861446A