An artificial intelligence driven information network system risk early warning method and system

CN122845264APending Publication Date: 2026-09-29HANGZHOU ZHISHUN TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611112565.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-24
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

这种匹配偏差使得安全运营人员看到错误的资产风险告警,执行不必要的处置动作,而实际风险资产持续面临攻击,造成防护盲区

Benefits of technology

[0017]相比于现有技术,本发明的有益效果为:本发明通过在资产标识历史链表中以标识有效期区间保留内部资产网络标识的完整变更历史,并以情报生成时间戳为锚点执行区间定位生成情报时刻资产标识快照,将现有技术中基于资产当前标识的静态匹配转变为基于历史状态区间的时序对齐匹配,从根源上切断了因云实例弹性伸缩、容器化部署和负载均衡导致的网络标识动态变更与情报采集时刻之间的时效性错位,规避了受影响资产识别结果与实际风险分布出现偏差的失准来源;在此基础上,以结构化命中标记和语义命中标记的双标记融合驱动情报置信度加权影响分值的三分支计算,并通过图神经网络推理模块在资产关联拓扑图上执行风险传播推理,使系统在情报与资产标识精确对应、情报与资产语义隐含关联以及受影响资产间接暴露三类场景下均具备识别能力,实现了动态网络环境中风险预警的全覆盖与高精度。本发明属于计算机数据处理领域,可广泛应用于企业资产安全风险分析。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122845264A_ABST
    Figure CN122845264A_ABST
Patent Text Reader

Abstract

The application relates to the technical field of network security risk early warning, and discloses an information network system risk early warning method and system driven by artificial intelligence, which comprises the following steps: constructing an asset identification history chain table and an intelligence time asset identification snapshot, generating an intelligence asset correlation pair through fusion of structured interval matching and semantic vector retrieval, calculating a comprehensive risk score and performing propagation enhancement through a graph neural network, and generating a graded early warning disposal suggestion; according to the application, the current state matching is replaced by historical state interval matching, the timeliness dislocation is cut off, the systematic missing matching of the pure field matching is made up, and the indirectly exposed assets are identified, so that the attack surface risk full life cycle monitoring and closed loop disposal are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of cybersecurity risk early warning technology, and in particular to an artificial intelligence-driven method and system for early warning of risks in information network systems. Background Technology

[0002] As cyber threats become increasingly complex, threat intelligence systems are widely used in enterprise security protection. These systems identify potentially affected assets by correlating and matching external threat intelligence with internal asset information. Existing systems typically rely on static identifier fields such as IP addresses, domain names, and file hashes carried in the intelligence to locate internal assets. However, in real-world enterprise deployment environments, the adoption of technologies such as cloud instance elastic scaling, containerization, and load balancing means that the network identifiers of internal assets are constantly changing, and static identifier matching methods cannot reflect this dynamic characteristic.

[0003] In actual security operations, it is common for intelligence matching results to fail to accurately pinpoint the actual risky assets. For example, a company uses cloud servers to host web services, and its public IP address is periodically reassigned by the cloud platform. An external threat intelligence platform may flag a particular IP address for active vulnerability exploitation at a certain time. However, by the time the intelligence access system completes parsing and performs asset matching, this IP address has already been reassigned to a test host unrelated to the web service, while the original web service has switched to the new IP. The intelligence matching module outputs that the test host is the affected asset, and security operations personnel initiate actions to address it. However, the original web service, which actually carries the risk, remains unidentified and continues to be exposed to threats. This matching discrepancy causes security operations personnel to see incorrect asset risk alerts and perform unnecessary actions, while the actual risky asset continues to face attacks, creating a protection blind spot. Therefore, reducing the timeliness and inaccuracy of matching results during the threat intelligence-internal asset correlation matching process is a pressing issue that needs to be addressed in this field. Summary of the Invention

[0004] To address the shortcomings of existing technologies, this invention aims to provide an AI-driven risk early warning method for information network systems. This method acquires raw external threat intelligence data and internal asset status data from multi-source data access channels, performs indicator field parsing, timestamp extraction, identifier change event tracking, and historical snapshot recording to generate intelligence indicator sequences and historical asset identifier lists. Based on the intelligence generation timestamp, it locates historical asset identifiers and generates snapshots. Semantic embedding is obtained by combining pre-trained language representation model encoding, followed by dual-path matching of structured interval matching and semantic vector retrieval. This results in the generation of intelligence-asset association pairs, followed by multi-dimensional association analysis to calculate the weighted impact score of intelligence confidence and the asset exposure score. These scores are then input into a graph neural network inference module for asset association risk propagation analysis, ultimately achieving tiered risk early warning and improving the accuracy of risk warnings and the efficiency of response.

[0005] To achieve the above objectives, the present invention adopts the following technical solution: In a first aspect, the present invention provides an artificial intelligence-driven risk early warning method for information network systems, comprising: The system obtains raw external threat intelligence data and internal asset status data from a multi-source data access channel. It performs indicator field parsing and timestamp extraction on the raw external threat intelligence data and performs identifier change event tracking and historical snapshot recording on the internal asset status data, generating an intelligence indicator sequence and an asset identifier historical linked list, respectively. For each intelligence indicator in the intelligence indicator sequence, the validity period range is located in the asset identifier history chain based on the intelligence generation timestamp, the set of historical asset identifiers corresponding to the identifier validity period range in which the intelligence generation timestamp falls is extracted, and an asset identifier snapshot at the intelligence moment is generated. Perform pre-trained language representation model encoding on all attack indicator fields in the intelligence indicator sequence and the asset description text in the asset identifier snapshot at the intelligence time, respectively, to generate an intelligence semantic embedding library and an asset semantic embedding library. Based on the intelligence indicator sequence, the intelligence moment asset identifier snapshot, the intelligence semantic embedding library, and the asset semantic embedding library, structured interval matching and semantic vector retrieval are performed respectively, and the two matching results are merged to generate a list of intelligence asset association pairs; Perform multidimensional association analysis on the intelligence asset association pair list, calculate the intelligence confidence weighted impact score and asset exposure score based on the combined state of structured hit tags and semantic hit tags, and generate an affected asset risk assessment table; The affected asset risk assessment form is input into the graph neural network inference module to perform asset-related risk propagation analysis, outputting a propagation risk enhancement score. Combined with the comprehensive risk score, a graded risk warning is executed, differentiated disposal suggestions are generated, and a risk warning disposal report is output.

[0006] Preferably, the step of parsing the indicator fields and extracting the timestamps from the original external threat intelligence data includes: obtaining the original external threat intelligence data from the threat intelligence platform interface, open intelligence subscription sources, and internal security device alarm channels; parsing the original external threat intelligence data to extract the attack indicator field and intelligence generation timestamp for each intelligence; binding the attack indicator field and the intelligence generation timestamp together and storing them as an intelligence indicator record; and arranging all the intelligence indicator records in ascending order according to the intelligence generation timestamp to generate an intelligence indicator sequence.

[0007] Preferably, the step of tracking and recording identifier change events and historical snapshots of the internal asset status data includes: retrieving the unique asset number, asset type description, deployment environment label, and business function label from the asset management platform; retrieving the public and private IP addresses currently bound to the cloud instance from the cloud platform API; and retrieving the domain name binding relationship and network endpoint information from the network configuration collection interface, summarizing them into internal asset status data; comparing the current retrieved value with the previous retrieved value for each asset's IP address, domain name binding, and network endpoint at a preset retrieval period to detect identifier change events; whenever an inconsistency is detected between the current retrieved value and the previous retrieved value for an asset, appending the identifier before the change, the identifier after the change, and the time of the change as a change record to the identifier change record of the current asset; and arranging all identifier change records for each asset in ascending order according to the time of the change to generate an asset identifier history linked list.

[0008] Preferably, the step of performing validity period range positioning in the asset identifier history chain based on the intelligence generation timestamp, extracting the set of historical asset identifiers corresponding to the identifier validity period range in which the intelligence generation timestamp falls, and generating an asset identifier snapshot at the intelligence time includes: reading the intelligence generation timestamp of each intelligence indicator in the intelligence indicator sequence, traversing the identifier change record sequence corresponding to each unique asset number in the asset identifier history chain; determining whether the intelligence generation timestamp falls within the identifier validity period range of a certain identifier change record; for identifier change records that meet the falling condition, extracting the changed identifier as the historical identifier held by the asset at the intelligence generation time, binding the historical identifier with the unique asset number to which it belongs, and writing it into the asset identifier snapshot at the intelligence time.

[0009] Preferably, the step of performing pre-trained language representation model encoding on all attack indicator fields in the intelligence indicator sequence and the asset description text in the asset identification snapshot at the intelligence moment includes: inputting the attack indicator field of each intelligence indicator in the intelligence indicator sequence into a fine-tuned pre-trained language representation model to obtain the semantic embedding vector of the attack indicator field; for each unique asset number in the asset identification snapshot at the intelligence moment, reading the asset type description, deployment environment label, and business function label corresponding to the unique asset number from the internal asset status data and concatenating them into asset description text in a fixed order; inputting the asset description text into the pre-trained language representation model to obtain the asset semantic embedding vector corresponding to the unique asset number.

[0010] Preferably, the steps of performing structured interval matching and semantic vector retrieval respectively include: constructing a hierarchical Bloom filter and performing structured interval matching; for intelligence indicators that pass the screening, performing value equality verification in the asset historical identifier auxiliary index of the asset identifier snapshot at the intelligence time to confirm the unique number of the asset to be matched, performing a binary search in the asset identifier historical chain list with the intelligence generation timestamp as the query parameter to confirm the structured valid association, and generating a structured association record; for each semantic embedding vector in the intelligence semantic embedding library, performing an approximate nearest neighbor search in the asset semantic embedding library, calculating the cosine similarity between the semantic embedding vector and the asset semantic embedding vector; retaining the unique number of the asset whose cosine similarity exceeds the semantic matching threshold, and generating a semantic association record.

[0011] Preferably, the construction of the hierarchical Bloom filter and the execution of structured interval matching includes: setting the first layer bit array of the hierarchical Bloom filter to correspond to the attack indicator type dimension, the second layer bit array to correspond to the time slice dimension, and the third layer bit array to correspond to the network identifier dimension; mapping the attack indicator type recorded in the asset identifier snapshot at the intelligence time using a hash function and setting the corresponding position of the first layer bit array to 1; mapping the concatenated string of the attack indicator type and the time slice number using a hash function and setting the corresponding position of the second layer bit array to 1; mapping the concatenated string of the attack indicator type, the time slice number, and the asset history identifier using a hash function and setting the corresponding position of the third layer bit array to 1, thereby obtaining a three-layer filtering state vector; performing bit operations to query each intelligence indicator in the intelligence indicator sequence on the three layers of the hierarchical Bloom filter in sequence, and terminating the matching if any layer returns no existence.

[0012] Preferably, the step of fusing the two matching results to generate an intelligence asset association pair list includes: merging all structured association records and semantic association records corresponding to each intelligence indicator identifier; setting both the structured hit flag and the semantic hit flag to 1 for asset unique numbers that appear simultaneously in the structured association records and the semantic association records; setting the structured hit flag to 1 and the semantic hit flag to 0 for asset unique numbers that appear only in the structured association records; setting the structured hit flag to 0 and the semantic hit flag to 1 for asset unique numbers that appear only in the semantic association records; and summarizing the intelligence indicator identifier, asset unique number, structured hit flag, semantic hit flag, hit identifier validity period interval, and cosine similarity into one association record to generate an intelligence asset association pair list.

[0013] Preferably, the step of calculating the weighted impact score of intelligence confidence and the asset exposure score based on the combined state of structured hit markers and semantic hit markers includes: when both the structured hit marker and the semantic hit marker are 1, calculating the weighted impact score of intelligence confidence based on the weighted sum of intelligence source credibility score, attack activity score, vulnerability exploitation maturity score, and semantic partial cosine similarity; when only the structured hit marker is 1, setting the cosine similarity weight to zero and calculating based on the weighted sum of the above three objective scores; when only the semantic hit marker is 1, calculating based on the cosine similarity multiplied by the corresponding weight and semantic decay coefficient; extracting the number of externally open ports, associated business level, and number of known weaknesses corresponding to each asset's unique number, and calculating the asset exposure score by a normalized weighted sum of the three.

[0014] Preferably, the step of inputting the affected asset risk assessment table into the graph neural network inference module to perform asset-related risk propagation analysis and output propagation risk enhancement score includes: constructing an asset association topology graph with network connectivity and business dependency relationships between assets as edges and the unique asset number as nodes; concatenating the comprehensive risk score, asset exposure score, structured hit marker, and semantic hit marker from the affected asset risk assessment table into a node feature vector and writing it into the corresponding node; inputting the asset association topology graph carrying the node feature vector into the graph neural network inference module to perform multi-round message passing and update the hidden layer representation of each node; and mapping the hidden layer representation of all nodes through the output head of the graph neural network inference module to obtain the propagation risk enhancement score for each unique asset number.

[0015] Preferably, the step of combining the comprehensive risk score to perform tiered risk warning and generate differentiated handling suggestions includes: calculating the final risk score by weighted summation of the comprehensive risk score and the propagation risk enhancement score, and dividing the risk into high-risk, medium-risk, and low-risk levels based on the first and second percentiles of the final risk score as dividing boundaries, wherein the sum of the first and second percentiles is 100%, and the first percentile is greater than the second percentile; for assets of high-risk and medium-risk levels, a unique identifier is assigned, and the changed identifier of the latest identifier change record is extracted from the asset identifier history chain as the current network identifier of the affected asset; based on the vulnerability exploitation maturity score of the corresponding intelligence indicators, the number of known weaknesses, and the combination state of structured hit tags and semantic hit tags, matching handling action entries are retrieved from a pre-set handling rule base, and differentiated handling suggestions including isolation operations, patch repair, and traffic filtering rules are generated in combination with the current network identifier of the affected asset.

[0016] Secondly, this invention provides an artificial intelligence-driven information network system risk early warning system, comprising: The multi-source data processing and feature encoding module is used to acquire raw external threat intelligence data and internal asset status data from multi-source data access channels. It performs indicator field parsing and timestamp extraction on the raw external threat intelligence data, and performs identifier change event tracking and historical snapshot recording on the internal asset status data, generating an intelligence indicator sequence and an asset identifier historical linked list, respectively. For each intelligence indicator in the intelligence indicator sequence, it performs validity period range positioning in the asset identifier historical linked list based on the intelligence generation timestamp, extracting the set of historical asset identifiers corresponding to the identifier validity period range falling within the intelligence generation timestamp, and generating an asset identifier snapshot at the intelligence time. It then performs pre-trained language representation model encoding on all attack indicator fields in the intelligence indicator sequence and the asset description text in the asset identifier snapshot at the intelligence time, respectively, generating an intelligence semantic embedding library and an asset semantic embedding library. The dual-path matching and fusion module is used to perform structured interval matching and semantic vector retrieval based on the intelligence indicator sequence, the intelligence time asset identifier snapshot, the intelligence semantic embedding library and the asset semantic embedding library, respectively, and fuse the two matching results to generate a list of intelligence asset association pairs; The correlation analysis and scoring calculation module is used to perform multidimensional correlation analysis on the intelligence asset correlation pair list, calculate the intelligence confidence weighted impact score and asset exposure score based on the combination state of structured hit tags and semantic hit tags, and generate an affected asset risk assessment table. The graph network propagation analysis and early warning module is used to input the affected asset risk assessment table into the graph neural network inference module to perform asset-related risk propagation analysis, output propagation risk enhancement score, combine the comprehensive risk score to perform graded risk early warning, generate differentiated disposal suggestions, and output risk early warning disposal report.

[0017] Compared to existing technologies, the advantages of this invention are as follows: This invention retains the complete change history of internal asset network identifiers within the identifier validity period in the asset identifier historical chain, and uses the intelligence generation timestamp as an anchor point to perform interval positioning to generate an asset identifier snapshot at the intelligence moment. This transforms the static matching based on the current asset identifier in existing technologies into time-series alignment matching based on historical state intervals. It fundamentally cuts off the timeliness misalignment between the dynamic changes of network identifiers and the intelligence collection moment caused by cloud instance elastic scaling, containerized deployment, and load balancing, avoiding the source of inaccuracy where the identification results of affected assets deviate from the actual risk distribution. Furthermore, it drives the three-branch calculation of the intelligence confidence weighted impact score with the fusion of structured hit tags and semantic hit tags, and performs risk propagation reasoning on the asset association topology graph through a graph neural network inference module. This enables the system to have identification capabilities in three scenarios: precise correspondence between intelligence and asset identifiers, implicit semantic association between intelligence and assets, and indirect exposure of affected assets. This achieves full coverage and high accuracy of risk warning in dynamic network environments. This invention belongs to the field of computer data processing and can be widely applied to enterprise asset security risk analysis. Attached Figure Description

[0018] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 This is a flowchart of an artificial intelligence-driven risk warning method for information network systems according to the present invention; Figure 2 This is a schematic diagram of the asset identification history linked list structure in an embodiment of the present invention; Figure 3 This is a schematic diagram illustrating the location of the snapshot interval of the asset identifier at the intelligence moment in an embodiment of the present invention; Figure 4 This is a schematic diagram of the three-layer structure of the layered Bloom filter in an embodiment of the present invention; Figure 5 This is a schematic diagram of a three-level progressive structured interval matching in an embodiment of the present invention; Figure 6This is a schematic diagram of dual-label fusion for intelligence asset association in an embodiment of the present invention; Figure 7 This is a schematic diagram illustrating the normalization of three indicators of asset exposure in an embodiment of the present invention; Figure 8 This is a schematic diagram of asset association topology graph node feature injection in an embodiment of the present invention; Figure 9 This is a schematic diagram illustrating the final risk score percentile classification in an embodiment of the present invention; Figure 10 This is a functional block diagram of an artificial intelligence-driven information network system risk early warning system according to the present invention. Detailed Implementation

[0020] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0021] Example 1: Please see Figure 1 As shown, this embodiment provides an artificial intelligence-driven risk early warning method for information network systems, including: This embodiment uses an enterprise attack surface management scenario as its application background. In this scenario, the enterprise's digital assets include elastically scalable cloud host instances, containerized deployment services, and web services carried by load balancing devices. The public IP addresses of the cloud instances are periodically reassigned by the cloud platform, container services correspond to different network endpoints at different times, and multiple physical hosts share the same external IP, causing the network identifier of internal assets to change dynamically over time. The risk warning system described in this embodiment is deployed on the enterprise security operations platform and interfaces with the threat intelligence platform, asset management platform, cloud platform, and network defense tools through multi-source data access channels, executing the following steps S1 to S4.

[0022] S1: Obtain raw external threat intelligence data and internal asset status data from the multi-source data access channel, perform indicator field parsing and timestamp extraction on the raw external threat intelligence data, and perform identifier change event tracking and historical snapshot recording on the internal asset status data, generating intelligence indicator sequences and asset identifier historical linked lists respectively.

[0023] In enterprise network environments where cloud instance elastic scaling, containerized deployment, and load balancing coexist, the network identifiers of internal assets are not static, while the attack indicators recorded by external threat intelligence are collected at a specific historical moment. If only the static identifier field at the time of intelligence collection is matched literally with the current asset identifier, a timeliness mismatch will occur due to the lack of verification correlation between the intelligence timestamp and the asset identifier's validity period. The targeted asset may not actually be the one bearing the risk. This step requires the structured preservation of both the time attributes of the intelligence indicators and the asset identifier's change history, laying the data foundation for subsequent matching based on historical state intervals instead of current state matching.

[0024] Further, step S1 includes: S11: Obtain raw external threat intelligence data from threat intelligence platform interfaces, open intelligence subscription sources, and internal security device alarm channels. Perform field parsing on the raw external threat intelligence data, extract attack indicator fields and intelligence generation timestamps, bind and store the attack indicator fields with the intelligence generation timestamps, and generate an intelligence indicator sequence.

[0025] The subsequent step S13 requires range positioning in the asset identification history chain with the intelligence generation timestamp as the anchor point. However, the original external threat intelligence data is accessed from multiple sources in the form of unstructured or semi-structured text, with attack indicators and time attributes mixed in the message fields. It does not yet have a structured form that can be aligned with the time sequence. Therefore, field parsing must be performed first to separate and bind the two types of attributes.

[0026] Specifically, raw external threat intelligence data is obtained from the threat intelligence platform interface in a structured threat information format, from open intelligence subscription sources in an Extensible Markup Language (XML) message format, and from internal security device alarm channels in a system log format.

[0027] The raw external threat intelligence data is parsed to extract the attack indicator field and intelligence generation timestamp for each piece of intelligence. The attack indicator field refers to the network identifier field used in external threat intelligence to identify the carrier of attack behavior, including four categories: IP address, domain name, file hash, and URL path. The intelligence generation timestamp refers to the Coordinated Universal Time (UTC) recorded by the external threat intelligence source when it confirms the existence of active threat behavior based on the attack indicator.

[0028] The attack indicator field is bound to the intelligence generation timestamp and stored as an intelligence indicator record. All intelligence indicator records are sorted in ascending order by intelligence generation timestamp to generate an intelligence indicator sequence. The intelligence indicator sequence is stored in an array structure. Each array element contains four components: intelligence indicator identifier, attack indicator field, attack indicator type, and intelligence generation timestamp. The intelligence indicator identifier is an integer number starting from 1 assigned by the system to each intelligence indicator according to the access order. The attack indicator type can be one of four: IP address, domain name, file hash, or URL path.

[0029] S12: Continuously pull internal asset status data from the asset management platform, cloud platform API and network configuration collection interface, perform identifier change event detection on the internal asset status data, append the change record to the identifier change record of the corresponding asset and arrange them in chronological order to generate an asset identifier history linked list.

[0030] Cloud instance elastic scaling and containerized deployment mean that the same asset corresponds to different network endpoints at different times, and the network identifier of the asset exhibits a discrete and volatile characteristic over time. This characteristic dictates that simply recording the current identifier of the asset is insufficient to support real-time intelligence matching; the entire historical state of the asset identifier must be completely preserved in a temporal structure in order to retrospectively query the identifier held by the asset at the time the intelligence was generated.

[0031] Specifically, the system retrieves unique asset identifiers, asset type descriptions, deployment environment tags, and business function annotations from the asset management platform; it retrieves the currently bound public and private IP addresses of the cloud instance from the cloud platform API; and it retrieves domain name binding relationships and network endpoint information from the network configuration collection interface, summarizing these into internal asset status data. The unique asset identifier is a globally unique and lifelong identifier assigned to each asset by the asset management platform when it is incorporated into the asset management system. It is independent of the asset's network identifier and does not change with changes to the network identifier.

[0032] The internal asset status data is subjected to identifier change event detection. The detection method is as follows: the IP address, domain name binding and network endpoint of each asset are compared with the previous retrieval value item by item at a preset retrieval period. The preset retrieval period is determined according to the minimum time interval for the cloud platform IP address reallocation. One-tenth of the minimum time interval is taken as the retrieval period. For example, in the scenario where the cloud instance reallocates the IP address every 24 hours, the retrieval period can be set to 2.4 hours.

[0033] Whenever an inconsistency is detected between the current fetch value and the previous fetch value for an asset's IP address, domain name binding, or network endpoint, the identifier before the change, the identifier after the change, and the time of the change are appended as a change record to the current asset's identifier change record. The time of the change is the fetch time at which the inconsistency was detected.

[0034] For each asset, all identification change records are sorted in ascending order by the time of the change, generating an asset identification history linked list. This asset identification history linked list is a composite data structure of hash table and linked list, using the asset's unique ID as the key and the sequence of identification change records arranged chronologically as the value. Each identification change record includes the original identification, the new identification, and the time of the change. In addition, each record contains an identification validity period range formed by the time of the previous change and the time of the next change record. The right endpoint of the validity period range for the last change record is taken as the current system time.

[0035] The asset identifier history chain retains the complete historical state of asset identifiers in a time-series chain structure, enabling asset identifiers at any historical moment to be accurately retrieved through their validity period. This fundamentally eliminates the timeliness misalignment introduced by replacing historical identifiers with current ones, providing a traceable identifier validity period boundary for the intelligence time interval positioning in S13. See also Figure 2 This is a schematic diagram of the asset identification history linked list structure provided in the embodiments of this application. For example... Figure 2 As shown, the diagram presents a composite data structure of hash tables and linked lists. On the left, the asset's unique identifier serves as the key, with lines connecting to a chronologically ordered sequence of identifier change records. Each node contains three components: the identifier before the change, the identifier after the change, and the time the change occurred. Nodes are connected by arrows indicating ascending order of change time. The change times of adjacent nodes form an identifier validity period interval, with the right endpoint of the validity period interval of the last change record pointing to the current system time via a dashed arrow. This structure, using a chronological chain structure, preserves the entire historical state of asset identifiers, allowing asset identifiers at any historical moment to be accurately retrieved through their validity period intervals. This fundamentally eliminates the timeliness misalignment introduced by replacing historical identifiers with current ones, providing a traceable identifier validity period boundary for subsequent intelligence time interval positioning.

[0036] S13: For each intelligence indicator in the intelligence indicator sequence, perform validity period range positioning in the asset identifier history chain based on the intelligence generation timestamp, extract the set of historical asset identifiers corresponding to the identifier validity period range that the intelligence generation timestamp falls into, and generate an asset identifier snapshot at the intelligence moment.

[0037] If this step is skipped and the intelligence indicator sequence is directly matched with the asset's current identifier, then when the intelligence generation timestamp is earlier than the asset's most recent identifier change time, the asset's current identifier is no longer the identifier at the time of intelligence collection. The matching result will point to an irrelevant asset whose identifier was reassigned after intelligence collection, resulting in the omission of truly affected assets. This step uses the intelligence generation timestamp as an anchor point to backtrack and locate within the identifier's validity period, transforming the matching from current state matching to historical state range matching.

[0038] Specifically, for each intelligence indicator in the intelligence indicator sequence, the intelligence generation timestamp of the intelligence indicator is read, the sequence of identifier change records corresponding to the unique number of each asset in the asset identifier history chain is traversed, and it is determined whether the intelligence generation timestamp falls within the identifier validity period range of a certain identifier change record, that is, the intelligence generation timestamp is greater than or equal to the left endpoint of the identifier validity period range and less than the right endpoint of the identifier validity period range.

[0039] For each identifier change record that meets the inclusion criteria, the changed identifier of the record is extracted as the historical identifier held by the asset at the time of intelligence generation. This historical identifier is then bound to the unique identifier of the asset and written into the asset identifier snapshot at the time of intelligence generation. See also... Figure 3 This is a schematic diagram illustrating the location of the snapshot range of the intelligence moment asset identifier provided in the embodiments of this application. For example... Figure 3 As shown, this graph uses a timeline as its main element, marking several validity periods of identifiers and their corresponding modified identifiers along the axis. The intelligence generation timestamp is indicated by a downward arrow falling into a specific identifier validity period, with the moment of falling into the interval marked by a solid dot. Modified identifiers falling into the interval are enclosed in dashed extraction boxes, and a broken line points to the asset historical identifiers in the asset identifier snapshot table at the intelligence time on the right. This graph uses the intelligence generation timestamp as an anchor point to trace back and locate within the identifier validity period, transforming the matching from current state matching to historical state interval matching. In scenarios where the public IP address of a cloud host is reassigned after intelligence collection, the intelligence indicator can still locate the unique asset number corresponding to the original Web service holding that IP address at the intelligence generation time, rather than the test host currently holding that IP address. This avoids the source of inaccuracies that cause deviations between the identified affected assets and the actual risk distribution at the data foundation level.

[0040] The snapshot of the asset identifier at the intelligence moment is stored in a table structure. Each record contains five components: intelligence indicator identifier, attack indicator type, time segment number to which the intelligence generation timestamp belongs, asset history identifier, and asset unique number. It records all internal asset identifiers and their associated asset unique numbers corresponding to each intelligence indicator at the intelligence generation moment. The time segment number will be defined and used uniformly in S211. An auxiliary index will be established using the asset history identifier as the key for subsequent retrieval of asset unique numbers by network identifier value.

[0041] The intelligence-time asset identifier snapshot uses the intelligence generation timestamp as an anchor to perform range positioning in the asset identifier historical chain. In the example scenario where the public IP address of the cloud host is reassigned to the test host after intelligence collection, the intelligence indicator can still locate the unique asset number corresponding to the original Web service that held the IP address at the time of intelligence generation, rather than the test host that currently holds the IP address. This avoids the source of inaccuracy that causes the identification results of affected assets to deviate from the actual risk distribution at the data foundation level.

[0042] S14: Perform pre-trained language representation model encoding on all attack indicator fields in the intelligence indicator sequence and the asset description text in the asset identifier snapshot at the intelligence moment, respectively, to generate an intelligence semantic embedding library and an asset semantic embedding library.

[0043] When external attackers describe their targets in threat intelligence, they don't always directly list network identifiers such as IP addresses. Instead, they may describe the targets using business semantics, such as referring to the attacked object by business function or service type. When the attack indicator field does not directly contain the internal asset network identifier, a pure field matching mechanism cannot establish a correspondence between intelligence and assets. This step maps the attack indicator field and asset description text to the same semantic vector space, enabling subsequent matching to capture the implicit correspondence between the attack target description and the asset's business function at the semantic level. This forms two independent matching bases in parallel with the structured matching links from S11 to S13.

[0044] Further, step S14 includes: S141: For each intelligence indicator in the intelligence indicator sequence, input the attack indicator field of the intelligence indicator into the pre-trained language representation model to obtain the semantic embedding vector of the attack indicator field.

[0045] Specifically, the pre-trained language representation model is based on the open-source Chinese sentence embedding model Sentence-BERT. The base network has been initially trained on a large-scale general Chinese corpus and has the ability to encode any text into a general text representation with a fixed-length semantic vector. The output vector has a dimension of 768.

[0046] Because the general Chinese corpus lacks samples of attack indicators and asset business descriptions in the cybersecurity field, the basic network's semantic distance between attack target descriptions and asset business functions in cybersecurity texts is inaccurate. Therefore, fine-tuning training was performed on the basic network specifically for cybersecurity text scenarios. The fine-tuning training data consisted of cybersecurity text pairs collected from historical archived intelligence from the threat intelligence platform and historical asset descriptions from the asset management platform. Each pair included one attack indicator field text and one confirmed associated asset description text as a positive sample, and randomly paired unassociated text pairs as negative samples. The data was labeled by qualified cybersecurity operators, and the dataset size was no less than 50,000 valid samples. The fine-tuning training froze the parameters of the first 6 encoding layers of the basic network, updating only the parameters of the last 6 encoding and pooling layers. The loss function used was the contrastive learning triplet loss function, and the optimizer was AdamW. The initial learning rate was set to 0.00002 and adjusted using a linear decay strategy. The training epochs were set to 5, and the batch size was set to 64. After fine-tuning, the pre-trained language representation model improves the semantic similarity accuracy of related text pairs in the cybersecurity field by no less than 15 percentage points compared to directly using the basic network.

[0047] The attack indicator field of each intelligence indicator is input into the pre-trained language representation model after fine-tuning, and the 768-dimensional vector output by the pooling layer is taken as the semantic embedding vector of the attack indicator field.

[0048] S142: For each unique asset number in the intelligence moment asset identification snapshot, read the asset type description, deployment environment label and business function annotation from the internal asset status data and concatenate them into an asset description text. Input the text into the pre-trained language representation model to obtain the asset semantic embedding vector corresponding to the unique asset number.

[0049] Specifically, for each unique asset number appearing in the asset identification snapshot at the intelligence moment, the asset type description, deployment environment label, and business function label corresponding to the unique asset number are read from the internal asset status data obtained in S12, and concatenated with separators in a fixed order of asset type description, deployment environment label, and business function label to form asset description text.

[0050] The asset description text is input into the pre-trained language representation model after fine-tuning in S141, and the 768-dimensional vector output by the pooling layer is taken as the asset semantic embedding vector corresponding to the unique asset number.

[0051] S143: Summarize all semantic embedding vectors obtained in S141 to generate an intelligence semantic embedding library, and summarize all asset semantic embedding vectors obtained in S142 to generate an asset semantic embedding library.

[0052] Specifically, the semantic embedding vector of each attack indicator field output by S141 is bound to the corresponding intelligence indicator identifier, and summarized into an intelligence semantic embedding library; the asset semantic embedding vector corresponding to each unique asset number output by S142 is bound to the unique asset number, and summarized into an asset semantic embedding library. Both the intelligence semantic embedding library and the asset semantic embedding library are stored in a vector index structure for S22 to perform approximate nearest neighbor retrieval.

[0053] The generation of the intelligence semantic embedding library and the asset semantic embedding library uses the business function annotations of the asset description text as an important input to the pre-trained language representation model, so that the attack target semantics expressed by the attack indicator field and the asset business function semantics are mapped to the same 768-dimensional vector space. In scenarios where external attackers use business semantics to describe attack targets instead of directly listing IP addresses, the structured matching link cannot hit any assets because the attack indicator field does not contain network identifiers. At this time, the semantic embedding vectors in the intelligence semantic embedding library and the asset semantic embedding vectors in the asset semantic embedding library that are highly consistent with the business functions are close in distance in the vector space, so that the system can still identify potentially affected assets. This semantic recognition capability is not the original design goal of semantic encoding. The original design goal of the pre-trained language representation model is only to convert text into comparable fixed-length vectors. However, the mathematical structure of mapping to the same vector space makes the intelligence and assets with consistent business semantics naturally cluster in the space. Without the need to introduce an independent business semantic parsing module, the ability to overcome the hidden defect of attackers using business semantics to covertly describe targets naturally emerges, making up for the inherent limitation of the pure field matching mechanism that produces systematic missed matches when the attack indicator field lacks network identifiers.

[0054] S2: Based on the intelligence indicator sequence, the intelligence moment asset identifier snapshot, the intelligence semantic embedding library, and the asset semantic embedding library, perform structured interval matching and semantic vector retrieval respectively, and merge the two matching results to generate a list of intelligence asset association pairs.

[0055] The snapshot of the asset identifier at the moment of intelligence output by S1 establishes a foundation for accurate time-series structured matching, while the intelligence semantic embedding library and asset semantic embedding library output by S1 establish a foundation for semantic-level matching. The outputs of these two basic links must be consumed through structured interval matching and semantic vector retrieval, respectively, and the two results must be fused in a source-distinguishing format to ensure that the final association result simultaneously covers both scenarios of precise identifier matching and implicit semantic association. When intelligence is accessed in batches, performing precise interval matching on each piece of intelligence may result in processing latency exceeding the asset identifier change cycle, thus introducing time-series misalignment again. Therefore, a low-overhead filtering mechanism must first be used to narrow down the candidate set.

[0056] Further, step S2 includes: S21: Construct a hierarchical Bloom filter and perform structured interval matching. Perform precise interval matching on the intelligence indicators that pass the screening in the asset identification history chain to generate structured association records.

[0057] Further, step S21 includes: S211: Construct a hierarchical Bloom filter. Write each record in the intelligence moment asset identifier snapshot into the corresponding layer of the hierarchical Bloom filter according to the attack indicator type, time fragment number and asset history identifier to obtain a three-layer filter state vector.

[0058] The subsequent step S213 performs precise interval matching by performing a binary search on the asset identifier history chain. The cost of a single search is proportional to the logarithm of the chain length. If a binary search is performed on each intelligence indicator during batch intelligence access, the cumulative cost will increase linearly with the number of intelligence indicators. This step uses a hierarchical Bloom filter to preemptively eliminate intelligence indicators that are unlikely to be matched, thus shrinking the size of the candidate set entering the binary search.

[0059] Specifically, the hierarchical Bloom filter comprises three layers, each being an independent bit array configured with a set of hash functions. The first layer corresponds to the attack indicator type dimension, with the attack indicator type as the input element; the second layer corresponds to the time shard dimension, with the string formed by concatenating the attack indicator type and the time shard number as the input element; and the third layer corresponds to the network identifier dimension, with the string formed by concatenating the attack indicator type, the time shard number, and the asset history identifier as the input element.

[0060] The time shard number is defined as follows: the time axis is divided equally by the duration of a single time shard. The duration of a single time shard is determined based on the minimum change cycle of the internal asset network identifier, taking half of the minimum change cycle. For example, in a scenario where the IP address of a cloud instance is reassigned every 24 hours, the duration of a single time shard can be set to 12 hours. Starting from the system deployment time, each time shard is assigned an integer number that increments from 0. The integer number corresponding to the time shard to which the intelligence generation timestamp falls is the time shard number.

[0061] The length of each bit array and the number of hash functions are determined according to the standard Bloom filter method, with the bit array length taken as... The number of hash functions is taken as follows ,in The expected number of elements to be written. Let be the expected false positive rate. For example, take . , At that time, the length of the bit array is approximately The number of bits and hash functions is approximately 7.

[0062] The attack indicator type, the concatenated string of attack indicator type and time segment number, and the concatenated string of attack indicator type, time segment number, and asset history identifier for each record in the intelligence moment asset identifier snapshot are mapped by each layer of hash functions. The corresponding bits in each bit array are then set to 1. After writing, a three-layer filtering state vector is obtained. The three-layer filtering state vector refers to the current bit state of the three-layer bit array. See also... Figure 4 This is a schematic diagram of the three-layer structure of the layered Bloom filter provided in an embodiment of this application. Figure 4 As shown, the diagram is divided into three layers from top to bottom. The first layer corresponds to the attack indicator type dimension, the second layer to the time slice number dimension, and the third layer to the asset history identifier dimension. The left side of each layer contains the input elements, which are pointed to by arrows along hash circles. The hash circles are then connected by lines to set the corresponding positions of the bit arrays to 1. The bit arrays are represented by a grid sequence, with black squares indicating the set state. The right side uses curly braces to uniformly label all three layers as bit arrays. This diagram writes each record in the asset identifier snapshot at the intelligence moment into the corresponding layer according to the attack indicator type, time slice number, and asset history identifier, resulting in a three-layer filtering state vector. A hierarchical Bloom filter is used to preemptively eliminate intelligence indicators that are unlikely to be matched, thus shrinking the size of the candidate set entering the binary search.

[0063] S212: For each intelligence indicator in the intelligence indicator sequence, perform bitwise operations to query the three layers of the hierarchical Bloom filter in turn. If any layer returns that it does not exist, terminate the matching. If all three layers return that it exists, proceed to S213.

[0064] Specifically, for each intelligence indicator in the intelligence indicator sequence, the attack indicator type, the time segment number to which the intelligence generation timestamp belongs, and the attack indicator field are extracted.

[0065] The layered Bloom filter is queried sequentially: the first layer queries whether all bits of the attack indicator type are 1 after being mapped by the hash function; the second layer queries whether all bits of the attack indicator type and time shard number are 1 after being concatenated into a string; and the third layer queries whether all bits of the attack indicator type, time shard number and attack indicator field are 1 after being concatenated into a string.

[0066] If any layer has a mapping bit of 0, it is determined that the intelligence indicator cannot be matched on the structured link, the matching of the intelligence indicator is terminated and skipped; if all three layers of mapping bits are 1, the intelligence indicator is transferred to S213 to perform precise matching.

[0067] S213: For intelligence indicators screened by the hierarchical Bloom filter, perform value equality verification in the asset history identifier auxiliary index of the asset identifier snapshot at the intelligence time to confirm the unique number of the asset to be matched, and then perform a binary search in the asset identifier history chain with the intelligence generation timestamp as the query parameter to confirm the structured valid association and generate a structured association record.

[0068] The Bloom filter has the potential for false positives. All three layers returned insufficient data to confirm that the attack indicator field value is truly equal to the asset's historical identifier and that the intelligence generation timestamp truly falls within the asset identifier's validity period. Therefore, value equality verification and precise timestamp verification must be performed sequentially.

[0069] Specifically, for each intelligence indicator screened by the hierarchical Bloom filter, a record that is identical to the attack indicator field of the intelligence indicator is retrieved from the asset history identifier auxiliary index of the asset identifier snapshot at the intelligence time. The unique asset number bound to the retrieved record is used as the unique asset number to be matched. If no record meets the equality condition, it is determined to be a Bloom filter misjudgment and discarded.

[0070] For intelligence indicators that confirm the unique number of the asset to be matched, read the sequence of identifier change records corresponding to the unique number of the asset to be matched in the asset identifier history chain. Since the sequence of identifier change records is arranged in ascending order according to the time of change, perform a binary search on the identifier validity period range of each identifier change record with the intelligence generation timestamp as the query parameter to determine whether the intelligence generation timestamp falls within the identifier validity period range of a certain identifier change record.

[0071] If the intelligence generation timestamp falls within the validity period range of an identifier change record, it is confirmed that the intelligence indicator and the unique identifier of the asset to be matched have a structured and valid association. The intelligence indicator identifier, the unique identifier of the asset, the matched identifier validity period range, and the intelligence generation timestamp are packaged into a structured association record. If it does not fall within any identifier validity period range, it is determined to be a Bloom filter misjudgment and discarded. The above operation is performed on all intelligence indicators that pass the screening, and a set of structured association records is generated.

[0072] S213's precise matching employs a three-tiered progressive mechanism: coarse screening (Bloom filter value matching) → exact value equality verification → timestamp binary search. This effectively reduces the size of the candidate set entering the high-overhead precise verification stage, ensuring stable overall matching throughput during batch intelligence access. (See also...) Figure 5 This is a schematic diagram of a three-level progressive structured interval matching provided in an embodiment of this application. For example... Figure 5As shown, the diagram uses three progressively smaller trapezoidal funnel blocks arranged from left to right, representing a hierarchical Bloom filter, value equality verification, and binary search, respectively. Adjacent funnel blocks are connected by arrows to indicate the progressive shrinking of the candidate set. The rightmost side displays structured association records with a rectangle indicating their validity period. Below each funnel block, a downward dashed arrow points to a misjudgment / discard rectangle, indicating that intelligence indicators that failed the verification were judged as misjudged by the Bloom filter and discarded. This diagram forms a three-tiered progressive mechanism: coarse screening, precise value equality verification, and timestamp binary search. This effectively shrinks the size of the candidate set entering the high-overhead precise verification stage, maintaining a stable overall matching throughput during batch intelligence access.

[0073] S22: For each semantic embedding vector in the intelligence semantic embedding library, perform an approximate nearest neighbor search in the asset semantic embedding library, retain the unique asset ID with a cosine similarity exceeding the semantic matching threshold, and generate a semantic association record.

[0074] The structured association records generated in S213 only cover cases where the attack indicator field and the asset network identifier correspond precisely within a historical period. They cannot establish associations when the attack indicator field describes the target using business semantics. This step utilizes the distance relationships in the vector space between the semantic embedding vectors generated in S14 to supplement the semantic-level associations.

[0075] Specifically, for each semantic embedding vector in the intelligence semantic embedding library, an approximate nearest neighbor retrieval is performed in the asset semantic embedding library. The approximate nearest neighbor retrieval uses a hierarchical navigable small world graph index as the retrieval structure, uses the semantic embedding vector in the intelligence semantic embedding library as the query vector, and returns several asset semantic embedding vectors in the asset semantic embedding library that are closest to the query vector and their corresponding unique asset numbers.

[0076] For each returned asset semantic embedding vector, the cosine similarity between the semantic embedding vector and the asset semantic embedding vector is calculated by dividing the inner product of the two vectors by the product of their magnitudes. ,in For semantic embedding vectors, This is the semantic embedding vector for the asset.

[0077] The unique asset IDs with cosine similarity exceeding the semantic matching threshold are retained. The semantic matching threshold is determined based on the cosine similarity statistics of historically confirmed associated text pairs. The threshold is calculated by subtracting one standard deviation from the mean of the cosine similarity of historically confirmed associated text pairs; for example, the semantic matching threshold can be set to 0.75. The intelligence indicator identifier, the unique asset ID, and the cosine similarity are packaged into a single semantic association record, and a set of semantic association records is generated.

[0078] S23: For each intelligence indicator identifier, merge all its corresponding structured association records and semantic association records, set the structured hit flag and semantic hit flag according to the hit source, and generate an intelligence asset association pair list.

[0079] S213 and S22 produce association results from the structured link and semantic link, respectively. The correspondence between the same intelligence indicator identifier and the unique asset number may appear simultaneously, individually, or without overlap in the two results. The intelligence confidence weighted impact score in subsequent S3 must select different calculation branches according to the hit source. Therefore, this step must record the hit status of each association pair in a double-label form that can distinguish the source.

[0080] Specifically, for each intelligence indicator identifier, all structured association records related to the intelligence indicator identifier in the structured association record set output by S213 and all semantic association records related to the intelligence indicator identifier in the semantic association record set output by S22 are merged.

[0081] For unique asset IDs that appear in both structured and semantic association records, both the structured hit flag and the semantic hit flag are set to 1; for unique asset IDs that appear only in structured association records, the structured hit flag is set to 1 and the semantic hit flag is set to 0; for unique asset IDs that appear only in semantic association records, the structured hit flag is set to 0 and the semantic hit flag is set to 1.

[0082] The intelligence indicator identifier, asset unique number, structured hit tag, semantic hit tag, hit tag validity period, and cosine similarity are aggregated into a single associated record. For associated records with only structured hits, the cosine similarity is set to null; for associated records with only semantic hits, the hit tag validity period is set to null. This process generates a list of intelligence asset associated pairs. See also... Figure 6 This is a schematic diagram of dual-label fusion of intelligence asset association provided in an embodiment of this application. For example... Figure 6As shown, the left side of the diagram contains two rectangles: a structured association record and a semantic association record. Arrows point to the intelligence asset association pair list table on the right. The table column names include the asset's unique ID, a structured hit marker, and a semantic hit marker, with the hit marker represented by numbers in small squares. The first row has both a structured hit marker and a semantic hit marker of 1, indicating that the pair appears in both structured and semantic association records. The second row has a structured hit marker of 1 and a semantic hit marker of 0, indicating that the pair appears only in structured association records. The third row has a structured hit marker of 0 and a semantic hit marker of 1, indicating that the pair appears only in semantic association records. This diagram records the hit status of each association pair using a dual-marker format that distinguishes the source, ensuring that the final association result covers both precise identification hits and implicit semantic associations, and providing a basis for selecting different calculation branches based on the hit source.

[0083] The structured interval matching links S211 to S213 and the semantic vector retrieval link S22 are executed in parallel and then fused in S23 using a dual-label approach. The three-dimensional filtering parameters of the hierarchical Bloom filter and the timestamp query for precise interval matching in S213 form parameter coupling within the structured link. Similarly, the cosine similarity of the semantic embedding vector and the semantic matching threshold in S22 form parameter coupling within the semantic link. The two links target non-overlapping association scenarios: the structured link only hits when the attack indicator field precisely matches the asset's historical network identifier, while the semantic link only hits when the attack target's semantics match the asset's business function. If only the structured link is retained without the semantic link, intelligence describing the target using business semantics will be completely unassociated with any asset because the attack indicator field does not contain a network identifier. Conversely, if only the semantic link is retained without the structured link, intelligence explicitly including the asset's historical IP address in the attack indicator field will be missed because the semantic vector may not fall within the nearest neighbor range, and semantic matching cannot provide temporal evidence of the validity period of the matched identifier. After the two links are fused with a combination of structured hit tags and semantic hit tags, the combined state further drives the three-branch calculation of the weighted influence score of S31 intelligence confidence, making the fusion result conditionally dependent across steps in the entire link. At the same time, the hierarchical Bloom filter's early shrinkage of the candidate set and the semantic matching threshold's truncation of the nearest neighbor results together control the candidate size entering exact matching and near nearest neighbor retrieval within a range that matches the asset identification change cycle, so that the overall matching throughput remains stable when intelligence is accessed in batches, avoiding the hidden defect of full-volume, one-by-one exact matching causing processing delays exceeding the asset identification change cycle and thus causing time sequence misalignment.

[0084] S3: Perform multidimensional association analysis on the list of intelligence asset associations, calculate the intelligence confidence weighted impact score and asset exposure score based on the combined state of structured hit markers and semantic hit markers, and generate a risk assessment table for affected assets.

[0085] The list of intelligence asset association pairs output by S23 records the source of each association pair with double tags, but the risk level of the association pairs has not yet been quantified. The risk reliability of the three types of association pairs—structured hits only, semantic hits only, and double hits—is different. Pure semantic hits may introduce false positives, so intelligence confidence must be quantified with differentiated calculation branches, and a comprehensive risk score must be calculated in combination with the asset's own exposure level.

[0086] Further, step S3 includes: S31: For each associated record in the list of intelligence asset association pairs, read the combined state of structured hit markers and semantic hit markers, and calculate the weighted impact score of intelligence confidence based on the combined state in three branches.

[0087] Specifically, for each associated record in the intelligence asset association list, the structured hit marker and semantic hit marker of the associated record are read. From the raw external threat intelligence data obtained in S11, the intelligence source credibility score, attack activity score, and vulnerability exploitation maturity score of the corresponding intelligence indicators of the associated record are read. All three scores are quantitative fields with values ​​ranging from 0 to 1 provided by the external threat intelligence source along with the intelligence.

[0088] When both the structured hit marker and the semantic hit marker are 1, the weighted impact score of intelligence confidence is calculated using the following formula: When only the structured hit marker is 1, the cosine similarity weight term is set to zero and calculated using the following formula: When only the semantic hit tag is 1, the cosine similarity is used as the main weight and multiplied by the semantic decay coefficient, and calculated according to the following formula: in Rate the credibility of intelligence sources. Rate the attack activity. Score the maturity of vulnerability exploitation. The cosine similarity output by S22. to For each weight, This represents the semantic attenuation coefficient. To ensure that objective intelligence evidence dominates while also considering the supplementary role of semantic matching, the weight... to The preferred values ​​are set to 0.35, 0.30, 0.20, and 0.15, respectively. Alternatively, the analytic hierarchy process (AHP) can be used to determine these values. A judgment matrix is ​​constructed by comparing four factors: intelligence source credibility, attack activity, vulnerability exploitation maturity, and semantic similarity, and then normalized. (Semantic attenuation coefficient) The value ranges from 0 to 1, and is determined based on the false alarm rate of historical pure semantic hit association pairs. The semantic decay coefficient is obtained by subtracting the false alarm rate of historical pure semantic hits from 1. For example, the semantic decay coefficient can be set to 0.6.

[0089] The design principle for calculating the weighted impact score of intelligence confidence is as follows: In the dual-hit branch, the objective scores from intelligence sources—intelligence source credibility, attack activity, and vulnerability exploitation maturity—are additively combined with semantic similarity to ensure that credible information from any dimension independently contributes to the intelligence confidence score. This additive combination guarantees that the remaining dimensions can still support the score even if a single dimension is missing. In the structured-only hit branch, since there is no semantic matching, the cosine similarity weight is set to zero to avoid introducing false similarity contributions when there is no semantic evidence. In the semantic-only hit branch, due to the lack of precise temporal correspondence evidence between intelligence indicators and asset historical identifiers, the correlation reliability is lower than that of structured hits. Therefore, a semantic decay coefficient is used to reduce the overall score, making the correlation pairs of pure semantic hits numerically lower than those of correlation pairs containing structured evidence. The upper limits of the scores for the three branches are as follows: , and The order of decreasing weights corresponds strictly and monotonically to the order of evidence reliability for double hits, structured hits only, and semantic hits only. This ensures that semantic corroboration receives a score bonus as additional evidence rather than being allocated to irrelevant items. The structured hits only branch does not perform normalization on the remaining weights, preserving... The corresponding semantic corroboration contributes to the score space only in the case of double hits, ensuring that the score ceiling for double hits is strictly higher than that for structured hits alone, which conforms to the engineering constraint that the more sufficient the evidence, the higher the confidence ceiling. The overall design goal of the three branches is to make the weighted influence score of intelligence confidence increase with the reliability of the hit source, while preserving the ability to discover semantic associations and suppressing the interference of false alarms introduced by pure semantic matching on the high-risk warning level.

[0090] For example, take , , , A certain related record has a double hit and , , , The weighted impact score of intelligence confidence is then... .

[0091] S32: Extract the number of externally accessible ports, associated business level, and number of known vulnerabilities corresponding to each asset's unique number from the internal asset status data, and calculate the asset exposure score.

[0092] Measuring risk solely based on intelligence confidence level is insufficient to reflect the attack surface of the attacked asset itself. Assets with more externally exposed ports, higher-level business operations, and more known weaknesses will suffer more severe consequences if attacked. This step quantifies the degree of exposure from the asset side, which, together with the confidence level from the intelligence side, constitutes two factors in the comprehensive risk score.

[0093] Specifically, the number of external open ports, associated business level, and number of known weaknesses of each asset corresponding to its unique number are read from the internal asset status data obtained from S12. The associated business level is an integer level from 1 to 5, which is assigned by the asset management platform according to the importance of the business carried by the asset.

[0094] Calculate the asset exposure score using the following formula: in The number of externally open ports is linearly normalized to the interval 0 to 1, with the denominator being the maximum number of externally open ports of all assets. The value obtained by dividing the associated business level by 5; The number of known weaknesses is linearly normalized to the interval 0 to 1, with the denominator being the maximum number of known weaknesses for all assets. , , As a weighted selection, considering that the impact of the business importance carried by the asset and the vulnerability of the system on the consequences of exposure is usually greater than the number of ports opened at the pure network layer, a weight is set. =0.2、 =0.4、 =0.4. The design principle of the asset exposure score is as follows: the three types of asset-side exposure indicators are normalized to have unified dimensions, and are combined additively so that an increase in any exposure dimension independently raises the asset exposure score, consistent with the physical law that the larger the attack surface, the more severe the exposure consequences. See also Figure 7 This is a normalized diagram of the three indicators of asset exposure provided in the embodiments of this application. For example... Figure 7 As shown in the figure, the left side of the graph contains three rectangles representing the number of open ports, the level of associated services, and the number of known vulnerabilities. Each rectangle is mapped to the interval 0 to 1 by a linearly normalized line segment, with 0 and 1 marked at the ends of the normalized line segment, respectively. The three normalized results are then converged and connected, with an arrow pointing to the asset exposure score rectangle on the right. This graph unifies the dimensions of the three types of asset-side exposure indicators after normalization, and uses additive combination so that an increase in any exposure dimension independently raises the asset exposure score, consistent with the physical law that the larger the attack surface, the more severe the exposure consequences.

[0095] S33: The comprehensive risk score is obtained by multiplying the weighted impact score of intelligence confidence level and the asset exposure score. All related records are sorted in descending order of comprehensive risk score to generate an affected asset risk assessment table.

[0096] Specifically, for each associated record in the list of intelligence asset associations, a comprehensive risk score is calculated using the following formula: in The weighted impact score is added to the confidence level of the intelligence output by S31. This is the asset exposure score output by S32. The comprehensive risk score is in a product form, such that when either the intelligence confidence level or the asset exposure level approaches 0, the comprehensive risk score also approaches 0, which conforms to the engineering constraint that risk does not exist when there is no credible threat or no exposure surface.

[0097] All associated records in the intelligence asset association list are sorted in descending order of comprehensive risk score. The unique asset number, intelligence indicator identifier, structured hit mark, semantic hit mark, intelligence confidence weighted impact score, asset exposure score, comprehensive risk score, and hit mark validity period are merged and output to generate an affected asset risk assessment table.

[0098] S4: Input the risk assessment form of the affected assets into the graph neural network inference module to perform asset-related risk propagation analysis, output the propagation risk enhancement score, combine the comprehensive risk score to perform graded risk warning, generate differentiated disposal suggestions, and output risk warning disposal report.

[0099] The affected asset risk assessment table only depicts the risks of assets directly identified by intelligence, without considering the risk propagation caused by network connectivity and business dependencies between assets. Assets that are not directly identified by any intelligence but are connected to or dependent on multiple high-risk assets are actually in a state of indirect exposure. This step uses a graph neural network to perform risk propagation reasoning on the asset association topology, identifies indirectly exposed assets, and drives the generation of tiered early warning and disposal recommendations by fusing propagation risks with direct risks.

[0100] Further, step S4 includes: S41: Construct an asset association topology graph and inject node feature vectors. Input the graph neural network inference module to perform multiple rounds of message passing and output the propagation risk enhancement score for each asset's unique number.

[0101] Further, step S41 includes: S411: Construct an asset association topology graph with the network connectivity and business dependencies between assets as edges and the unique asset number as nodes, and concatenate the scores and labels in the affected asset risk assessment table into a node feature vector and write it into the corresponding node.

[0102] Specifically, the network connectivity and business dependencies between assets are read from the internal asset status data obtained in S12. Using the unique asset ID as a node, and the connection between two unique asset IDs that have a network connectivity or business dependency relationship as an edge, an asset association topology graph is constructed. The edge weight is the weighted sum of the network connectivity and business dependencies; it is 0.5 when a network connectivity relationship exists, 0.5 when a business dependency relationship exists, and 1 when both exist simultaneously.

[0103] For each node corresponding to a unique asset ID in the asset association topology graph, the comprehensive risk score, asset exposure score, structured hit marker, and semantic hit marker of the unique asset ID are read from the affected asset risk assessment table output by S33, concatenated into a 4-dimensional node feature vector, and written to the corresponding node; for nodes corresponding to unique asset IDs that do not appear in the affected asset risk assessment table, all four components of the node feature vector are set to 0. See also Figure 8 This is a schematic diagram of node feature injection in the asset association topology graph provided in the embodiments of this application. For example... Figure 8 As shown, this graph uses unique asset identifiers as nodes and connections between two unique asset identifiers with network connectivity or business dependencies as edges to construct an asset association topology. The edges are labeled with network connectivity, business dependencies, and edge weights. On the left, four-cell node feature vector boxes represent four components: comprehensive risk score, asset exposure score, structured hit label, and semantic hit label. Dashed lines connect these components to nodes in the topology, indicating that the node feature vector is written to the corresponding node. This graph concatenates the scores and labels from the affected asset risk assessment table into a node feature vector, which is then written to the corresponding node. This ensures that the risk source type of neighboring nodes is carried and transmitted during message transmission, providing a basis for parameter coupling between the enhanced score for the propagation risk of indirectly exposed assets and the reliability of the risk source.

[0104] S412: Input the asset association topology graph carrying node feature vectors into the graph neural network inference module, perform multiple rounds of message passing, and update the hidden layer representation of each node.

[0105] Specifically, the graph neural network inference module is a graph convolutional neural network, containing three graph convolutional layers. The initial hidden layer representation of each node is taken from the 4-dimensional node feature vector written in S411. In each round of message passing, each node aggregates the sum of the products of the hidden layer representations of its neighboring nodes and the corresponding edge weights, adds this sum to its own hidden layer representation, multiplies it by the weight matrix of the current graph convolutional layer, adds the bias vector, and then processes it through a non-linear activation function to obtain the updated hidden layer representation, executed according to the following formula: in For nodes In the Hidden layer representation of a layer, For nodes The set of neighboring nodes, For nodes With nodes Edge weights between them For the first Layer weight matrix, For the first Layer bias vector, This is a linear rectified activation function. The weight matrix... With bias vector All parameters are learnable parameters obtained from training data. The hidden layer representations of the first and second layers have a dimension of 16, and the output hidden layer representation of the third layer has a dimension of 8.

[0106] The training of graph neural network inference uses historically reviewed asset risk propagation events as samples, with the binary label of whether each node actually suffered indirect risk propagation in the historical events as the target variable. The loss function adopts the binary cross-entropy loss function, the optimizer adopts Adam, and the initial learning rate is set to 0.001. Training continues until the loss value converges. The convergence of the loss value means that the absolute value of the difference between the loss values ​​of two adjacent iterations is less than the convergence threshold, which is preferably 0.0001.

[0107] S413: Take the hidden layer representation of all nodes in the graph neural network inference module as input and output the propagation risk enhancement score of each asset's unique number.

[0108] Specifically, the 8-dimensional hidden layer representation of each node output from layer 3 of S412 is used as the output head of the input graph neural network inference module. This output head is an independent linear transformation layer. The 8-dimensional hidden layer representation is multiplied by the output head weight matrix and then added to the output head bias vector. This result is mapped to the interval 0 to 1 using a Sigmoid activation function, yielding the propagation risk enhancement score for the unique asset ID corresponding to that node. The output head weight matrix and output head bias vector are learnable parameters obtained through training data. The propagation risk enhancement score reflects the accumulated indirect risk borne by this asset in the asset association topology graph due to the diffusion of risk states from neighboring assets.

[0109] S411 to S413 inject the comprehensive risk score from the affected asset risk assessment table into the asset association topology graph as node feature vectors. Then, through multiple rounds of message passing by the graph neural network inference module, risk propagation inference is performed on the topology graph. The node feature vector contains both structured hit tags and semantic hit tags, so that the risk source type of neighboring nodes is carried and passed on during message passing. During training, the weight matrix learns different strengths of propagation weights for structured hit neighbors and semantically hit neighbors, so that the propagation risk enhancement score and the reliability of risk source are parametrically coupled. This step addresses the systematic underreporting of indirectly exposed assets when relying solely on direct intelligence hits to identify affected assets: If only direct intelligence hits are retained without risk propagation inference, assets that are not hit by any intelligence but are connected to multiple high-confidence hit neighbors will have all four components of their node feature vector being 0, and will be completely missed. If risk propagation inference is performed but the node feature vector does not contain a hit marker, message passing cannot distinguish the source of neighbor risk. Structured hit neighbors and semantic hit neighbors propagate with the same weight, causing false positives that may be introduced by semantic hit neighbors to spread indiscriminately along the topology, and the propagation risk enhancement score of indirectly exposed assets loses its source reliability basis. After combining the two mechanisms, indirectly exposed assets are identified by obtaining a non-zero propagation risk enhancement score through risk propagation from multiple high-confidence structured hit neighbors, and the propagation contribution of semantic hit neighbors is suppressed due to the learned lower propagation weight. Without either of these two mechanisms, the ability to identify indirectly exposed assets or distinguish risk sources is lost, and the propagation risk enhancement score loses its physical basis for depicting the actual indirect exposure state.

[0110] S42: For all assets in the affected asset risk assessment table, use a unique number to calculate the final risk score by weighting the comprehensive risk score and the risk enhancement score, and classify the warning level according to the final risk score.

[0111] The comprehensive risk score output by S33 only reflects direct risk, while the enhanced risk propagation score output by S413 only reflects indirect propagation risk. Both must be combined to comprehensively depict the actual risk of the asset. This step weights and merges the two scores and classifies them accordingly, ensuring that the warning level is affected by both direct impact and indirect propagation.

[0112] Specifically, for all assets with unique identifiers in the affected asset risk assessment table, the final risk score is calculated using the following formula: in The overall risk score output for S33. The score for enhanced propagation risk output by S413. and Given weights that sum to 1, an example is taken as follows: , .

[0113] Warning levels are determined based on the final risk score: The final risk score of all unique asset IDs is calculated. Using percentile 1 and percentile 2 as the dividing line, unique asset IDs with a final risk score greater than or equal to percentile 1 are classified as high-risk; those with a final risk score greater than or equal to percentile 2 but less than percentile 1 are classified as medium-risk; and those with a final risk score less than percentile 2 are classified as low-risk. The sum of percentile 1 and percentile 2 is 100%, and percentile 1 is greater than percentile 2. For example, percentile 1 can be the 67th percentile, and percentile 2 can be the 33rd percentile. When the number of affected assets is less than the preset minimum statistical sample size (e.g., 30), percentile grading is abandoned, and a preset absolute threshold is used for direct classification: a final risk score higher than the high-risk absolute threshold by 0.7 is classified as high-risk, higher than the medium-risk absolute threshold by 0.4 is classified as medium-risk, and the rest are low-risk. The absolute thresholds can be determined by security operations personnel based on historical handling experience and stored in the configuration items.

[0114] See Figure 9 This is a schematic diagram illustrating the percentile grading of the final risk score provided in the embodiments of this application. For example... Figure 9As shown, the left side of the graph displays the final risk score axis, indicated by upward arrows, with the lower end labeled "low" and the upper end "high." The middle section features a tiered strip, divided into three segments by two dotted lines. The upper dotted line corresponds to the 67th percentile, and the lower dotted line corresponds to the 33rd percentile. Both dotted lines are led out by text labels. The strip, from top to bottom, labels high-risk, medium-risk, and low-risk levels. This graph categorizes asset unique IDs with a final risk score greater than or equal to the 67th percentile as high-risk, those greater than or equal to the 33rd percentile but less than the 67th percentile as medium-risk, and those less than the 33rd percentile as low-risk, ensuring that the warning level is influenced by both direct impact and indirect transmission.

[0115] S43: Identify high-risk and medium-risk assets with unique numbers, confirm the current network identifier of the affected assets, search the disposal rule base, and generate differentiated disposal suggestions.

[0116] Actions must be issued to the current network identifier of the affected asset to take effect. However, intelligence matching is based on the asset's historical identifier, so the current identifier must be confirmed first. This step retrieves matching actions from the action rule base based on risk characteristics and outputs executable, differentiated action recommendations.

[0117] Specifically, for the unique identifiers of high-risk and medium-risk assets, the latest identifier change record in the identifier change record sequence corresponding to the unique asset identifier is extracted from the asset identifier history chain generated in S12, and the changed identifier of the identifier change record is taken as the current network identifier of the affected asset. The department, person in charge, and network isolation capability identifier of the asset corresponding to the unique asset identifier are read from the internal asset status data obtained in S12.

[0118] Based on the vulnerability exploitation maturity score, the number of known weaknesses, and the combination status of structured hit tags and semantic hit tags corresponding to the unique asset ID, matching action entries are retrieved from a pre-set action rule base. The action rule base is a mapping table with vulnerability exploitation maturity score ranges, known weakness number ranges, and hit tag combination statuses as keys, and a priority-sorted list of action entries as values. Actions include three categories: isolation operations, patching, and traffic filtering rules. The initial values ​​of the mapping table are determined by security operations personnel based on historical action experience. New risk feature combinations are added with corresponding action entries by security operations personnel upon their first appearance. After each action review, the system updates the priority of action entries in the mapping table based on the action effectiveness feedback. The retrieved action entries are combined with the current network identifier and network isolation capability identifier of the affected asset to generate differentiated action suggestions including isolation operations, patching, and traffic filtering rules.

[0119] S44: Summarize and package the warning level, current network identifier of affected assets, final risk score, intelligence indicator identifier, enhanced spread risk score, differentiated handling suggestions, and intelligence generation timestamp to generate a risk warning and handling report and distribute it according to the push strategy.

[0120] Specifically, the warning level output by S42, the current network identifier and differentiated handling suggestions of the affected assets output by S43, the final risk score output by S42, the propagation risk enhancement score output by S413, and the corresponding intelligence indicator identifiers and intelligence generation timestamps are summarized and packaged to generate a risk warning and handling report. This risk warning and handling report is output in a structured document format and distributed according to the push strategy corresponding to the warning level: high-risk levels are distributed to the security operations platform and the head of the asset-owning department in the form of real-time alerts; medium-risk levels are distributed in the form of periodic summaries. This enables security operations personnel to initiate handling of assets actually bearing risk based on the current network identifier of the affected assets, and to issue corresponding isolation operations, patch repairs, and traffic filtering rules to vulnerability scanners, endpoint protection platforms, and external attack surface management tools through industrial communication interfaces, achieving full lifecycle monitoring and closed-loop handling of attack surface risks.

[0121] Through step-by-step processing from S1 to S4, the output parameters of the structured temporal matching link, the semantic embedding matching link, and the graph neural network risk propagation link are continuously coupled between each step: the asset identification historical linked list in S1 and the asset identification snapshot at the intelligence moment establish a temporally accurate structured matching foundation; the pre-trained language representation model encoding in S1 establishes a semantic matching foundation; the two foundations are consumed by hierarchical Bloom filters and near nearest neighbor retrieval in S2 and fused with dual labels; the fused combined state drives the three-branch calculation of the intelligence confidence weighted influence score in S3; the comprehensive risk score in S3 is injected into the asset association topology graph in S4 as a node feature vector; the propagation risk enhancement score of the graph neural network inference module is weighted with the comprehensive risk score to obtain the final risk score; the final risk score drives the generation of graded early warning and differentiated disposal suggestions, enabling this embodiment to have the ability to identify in three scenarios: accurate correspondence between intelligence and asset identification, implicit semantic association between intelligence and assets, and indirect exposure of affected assets.

[0122] Example 2: This embodiment, based on Embodiment 1, provides an artificial intelligence-driven information network system risk early warning system, such as... Figure 10 As shown, it includes: The multi-source data processing and feature encoding module is used to acquire raw external threat intelligence data and internal asset status data from multi-source data access channels. It performs indicator field parsing and timestamp extraction on the raw external threat intelligence data, and performs identifier change event tracking and historical snapshot recording on the internal asset status data, generating an intelligence indicator sequence and an asset identifier historical linked list, respectively. For each intelligence indicator in the intelligence indicator sequence, it performs validity period range positioning in the asset identifier historical linked list based on the intelligence generation timestamp, extracting the set of historical asset identifiers corresponding to the identifier validity period range falling within the intelligence generation timestamp, and generating an asset identifier snapshot at the intelligence time. It then performs pre-trained language representation model encoding on all attack indicator fields in the intelligence indicator sequence and the asset description text in the asset identifier snapshot at the intelligence time, respectively, generating an intelligence semantic embedding library and an asset semantic embedding library. The dual-path matching and fusion module is used to perform structured interval matching and semantic vector retrieval based on the intelligence indicator sequence, the intelligence time asset identifier snapshot, the intelligence semantic embedding library and the asset semantic embedding library, respectively, and fuse the two matching results to generate a list of intelligence asset association pairs; The correlation analysis and scoring calculation module is used to perform multidimensional correlation analysis on the intelligence asset correlation pair list, calculate the intelligence confidence weighted impact score and asset exposure score based on the combination state of structured hit tags and semantic hit tags, and generate an affected asset risk assessment table. The graph network propagation analysis and early warning module is used to input the affected asset risk assessment table into the graph neural network inference module to perform asset-related risk propagation analysis, output propagation risk enhancement score, combine the comprehensive risk score to perform graded risk early warning, generate differentiated disposal suggestions, and output risk early warning disposal report.

Claims

1. A risk early warning method for an information network system driven by artificial intelligence, characterized in that, include: Raw external threat intelligence data and internal asset status data are obtained from multi-source data access channels. The raw external threat intelligence data undergoes indicator field parsing and timestamp extraction, while the internal asset status data undergoes identifier change event tracking and historical snapshot recording, generating an intelligence indicator sequence and an asset identifier historical linked list, respectively. For each intelligence indicator in the intelligence indicator sequence, validity period range positioning is performed in the asset identifier historical linked list based on the intelligence generation timestamp. The corresponding set of historical asset identifiers within the identifier validity period range that the intelligence generation timestamp falls into is extracted, generating an asset identifier snapshot at the intelligence moment. Pre-trained language representation models are used to encode all attack indicator fields in the intelligence indicator sequence and the asset description text in the asset identifier snapshot at the intelligence moment, generating an intelligence semantic embedding library and an asset semantic embedding library, respectively. Based on the intelligence indicator sequence, the intelligence moment asset identifier snapshot, the intelligence semantic embedding library, and the asset semantic embedding library, structured interval matching and semantic vector retrieval are performed respectively, and the two matching results are merged to generate a list of intelligence asset association pairs; Perform multidimensional association analysis on the intelligence asset association pair list, calculate the intelligence confidence weighted impact score and asset exposure score based on the combined state of structured hit tags and semantic hit tags, and generate an affected asset risk assessment table; The affected asset risk assessment form is input into the graph neural network inference module to perform asset-related risk propagation analysis, outputting a propagation risk enhancement score. Combined with the comprehensive risk score, a graded risk warning is executed, differentiated disposal suggestions are generated, and a risk warning disposal report is output.

2. The AI-driven risk early warning method for information network systems according to claim 1, characterized in that, The process of parsing the indicator fields and extracting the timestamps from the original external threat intelligence data includes: Raw external threat intelligence data is obtained from threat intelligence platform interfaces, open intelligence subscription sources, and internal security device alert channels; Perform field parsing on the raw external threat intelligence data to extract attack indicator fields and intelligence generation timestamps; The attack indicator field is bound to the intelligence generation timestamp and stored to generate the intelligence indicator sequence.

3. The AI-driven risk early warning method for information network systems according to claim 1, characterized in that, The generation of intelligence indicator sequences and asset identifier historical linked lists respectively includes: Continuously pull internal asset status data from the asset management platform, cloud platform API, and network configuration collection interface; The internal asset status data is subjected to identifier change event detection. The detection method is to compare the current retrieved value with the previous retrieved value for each asset's IP address, domain name binding and network endpoint at a preset retrieval period. Whenever a discrepancy is detected between the current fetch value and the previous fetch value of an asset, the identifier before the change, the identifier after the change, and the time of the change are added as a change record to the identifier change record of the corresponding asset. All identification change records for each asset are sorted in ascending order by the time the change occurred, generating an asset identification history linked list.

4. The AI-driven risk early warning method for information network systems according to claim 1, characterized in that, The generated intelligence semantic embedding library and the asset semantic embedding library include: For each intelligence indicator in the intelligence indicator sequence, the attack indicator field of the intelligence indicator is input into a pre-trained language representation model to obtain the semantic embedding vector of the attack indicator field. For each asset in the intelligence moment asset identification snapshot, a unique number is assigned. The asset type description, deployment environment label, and business function label are read from the internal asset status data and concatenated into an asset description text. This text is then input into the pre-trained language representation model to obtain the corresponding asset semantic embedding vector. All semantic embedding vectors are aggregated to generate an intelligence semantic embedding library, and all asset semantic embedding vectors are aggregated to generate an asset semantic embedding library.

5. The AI-driven risk early warning method for information network systems according to claim 1, characterized in that, The separate execution of structured interval matching and semantic vector retrieval includes: Construct a hierarchical Bloom filter, and write each record in the intelligence moment asset identifier snapshot into the corresponding layer of the hierarchical Bloom filter according to the attack indicator type, time fragment number and asset history identifier, to obtain a three-layer filter state vector; For each intelligence indicator in the intelligence indicator sequence, perform bitwise operations on the three layers of the hierarchical Bloom filter in sequence; For the intelligence indicators screened by the hierarchical Bloom filter, an equality check is performed in the asset history identifier auxiliary index of the asset identifier snapshot at the intelligence time to confirm the unique number of the asset to be matched. A binary search is performed in the asset identifier history chain with the intelligence generation timestamp as the query parameter to confirm the valid structured association and generate a structured association record.

6. The AI-driven risk early warning method for information network systems according to claim 5, characterized in that, The sequential bitwise operation query performed on the three layers of the hierarchical Bloom filter includes: Extract the attack indicator type, the time segment number to which the intelligence generation timestamp belongs, and the attack indicator field from the intelligence indicator; The layered Bloom filter is queried sequentially. The first layer queries whether all the corresponding bits of the attack indicator type are 1 after being mapped by the hash function. The second layer queries whether all the corresponding bits of the string concatenated with the attack indicator type and the time slice number are 1. The third layer queries whether all the corresponding bits of the string concatenated with the attack indicator type, the time slice number and the attack indicator field are 1. If any layer has a mapping bit of 0, the intelligence indicator is determined to be impossible to match; if all three layers have mapping bits of 1, the intelligence indicator is transferred to perform precise matching.

7. The artificial intelligence-driven information network system risk early warning method according to claim 1, characterized in that, The process of fusing the two matching results to generate a list of intelligence asset association pairs includes: For each intelligence indicator identifier, merge all its corresponding structured association records and semantic association records; For asset unique IDs that appear in both structured association records and semantic association records, both the structured hit flag and the semantic hit flag are set to 1; For unique asset IDs that appear only in structured association records, set the structured hit flag to 1 and the semantic hit flag to 0; For asset unique IDs that only appear in semantic association records, set the structured hit flag to 0 and the semantic hit flag to 1; The intelligence indicator identifier, asset unique number, structured hit mark, semantic hit mark, hit mark validity period range, and cosine similarity are summarized into a single association record, and a list of intelligence asset association pairs is generated.

8. The artificial intelligence-driven information network system risk early warning method according to claim 1, characterized in that, The calculation of the intelligence confidence weighted impact score and asset exposure score based on the combined state of structured hit tags and semantic hit tags includes: The system reads the intelligence source credibility score, attack activity score, vulnerability exploitation maturity score, and cosine similarity score from the associated records. When both the structured hit marker and the semantic hit marker are 1, the weighted sum of the four scores is used as the weighted impact score of the intelligence confidence. When only the structured hit is marked as 1, the cosine similarity term is set to zero and then summed as the weighted impact score of the intelligence confidence. When the semantic hit marker is only 1, the weighted impact score of the intelligence confidence is calculated by multiplying the cosine similarity by the semantic decay coefficient. The number of externally accessible ports, associated business levels, and known vulnerabilities of the corresponding assets are extracted from the internal asset status data. After normalization, the data is weighted and summed to obtain the asset exposure score.

9. The artificial intelligence-driven risk early warning method for information network systems according to claim 1, characterized in that, The generation of the affected asset risk assessment table includes: For each associated record in the intelligence asset association pair list, the comprehensive risk score is the product of the intelligence confidence weighted impact score and the asset exposure score. All associated records in the intelligence asset association list are sorted in descending order of comprehensive risk score; The system combines the unique asset number, intelligence indicator identifier, structured hit marker, semantic hit marker, intelligence confidence weighted impact score, asset exposure score, comprehensive risk score, and hit marker validity period range to generate an affected asset risk assessment table.

10. The artificial intelligence-driven risk early warning method for information network systems according to claim 1, characterized in that, The step of inputting the affected asset risk assessment form into the graph neural network inference module to perform asset-related risk propagation analysis and outputting a propagation risk enhancement score includes: Construct an asset association topology graph with network connectivity and business dependencies between assets as edges and unique asset numbers as nodes; The comprehensive risk score, asset exposure score, structured hit marker, and semantic hit marker in the affected asset risk assessment table are concatenated into a node feature vector and written into the corresponding node; The asset association topology graph carrying node feature vectors is input into the graph neural network inference module, which performs multiple rounds of message passing and updates the hidden layer representation of each node. Using the hidden layer representations of all nodes in the graph neural network inference module as input, the propagation risk enhancement score of each asset's unique identifier is obtained through linear transformation layers and activation function mapping.

11. The AI-driven risk early warning method for information network systems according to claim 1, characterized in that, The method of combining comprehensive risk scores to implement tiered risk warnings and generate differentiated handling recommendations includes: The final risk score is calculated by weighting and summing the overall risk score and the risk enhancement score. Based on the first and second percentiles of the final risk score as the dividing line, the unique asset number is divided into high-risk, medium-risk, or low-risk levels; the sum of the first and second percentiles is 100%, and the first percentile is greater than the second percentile. For assets with high and medium risk levels, a unique identifier is assigned, and the latest changed identifier in the identifier change record sequence is extracted as the current network identifier of the affected asset. Based on the vulnerability exploitation maturity score, the number of known weaknesses, and the hit tag combination status of the corresponding intelligence indicators, the system retrieves matching action entries from the pre-set action rule base and generates differentiated action suggestions.

12. An artificial intelligence-driven information network system risk early warning system, used to implement the artificial intelligence-driven information network system risk early warning method according to any one of claims 1 to 11, characterized in that, The system includes: The multi-source data processing and feature encoding module is used to acquire raw external threat intelligence data and internal asset status data from multi-source data access channels. It performs indicator field parsing and timestamp extraction on the raw external threat intelligence data, and performs identifier change event tracking and historical snapshot recording on the internal asset status data, generating an intelligence indicator sequence and an asset identifier historical linked list, respectively. For each intelligence indicator in the intelligence indicator sequence, it performs validity period range positioning in the asset identifier historical linked list based on the intelligence generation timestamp, extracting the set of historical asset identifiers corresponding to the identifier validity period range falling within the intelligence generation timestamp, and generating an asset identifier snapshot at the intelligence time. It then performs pre-trained language representation model encoding on all attack indicator fields in the intelligence indicator sequence and the asset description text in the asset identifier snapshot at the intelligence time, respectively, generating an intelligence semantic embedding library and an asset semantic embedding library. The dual-path matching and fusion module is used to perform structured interval matching and semantic vector retrieval based on the intelligence indicator sequence, the intelligence time asset identifier snapshot, the intelligence semantic embedding library and the asset semantic embedding library, respectively, and fuse the two matching results to generate a list of intelligence asset association pairs; The correlation analysis and scoring calculation module is used to perform multidimensional correlation analysis on the intelligence asset correlation pair list, calculate the intelligence confidence weighted impact score and asset exposure score based on the combination state of structured hit tags and semantic hit tags, and generate an affected asset risk assessment table. The graph network propagation analysis and early warning module is used to input the affected asset risk assessment table into the graph neural network inference module to perform asset-related risk propagation analysis, output propagation risk enhancement score, combine the comprehensive risk score to perform graded risk early warning, generate differentiated disposal suggestions, and output risk early warning disposal report.