Network visitor reputation rating system based on public intelligence clues
Through the network visitor reputation rating system based on public intelligence clues, the traditional system's shortcomings in dynamic correlation data and adapting to complex modes is solved, and the accurate assessment and dynamic adaptation of visitor risk levels are achieved, which meets the needs of precise prevention and control.
Patent Information
- Application Number
- CN202510415502.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-06-27
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The traditional network visitor reputation rating system lacks the ability to dynamically correlate multiple heterogeneous data, and it is difficult to integrate dispersed information, resulting in one-sided risk assessment; at the same time, the scoring mechanism relies too much on static rules or manual thresholds, lacks the ability to capture complex patterns, and cannot dynamically adapt to new attack methods.
The network visitor reputation rating system based on public intelligence clues is adopted. Public intelligence data is collected from OSINT data sources through the data acquisition module. The data cleaning module cleanses the data. The data analysis module performs network visitor portrait analysis to generate a joint coded vector of public clues. Finally, the reputation rating label generation module automatically generates a reputation rating label based on these vectors.
It realizes a more accurate assessment of the visitor's true risk level, avoids one-sided risk assessment, can dynamically adapt to new attack methods, and meets the needs of precise prevention and control.
Smart Images

Figure CN120223404A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of network security, and more specifically, to a network visitor credit rating system based on public intelligence clues. Background Art
[0002] In a digital network environment, network visitor credit rating is a core link in defending against network attacks and protecting system security. Its core value lies in quantifying the risk level of access behaviors to prevent threats such as malicious crawlers, account theft, and fraud attacks, and to ensure the security of data assets and the availability of services.
[0003] However, traditional network visitor credit rating systems have the following significant defects: on the one hand, traditional systems lack the ability to dynamically associate multiple heterogeneous data, making it difficult to integrate scattered information into a unified visitor profile, resulting in one-sided risk assessment; on the other hand, in the scoring mechanism, traditional systems overly rely on static rules or manual threshold settings, lacking the ability to capture complex patterns and unable to dynamically adapt to new attack methods. These problems together cause traditional systems to be difficult to accurately evaluate the true risk level of visitors in a complex network environment. Especially when facing covert attacks using legitimate infrastructure, there are significant identification blind spots, seriously restricting the refinement and intelligence level of network security protection and unable to meet the precise prevention and control requirements for access behavior risks in the digital age.
[0004] Therefore, an optimized network visitor credit rating system is expected. Summary of the Invention
[0005] To solve the above technical problems, this application is proposed. Embodiments of this application provide a network visitor credit rating system based on public intelligence clues.
[0006] According to one aspect of this application, there is provided a network visitor credit rating system based on public intelligence clues, which includes:
[0007] A data collection module for collecting public intelligence data of a target network visitor from an OSINT data source, where the public intelligence data includes IP address source, historical record of the IP address, user agent string, and email domain name reputation;
[0008] A data cleaning module for cleaning the public intelligence data of the target network visitor to obtain cleaned public intelligence data;
[0009] A data analysis module, configured to perform network visitor portrait analysis on the cleaned public intelligence data to obtain a target network visitor public clue joint coding vector, wherein the data analysis module is configured to: perform network visitor portrait analysis based on public clue collaborative coding on the cleaned public intelligence data to obtain the target network visitor public clue joint coding vector;
[0010] A reputation rating label generation module, configured to obtain a reputation rating label of the target network visitor based on the target network visitor public clue joint coding vector.
[0011] Compared with the prior art, the network visitor reputation rating system based on public intelligence clues provided by the present application uses artificial intelligence-based data processing technology to perform data cleaning and structured mapping on the obtained public intelligence data of the target network visitor to obtain multiple public intelligence data features, and then performs IP address activity record semantic dynamic transmission on the time series of the semantic embedded coding features of the IP address activity record among them to obtain IP address activity time series pattern features. Finally, based on the joint representation between the IP address activity time series pattern features and other public intelligence data features, a reputation rating label of the target network visitor is automatically obtained. In this way, the true risk level of the visitor can be more accurately evaluated, and thus the precise prevention and control requirements can be met. Brief Description of the Drawings
[0012] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts. In the drawings:
[0013] Figure 1 It is a system block diagram of a network visitor reputation rating system based on public intelligence clues according to an embodiment of the present application.
[0014] Figure 2 It is a schematic diagram of data flow of a network visitor reputation rating system based on public intelligence clues according to an embodiment of the present application.
[0015] Figure 3 It is a block diagram of a data analysis module in a network visitor reputation rating system based on public intelligence clues according to an embodiment of the present application.
[0016] Figure 4 It is a block diagram of an IP address activity time series pattern feature extraction unit in a network visitor reputation rating system based on public intelligence clues according to an embodiment of the present application. Detailed Description of the Embodiment
[0017] Next, exemplary embodiments according to the present application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. It should be understood that the present application is not limited by the exemplary embodiments described herein.
[0018] In a digital network environment, the network visitor credit rating mechanism, as a key link in resisting network attacks and maintaining system security, its core value lies in effectively identifying and preventing security threats such as malicious crawlers, account theft, and fraud attacks through quantitative risk assessment means, thereby ensuring the security of data assets and the continuous availability of services.
[0019] However, the existing network visitor credit rating systems have obvious limitations: First, when dealing with various heterogeneous data, traditional systems lack the ability of dynamic association and are difficult to integrate scattered data information into a comprehensive and unified visitor profile, resulting in one-sided risk assessment results; Second, traditional scoring mechanisms mainly rely on static rules or manually set thresholds and lack the ability to identify complex behavior patterns and cannot respond in a timely manner to the evolution of new attack methods. These defects make it difficult for traditional systems to accurately assess the true risk level of visitors in a complex network environment, especially when facing hidden attacks using legitimate infrastructure, there are often blind spots in identification, seriously affecting the refinement and intelligence level of network security protection and making it difficult to meet the needs of precise prevention and control of access behavior risks in the digital age.
[0020] It should be understood that public intelligence clues refer to various information that can be obtained through public channels such as the Internet, including but not limited to data such as IP addresses, domain names, email addresses, user agent information, and user behavior patterns. These information can objectively reflect the historical activities and potential characteristics of network visitors. Conducting credit rating based on such public intelligence clues can comprehensively evaluate the credibility and risk level of visitors by analyzing public data. Compared with the traditional evaluation method that relies on a single internal log or a closed database, public intelligence clues have the characteristics of wide source, strong timeliness, and dynamic update, and can capture malicious behavior patterns (such as abnormal login frequency, association with known blacklist IPs, etc.) more comprehensively, thereby helping websites or applications to achieve accurate risk prediction and access control. This not only reduces the misjudgment probability caused by information silos, but also can enhance the initiative and adaptability of overall security defense through the correlation analysis of public data without the need for users to actively provide sensitive information, and can provide data-driven decision support for security teams to optimize security strategies in a complex network environment.
[0021] Based on this, the present application proposes a network visitor reputation rating system based on public intelligence clues. First, it collects public intelligence data of target network visitors from OSINT data sources. Then, it uses artificial intelligence-based data processing techniques to clean and structurally map the public intelligence data of target network visitors to obtain multiple public intelligence data features. Subsequently, it performs semantic dynamic transmission on the time series of the semantic embedding coding features of the IP address activity records to obtain IP address activity time series pattern features. Finally, based on the joint representation between the IP address activity time series pattern features and other public intelligence data features, it automatically obtains the reputation rating label of the target network visitor. In this way, by dynamically capturing the associations of multiple heterogeneous data, the one-sidedness of risk assessment can be effectively avoided. At the same time, the ability of the machine learning model to capture complex patterns can be utilized to dynamically rate the visitors' reputation, achieving an accurate assessment of the true risk level of visitors in a complex network environment, and thus meeting the needs of precise prevention and control.
[0022] Figure 1 FIG. is a system block diagram of a network visitor reputation rating system based on public intelligence clues according to an embodiment of the present application. Figure 2 FIG. is a schematic diagram of data flow of a network visitor reputation rating system based on public intelligence clues according to an embodiment of the present application. As Figure 1 and Figure 2 shown, in the network visitor reputation rating system 100 based on public intelligence clues, it includes: a data collection module 110, configured to collect public intelligence data of target network visitors from OSINT data sources, where the public intelligence data includes IP address source, historical records of the IP address, user agent string, and email domain name reputation; a data cleaning module 120, configured to clean the public intelligence data of the target network visitors to obtain the cleaned public intelligence data; a data analysis module 130, configured to perform network visitor portrait analysis on the cleaned public intelligence data to obtain a target network visitor public clue joint coding vector; and a reputation rating label generation module 140, configured to obtain the reputation rating label of the target network visitor based on the target network visitor public clue joint coding vector.
[0023] In the embodiment of the present application, the data collection module 110 is used to collect public intelligence data of target network visitors from OSINT data sources. The public intelligence data includes IP address source, historical records of IP addresses, user agent strings, and email domain name reputations. It should be understood that the IP address source includes information such as the geographical location of the IP address, the affiliated organization (ASN - Autonomous System Number), whether it is a known proxy IP or VPN, and whether it is on the malicious IP blacklist. These information can determine the approximate location and affiliated organization of the target network visitor, judge whether it uses a proxy or VPN to hide its true identity, and whether it comes from a known malicious IP address, thereby helping to evaluate potential risks. The historical records of IP addresses refer to the network activities that the IP address has participated in. This information helps to discover whether the IP has a history of malicious behavior and can provide a reference for judging the reputation of the target network visitor. The user agent string contains user agent analysis information (judging whether it is a common browser, crawler, malicious robot, or automated tool, detecting whether the user agent is abnormal or forged) and browser fingerprint information. Obtaining this information can judge whether the visitor is a normal user or a malicious program. An abnormal or forged user agent may imply malicious access. Email is an important way of network activities. Understanding the email domain name reputation can judge whether the target network visitor is associated with malicious activities such as spam and phishing. Generally speaking, the above public intelligence data represents the behavior and background of the target network visitor from multiple dimensions, and they jointly form a comprehensive network visitor profile. The IP address source and historical records can reflect its physical location and the legality of past behaviors; the user agent string can identify the nature of the visitor's tool and judge whether there are abnormalities; the email domain name reputation evaluates potential risks from the perspective of communication. These information corroborate each other and can provide a comprehensive basis for determining the reputation rating of the target network visitor, thereby effectively preventing potential malicious access and network threats.
[0024] The following is a detailed elaboration of a specific implementation process of "collecting public intelligence data of target network visitors from OSINT data sources":
[0025] First, for the collection of IP address sources, the system retrieves relevant information about the target IP address from multiple public data sources. Specifically, the system first inputs the target IP address into IP geolocation databases (such as MaxMind GeoIP, IP2Location) to obtain its geographical location information. Then, the system queries the Autonomous System Number (ASN) database to obtain information about the organization to which the IP address belongs, such as whether the IP address is owned by an Internet Service Provider (ISP) or an enterprise. In addition, the system also searches for the IP address in threat intelligence platforms (such as VirusTotal, AbuseIPDB) to check whether it is marked as a malicious IP and whether it is associated with known attack activities.
[0026] Secondly, the historical record of the IP address is also an important source of information for evaluating the credibility of network visitors. The historical record of the IP address includes network activities the IP address has participated in, open ports, and service types, etc. To obtain this information, the system queries from multiple OSINT data sources. For example, the system queries the target IP address in threat intelligence platforms (such as AlienVault OTX, Cisco Talos) to obtain its historical activity records, including whether it has participated in malicious activities such as DDoS attacks and port scans. In addition, the system uses network traffic analysis platforms (such as Censys, Shodan) to obtain the open ports, service types, and historical activity patterns of the IP address, such as whether it frequently initiates abnormal requests.
[0027] Next, the collection and analysis of the User-Agent string are also important steps in evaluating the credibility of network visitors. The User-Agent string contains information such as the browser type, operating system, and device type used by the visitor, which can help determine whether the visitor is a normal user or a malicious program. To obtain this information, the system inputs the target User-Agent string into user-agent analysis tools (such as UserAgentString.com, WhatIsMyBrowser) to parse detailed information such as its browser type, operating system, and device type. In addition, the system also searches for the User-Agent string in threat intelligence platforms (such as VirusTotal, Hybrid Analysis) to check whether it is associated with known malicious tools (such as crawlers, automated attack tools).
[0028] Finally, the collection and analysis of email domain reputation are also an important part of evaluating the reputation of network visitors. Email is an important means of communication in network activities. Understanding the reputation of email domains can help determine whether the target network visitors are associated with malicious activities such as spam and phishing attacks. To obtain this information, the system extracts the domain name part from the target email address and queries the reputation score of the domain name in the domain reputation database (such as Spamhaus, MXToolbox) to determine whether it is marked as a spam source or associated with malicious activities. In addition, the system also uses DNS record query tools (such as DNSChecker, Whois) to check the SPF, DKIM, and DMARC configurations of the domain name to determine whether it has the legitimate ability to send emails. At the same time, the system searches for the domain name in the threat intelligence platform (such as VirusTotal, AbuseIPDB) to check whether it is associated with known malicious activities (such as phishing attacks, malware distribution).
[0029] After completing the collection of the above data, the system will summarize all the obtained information and store these data in the database for subsequent data processing operations. At the same time, the system will ensure the traceability of the data sources, record the sources of each data item (such as data source name, query time), so as to verify and update when needed.
[0030] In an embodiment of the present application, the data cleaning module 120 is used to clean the public intelligence data of the target network visitor to obtain the cleaned public intelligence data. Accordingly, considering the public intelligence data collected from the OSINT data source, there may be problems such as incomplete data (such as missing some fields of IP address history records), duplication (the same redundant information is collected from different channels), confusing formats (such as inconsistent user agent string formats) or errors (such as errors in email domain name reputation data records). If these "dirty data" are used directly, subsequent analysis will be biased, affecting the accuracy of reputation ratings. Therefore, it is necessary to clean the public intelligence data of the target network visitor to eliminate interference factors and ensure data quality. Specifically, for repeated or conflicting records (such as redundant markings or status conflicts of the same IP in different blacklists), entity alignment and cross-source calibration mechanisms are used, combined with timestamps and confidence weights, to dynamically determine the latest valid information (such as when an IP is marked as "malicious" in the threat intelligence platform but the private log shows that it has been recycled as a normal business IP, the status coverage is preferentially based on the recent verification results). Secondly, the invalid or erroneous content in the original data is formatted and logically checked, such as removing reserved addresses (such as 192.168.0.0 / 16), repairing incomplete or obviously forged user agent strings (such as parsing "Mozilla / 5.0(;;;;;)" as an abnormal UA mode), and unifying the geographic location encoding standard to ensure the consistency of cross-field semantics. For missing key fields (such as domain name SPF configuration status or IP activity time series breaks), reasonable completion is carried out through context association analysis (such as filling missing values according to the median activity of similar IPs) or business logic inference (such as marking domain names with no SPF records but normal mail traffic as "configuration defects" instead of directly classifying them as malicious). After the above cleaning process, the noise, redundancy and contradictions in the original data are effectively eliminated, forming highly consistent and low-error cleaned public intelligence data, which can enable the subsequent data analysis process to more accurately capture risk signals (such as accurately identifying malicious crawlers disguised as legitimate browsers), and ultimately support the reputation rating system to achieve highly robust risk decision-making in a complex network environment.
[0031] In the embodiment of the present application, the data analysis module 130 is used to perform a network visitor profile analysis on the cleaned public intelligence data to obtain a target network visitor public clue joint coding vector. Specifically, in the embodiment of the present application, the data analysis module is used to perform a network visitor profile analysis based on public clue collaborative coding on the cleaned public intelligence data to obtain a target network visitor public clue joint coding vector. More specifically, Figure 3 FIG. 1 is a block diagram of a data analysis module in a network visitor reputation rating system based on public intelligence clues according to an embodiment of the present application.Figure 3 As shown, the data analysis module 130 includes: an intelligence data mapping unit 131, configured to perform structured mapping encoding on the cleaned public intelligence data to obtain a time series of an IP address source structured encoding vector, an IP address activity record semantic embedding encoding vector, a user agent string semantic embedding encoding vector, and an email domain name reputation semantic embedding encoding vector; an IP address activity time series pattern feature extraction unit 132, configured to perform IP address activity record semantic dynamic transfer on the time series of the IP address activity record semantic embedding encoding vector to obtain an IP address activity time series pattern feature encoding vector; and an intelligence data joint encoding unit 133, configured to perform public clue joint encoding on the IP address source structured encoding vector, the IP address activity time series pattern feature encoding vector, the user agent string semantic embedding encoding vector, and the email domain name reputation semantic embedding encoding vector to obtain the target network visitor public clue joint encoding vector.
[0032] In an embodiment of the present application, the intelligence data mapping unit 131 is configured to perform structured mapping encoding on the cleaned public intelligence data to obtain a time series of an IP address source structured encoding vector, an IP address activity record semantic embedding encoding vector, a user agent string semantic embedding encoding vector, and an email domain name reputation semantic embedding encoding vector. Specifically, in an embodiment of the present application, the intelligence data mapping unit is configured to: perform structured mapping encoding on the IP address source using one-hot encoding to obtain the IP address source structured encoding vector; use an embedding layer to perform structured mapping encoding on the historical record of the IP address, the user agent string, and the email domain name reputation respectively to obtain the time series of the IP address activity record semantic embedding encoding vector, the user agent string semantic embedding encoding vector, and the email domain name reputation semantic embedding encoding vector.
[0033] It should be understood that after cleaning, the public intelligence data still exists in the original format (such as text, discrete category information), and computers cannot directly understand and process this unstructured data. Through structured mapping encoding, the data can be converted into a vector form recognizable by a computer (such as a numerical vector), and then the data features can be effectively extracted, providing a basis for more detailed analysis by subsequent models. In particular, in a specific embodiment of the present application, one-hot encoding is used to perform structured mapping encoding on the source of the IP address to obtain the structured encoding vector of the IP address source, and an embedding layer is used to perform structured mapping encoding on the historical record of the IP address, the user agent string, and the email domain name reputation respectively to obtain the time series of the semantic embedding encoding vector of the IP address activity record, the semantic embedding encoding vector of the user agent string, and the semantic embedding encoding vector of the email domain name reputation. Specifically, one-hot encoding is to convert discrete categorical variables into binary vectors, with each category corresponding to one dimension in the vector, and only one position is 1 (the rest are 0), which is suitable for processing data with a small number of categories and no obvious semantic associations. The embedding layer is usually used to map high-dimensional sparse raw data (such as text, long sequences) to a low-dimensional dense vector space, while retaining the semantic information of the data when compressing the dimensions, and is suitable for processing data that needs to mine semantic associations. Considering that the attributes of the IP address source (such as whether it is a proxy IP, the type of affiliated organization) are usually limited in categories (such as "yes / no", "a small number of organization categories"), and there is no complex semantic association between categories, using one-hot encoding to process it can clearly distinguish each category and simply and directly express the characteristics of the IP address source. The IP address activity record contains time series information of activity behaviors and there are temporal semantic associations; the text content of the user agent string is complex, and implicit semantics such as "browser type", "whether it is a malicious crawler" need to be mined; the email domain name reputation has diverse domain name categories, and deep information such as "whether the domain name is associated with spam" needs to be captured through semantic embedding. By using the embedding layer to process them respectively, the original semantic information can be retained while converting the data representation format, which is beneficial to improving the accuracy of feature representation.
[0034] In the embodiment of the present application, the IP address activity time series pattern feature extraction unit 132 is used to perform IP address activity record semantic dynamic transfer on the time series of the semantic embedding encoding vector of the IP address activity record to obtain the IP address activity time series pattern feature encoding vector. Specifically, Figure 4 It is a block diagram of the IP address activity time series pattern feature extraction unit in the network visitor reputation rating system based on public intelligence clues according to the embodiment of the present application. As Figure 4As shown, the IP address activity timing pattern feature extraction unit 132 includes: a semantic transfer terminal axial adjustment factor calculation subunit 1321, configured to calculate the IP address activity record semantic transfer terminal adjustment factor and the IP address activity record semantic transfer axial adjustment factor of each IP address activity record semantic embedding coding vector based on the IP address activity record semantic transfer terminal reference positioning coding vector and the IP address activity record semantic transfer axial reference positioning coding vector of the time series of the IP address activity record semantic embedding coding vector; a semantic transfer adjustment factor calculation subunit 1322, configured to calculate the IP address activity record semantic transfer adjustment factor of each IP address activity record semantic embedding coding vector based on the IP address activity record semantic transfer terminal adjustment factor and the IP address activity record semantic transfer axial adjustment factor of each IP address activity record semantic embedding coding vector in the time series of the IP address activity record semantic embedding coding vector; an IP address activity timing pattern feature generation subunit 1323, configured to perform IP address activity record semantic dynamic constraint transfer coding on the time series of the IP address activity record semantic embedding coding vector based on the IP address activity record semantic transfer adjustment factor of each IP address activity record semantic embedding coding vector to obtain the IP address activity timing pattern feature coding vector.
[0035] It should be understood that the time series of the semantic embedding coding vectors of the IP address activity records contains various activity information of the IP at different time points. In this sequence, the early activity information may have an important impact on the current reputation assessment, and there are long-range dependencies. For example, an IP address launched malicious attacks frequently several months ago. Although it has performed normally recently, there may still be potential threats. Traditional data analysis methods often do not work well when dealing with such long-range dependencies. As the sequence length increases, the transmission of information will gradually decay or be lost. And when processing, it is also necessary to consider the global structure information of the sequence data (the overall distribution pattern of the IP address activities, such as whether the IP has frequent abnormal activities within a specific time period) and the local structure information (reflecting the short-term patterns in the sequence, such as consecutive access behaviors within a certain time period). Based on this, in this application, it is necessary to perform semantic dynamic transmission on the time series of the semantic embedding coding vectors of the IP address activity records to obtain the IP address activity time series pattern feature coding vectors. In this way, on the basis of retaining the local structure information, it can be integrated into the global time series, so as to better balance the two, so that the finally generated IP address activity time series pattern feature coding vectors can contain both short-term activity patterns and reflect long-term activity trends. At the same time, the long-range dependencies of the IP address activity records can be effectively captured during this coding process, which helps to more comprehensively reflect the activity patterns of the IP address and provide richer information for accurately evaluating the reputation of network visitors.
[0036] Specifically, in the embodiment of this application, the semantic transfer terminal axial adjustment factor calculation sub-unit 1321 is used to: extract the last semantic embedding coding vector of the IP address activity record from the time series of the semantic embedding coding vectors of the IP address activity records as the semantic transfer terminal reference positioning coding vector of the IP address activity record. This process can be represented by the formula:
[0037] O={v1,v2,...,v i ,...,v t}
[0038] v tail =v t
[0039] Where O is the time series of the semantic embedding coding vectors of the IP address activity records, v1, v2, v i and v t are the 1st, 2nd, i-th and t-th semantic embedding coding vectors of the IP address activity records in the time series of the semantic embedding coding vectors of the IP address activity records respectively, and v tail is the semantic transfer terminal reference positioning coding vector of the IP address activity record;
[0040] Perform clustering analysis on the time series of the semantic embedding coding vectors of the IP address activity records to obtain the semantic transfer axial reference positioning coding vectors of the IP address activity records. This process can be expressed by the formula:
[0041]
[0042] where Cluster is the clustering analysis operation, max(v i ) and min(v i ) are respectively the maximum and minimum values of v i , η is the adjustment parameter, a i is the i-th semantic transfer axial reference positioning value in the time series of the semantic transfer axial reference positioning values of the IP address activity records, softmax is the normalization function, e i is the i-th semantic transfer axial reference positioning weight value in the time series of the semantic transfer axial reference positioning weight values of the IP address activity records, t is the number of vectors in O, and v axis is the semantic transfer axial reference positioning coding vector of the IP address activity records;
[0043] Calculate the semantic transfer terminal adjustment factors of each semantic embedding coding vector in the time series of the semantic embedding coding vectors of the IP address activity records relative to the semantic transfer terminal reference positioning coding vectors of the IP address activity records. This process can be expressed by the formula:
[0044]
[0045] where f(v i , v tail ) is to calculate the terminal adjustment factor between v i and v tail , v ij is the eigenvalue at the j-th position of v i , v tailj is the eigenvalue at the j-th position of v tail , log2 is the logarithmic function value with base 2, n is the number of eigenvalues in v i , exp is the exponential function value with base e (the natural constant), and α i is the semantic transfer terminal adjustment factor corresponding to v i ;
[0046] The IP address activity record semantic transfer axial adjustment factor of each IP address activity record semantic embedding encoding vector in the sequence distribution of the IP address activity record semantic embedding encoding vector relative to the IP address activity record semantic transfer axial reference positioning encoding vector is calculated. The process can be expressed by the formula:
[0047]
[0048] Among them, f(v i ,v axis ) is used to calculate v i and v axis The axial adjustment factor between ||·||2 is the Euclidean norm of the calculation vector, arccosh is the inverse hyperbolic cosine function, β i Yes i The corresponding IP address activity record semantic delivery axial adjustment factor.
[0049] It should be understandable that the last embedding vector of the time series of the semantic embedding coding vector of the IP address activity record represents the semantic state at the current moment and is the end point of the sequence evolution. Traditional methods are prone to loss of early key activity patterns due to information attenuation when processing long sequences, while the IP address activity record semantic transmission terminal reference positioning coding vector can constrain the convergence direction of information transmission, ensure that information transmission is always anchored to the current state, and avoid semantic deviation caused by too long a sequence. For example, even if an IP has performed normally recently, by using the terminal vector as a fixed reference, the model can still trace back to malicious attack patterns from several months ago. In other words, by introducing the time series end point as a positioning reference, the model not only enhances the sensitivity of the model to the final state of the sequence, but also can encode long-term dependencies into the current features through the back-propagation mechanism, which can effectively improve the stability of the encoding process and avoid feature drift.
[0050] Accordingly, considering that the global pattern of IP activities (such as periodic anomalies) is implicit in the sequence distribution, it is difficult for traditional data analysis methods to model explicitly. By clustering the time series of the semantic embedding coding vector of the IP address activity record and extracting the principal components of the original feature sequence data (such as frequent attack clusters and normal access clusters), the common structural features across time periods can be captured. That is, the generated IP address activity record semantic transfer axial reference positioning coding vector is used as a distribution skeleton to provide regularization constraints for the temporal transmission of IP address activity information, forcing the encoding process to proceed along the main direction of the data manifold. This effectively suppresses noise interference and ensures that long-term trends (such as slow changes in attack frequency) are not overwhelmed by local fluctuations, which is conducive to improving the model's generalization recognition ability for overall activity patterns.
[0051] It should be understood that the similarity difference between the semantic embedding coding vector of the IP address activity record at each moment and the semantic transmission terminal reference positioning coding vector of the IP address activity record can reflect the correlation strength with the current state. By calculating the semantic transmission terminal adjustment factor of the IP address activity record through the exponential deformation formula of the KL divergence, it is essentially constructing a time decay attention mechanism - the features closer to the terminal moment obtain higher weights, and the early features need to prove their causal association with the current state to be retained. For example, if the abnormal login of a certain IP three months ago has pattern coherence with the current suspicious behavior, its information transmission is strengthened through the adjustment factor. That is, through the calculation of the terminal adjustment factor, it not only allows important historical features to break through the time decay constraint, but also suppresses the propagation of irrelevant noise, so as to achieve the balance between "key event memory enhancement" and "redundant information filtering".
[0052] Correspondingly, considering that the global structural features of IP address activities (such as periodic patterns, normal behavior baselines) need to be explicitly modeled through the backbone patterns of data distributions. Since the semantic transmission axial reference positioning coding vectors of the IP address activity records obtained by clustering are essentially the clustering centers in the sequence feature space, which carry the common semantics across time periods (such as the active patterns during working hours of office IPs, the high-frequency probing at random times of attack IPs). By calculating the similarity between the semantic embedding coding vector of the IP address activity record at each time point and the axial reference positioning coding vector, the deviation degree of the activity behavior at that moment from the global pattern can be quantified. That is, the generated semantic transmission axial adjustment factor of the IP address activity record is used as the global constraint weight to project the local features at each time point into the semantic subspace spanned by the clustering centers. This is equivalent to correcting the distribution of abnormal segments (such as intensive logins during low-probability early morning hours), weakening the interference of their statistical deviation from the main pattern on the coding process. At the same time, this factor ensures that the long-term behavior trends (such as the monthly activity decay) obtain higher weights during coding by strengthening the time points with high similarity to the axial vector (such as legal accesses that conform to the 9-18 o'clock periodicity), thereby improving the model's sensitivity to persistent threats (such as periodic low-frequency scans).
[0053] Specifically, in the embodiment of the present application, the semantic transmission adjustment factor calculation subunit 1322 is used to calculate the semantic transmission adjustment factor of each IP address activity record semantic embedding coding vector based on the semantic transmission terminal adjustment factor and the semantic transmission axial adjustment factor of each IP address activity record semantic embedding coding vector in the time series of the IP address activity record semantic embedding coding vector.
[0054] More specifically, in the embodiments of the present application, the semantic transfer adjustment factor calculation subunit 1322 is configured to: based on the canonical space constraint matrix, perform orthogonal convergence covariance on the IP address activity record semantic transfer terminal adjustment factor and the IP address activity record semantic transfer axial adjustment factor of the IP address activity record semantic embedding encoding vector to obtain the IP address activity record semantic transfer terminal convergence covariance adjustment factor and the IP address activity record semantic transfer axial convergence covariance adjustment factor; based on the IP address activity record semantic transfer terminal convergence covariance adjustment factor and the IP address activity record semantic transfer axial convergence covariance adjustment factor, perform convergence constraint based on the component rotation matrix on the IP address activity record semantic transfer terminal adjustment factor and the IP address activity record semantic transfer axial adjustment factor to obtain the optimized IP address activity record semantic transfer terminal adjustment factor and the optimized IP address activity record semantic transfer axial adjustment factor; perform weighted processing based on the sigmoid function on the optimized IP address activity record semantic transfer terminal adjustment factor and the optimized IP address activity record semantic transfer axial adjustment factor to obtain the IP address activity record semantic transfer adjustment factor of the IP address activity record semantic embedding encoding vector. It can be expressed by the following formula:
[0055]
[0056] y i = Sigmoid(ω1·α i '+ ω2·β i ')
[0057] where δ i and ε i are respectively the IP address activity record semantic transfer terminal convergence covariance adjustment factor and the IP address activity record semantic transfer axial convergence covariance adjustment factor after orthogonal convergence covariance of α i and β i , T represents the transpose operation, |·| is the absolute value operation, α i ' and β i ' are respectively the optimized IP address activity record semantic transfer terminal adjustment factor and the optimized IP address activity record semantic transfer axial adjustment factor corresponding to v i , ω1 and ω2 are respectively weighted hyperparameters, Sigmoid is the weight mapping function, and y i is the IP address activity record semantic transfer adjustment factor corresponding to v i .
[0058] It should be understood that the IP address activity record semantic transfer terminal regulator and the IP address activity record semantic transfer axial regulator impose constraints from two orthogonal perspectives, namely the local relevance in the time dimension and the global distribution in the feature space respectively. However, using them independently may lead to information imbalance. For example, over-reliance on the terminal factor will cause the model to only focus on recent activities and ignore historical global patterns. On the contrary, it may cause the current key signals to be overwhelmed by long-term trends. The IP address activity record semantic transfer regulator obtained by weighted fusion can form a multi-constraint joint optimization space, enabling the feature vector at each time point to simultaneously satisfy the semantic coherence with the current terminal state (such as the contradiction between recent consecutive normal accesses and historical attack behaviors being reasonably weighed) and the distribution consistency of the global axial structure (such as whether the current behavior deviates from the typical active pattern of this IP) during the information transfer process. This dual constraint forces the model to capture the local details of short-term emergencies (such as an abnormal access to a certain high-risk port) during encoding and also place it in the long-term activity profile (such as the overall threat level trend in the past six months) for context calibration, thus achieving a dynamic balance between local sensitivity and global robustness.
[0059] In particular, during the semantic encoding process of IP address activity records, in view of the possible orthogonal diffusion phenomenon between the IP address activity record semantic transmission terminal regulator and the IP address activity record semantic transmission axial regulator in the feature vector information flow transmission field and their regularization contradiction with the global constraint, it is necessary to achieve covariant integration through the orthogonal convergence mechanism of the information transmission field. Specifically, first, a scalar gauge space constraint matrix is constructed to impose orthogonal projection on the two types of regulators, mapping them to mutually independent and complete vector subspaces, thereby eliminating the distortion of the feature manifold caused by non-orthogonal coupling. During this process, the temporal end-point correlation information carried by the IP address activity record semantic transmission terminal regulator and the global distribution main mode features represented by the IP address activity record semantic transmission axial regulator are orthogonally decoupled into a time-dimensional local sensitive field and a feature space global distribution field through tensor decomposition, forming a two-channel constraint basis. Subsequently, a component rotation matrix based on the reversible space metric is introduced to perform a rotation transformation on the feature components in the orthogonal subspace, so that while maintaining orthogonality, they re-calibrate the projection direction along the geometric structure of the data manifold. This operation transforms the linear superposition relationship of the original regulators into a non-linear cooperative effect in the manifold space through finite-dimensional spinor convergence constraints, enabling the backpropagation path of the IP address activity record semantic transmission terminal factor and the regularization constraint trajectory of the axial factor to form a dynamic balance during the information transmission process. This covariant integration mechanism within the orthogonal structure field not only retains the anchoring effect of the terminal reference positioning on the current semantic state but also strengthens the skeleton support of the axial reference positioning on the global mode distribution. At the same time, it eliminates the overfitting risk during the regularization process through the spatial rotation of the reversible metric, ultimately achieving the unified optimization of long-term dependence relationship encoding and local burst pattern detection, enabling the model to accurately capture the pattern mutations of recent attack links and stably trace the weak correlation signals of cross-period latent threats when facing complex IP activity scenarios.
[0060] Specifically, in the embodiment of the present application, the IP address activity time-series pattern feature generation sub-unit 1323 is configured to perform IP address activity record semantic dynamic constraint transmission encoding on the time series of the IP address activity record semantic embedding encoding vectors based on the IP address activity record semantic transmission regulators of the IP address activity record semantic embedding encoding vectors, so as to obtain the IP address activity time-series pattern feature encoding vectors, which can be expressed by the following formula:
[0061]
[0062] where z is the IP address activity time-series pattern feature encoding vector.
[0063] It should be understood that by performing IP address activity record semantic dynamic constraint transfer encoding on the time series of the IP address activity record semantic embedding encoding vector through the IP address activity record semantic transfer regulator, it is essentially performing a dynamic weighted aggregation operation, which can make the finally aggregated IP address activity time series pattern feature encoding vector retain both local short-term behaviors (such as high-frequency requests in the recent week) and integrate global long-term patterns (such as quarterly threat trends), and explicitly model long-range dependencies (such as the contribution of early malicious behaviors to the current risk). This helps to more comprehensively reflect the activity patterns of IP addresses, and thus can provide richer information for accurately evaluating the reputation of network visitors.
[0064] In the embodiment of the present application, the intelligence data joint encoding unit 133 is configured to perform common clue joint encoding on the IP address source structured encoding vector, the IP address activity time series pattern feature encoding vector, the user agent string semantic embedding encoding vector, and the email domain name reputation semantic embedding encoding vector to obtain the target network visitor common clue joint encoding vector. Specifically, in the embodiment of the present application, the intelligence data joint encoding unit is configured to: use a joint encoder based on the Transformer model to perform common clue joint encoding on the IP address source structured encoding vector, the IP address activity time series pattern feature encoding vector, the user agent string semantic embedding encoding vector, and the email domain name reputation semantic embedding encoding vector to obtain the target network visitor common clue joint encoding vector.
[0065] It should be understood that information such as the source of the IP address, the temporal pattern of IP address activities, the user - agent string, and the reputation of the email domain name obtained from different dimensions, although able to characterize certain aspects of the visitor after their respective encoding processes, are relatively scattered. To integrate the characteristics of multi - source heterogeneous data and fuse the characteristics of different aspects of the visitor into a unified vector, thereby forming a comprehensive and complete network visitor profile, in this application, it is necessary to perform joint encoding of common clues on the structured encoding vector of the IP address source, the feature encoding vector of the IP address activity temporal pattern, the semantic embedding encoding vector of the user - agent string, and the semantic embedding encoding vector of the email domain name reputation to obtain the target network visitor common - clue joint encoding vector. The generated target network visitor common - clue joint encoding vector can comprehensively reflect various behavior patterns and attributes of the visitor in network activities, can provide richer and more representative inputs for subsequent risk scoring, making the reputation rating more accurate and reliable, and effectively identifying potential malicious visitors. In particular, in a specific embodiment of this application, a joint encoder based on the Transformer model is used to perform joint encoding of common clues on the structured encoding vector of the IP address source, the feature encoding vector of the IP address activity temporal pattern, the semantic embedding encoding vector of the user - agent string, and the semantic embedding encoding vector of the email domain name reputation to obtain the target network visitor common - clue joint encoding vector. Those of ordinary skill in the art should know that the core of the Transformer model is the self - attention mechanism, which can automatically focus on the relationships between elements at different positions in the input sequence and assign different weights to each element according to these relationships. During the joint - encoding process, different types of encoding vectors can be regarded as a sequence, and the self - attention mechanism can help the model dynamically capture the important associations between various vectors. For example, when evaluating the visitor's reputation, in some cases, the IP address activity temporal pattern feature may be more critical, and the self - attention mechanism can assign a higher weight to this vector, thereby highlighting important information and improving the accuracy of encoding. In addition, the Transformer model can process the input sequence in parallel, without the need to process elements one by one in sequence like an RNN. This enables a significant improvement in computational efficiency when processing the joint encoding of multiple encoding vectors.
[0066] In the embodiment of the present application, the credit rating label generation module 140 is configured to obtain the credit rating label of the target network visitor based on the joint encoding vector of the public clues of the target network visitor. Specifically, in the embodiment of the present application, the credit rating label generation module is configured to: input the joint encoding vector of the public clues of the target network visitor into a risk score calculator based on a machine learning model to obtain the credit rating label of the target network visitor. It should be understood that the joint encoding vector of the public clues of the target network visitor integrates multi-dimensional features, but it is still a high-dimensional dense vector and cannot directly represent the specific credit rating. It needs to be input into the risk score calculator. Specifically, the risk score calculator is a model constructed based on machine learning algorithms. It can learn the mapping relationship of "data - features - risk scores - credit rating labels" through a large amount of training data. When the newly processed joint encoding vector of the public clues of the target network visitor is input into the trained risk score calculator, it will output the credit rating label result of the target network visitor according to the learned mapping pattern. The credit rating label is a quantitative identification of the risk level of the target network visitor, in the form of a level label (such as "high risk", "medium risk", "low risk"). According to the obtained credit rating label result, differential security decisions can be implemented, which helps to achieve precise protection. For example, when the risk score is greater than 0.8, the credit rating label is "high risk", and the system automatically adds its IP to the temporary blacklist. When the risk score is greater than 0.6, the credit rating label is "medium risk", and the system implements additional verification measures (such as two-factor authentication or behavioral verification codes). When the risk score is less than 0.6, the credit rating label is "low risk", and the system allows normal access.
[0067] In a specific embodiment of the present application, a logistic regression model is used as the infrastructure of the risk score calculator. The following is a detailed description of a specific implementation process of "inputting the joint encoding vector of the public clues of the target network visitor into a risk score calculator based on a machine learning model to obtain the credit rating label of the target network visitor":
[0068] The first step of the implementation is to prepare the training data. This requires widely collecting samples from a large amount of network visitor data. These samples must be comprehensive, covering various situations such as normal visitors and abnormal visitors with different risk levels. Each sample needs to clearly define the corresponding true credit rating label. For example, the sample of a malicious crawler visitor is labeled as "high risk", the one with occasional abnormal behavior is labeled as "medium risk", and the one with long-term normal access is labeled as "low risk". After collecting the sample data, it needs to be reasonably divided. Usually, the data is divided into a training set and a test set according to a ratio of 70% - 30%. The training set is used to train the logistic regression model to let the model learn the rules in the data; the test set is used to evaluate the performance of the model and test the performance of the model on unknown data.
[0069] After the data is prepared, it enters the stage of training the logistic regression model. First, the model parameters need to be initialized, setting the weights and biases. These initial values will be continuously optimized and adjusted during subsequent training. Then, the jointly encoded vector of the common clues of network visitors obtained by encoding in the training set is input into the model. The model calculates the predicted risk probability value through the sigmoid function. At this time, the logarithmic loss function is used to measure the difference between the predicted value and the true label. The logarithmic loss function can accurately reflect the accuracy of the model prediction. The smaller its value, the better the prediction effect of the model. To make the model's prediction closer to the true label, optimization algorithms such as gradient descent are used to update the weights and biases of the model according to the gradient of the loss function. The gradient descent algorithm gradually adjusts the parameters along the direction where the loss function decreases. This process will be iterated multiple times until the loss function converges to a smaller value, indicating that the model has learned the patterns in the data well.
[0070] After training is completed, the performance of the model needs to be evaluated. The jointly encoded vector of the common clues of network visitors obtained by encoding in the test set is input into the trained logistic regression model, and the model will output the corresponding risk prediction results. Then, performance metrics such as accuracy, recall rate, and F1 value are used to comprehensively evaluate the performance of the model. Accuracy reflects the proportion of the number of samples correctly predicted by the model in the total number of samples. The recall rate reflects the proportion of the number of positive samples (such as high-risk samples) correctly predicted by the model in the actual number of positive samples. The F1 value, as the harmonic mean of accuracy and recall rate, can comprehensively reflect the performance of the model. Through these metrics, the accuracy and reliability of the model in identifying network visitors with different risk levels can be accurately judged. If the performance metrics of the model are not ideal, the model needs to be optimized. Possible operations include adjusting the model parameters, increasing the amount of training data, or trying other more suitable machine learning models.
[0071] Finally, it is the link of generating credit rating labels. When the logistic regression model is trained and passes the performance evaluation, the jointly encoded vector of the common clues of the target network visitor to be parsed is input into the model. The model will calculate the risk probability value of this visitor, that is, the risk score, based on the learned parameters. For example, if the calculated risk probability value is 0.8, it indicates that this visitor has an 80% probability of belonging to the high-risk category. Then, the credit rating label is determined according to the pre-set risk score threshold. Suppose it is set that a risk score greater than 0.8 is "high risk", between 0.6 - 0.8 is "medium risk", and less than 0.6 is "low risk". When the risk score output by the model is 0.8, the credit rating label of this target network visitor is determined as "high risk". In this way, the risk score can be accurately mapped to discrete credit rating labels, realizing the quantitative identification of the risk level of network visitors.
[0072] In summary, the network visitor credit rating system 100 based on public intelligence clues according to the embodiments of the present application is elucidated. It uses artificial intelligence-based data processing technology to perform data cleaning and structured mapping on the obtained public intelligence data of the target network visitor to obtain multiple public intelligence data features. Subsequently, time series of semantic embedding coding features of IP address activity records are subjected to semantic dynamic transmission of IP address activity records to obtain IP address activity time series pattern features. Finally, based on the joint representation between the IP address activity time series pattern features and other public intelligence data features, the credit rating label of the target network visitor is automatically obtained. In this way, the true risk level of the visitor can be evaluated more accurately, and thus the precise prevention and control requirements can be met.
Claims
1. A network visitor reputation rating system based on public intelligence clues, characterized in that: include: A data collection module for collecting public intelligence data of target network visitors from OSINT data sources, wherein the public intelligence data includes IP address sources, IP address history, user agent strings, and email domain reputation; A data cleaning module, used for cleaning the public intelligence data of the target network visitor to obtain cleaned public intelligence data; A data analysis module is used to perform a network visitor portrait analysis on the cleaned public intelligence data to obtain a target network visitor public clue joint coding vector, wherein the data analysis module is used to: perform a network visitor portrait analysis based on public clue collaborative coding on the cleaned public intelligence data to obtain the target network visitor public clue joint coding vector; The reputation rating label generation module is used to obtain the reputation rating label of the target network visitor based on the target network visitor public clue joint encoding vector.
2. The network visitor reputation rating system based on public intelligence clues according to claim 1 is characterized in that: The data analysis module comprises: An intelligence data mapping unit, used for performing structured mapping encoding on the cleaned public intelligence data to obtain an IP address source structured encoding vector, a time series of an IP address activity record semantic embedding encoding vector, a user agent string semantic embedding encoding vector, and an email domain name reputation semantic embedding encoding vector; An IP address activity time series pattern feature extraction unit is used to perform IP address activity record semantic dynamic transfer on the time series of the IP address activity record semantic embedded coding vector to obtain the IP address activity time series pattern feature coding vector; The intelligence data joint encoding unit is used to perform public clue joint encoding on the IP address source structured encoding vector, the IP address activity time series pattern feature encoding vector, the user agent string semantic embedding encoding vector and the email domain name reputation semantic embedding encoding vector to obtain the target network visitor public clue joint encoding vector.
3. The network visitor reputation rating system based on public intelligence clues according to claim 2 is characterized in that: The intelligence data mapping unit is used to: Performing structured mapping encoding on the IP address source using one-hot encoding to obtain a structured encoding vector of the IP address source; The history record of the IP address, the user agent string and the email domain reputation are respectively structured mapped and encoded using an embedding layer to obtain a time series of the IP address activity record semantic embedding encoding vector, the user agent string semantic embedding encoding vector and the email domain reputation semantic embedding encoding vector.
4. The network visitor reputation rating system based on public intelligence clues according to claim 2 is characterized in that: The IP address activity timing pattern feature extraction unit comprises: A semantic transfer terminal axial adjustment factor calculation subunit is used to calculate the IP address activity record semantic transfer terminal adjustment factor and the IP address activity record semantic transfer axial adjustment factor of each IP address activity record semantic embedding coding vector based on the IP address activity record semantic transfer terminal reference positioning coding vector and the IP address activity record semantic transfer axial reference positioning coding vector of the time series of the IP address activity record semantic embedding coding vector; A semantic transfer adjustment factor calculation subunit, used to calculate the IP address activity record semantic transfer adjustment factor of each IP address activity record semantic embedding coding vector based on the IP address activity record semantic transfer terminal adjustment factor and the IP address activity record semantic transfer axial adjustment factor of each IP address activity record semantic embedding coding vector in the time series of the IP address activity record semantic embedding coding vector; The IP address activity timing pattern feature generation subunit is used to perform IP address activity record semantic dynamic constraint transfer encoding on the time series of the IP address activity record semantic embedded coding vector based on the IP address activity record semantic transfer adjustment factor of each IP address activity record semantic embedded coding vector to obtain the IP address activity timing pattern feature coding vector.
5. The network visitor reputation rating system based on public intelligence clues according to claim 4 is characterized in that: The semantic transfer terminal axial adjustment factor calculation subunit is used to: Extracting the last IP address activity record semantic embedding coding vector from the time series of the IP address activity record semantic embedding coding vector as the IP address activity record semantic transmission terminal reference positioning coding vector; Performing cluster analysis on the time series of the semantic embedding coding vector of the IP address activity record to obtain the axial reference positioning coding vector of the semantic transfer of the IP address activity record; Calculate the IP address activity record semantics delivery terminal adjustment factor of each IP address activity record semantics embedding coding vector in the time series of the IP address activity record semantics embedding coding vector relative to the IP address activity record semantics delivery terminal reference positioning coding vector; Calculate the IP address activity record semantic transfer axial adjustment factor of each IP address activity record semantic embedding coding vector in the sequence distribution of the IP address activity record semantic embedding coding vector relative to the IP address activity record semantic transfer axial reference positioning coding vector.
6. The network visitor reputation rating system based on public intelligence clues according to claim 5 is characterized in that: The semantic transfer adjustment factor calculation subunit is used to: Based on the canonical space constraint matrix, the IP address activity record semantic transfer terminal adjustment factor and the IP address activity record semantic transfer axial adjustment factor of the IP address activity record semantic embedding coding vector are subjected to orthogonal convergence covariation to obtain the IP address activity record semantic transfer terminal convergence covariation adjustment factor and the IP address activity record semantic transfer axial convergence covariation adjustment factor; Based on the IP address activity record semantic transfer terminal convergence covariation adjustment factor and the IP address activity record semantic transfer axial convergence covariation adjustment factor, the IP address activity record semantic transfer terminal adjustment factor and the IP address activity record semantic transfer axial adjustment factor are subjected to convergence constraints based on component rotation matrices to obtain optimized IP address activity record semantic transfer terminal adjustment factors and optimized IP address activity record semantic transfer axial adjustment factors; The optimized IP address activity record semantic transfer terminal adjustment factor and the optimized IP address activity record semantic transfer axial adjustment factor are weighted based on the sigmoid function to obtain the IP address activity record semantic transfer adjustment factor of the IP address activity record semantic embedding coding vector.
7. The network visitor reputation rating system based on public intelligence clues according to claim 2 is characterized in that: The intelligence data joint encoding unit is used to use a joint encoder based on a Transformer model to perform public clue joint encoding on the IP address source structured encoding vector, the IP address activity time series pattern feature encoding vector, the user agent string semantic embedding encoding vector and the email domain name reputation semantic embedding encoding vector to obtain the target network visitor public clue joint encoding vector.
8. The network visitor reputation rating system based on public intelligence clues according to claim 1 is characterized in that: The reputation rating label generation module is used to: input the public clue joint encoding vector of the target network visitor into a risk scorer based on a machine learning model to obtain the reputation rating label of the target network visitor.