A computer remote login recognition system
The remote login identification system uses packet and memory analysis with statistical methods to detect behavioral anomalies, addressing the limitations of traditional systems by enhancing the detection of abnormal remote login activities and improving security.
Patent Information
- Application Number
- CN202510562750.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-04-30
AI Technical Summary
The existing computer remote login identification system cannot effectively identify abnormal behaviors during the remote login process, especially the inability to detect dynamic behavior abnormalities and remote environment changes, resulting in attackers being able to bypass authentication and perform malicious operations, and there is a risk of identity credentials being tampered with or forged during transmission.
Through the traffic monitoring module, memory paging monitoring module, environment abnormality detection module and command sequence analysis module, packet data, memory paging data and user input commands are obtained and analyzed separately, and probability distribution and sequence library are constructed. Combined with multi-factor abnormality analysis methods, the overall abnormality degree of remote login behavior is evaluated.
It realizes accurate identification of packet behavior, memory access and user behavior, improves remote login security, can distinguish between normal login and abnormal attack behavior, reduces false alarms and missed alarm rates, and adapts to dynamic adjustments in different usage scenarios.
Smart Images

Figure CN120090877B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of remote identity authentication, and particularly to a computer remote login recognition system. Background Art
[0002] A computer remote login recognition system refers to a system used to perform identity authentication during the computer remote login process to ensure that the remote user is a legitimate authorized entity. The system involves the acquisition, transmission, and comparison of user identity information, and usually adopts methods such as password input, biometric-based identity authentication, hardware token verification, or dynamic key-based identity confirmation to effectively identify the identity of remote users. During the identity information acquisition process, the system may involve specific means such as keyboard input, fingerprint scanning, iris recognition, or face recognition, and combines the transmission and storage mechanisms of identity credentials to prevent data tampering or forgery. In the identity verification link, the system can apply a one-time password generation mechanism triggered by time or events, or combine a public-private key encryption system to achieve accurate matching and authentication of the identity of remote users.
[0003] Traditional remote login recognition systems mainly rely on identity authentication methods for security determination. Although they can ensure the legitimacy of the login entity, they cannot effectively identify abnormal behaviors during the remote login process. Identity authentication mechanisms usually rely on static passwords, biometrics, or hardware tokens, but these methods cannot detect dynamic behavior anomalies after remote login, resulting in the system being difficult to further distinguish abnormal operations once an attacker bypasses the authentication. There is a risk that the identity credentials of remote users can be tampered with or forged during transmission. Even with encryption mechanisms, it is difficult to completely prevent credential leakage problems. At the same time, the identity authentication process is usually independent of subsequent interaction behavior analysis, enabling attackers to perform malicious operations after authentication by hijacking legitimate credentials, and the system lacks effective monitoring of subsequent behaviors. Traditional methods mainly rely on a dynamic key mechanism triggered by time or events, which can only prevent the direct reuse of credentials but cannot identify changes in the remote environment. For example, when an attacker uses remote tools to operate the target system, execute abnormal commands, or hijack normal processes, the system is difficult to determine whether abnormal behaviors exist. Due to the lack of joint analysis of multi-dimensional data such as packets, memory, and command sequences, traditional methods have limited recognition capabilities when facing complex remote attacks and are easily circumvented by means such as forged traffic and hidden operations of remote tools, resulting in the inability to accurately distinguish normal remote logins from abnormal remote attack behaviors. Summary of the Invention
[0004] The purpose of the present invention is to solve the deficiencies existing in the prior art and propose a computer remote login recognition system.
[0005] To achieve the above purpose, the present invention adopts the following technical solutions: A computer remote login recognition system includes:
[0006] The traffic monitoring module obtains the packet data during the user's current remote login interaction, counts the occurrence probability of each packet size category in the packet data, and generates the probability distribution of the current packet size.
[0007] The packet analysis module queries the historical occurrence probability of each packet size category, constructs the probability distribution of the historical packet size, determines the abnormal packet behavior based on the probability distribution of the current packet size and the probability distribution of the historical packet size, and generates the packet behavior analysis result.
[0008] The memory paging monitoring module obtains the memory paging data during the user's remote login interaction and divides it through a time window to generate a memory paging access sequence.
[0009] The environment anomaly detection module regards each type of memory paging data in the memory paging access sequence as an independent data point, and determines the memory access abnormal behavior by calculating the density distribution of each type of data point and adjacent data points, and generates the remote environment analysis result.
[0010] The command sequence analysis module obtains the historical and current user input commands during the remote login interaction, constructs the corresponding historical and current command sequence libraries, determines the abnormal command usage behavior by calculating the similarity of the historical and current command sequence libraries, and generates the user behavior analysis result.
[0011] The security determination module performs remote session identification based on the overall abnormal degree of the packet behavior analysis result, the remote environment analysis result, and the user behavior analysis result, and obtains the remote login behavior identification result.
[0012] As a further solution of the present invention, the specific steps for obtaining the probability distribution of the current packet size are as follows:
[0013] Collect the packet data during the user's current remote login interaction, parse the packet size information from the packet data and classify it according to different byte ranges to generate the packet data classification result.
[0014] Based on the packet data classification result, count the occurrence probability of each packet size category after classification to generate the probability distribution of the current packet size.
[0015] As a further solution of the present invention, the specific steps for obtaining the packet behavior analysis result are as follows:
[0016] Query each packet size category and the corresponding historical occurrence probability during the user's historical remote login interaction, construct the probability distribution of the historical packet size, and calculate the divergence value between the probability distribution of the current packet size and the probability distribution of the historical packet size using KL divergence.
[0017] Compare the divergence value with a preset divergence threshold. If the divergence value exceeds the divergence threshold, it is determined that the distribution of the corresponding packet size category is abnormal, and the abnormal packet size category and its packet behavior characteristics are extracted. The packet behavior characteristics include the packet size change trend, the abnormal packet ratio, and the packet interaction mode, and a packet behavior analysis result is generated.
[0018] As a further solution of the present invention, the step of obtaining the memory paging access sequence is specifically as follows:
[0019] Obtain the memory paging data during the user's remote login interaction, including the memory paging access times, the number of page faults, and the number of memory pages occupied by each process. Remove the non-interactive paging activities generated by background processes. The non-interactive paging activities include background running processes, log writing processes, and cache management processes, so as to retain the paging data related to user interaction and generate the processed memory paging data;
[0020] Divide the processed memory paging data of different categories through a time window, including normal access paging, page fault access paging, and background process paging, and generate a memory paging access sequence.
[0021] As a further solution of the present invention, the step of obtaining the remote environment analysis result is specifically as follows:
[0022] Regard each type of memory paging data in the memory paging access sequence as an independent data point, calculate the density distribution of each type of data point under different time windows, and integrate to obtain the density distribution of each type of data point;
[0023] Based on the density distribution of each type of data point, use the local outlier factor algorithm to compare the density of each data point with the data points in its adjacent time window, and calculate the local outlier factor of the corresponding data point relative to the adjacent data points;
[0024] Compare the local outlier factor with a preset access density range to determine whether there are memory access abnormal behaviors, including process behavior abnormalities, memory access mode mutations, and remote control tool interventions, and generate a remote environment analysis result.
[0025] As a further solution of the present invention, the step of obtaining the user behavior analysis result is specifically as follows:
[0026] Query the historical user input commands during the remote login interaction, including the command execution order and the command parameter combination method. Remove the non-interactive commands such as repeated input, auto-completion, and format adjustment to retain the actual user interaction operations, construct a historical command sequence library, extract the current user input commands during the remote login interaction, and construct a current command sequence library;
[0027] Analyze the length of the longest common subsequence of the command sequences in the current command sequence library and the historical command sequence library using the longest common subsequence algorithm, judge their similarity, compare with the preset similarity range, determine whether there is abnormal command usage behavior, and generate user behavior analysis results.
[0028] As a further solution of the present invention, the specific steps for obtaining the remote login behavior recognition result are as follows:
[0029] Based on the packet behavior analysis result, remote environment analysis result and user behavior analysis result, evaluate the overall abnormal degree of the remote login interaction through a multi-factor abnormal analysis method to obtain an abnormal degree evaluation result;
[0030] Perform remote session recognition according to the abnormal degree evaluation result, including recognizing normal remote login behavior and abnormal remote attack behavior, and obtain the remote login behavior recognition result.
[0031] Compared with the prior art, the advantages and positive effects of the present invention are as follows:
[0032] In the present invention, in remote login interaction, the statistical analysis of packet data can obtain the probability distribution of packet sizes, and detect abnormal traffic through comparison with historical data to prevent potential risks caused by large-scale data transmission or traffic mutations. By calculating the divergence value of the packet size distribution, the abnormal degree of packet behavior can be accurately identified to ensure the rationality of the packet interaction mode. Combining the analysis of the memory paging access sequence, the memory access mode is divided by time windows to exclude the influence of background processes, making the detection of memory access anomalies more accurate. The application of the density distribution calculation method enables the judgment of environmental anomalies not to be limited to a single data point, but to combine the access density changes in adjacent time windows, enhancing the ability to identify sudden changes in process behavior and remote control interventions. The command sequence comparison uses the longest common subsequence algorithm, making the detection of abnormal command usage patterns more refined, capable of identifying deviations in command execution logic rather than simply matching the commands themselves. The introduction of the multi-factor abnormal analysis method realizes the comprehensive evaluation of packets, environment, and user behavior, improves the accuracy of remote session recognition, and avoids false positives and false negatives caused by single-index judgment. When comprehensively evaluating remote login behavior, different abnormal scoring weights can be dynamically adjusted according to historical statistical data, making the abnormal determination more flexible and capable of adapting to different usage scenarios. By comprehensively analyzing packet characteristics, memory access behavior, and command input patterns, normal remote logins and abnormal remote attack behaviors can be accurately identified, enhancing the security of remote logins. Brief Description of the Drawings
[0033] Figure 1 is the system flow chart of the present invention;
[0034] Figure 2 is the flow chart of the traffic monitoring module of the present invention;
[0035] Figure 3 is the flowchart of the packet analysis module of the present invention;
[0036] Figure 4 is the flowchart of the memory paging monitoring module of the present invention;
[0037] Figure 5 is the flowchart of the environmental anomaly detection module of the present invention;
[0038] Figure 6 is the flowchart of the command sequence analysis module of the present invention;
[0039] Figure 7 is the flowchart of the security determination module of the present invention. Detailed implementation manners
[0040] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0041] Please refer to Figure 1 , a computer remote login identification system includes:
[0042] The traffic monitoring module obtains packet data during the current remote login interaction of the user, counts the occurrence probability of each packet size category in the packet data, and generates the probability distribution of the current packet size;
[0043] The packet analysis module queries the historical occurrence probability of each packet size category, constructs the probability distribution of the historical packet size, determines abnormal packet behavior according to the probability distribution of the current packet size and the probability distribution of the historical packet size, and generates the packet behavior analysis result;
[0044] The memory paging monitoring module obtains the memory paging data during the user's remote login interaction and divides it through a time window to generate a memory paging access sequence;
[0045] The environmental anomaly detection module regards each type of memory paging data in the memory paging access sequence as an independent data point, and determines the memory access abnormal behavior by calculating the density distribution of each type of data point and adjacent data points, and generates the remote environment analysis result;
[0046] The command sequence analysis module obtains the historical and current user input commands during the remote login interaction, constructs the corresponding historical and current command sequence libraries, determines the abnormal command usage behavior by calculating the similarity of the historical and current command sequence libraries, and generates the user behavior analysis result;
[0047] The security determination module performs remote session identification based on the overall abnormality degree of the packet behavior analysis result, the remote environment analysis result, and the user behavior analysis result, and obtains the remote login behavior identification result.
[0048] Please refer to Figure 2 , the steps to obtain the probability distribution of the current packet size are specifically as follows:
[0049] Collect the packet data during the user's current remote login interaction, parse the packet size information from the packet data and classify it according to different byte ranges to generate the packet data classification result;
[0050] During the remote login interaction, the traffic monitoring module captures the network packet data in real time, intercepts the TCP / IP packets in transmission using the data stream listening technology, caches the packet data in the local storage area, extracts the basic information of the packets through the protocol parsing tool, including but not limited to the source IP address, destination IP address, port number, protocol type, packet size, etc. For the packet size information, classify and store the packets according to the preset byte interval threshold. For example, divide the packet size into multiple intervals such as 0-64 bytes, 65-512 bytes, 513-1024 bytes, 1025-2048 bytes, etc. At the same time, establish an index structure during the data storage process for subsequent efficient query and statistics. For the packet size classification and statistics process, use the cumulative counting method to record the occurrence times of each type of packet, that is, for each packet arrival, the system determines the interval it belongs to according to its size and increments the counter in the corresponding interval. The statistical data can be stored in the local database or stored in real time using an efficient data structure (such as a hash table) and can be periodically refreshed to the file system.
[0051] Based on the packet data classification result, use the formula:
[0052]
[0053] Calculate the current occurrence probability of the packet size of the th class in the packet data classification result , and count the occurrence probability of each packet size category after classification to generate the probability distribution of the current packet size;
[0054] Among them: represents the total number of occurrences of the packet size of the th class within the statistical time period in the packet data classification result, which is obtained by classifying and counting according to the packet size after real-time capturing of the packets, represents the total number of all packets within the statistical time period in the packet data classification result, which is equal to the sum of the occurrence times of all packet categories.
[0055] Set a statistical time period, for example, 10 minutes. During this time, capture all packet data generated by remote login interactions, parse the packet size information, and classify and count according to the set byte range, recording the occurrence times of each packet size category and the total number of packets , assuming different packet size ranges and counts: 0 - 64 bytes: , 65 - 512 bytes: , 513 - 1024 bytes: , 1025 - 2048 bytes: .
[0056] Total number of packets: .
[0057] Calculate the probability of each packet category: , , , .
[0058] The result shows the probability distribution of packets of different sizes during the remote login interaction. Among them, the packets in the range of 65 - 512 bytes have the highest proportion, reaching 48.61%. Followed by 513 - 1024 bytes, accounting for 25%. The smallest category is 1025 - 2048 bytes, only accounting for 9.72%. This data can be used to analyze the packet characteristics of remote login behavior and provide support for subsequent anomaly detection and optimization strategies
[0059] Through the statistical analysis of the packet size during the remote login interaction, the occurrence probability of different categories of packets can be accurately obtained, forming the probability distribution of the packet size, so as to identify abnormal packet behaviors. Real-time capture of packet data through traffic monitoring technology, and the use of protocol parsing tools to extract key packet information to ensure the accuracy and integrity of packet classification. The preset byte range threshold enables packets to be classified and stored according to a unified standard, and the establishment of the index structure improves the efficiency of data query and statistics. The cumulative counting method is used for packet size statistics, so that the distribution of packet categories can be updated at any time, and large-scale data processing can be completed in a short time. Through probability calculation, the distribution of each packet category during the remote login interaction can be intuitively presented, so as to judge the normal mode of packet traffic
[0060] Please refer to Figure 3 , the specific steps to obtain the packet behavior analysis results are as follows:
[0061] Query each packet size category and the corresponding historical occurrence probability in the user's historical remote login interaction process, and construct the probability distribution of the historical packet size
[0062] Extract historical packet data from remote login interaction logs, query the packet size category information during the user's past interactions in a hierarchical indexing manner, and obtain the corresponding historical occurrence probabilities through database retrieval or time series storage structures to construct a historical packet size probability distribution. During the query process, first retrieve the remote login logs according to the user's unique identifier, extract the size information of each packet in the historical interaction records, and classify and store them according to the set packet size intervals. The historical probability value of each category is obtained by calculating the ratio of the number of packets of each category appearing in the historical data to the total number of packets. The formula is:
[0063]
[0064] Among them, is the historical occurrence probability of the category of packet size, is the total number of times the category appears in the historical data, is the total number of all packets within the historical time period. During the statistical process, to ensure data integrity, a sliding time window is used to screen the historical data. For example, the most recent 30 days are set as the reference time range, and the window length is adjusted according to the user's historical behavior pattern. Finally, a complete historical packet size probability distribution is obtained.
[0065] Use the KL divergence formula:
[0066] ;
[0067] Calculate the divergence value between the probability distribution of the current packet size and the probability distribution of the historical packet size;
[0068] Among them, is the current occurrence probability of the category of packet size in the probability distribution of the current packet size, is the historical occurrence probability of the category of packet size in the probability distribution of the historical packet size, is the adjustment parameter set according to and to reflect the overall deviation degree between the current and historical distributions. The calculation formula is: , by introducing the parameter , the sensitivity of the KL divergence calculation can be dynamically adjusted, weighted according to the deviation degree of the overall probability distribution, making the anomaly detection more sensitive during large-scale changes, and reducing the false positive rate and improving the stability of the algorithm under normal fluctuations. represents the summation calculation for all packet size categories.
[0069] Assume the current probability and the historical probability are as follows: 0 - 64 bytes: , (difference 0.03), 65 - 512 bytes: , (difference 0.04), 513 - 1024 bytes: , (difference 0.02), 1025 - 2048 bytes: , (difference 0.01).
[0070] Calculate : ; ; .
[0071] Calculate the divergence value: ;
[0072] First calculate the term: , , , ;
[0073] Calculate the product of each term: , , , .
[0074] Finally sum up: .
[0075] This result indicates that the divergence value between the current packet size distribution and the historical packet size distribution is 0.00275.
[0076] Compare the divergence value with the preset divergence threshold. If the divergence value exceeds the divergence threshold, it is determined that the distribution of the corresponding packet size category is abnormal, and the abnormal packet size category and its packet behavior characteristics are extracted. The packet behavior characteristics include the packet size change trend, the abnormal packet ratio, and the packet interaction mode, and the packet behavior analysis result is generated;
[0077] Compare the calculated divergence value with the preset divergence threshold . Among them, the threshold Obtained by the system based on statistical analysis of historical packet traffic data. The specific calculation method is as follows: Construct the KL divergence distribution of normal packet traffic. Select multiple time windows (for example, the past 30 days, calculate the KL divergence value once a day), and obtain a sequence of historical KL divergence values. Calculate the mean and standard deviation of this set of historical KL divergence values. Determine the divergence threshold : Use statistical analysis methods to set the threshold. Usually, the mean plus a certain multiple of the standard deviation is used as the boundary. The selection of the adjustment coefficient is usually 2 or 3, specifically based on the normal distribution principle in statistics. When the adjustment coefficient is 2, the threshold range covers approximately 95.4% of the normal data, and values outside this range may be abnormal. When the adjustment coefficient is 3, the threshold range covers approximately 99.7% of the normal data, and values outside this range are very likely to be abnormal situations. The setting range is generally between 1.5 and 3.0, and the specific value can be adjusted according to the stability of historical data. When the data fluctuation is small, a smaller value is taken; when the data fluctuation is large, a larger value is taken. If exceeds the threshold , it is determined that the packet size category distribution is abnormal; if , it is considered that the packet size category distribution is normal.
[0078] In this calculation, 0.00275 is greater than the threshold 0.0025, so the packet size category of this remote login interaction is determined to be abnormal. After anomaly detection, extract the packet behavior characteristics of the abnormal category, including the change trend of packet size, the proportion of abnormal packets, and the packet interaction mode: Change trend of packet size: Statistically calculate the sliding average of the packet size within a fixed time window, calculate the change between the current window mean and the previous window mean, and analyze whether there is a sudden increase or decrease. Proportion of abnormal packets: Calculate the proportion of the number of abnormal category packets, which represents the proportion of the number of abnormal packets in the total number of packets. If the proportion exceeds the set abnormal proportion threshold, further analyze the traffic pattern of this category of packets. Packet interaction mode analysis: Extract information such as the source IP, destination IP, protocol type, and packet interval time of the packet, and perform pattern matching to detect whether there is abnormal communication behavior. Finally, the packet analysis module generates the packet behavior analysis result, and can further combine historical anomaly records to evaluate whether the abnormal behavior has persistence or potential security risks.
[0079] By comparing the current and historical packet size distributions, abnormal changes in packet traffic can be accurately identified, avoiding misjudgments based solely on single - data fluctuations. A sliding time window is used to filter historical data, enabling anomaly detection to adapt to the long - term usage habits of different users and improving the stability of detection. The KL divergence is introduced to calculate the deviation degree of packet distribution, making packet anomaly detection no longer limited to a fixed threshold but dynamically adjusted in combination with the statistical characteristics of the overall packet traffic, thereby enhancing the sensitivity to abnormal behaviors and reducing the false - alarm rate caused by normal fluctuations. By setting statistical methods to determine the divergence threshold, anomaly determination becomes more scientific, reducing interference caused by environmental changes or short - term traffic fluctuations. In addition, the extraction of abnormal packet behavior characteristics, including the trend of packet size changes, the proportion of abnormal packets, and packet interaction patterns, makes the detection of abnormal traffic not only limited to the data - statistical level but also combined with traffic behavior characteristics, further improving the accuracy of abnormal packet analysis and providing more reliable data support for subsequent remote - environment analysis and user - behavior analysis.
[0080] Please refer to Figure 4 , the steps to obtain the memory paging access sequence are specifically as follows:
[0081] Obtain the memory paging data during the user's remote - login interaction, including the number of memory paging accesses, the number of paging faults, and the number of memory pages occupied by each process. Remove non - interactive paging activities generated by background processes. Non - interactive paging activities include background - running processes, log - writing processes, and cache - management processes, so as to retain the paging data related to user interaction and generate processed memory paging data;
[0082] First, the memory - management unit (MMU) records the memory paging access situation of each process, including the number of paging accesses, the number of paging faults, and the number of memory pages occupied by each process. For the number of memory paging accesses, the system obtains the access frequency of each page by periodically scanning the access - bit markers of page - table entries and stores the access records in the kernel data structure. For paging faults, the system counts the number of page faults through the page - fault exception - handling mechanism of the processor and differentiates the types of page faults, such as main - memory page faults and disk page faults. At the same time, to avoid the interference of background processes on interactive paging activities, the system filters the paging data generated by background processes, including non - interactive processes (such as log - writing processes and cache - management processes). The specific method is to record the active - process list through the process scheduler and determine whether a process is a background task according to the process priority and scheduling policy. If the process running state is in the background mode or there has been no user interaction for a long time, the paging activities generated by this process are excluded from the monitoring data. Finally, ensure that the memory paging monitoring module only retains the paging data related to user interaction for subsequent analysis.
[0083] Divide the memory paging data processed for different categories through a time window, including normal access paging, page fault access paging, and background process paging, to generate a memory paging access sequence;
[0084] Use a time window mechanism to divide memory paging data of different categories. First, set the time window length, such as 1 second, 10 seconds, or 1 minute, and count the paging access situations of each process within this window. Record the paging operations occurring within the time window as time series data. The division of the time window is triggered by the CPU clock interrupt signal. At the end of each time slice, the system collects the paging data within the current time window and stores it in the paging access log. At the same time, for paging data of different categories, the system classifies and stores them according to the paging access type, such as normal access paging, page fault access paging, and background process paging, etc. The classification standard is determined based on the trigger source of the paging event. For example, if the paging access is actively triggered by a user-mode process, it is classified as normal access paging; if the paging access is triggered by a page fault exception, it is classified as page fault access paging; background process paging is identified and excluded through process scheduling information. Finally, the paging data within all time windows is integrated into a memory paging access sequence for subsequent analysis and optimization of the memory access pattern.
[0085] By recording the memory paging access situation and excluding the non-interactive paging activities of background processes, the analysis data becomes more accurate, avoiding misjudgment caused by non-user operations. Using a time window to divide paging data of different categories enables the memory access pattern to be analyzed in a serialized manner, improving the traceability and temporal correlation of the data. Through classified storage by paging type, it is possible to distinguish normal access, page fault access, and background process paging, providing a finer-grained monitoring basis for anomaly detection, enabling the system to accurately identify abnormal memory access behaviors, and combined with remote environment analysis, improving the reliability of overall anomaly detection.
[0086] Please refer to Figure 5 , and the steps to obtain the remote environment analysis results are specifically as follows:
[0087] Regard each type of memory paging data in the memory paging access sequence as an independent data point, and use the formula:
[0088] ;
[0089] Calculate the density of the th data point in the time window in the memory paging access sequence, and integrate to obtain the density distribution of various data points;
[0090] Among them, is the The number of accesses to a data point within the corresponding time window can be obtained by sampling the memory access logs, recording the page access counts of each process within each time window, and then counting all the access requests. is the time window of the memory page access sequence. For example, 10 seconds, 30 seconds, or 60 seconds, which determines the range of density calculation. The window length can be selected through experimental analysis. For example, by monitoring the memory page trends of multiple time windows and choosing the window with the most obvious changes as the analysis benchmark. represents the time window all data points within the sum of the number of accesses, representing the overall access volume within this time window, which is obtained by accumulating all the page requests within the time window.
[0091] Suppose the set time window is 10 seconds, and multiple memory page access data points are recorded within each window.
[0092] The number of accesses to the 1st data point , the number of accesses to the 2nd data point , the number of accesses to the 3rd data point , the number of accesses to the 4th data point , the number of accesses to the 5th data point .
[0093] Calculate the total number of accesses: .
[0094] Calculate the density of each data point: , , , , .
[0095] The results show that the density of the 1st data point is 0.1951, indicating that its access frequency is at a medium level within this time window. The density of the 2nd data point is 0.2439, which is the data point with the highest access frequency within this window, indicating that it has the most page accesses. The density of the 3rd data point is 0.2195, second only to the 2nd data point, and still belongs to the data points with a relatively high access frequency. The density of the 4th data point is 0.1789, lower than the previous data points, indicating that its access frequency is relatively low. The density of the 5th data point is 0.1626, which is the data point with the lowest access frequency within this window, indicating that it has the fewest page accesses.
[0096] Based on the density distribution of various data points, the Local Outlier Factor algorithm is used to compare the density of each data point with that of the data points within its adjacent time windows, using the formula:
[0097] ;
[0098] Calculate the local anomaly degree of the th data point in the density distribution of various data points , and obtain the local anomaly factor of the corresponding data point relative to adjacent data points;
[0099] Among them, is the density of the th data point within the time window , is the density of the th data point adjacent to the th data point within the time window , represents the total number of nearest neighbors selected for the th data point within the time window, generally screened by the similarity of access behaviors within the time window. For example, if and , then .
[0100] For example, the density of the 1st data point is 0.1951, indicating that its access frequency is at a medium level within this time window. The density of the 2nd data point is 0.2439, which is the data point with the highest access frequency within this window, indicating that it has the most paging accesses. The density of the 3rd data point is 0.2195, second only to the 2nd data point, and still belongs to the data points with relatively high access frequencies. The density of the 4th data point is 0.1789, lower than the previous several data points, indicating that its access frequency is relatively low. The density of the 5th data point is 0.1626. Assume that for each data point, 4 of its nearest neighbor data points are selected for calculation.
[0101] Calculate the LOF of the 1st data point (neighbors: the 2nd, 3rd, 4th, and 5th data points):
[0102] ;
[0103] Calculate the LOF of the th data point (neighbors: the 1st, 3rd, 4th, and 5th data points):
[0104] ;
[0105] Calculate the LOF of the th data point (neighbors: the 1st, 2nd, 4th, and 5th data points):
[0106] ;
[0107] Calculate the LOF of the th data point (neighbors: the 1st, 2nd, 3rd, and 5th data points):
[0108] ;
[0109] Calculate the LOF of the th data point (neighbors: the 1st, 2nd, 3rd, and 4th data points):
[0110] ;
[0111] The results show that the local outlier factor of the 1st data point is 1.031, the local outlier factor of the 2nd data point is 0.775, the local outlier factor of the 3rd data point is 0.889, the local outlier factor of the 4th data point is 1.147, and the local outlier factor of the 5th data point is 1.288.
[0112] Compare the local outlier factor with the preset access density range to determine whether there are memory access abnormal behaviors, including process behavior anomalies, memory access pattern mutations, and remote control tool interventions, and generate remote environment analysis results;
[0113] Perform a threshold judgment on the calculated local outlier factor The preset access density range is set based on statistical analysis methods. The specific execution process is as follows: Statistically analyze the LOF values of normal data points within the historical time window, and calculate the mean and standard deviation. For example, the normal data points in the past 30 days The mean is 1.02 and the standard deviation is 0.03. These data are obtained by recording and statistically analyzing the paging access behaviors during normal system operation, mainly analyzing the distribution of local outlier factors of historical data. Set the threshold, and use the mean + adjustment coefficient × standard deviation as the judgment basis. The selection of the adjustment coefficient also depends on the normal distribution principle in statistics and the fluctuation range of historical data: When the adjustment coefficient is taken as 2, 95% of the normal data is covered, that is: , if the data fluctuates greatly, the adjustment coefficient can be 3 to cover 99.7% of the normal data: , The specific selection of the adjustment coefficient is based on the statistical distribution of historical data, and can be optimized by observing the false positive rate and false negative rate of anomaly detection.
[0114] Assume different levels of anomaly determination thresholds:
[0115] Normal range: (mean + 2 × standard deviation), Slight anomaly range: (mean + 2 - 3 × standard deviation), Obvious anomaly range: (above mean + 3 × standard deviation).
[0116] Compare the calculated LOF value with the threshold, the 1st data point: , Compare with the threshold: , It belongs to the normal range and no further processing is required.
[0117] The second data point: , comparison threshold: , within the normal range, no further processing is required.
[0118] The third data point: , comparison threshold: , within the normal range, no further processing is required.
[0119] The fourth data point: , comparison threshold: , within the slightly abnormal range, it is necessary to further analyze the memory access behavior of this data point to determine whether it is a transient fluctuation or a persistent anomaly.
[0120] The fifth data point: , comparison threshold: , within the significantly abnormal range, it is necessary to focus on monitoring the process to which this data point belongs, and combine the CPU load, memory occupancy, and remote connection status to determine whether there are abnormal activities or remote intervention behaviors.
[0121] For data points, it is necessary to check the behavior patterns of the processes to which they belong, such as whether there are abnormal memory occupancies, abnormal CPU loads, or high-frequency system calls, and further analyze whether there are malicious behaviors or program errors causing the anomalies. Mutation in memory access pattern: If the access behavior of a data point within a certain time window changes suddenly compared to historical data, such as a sharp increase or decrease in access density within a short period, it may be necessary to further monitor whether there are abnormal paging patterns in the process where this data point is located, such as memory leaks or malicious code execution. Intervention by remote control tools: If the process to which an abnormal data point belongs has remote connection behaviors, such as remote desktop, remote management tools, etc., it is necessary to combine network traffic analysis to check whether there are abnormal remote instruction executions or remote data transmission behaviors. Finally, generate the remote environment analysis results, output the list of abnormal data points according to the determination results, record the corresponding processes, access patterns, and possible risk types, and perform trend analysis in combination with historical data to evaluate the security of system operation.
[0122] By calculating the density distribution of data accessed by memory paging and combining with the local outlier factor algorithm, it is possible to accurately identify abnormal access behaviors in a remote environment, improving the detection ability for process behavior mutations and interventions by remote control tools. The time window mechanism is adopted to enable the dynamic capture of memory access patterns, avoiding misjudgments caused by short-term fluctuations. At the same time, by statistically setting the outlier threshold based on historical LOF values, the stability and adaptability of anomaly detection are ensured. After classifying the anomaly levels, monitoring can be carried out for minor anomalies, and in-depth analysis can be conducted for obvious anomalies. For example, by combining information such as CPU load and remote connections, further identify whether there is malicious behavior or remote intrusion, improving the accuracy of remote environment security analysis.
[0123] Please refer to Figure 6 , the steps to obtain the user behavior analysis results are specifically as follows:
[0124] Query the historical user input commands during remote login interactions, including the command execution order and the combination method of command parameters. Remove duplicate inputs, auto-completion, and non-interactive commands for format adjustment to retain the actual user interaction operations, construct a historical command sequence library, extract the current user input commands during remote login interactions, and construct the current command sequence library;
[0125] First, it is necessary to extract the complete user input command sequence from the system log or terminal session record, filter out non-interactive input content caused by auto-completion, command format adjustment, etc., to ensure that the extracted commands only reflect the actual user interaction operations. Then, compare the list of commands input by the user, remove duplicate commands, use the timestamp sorting method to ensure the consistency of the command sequence execution order, and store them separately according to different users or different remote login sessions to construct a historical command sequence library. During the remote login interaction, the commands currently input by the user are extracted in real-time, and they are also standardized according to the execution order and parameter combination method, removing irrelevant information and storing them in the current command sequence library to ensure that the historical command sequence library and the current command sequence library have the same structure.
[0126] Use the longest common subsequence algorithm to analyze the length of the longest common subsequence of the command sequences in the current command sequence library and the historical command sequence library and judge their similarity. Compare the preset similarity range to judge whether there are abnormal command usage behaviors and generate user behavior analysis results;
[0127] Through recursive formula, analyze the current command sequence library and the historical command sequence library to obtain the length of the longest common subsequence of the command sequences , where, is the th command in the current command sequence library , Is the historical command sequence library The th command in Is to remove the th command from the current command sequence library Subsequence of Is to remove the th command from the historical command sequence library Subsequence of And Indicates that if the current command And Do not match, then calculate the length of the longest common subsequence after removing one element from the current sequence library Or the length of the longest common subsequence after removing one element from the historical sequence library And take the larger value among them, Is the maximum value function, used to select the maximum value, Indicates the th command of the current command sequence Is the same as the th command of the historical command sequence At this time, the length of the longest common subsequence increases by 1, Is when The longest common subsequence may come from two cases: 1. Ignore the th command of the current command sequence That is, calculate , 2. Ignore the th command of the historical command sequence That is, calculate , Take the larger value of the two as the current Final result.
[0128] Assume that the historical command sequence library Is:
[0129] ls - l, cd / var / log, cat auth.log, grep "error" auth.log, exit;
[0130] Assume that the current command sequence library Is:
[0131] ls - l, cd / var / log, cat auth.log, grep "warning" auth.log, exit;
[0132] Longest common subsequence calculation:
[0133] Compare "exit" and "exit", match ;
[0134] Compare "grep warning auth.log" and "grep error auth.log", do not match ;
[0135] Compare "cat auth.log" and "cat auth.log", match ;
[0136] Compare "cd / var / log" and "cd / var / log", match ;
[0137] Compare "ls -l" and "ls -l", match ;
[0138] The final calculation result is .
[0139] After calculating the length of the longest common subsequence the system compares the preset similarity range to determine whether the command usage is abnormal, using the formula:
[0140] ;
[0141] Calculate the similarity ;
[0142] where is the length of the current command sequence library, is the length of the historical command sequence library.
[0143] If the length of the current command sequence library: , the length of the historical command sequence library: , the length of the longest common subsequence: , then .
[0144] Similarity threshold setting includes statistical analysis based on historical data, analyzing the command input patterns of different users during remote login, calculating the command similarity distribution under normal usage of each user, and finding the similarity interval for normal operations. It also includes the possible reasonable minor changes during the command input process, such as certain parameter changes, path adjustments, etc. Therefore, an appropriate range needs to be set to ensure that normal operations are not misjudged as abnormal. It also includes adjusting the similarity threshold according to the security policies set by the system administrator to ensure that abnormal behaviors can be identified while not triggering too many false alarms due to minor operation differences. It also includes using the past remote login command logs for regression analysis, testing the false alarm rate and missed alarm rate under different threshold settings, and selecting the optimal similarity threshold.
[0145] Assume that the similarity threshold is set as follows: for normal behavior: (The similarity of most normal user behaviors is above 0.75), for minor anomalies: (Larger behavior changes, but still having a certain similarity with the historical command sequence), for obvious anomalies: (Greatly different from the historical operation mode, may be abnormal behavior), the current similarity is 0.8, greater than 0.75, determined as normal behavior, no further processing is required.
[0146] Among them, normal behavior means that the current input command sequence has a high similarity with the historical command sequence and the execution logic has not changed significantly. Minor anomalies mean that if the similarity decreases but does not completely deviate from the historical behavior, it is necessary to analyze whether it is a new task scenario or a short-term behavior change of the user, such as an adjustment of a certain command parameter but the overall command flow has not changed significantly. Obvious anomalies mean that if the similarity is too low, it indicates that the commands executed by the user are quite different from the historical behavior, and there may be abnormal usage behaviors or potential security risks, such as entering a completely different command set or having suspicious remote control instructions. For minor anomalies: The system records this behavior and compares it with the user's past command input patterns. If this behavior persists in multiple time windows, it may indicate a change in the user's operation habits and further monitoring is required. For obvious anomalies: Trigger a security alarm, analyze whether there are abnormal processes running, check for unknown remote access, and conduct further investigations in combination with other security logs (such as IP changes, geographical location anomalies). Finally, generate the user behavior analysis results, record the similarity of the current command sequence, anomaly classification, and provide further monitoring suggestions to ensure that the remote login behavior conforms to the historical operation mode and prevent abnormal behaviors or potential security threats.
[0147] By constructing a library of historical and current command sequences and using the longest common subsequence algorithm to calculate command similarity, the system can accurately identify abnormal situations in the command usage patterns during the remote login process. Compared with simply matching command strings, this method can consider factors such as command order and parameter combination methods, reduce false alarms, and at the same time can detect malicious tampering or abnormal command execution. By setting a similarity threshold based on historical data, the system can adapt to changes in the normal behavior of users, effectively distinguish normal, slightly abnormal, and significantly abnormal situations, and improve the detection accuracy of remote login abnormal behaviors. In addition, for slightly abnormal situations, the system can conduct long-term monitoring, while for significantly abnormal situations, it triggers a security alert and combines with other security logs for further analysis, thereby enhancing the security protection ability of remote login.
[0148] Please refer to Figure 7 , and the steps to obtain the remote login behavior recognition result are specifically as follows:
[0149] Based on the packet behavior analysis result, remote environment analysis result, and user behavior analysis result, through a multi-factor anomaly analysis method, evaluate the overall anomaly degree of the remote login interaction to obtain an anomaly degree evaluation result;
[0150] First, the packet behavior analysis result provides data such as packet size, interaction mode, and abnormal packet ratio. The remote environment analysis result provides memory paging access characteristics, process behavior patterns, and the possibility of remote control tool intervention. The user behavior analysis result provides features such as command sequence similarity and command execution mode. Standardize these data, remove outliers, and ensure the dimensional consistency of different types of data. Then, perform individual scoring on each analysis result. For example, the abnormal score of packet behavior is calculated based on the deviation degree of packet size distribution, abnormal packet ratio, etc. The abnormal score of the remote environment is measured based on the running characteristics of the process, paging access mode, and historical behavior differences. The abnormal score of user behavior is analyzed based on factors such as command similarity, change in execution order, and abnormal ratio of key operations. Each score ranges from low to high, and the higher the value, the more serious the abnormal degree. After that, sum up the weighted scores of each item, and the weights are determined by the statistical results of historical data. For example, the weight increases when the packet behavior contributes more in a high-traffic attack, and the weight is adjusted when the user behavior contributes more in a specific abnormal login event. Finally, comprehensively calculate the overall abnormal score as the basis for evaluating whether the remote login interaction is abnormal. The combination of packet behavior, remote environment, and user behavior can improve the accuracy of anomaly detection. For example, relying solely on packet characteristics may lead to misjudgment because some legitimate high-traffic operations may also trigger alarms, while combining user behavior analysis can verify whether the operation conforms to user habits. At the same time, remote environment analysis can provide supporting information on underlying processes and memory activities to help distinguish normal and abnormal sessions. Therefore, multi-factor comprehensive analysis can reduce false alarms and enhance the robustness of detection.
[0151] Remote session identification is performed based on the abnormal degree evaluation result, including identifying normal remote login behavior and abnormal remote attack behavior, and obtaining the remote login behavior identification result;
[0152] First, set the abnormal score threshold for remote session identification. Based on historical data analysis, determine the remote login behavior categories corresponding to different abnormal score ranges. For example, when the abnormal score is low, it is determined as normal; when the abnormal score is in the medium range, it is determined as suspicious behavior; when the abnormal score is high, it is determined as abnormal remote attack behavior. Compare the abnormal score of the current remote login interaction with the threshold. If the score is lower than the threshold, it is determined as normal remote login behavior; otherwise, enter further analysis. Analyze the source of abnormal data to determine whether the abnormality is mainly caused by packet behavior, remote environment, or user behavior, and verify it in combination with log data, system calls, process behavior, etc. For example, if the abnormal score of packet behavior is high, it may mean large-scale data transmission or scanning behavior; if the abnormal score of the remote environment is high, it may mean process abnormality or remote tool intervention; if the abnormal score of user behavior is high, it may mean that the input command mode deviates from historical behavior. Finally, the system classifies the remote session as a normal session, a suspicious session, or an abnormal session according to the comprehensive analysis result, records the judgment basis, and outputs the remote login behavior identification result.
[0153] Through the comprehensive analysis of packet, remote environment, and user behavior, the accuracy of remote login anomaly detection is improved, and false positives or false negatives caused by single-factor judgment are avoided. Through standardization processing and weighted scoring, different types of abnormal data can be evaluated under the same system, ensuring that the overall measurement of the abnormal degree is more objective and reasonable. Set the abnormal score threshold based on historical data statistics, enabling the system to dynamically adapt to different users' remote login behavior patterns and improving the ability to identify suspicious behavior. Combine log data, system calls, and process behavior to further verify the source of the anomaly, ensure the reliability of the anomaly determination, and provide clear classification results, so that the remote session can be accurately classified as normal, suspicious, or abnormal, providing strong support for security policy adjustment and intrusion prevention.
[0154] The above is only a preferred embodiment of the present invention and does not impose other forms of limitations on the present invention. Any person skilled in the art may use the disclosed technical content to make changes or modifications into equivalent embodiments with equivalent changes and apply them to other fields. However, as long as it does not depart from the technical content of the technical solution of the present invention, any simple modification, equivalent change, and modification made to the above embodiments based on the technical essence of the present invention still fall within the protection scope of the technical solution of the present invention.
Claims
1. A computer remote login recognition system, characterized in that, The system includes: The traffic monitoring module obtains packet data during the user's current remote login interaction, counts the occurrence probabilities of each packet size category in the packet data, and generates the probability distribution of the current packet size. The packet analysis module queries the historical occurrence probabilities of each packet size category, constructs the probability distribution of the historical packet size, determines abnormal packet behavior based on the probability distribution of the current packet size and the probability distribution of the historical packet size, and generates the packet behavior analysis result. The memory paging monitoring module obtains the memory paging data during the user's remote login interaction and divides it through a time window to generate a memory paging access sequence. The environment anomaly detection module regards each type of memory paging data in the memory paging access sequence as an independent data point, determines the memory access abnormal behavior by calculating the density distribution of each type of data point and its adjacent data points, and generates the remote environment analysis result. The command sequence analysis module obtains the historical and current user input commands during the remote login interaction, constructs the corresponding historical and current command sequence libraries, determines the abnormal command usage behavior by calculating the similarity of the historical and current command sequence libraries, and generates the user behavior analysis result. The security determination module performs remote session identification based on the overall anomaly degree of the packet behavior analysis result, the remote environment analysis result, and the user behavior analysis result to obtain the remote login behavior identification result. The specific steps for obtaining the packet behavior analysis result are as follows: Query each packet size category and its corresponding historical occurrence probability in the user's historical remote login interaction process, construct the probability distribution of the historical packet size, and calculate the divergence value between the probability distribution of the current packet size and the probability distribution of the historical packet size using KL divergence. Compare the divergence value with a preset divergence threshold. If the divergence value exceeds the divergence threshold, it is determined that the distribution of the corresponding packet size category is abnormal, and the abnormal packet size category and its packet behavior characteristics are extracted. The packet behavior characteristics include the packet size change trend, the abnormal packet ratio, and the packet interaction mode, and the packet behavior analysis result is generated. The divergence value between the probability distribution of the current packet size and the probability distribution of the historical packet sizes , and the calculation formula is as follows: ; Among them, is the current occurrence probability of the i-th type of packet size in the probability distribution of the current packet size, is the historical occurrence probability of the i-th type of packet size in the probability distribution of the historical packet size, is based on and The set adjustment parameter.
2. The computer remote login identification system according to claim 1, characterized in that The specific steps for obtaining the probability distribution of the current packet size are as follows: Collect the packet data during the user's current remote login interaction, parse the packet size information from the packet data and classify it according to different byte ranges to generate the packet data classification result. Based on the packet data classification result, count the occurrence probabilities of each packet size category after classification to generate the probability distribution of the current packet size.
3. The computer remote login identification system according to claim 1, characterized in that, The specific steps for obtaining the memory paging access sequence are as follows: Obtain the memory paging data during the user's remote login interaction, including the memory paging access times, the number of page faults, and the number of memory pages occupied by each process, and remove the non-interactive paging activities generated by background processes. The non-interactive paging activities include background running processes, log writing processes, and cache management processes, so as to retain the paging data related to user interaction and generate the processed memory paging data. Divide the processed memory paging data of different categories through a time window, including normal access paging, page fault access paging, and background process paging, to generate a memory paging access sequence.
4. The computer remote login identification system according to claim 1, characterized in that, The specific steps for obtaining the remote environment analysis result are as follows: Treat each type of memory paging data in the memory paging access sequence as an independent data point, calculate the density distribution of each type of data point under different time windows, and integrate to obtain the density distribution of each type of data point; Based on the density distribution of each type of data point, use the Local Outlier Factor algorithm to compare the density of each data point with the data points within its adjacent time window, and calculate the local outlier factor of the corresponding data point relative to the adjacent data points; Compare the local outlier factor with the preset access density range to determine whether there are memory access abnormal behaviors, including abnormal process behaviors, sudden changes in memory access patterns, and interventions by remote control tools, and generate remote environment analysis results.
5. The computer remote login identification system according to claim 1, characterized in that, The specific steps for obtaining the user behavior analysis results are as follows: Query the historical user input commands during the remote login interaction, including the command execution order and the command parameter combination method, remove the non-interactive commands such as repeated input, auto-completion, and format adjustment, so as to retain the actual user interaction operations, construct a historical command sequence library, extract the current user input commands during the remote login interaction, and construct the current command sequence library; Use the Longest Common Subsequence algorithm to analyze the length of the longest common subsequence of the command sequences in the current command sequence library and the historical command sequence library and judge their similarity, compare with the preset similarity range, determine whether there are abnormal command usage behaviors, and generate user behavior analysis results.
6. The computer remote login identification system according to claim 1, characterized in that, The specific steps for obtaining the remote login behavior recognition results are as follows: Based on the packet behavior analysis results, remote environment analysis results, and user behavior analysis results, use the multi-factor anomaly analysis method to evaluate the overall anomaly degree of the remote login interaction and obtain the anomaly degree evaluation results; Perform remote session recognition according to the anomaly degree evaluation results, including recognizing normal remote login behaviors and abnormal remote attack behaviors, and obtain remote login behavior recognition results.
Citation Information
Patent Citations
Internet data security protection method and system based on intelligent algorithm
CN119272339A
Computer network big data security protection method and system
CN119892504A