Computer remote login identification system

By integrating traffic monitoring, packet analysis, memory paging monitoring and command sequence analysis modules in the remote login recognition system, analyzing packet behavior, memory access and user command sequences during the remote login process, the problem of difficulty in identifying abnormal behavior in traditional systems is solved, and accurate identification and security improvement of remote login behavior is achieved.

CN120090877AActive Publication Date: 2025-06-03JIANGSU SHEHUITONG INTELLIGENT TECH CO LTD +1
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510562750.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-06-03
Estimated Expiration
2045-04-30

AI Technical Summary

Technical Problem

Traditional remote login identification systems are difficult to effectively identify abnormal behaviors during remote login, cannot detect dynamic behavior abnormalities, and there is a risk of identity credentials being tampered with or forged during transmission, making it difficult to identify changes in the remote environment.

Method used

A computer remote login recognition system is adopted to obtain and analyze packet data, memory paging data and user command sequence through traffic monitoring module, packet analysis module, memory paging data and user command sequence analysis module, and determine the abnormality of packet behavior, memory access and user behavior, and generate remote login behavior recognition results.

Benefits of technology

It realizes multi-dimensional analysis of remote login behavior, accurately identify normal remote login and abnormal remote attack behavior, improves remote login security, and avoids false alarms and missed reports caused by single indicator judgment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120090877A_ABST
    Figure CN120090877A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of remote identity authentication, in particular to a computer remote login recognition system. According to the invention, in the remote login interaction, the probability distribution of the packet size can be obtained through the statistical analysis of the packet data, and the abnormal traffic is detected through the comparison of historical data, so that the potential risk caused by large-scale data transmission or traffic mutation is prevented. By calculating the divergence value of the packet size distribution, the abnormal degree of the packet behavior can be accurately identified, and the rationality of the packet interaction mode is ensured. In combination with the analysis of the memory paging access sequence, the memory access mode is divided through the time window, and the influence of the background process is eliminated, so that the detection of the memory access abnormality is more accurate. Through application of a density distribution calculation method, environment abnormity judgment is not limited to a single data point, and the recognition capacity for process behavior mutation and remote control intervention is enhanced by combining access density changes of adjacent time windows.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of remote identity authentication, and particularly to a computer remote login recognition system. Background Art

[0002] A computer remote login recognition system refers to a system used to perform identity authentication during the computer remote login process to ensure that the remote user is a legally authorized entity. The system involves the acquisition, transmission, and comparison of user identity information, and usually adopts methods such as password input, biometric-based identity authentication, hardware token verification, or dynamic key-based identity confirmation to effectively identify the identity of remote users. During the identity information acquisition process, the system may involve specific means such as keyboard input, fingerprint scanning, iris recognition, or face recognition, and combines the transmission and storage mechanisms of identity credentials to prevent data tampering or forgery. In the identity verification link, the system can apply a one-time password generation mechanism triggered by time or events, or combine a public-private key encryption system to achieve accurate matching and authentication of the identity of remote users.

[0003] Traditional remote login recognition systems mainly rely on identity authentication methods for security determination. Although they can ensure the legality of the login entity, they cannot effectively identify abnormal behaviors during the remote login process. The identity authentication mechanism usually relies on static passwords, biometrics, or hardware tokens, but these methods cannot detect dynamic behavior abnormalities after remote login, resulting in the system being difficult to further distinguish abnormal operations once an attacker bypasses the authentication. There is a risk that the identity credentials of remote users may be tampered with or forged during transmission. Even with encryption mechanisms, it is difficult to completely prevent credential leakage problems. At the same time, the identity authentication process is usually independent of subsequent interaction behavior analysis, enabling attackers to execute malicious operations after passing the authentication by hijacking legitimate credentials, and the system lacks effective monitoring of subsequent behaviors. Traditional methods mainly rely on time-triggered or event-triggered dynamic key mechanisms, which can only prevent the direct reuse of credentials, but cannot identify changes in the remote environment. For example, when an attacker uses remote tools to operate the target system, execute abnormal commands, or hijack normal processes, the system is difficult to determine whether abnormal behaviors exist. Due to the lack of joint analysis of multi-dimensional data such as packets, memory, and command sequences, traditional methods have limited recognition capabilities in the face of complex remote attacks and are easily evaded by means such as forged traffic and hidden operations of remote tools, resulting in the inability to accurately distinguish normal remote logins from abnormal remote attack behaviors. Summary of the Invention

[0004] The purpose of the present invention is to solve the drawbacks existing in the prior art and propose a computer remote login recognition system.

[0005] To achieve the above purpose, the present invention adopts the following technical solutions: A computer remote login recognition system includes: The traffic monitoring module obtains the packet data during the user's current remote login interaction, counts the occurrence probabilities of each packet size category in the packet data, and generates the probability distribution of the current packet size; The packet analysis module queries the historical occurrence probabilities of each packet size category, constructs the probability distribution of the historical packet size, judges the abnormal packet behavior according to the probability distribution of the current packet size and the probability distribution of the historical packet size, and generates the packet behavior analysis result; The memory paging monitoring module obtains the memory paging data during the user's remote login interaction and divides it through a time window to generate a memory paging access sequence; The environment anomaly detection module regards each type of memory paging data in the memory paging access sequence as an independent data point, judges the memory access abnormal behavior by calculating the density distribution of each data point and its adjacent data points, and generates the remote environment analysis result; The command sequence analysis module obtains the historical and current user input commands during the remote login interaction, constructs the corresponding historical and current command sequence libraries, judges the abnormal command usage behavior by calculating the similarity of the historical and current command sequence libraries, and generates the user behavior analysis result; The security determination module performs remote session identification based on the overall anomaly degree of the packet behavior analysis result, the remote environment analysis result and the user behavior analysis result, and obtains the remote login behavior identification result.

[0006] As a further solution of the present invention, the steps for obtaining the probability distribution of the current packet size are specifically as follows: Collect the packet data during the user's current remote login interaction, parse the packet size information from the packet data and classify it according to different byte ranges to generate the packet data classification result; Based on the packet data classification result, count the occurrence probabilities of each packet size category after classification, and generate the probability distribution of the current packet size.

[0007] As a further solution of the present invention, the steps for obtaining the packet behavior analysis result are specifically as follows: Query each packet size category and its corresponding historical occurrence probability during the user's historical remote login interaction, construct the probability distribution of the historical packet size, and calculate the divergence value between the probability distribution of the current packet size and the probability distribution of the historical packet size by using KL divergence; Compare the divergence value with a preset divergence threshold. If the divergence value exceeds the divergence threshold, it is determined that the distribution of the corresponding packet size category is abnormal, and the abnormal packet size category and its packet behavior characteristics are extracted. The packet behavior characteristics include the packet size change trend, the abnormal packet ratio, and the packet interaction mode, and the packet behavior analysis result is generated.

[0008] As a further solution of the present invention, the step of obtaining the memory paging access sequence is specifically as follows: Obtain the memory paging data during the user's remote login interaction, including the memory paging access times, the number of page fault occurrences, and the number of memory pages occupied by each process, and remove the non-interactive paging activities generated by background processes. The non-interactive paging activities include background running processes, log writing processes, and cache management processes, so as to retain the paging data related to user interaction and generate the processed memory paging data; Divide the processed memory paging data of different categories through a time window, including normal access paging, page fault access paging, and background process paging, to generate a memory paging access sequence.

[0009] As a further solution of the present invention, the step of obtaining the remote environment analysis result is specifically as follows: Regard each type of memory paging data in the memory paging access sequence as an independent data point, calculate the density distribution of each type of data point under different time windows, and integrate to obtain the density distribution of each type of data point; Based on the density distribution of each type of data point, use the local outlier factor algorithm to compare the density of each data point with the data points within its adjacent time window, and calculate the local outlier factor of the corresponding data point relative to the adjacent data points; Compare the local outlier factor with the preset access density range to determine whether there are memory access abnormal behaviors, including abnormal process behaviors, sudden changes in memory access patterns, and interventions by remote control tools, and generate a remote environment analysis result.

[0010] As a further solution of the present invention, the step of obtaining the user behavior analysis result is specifically as follows: Query the historical user input commands during the remote login interaction, including the command execution order and the command parameter combination method, remove the non-interactive commands such as repeated input, auto-completion, and format adjustment, so as to retain the actual user interaction operations, construct a historical command sequence library, extract the current user input commands during the remote login interaction, and construct a current command sequence library; Use the longest common subsequence algorithm to analyze the length of the longest common subsequence of the command sequences in the current command sequence library and the historical command sequence library and judge their similarity, compare with the preset similarity range, and determine whether there are abnormal command usage behaviors to generate a user behavior analysis result.

[0011] As a further solution of the present invention, the step of obtaining the remote login behavior recognition result is specifically as follows: Based on the packet behavior analysis result, the remote environment analysis result, and the user behavior analysis result, evaluate the overall abnormal degree of the remote login interaction through a multi-factor abnormal analysis method to obtain an abnormal degree evaluation result; Remote session recognition is performed according to the evaluation result of the anomaly degree, including recognizing normal remote login behavior and abnormal remote attack behavior, and obtaining the remote login behavior recognition result.

[0012] Compared with the prior art, the advantages and positive effects of the present invention are as follows: In the present invention, in remote login interaction, the statistical analysis of packet data can obtain the probability distribution of packet sizes, and detect abnormal traffic through comparison with historical data, preventing potential risks caused by large-scale data transmission or traffic mutation. By calculating the divergence value of the packet size distribution, the anomaly degree of packet behavior can be accurately identified, ensuring the rationality of the packet interaction mode. Combining the analysis of the memory paging access sequence, the memory access mode is divided by time windows, excluding the influence of background processes, making the detection of memory access anomalies more accurate. The application of the density distribution calculation method enables the judgment of environmental anomalies not to be limited to individual data points, but to combine the changes in access density in adjacent time windows, enhancing the ability to identify process behavior mutations and remote control interventions. The longest common subsequence algorithm is used for command sequence comparison, making the detection of abnormal command usage patterns more refined, and capable of identifying deviations in command execution logic rather than simply matching the commands themselves. The introduction of the multi-factor anomaly analysis method realizes the comprehensive evaluation of packets, environment, and user behavior, improves the accuracy of remote session recognition, and avoids false positives and false negatives caused by single-index judgment. When comprehensively evaluating remote login behavior, different anomaly scoring weights can be dynamically adjusted according to historical statistical data, making the anomaly determination more flexible and capable of adapting to different usage scenarios. By comprehensively analyzing packet characteristics, memory access behavior, and command input patterns, normal remote login and abnormal remote attack behaviors can be accurately identified, enhancing the security of remote login. Brief Description of the Drawings

[0013] Figure 1 It is the system flow chart of the present invention; Figure 2 It is the flow chart of the traffic monitoring module of the present invention; Figure 3 It is the flow chart of the packet analysis module of the present invention; Figure 4 It is the flow chart of the memory paging monitoring module of the present invention; Figure 5 It is the flow chart of the environmental anomaly detection module of the present invention; Figure 6 It is the flow chart of the command sequence analysis module of the present invention; Figure 7 It is the flow chart of the security determination module of the present invention. Detailed Embodiment

[0014] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0015] Please refer to Figure 1 , a computer remote login recognition system includes: The traffic monitoring module obtains packet data during the current remote login interaction of the user, counts the occurrence probability of each packet size category in the packet data, and generates the probability distribution of the current packet size. The packet analysis module queries the historical occurrence probability of each packet size category, constructs the probability distribution of the historical packet size, judges the abnormal packet behavior according to the probability distribution of the current packet size and the probability distribution of the historical packet size, and generates the packet behavior analysis result. The memory paging monitoring module obtains the memory paging data during the user's remote login interaction and divides it through a time window to generate a memory paging access sequence. The environment anomaly detection module regards each type of memory paging data in the memory paging access sequence as an independent data point, and judges the memory access abnormal behavior by calculating the density distribution of each type of data point and adjacent data points, and generates the remote environment analysis result. The command sequence analysis module obtains the historical and current user input commands during the remote login interaction, constructs the corresponding historical and current command sequence libraries, judges the abnormal command usage behavior by calculating the similarity of the historical and current command sequence libraries, and generates the user behavior analysis result. The security determination module performs remote session recognition based on the overall anomaly degree of the packet behavior analysis result, the remote environment analysis result and the user behavior analysis result, and obtains the remote login behavior recognition result.

[0016] Please refer to Figure 2 , the steps of obtaining the probability distribution of the current packet size are specifically as follows: Collect the packet data during the current remote login interaction of the user, parse the packet size information from the packet data and classify it according to different byte ranges to generate the packet data classification result. During the remote login interaction process, the traffic monitoring module captures network packet data in real time, intercepts TCP / IP packets in transmission using data stream listening technology, caches the packet data in the local storage area, extracts the basic information of the packets through a protocol parsing tool, including but not limited to source IP address, destination IP address, port number, protocol type, packet size, etc. For the packet size information, the packets are classified and stored according to a preset byte range threshold. For example, the packet size is divided into multiple ranges such as 0 - 64 bytes, 65 - 512 bytes, 513 - 1024 bytes, 1025 - 2048 bytes, etc. At the same time, an index structure is established during the data storage process for efficient subsequent query and statistics. For the packet size classification and statistics process, an accumulative counting method is used to record the occurrence times of packets in each category. That is, when each packet arrives, the system determines the interval it belongs to according to its size and increments the counter in the corresponding interval. The statistical data can be stored in the local database or stored in real time using an efficient data structure (such as a hash table) and can be periodically refreshed to the file system.

[0017] Based on the packet data classification results, use the formula:

[0018] Calculate the current occurrence probability of the packet size of the th category in the packet data classification results , and count the occurrence probabilities of each packet size category after classification to generate the probability distribution of the current packet size; Where: represents the total number of occurrences of the packet size of the th category within the statistical time period in the packet data classification results, obtained by classifying and counting according to the packet size after real-time packet capture, represents the total number of all packets within the statistical time period in the packet data classification results, which is equal to the sum of the occurrence times of all packet categories.

[0019] Set a statistical time period, such as 10 minutes. During this time, capture all the packet data generated by remote login interactions, parse the packet size information, and classify and count according to the set byte range, recording the occurrence times of each packet size category and the total number of packets . Assume different packet size intervals and counts: 0 - 64 bytes: , 65 - 512 bytes: , 513 - 1024 bytes: .

[0020] Total number of packets: .

[0021] Calculate the probability of each packet category: , , , .

[0022] The result shows the probability distribution of packets of different sizes during the Telnet interaction. Among them, the packets with a size of 65 - 512 bytes have the highest proportion, reaching 48.61%. Followed by 513 - 1024 bytes, accounting for 25%. The smallest category is 1025 - 2048 bytes, only accounting for 9.72%. This data can be used to analyze the packet characteristics of Telnet behavior and provide support for subsequent anomaly detection and optimization strategies.

[0023] Through the statistical analysis of the packet sizes during the Telnet interaction, the occurrence probability of different categories of packets can be accurately obtained, forming the probability distribution of packet sizes to identify abnormal packet behaviors. Real-time capture of packet data through traffic monitoring technology and the use of protocol parsing tools to extract key packet information ensure the accuracy and integrity of packet classification. The preset byte range threshold enables packets to be classified and stored according to a unified standard, while the establishment of an index structure improves the efficiency of data query and statistics. The cumulative counting method is used for packet size statistics, enabling the distribution of packet categories to be updated at any time and large-scale data processing to be completed in a short time. Through probability calculation, the distribution of each packet category during the Telnet interaction can be visually presented to judge the normal mode of packet traffic.

[0024] Please refer to Figure 3 to obtain the specific steps for the analysis result of packet behavior: Query each packet size category and its corresponding historical occurrence probability in the user's historical Telnet interaction process to construct the probability distribution of historical packet sizes; Extract historical packet data from the Telnet interaction log, query the packet size category information in the user's past interactions in a hierarchical index manner, and obtain the corresponding historical occurrence probability through database retrieval or time series storage structure to construct the probability distribution of historical packet sizes. During the query process, first retrieve the Telnet log according to the user's unique identifier, extract the size information of each packet in the historical interaction record, and classify and store them according to the set packet size interval. The historical probability value of each category is obtained by statistically calculating the ratio of the number of packets of each category appearing in the historical data to the total number of packets. The formula is:

[0025] where is the historical occurrence probability of the packet size of the rd category, is the The total number of occurrences of a class in historical data, is the total number of all packets within a historical time period. During the statistical process, to ensure data integrity, a sliding time window is used to filter historical data. For example, the most recent 30 days are set as the reference time range, and the window length is adjusted according to the user's historical behavior pattern. Finally, a complete probability distribution of historical packet sizes is obtained.

[0026] The KL divergence formula is adopted: ; Calculate the probability distribution of the current packet size and the probability distribution of historical packet sizes to obtain the divergence value ; Among them, is the current occurrence probability of the packet size of the th class in the probability distribution of the current packet size, is the historical occurrence probability of the packet size of the th class in the probability distribution of historical packet sizes, is the adjustment parameter set according to and to reflect the overall deviation degree between the current and historical distributions. The calculation formula is: . By introducing the parameter , the sensitivity of the KL divergence calculation can be dynamically adjusted, weighted according to the deviation degree of the overall probability distribution, making the anomaly detection more sensitive during large-scale changes, and reducing the false positive rate and improving the stability of the algorithm under normal fluctuations. represents the summation calculation for all packet size categories.

[0027] Assume the current probability and the historical probability are as follows: 0 - 64 bytes: , (difference 0.03), 65 - 512 bytes: , (difference 0.04), 513 - 1024 bytes: , (difference 0.02), 1025 - 2048 bytes: , (difference 0.01).

[0028] Calculate : ; ; .

[0029] Calculate the divergence value: ; First, calculate terms: , , , ; Calculate the product of each term: , , , .

[0030] Finally, sum up: .

[0031] This result indicates that the divergence value between the current packet size distribution and the historical packet size distribution is 0.00275.

[0032] Compare the divergence value with the preset divergence threshold. If the divergence value exceeds the divergence threshold, it is determined that the distribution of the corresponding packet size category is abnormal, and the abnormal packet size category and its packet behavior characteristics are extracted. The packet behavior characteristics include the change trend of the packet size, the proportion of abnormal packets, and the packet interaction mode, and a packet behavior analysis result is generated; Compare the calculated divergence value , with the preset divergence threshold . Among them, the threshold is statistically obtained by the system based on historical packet traffic data. The specific calculation method is as follows: Construct the KL divergence distribution of normal packet traffic: Select multiple time windows (for example, the past 30 days, calculate the KL divergence value once a day), and obtain a sequence of historical KL divergence values; Calculate the mean and standard deviation of this group of historical KL divergence values. Determine the divergence threshold : Use statistical analysis methods to set the threshold. Usually, the mean plus a certain multiple of the standard deviation is used as the boundary; The choice of the adjustment coefficient is usually 2 or 3, specifically based on the normal distribution principle in statistics: When the adjustment coefficient is 2, the threshold range covers approximately 95.4% of the normal data, and values outside this range may be abnormal; When the adjustment coefficient is 3, the threshold range covers approximately 99.7% of the normal data, and values outside this range are very likely to be abnormal; The setting range is generally between 1.5 and 3.0, and the specific value can be adjusted according to the stability of the historical data. When the data fluctuation is small, a smaller value is taken, and when the data fluctuation is large, a larger value is taken. If exceeds the threshold , it is determined that the packet size category distribution is abnormal; If , it is considered that the packet size category distribution is normal.

[0033] In this calculation, 0.00275 is greater than the threshold value of 0.0025. Therefore, the packet size category of this remote login interaction is determined to be abnormal. After anomaly detection, the packet behavior characteristics of the abnormal category are extracted, including the packet size change trend, the abnormal packet ratio, and the packet interaction pattern: Packet size change trend: The sliding average value of the packet size within a fixed time window is statistically calculated, the change between the current window mean value and the previous window mean value is calculated, and whether there is a sudden increase or decrease is analyzed; Abnormal packet ratio: The proportion of the number of abnormal category packets is calculated, indicating the proportion of the number of abnormal packets in the total number of packets. If the proportion exceeds the set abnormal proportion threshold, the traffic pattern of this category of packets is further analyzed; Packet interaction pattern analysis: Information such as the source IP, destination IP, protocol type, and packet interval time of the packet is extracted, and pattern matching is performed to detect whether there is abnormal communication behavior. Finally, the packet analysis module generates the packet behavior analysis result, and can further combine the historical anomaly records to evaluate whether the abnormal behavior has persistence or potential security risks.

[0034] By comparing the current and historical packet size distributions, the abnormal changes in packet traffic can be accurately identified, avoiding misjudgments based solely on single data fluctuations. The sliding time window is used to filter historical data, enabling anomaly detection to adapt to the long-term usage habits of different users and improving the stability of detection. The KL divergence is introduced to calculate the deviation degree of the packet distribution, making the packet anomaly detection no longer limited to a fixed threshold, but dynamically adjusted in combination with the statistical characteristics of the overall packet traffic, thereby enhancing the sensitivity to abnormal behaviors and reducing the false alarm rate caused by normal fluctuations. By setting statistical methods to determine the divergence threshold, the anomaly determination is made more scientific, reducing the interference caused by environmental changes or short-term traffic fluctuations. In addition, the extraction of the abnormal packet behavior characteristics, including the packet size change trend, the abnormal packet ratio, and the packet interaction pattern, makes the detection of abnormal traffic not only limited to the data statistics level, but also combines the traffic behavior characteristics, further improving the accuracy of abnormal packet analysis and providing more reliable data support for subsequent remote environment analysis and user behavior analysis.

[0035] Please refer to Figure 4 , and the steps to obtain the memory paging access sequence are specifically as follows: Obtain the memory paging data during the user's remote login interaction, including the memory paging access times, the number of paging faults, and the number of memory pages occupied by each process. Remove the non-interactive paging activities generated by background processes. The non-interactive paging activities include background running processes, log writing processes, and cache management processes, so as to retain the paging data related to user interaction and generate the processed memory paging data; First, the Memory Management Unit (MMU) records the memory paging access situation of each process, including the number of paging accesses, the number of paging faults, and the number of memory pages occupied by each process. Regarding the number of paging accesses, the system obtains the access frequency of each page by periodically scanning the access bit flags of page table entries and stores the access records in the kernel data structure. For paging faults, the system counts the number of page faults through the page fault exception handling mechanism of the processor and differentiates the types of page faults, such as main memory page faults and disk page faults. At the same time, to avoid the interference of background processes on interactive paging activities, the system filters the paging data generated by background processes, including non-interactive processes (such as log writing processes, cache management processes). The specific method is to record the list of active processes through the process scheduler and determine whether a process is a background task based on the process priority and scheduling policy. If the process running state is in the background mode or there has been no user interaction for a long time, the paging activities generated by this process are excluded from the monitoring data. Finally, it is ensured that the memory paging monitoring module only retains the paging data related to user interaction for subsequent analysis.

[0036] The processed memory paging data of different categories is divided through time windows, including normal access paging, page fault access paging, and background process paging, to generate a memory paging access sequence. The time window mechanism is used to divide the memory paging data of different categories. First, the time window length is set, such as 1 second, 10 seconds, or 1 minute, and the paging access situation of each process is counted within this window. The paging operations occurring within the time window are recorded as time series data. The division of the time window is triggered by the CPU clock interrupt signal. At the end of each time slice, the system collects the paging data within the current time window and stores it in the paging access log. At the same time, for different categories of paging data, the system classifies and stores them according to the paging access type, such as normal access paging, page fault access paging, and background process paging, etc. The classification standard is determined based on the trigger source of the paging event. For example, if the paging access is actively triggered by a user-mode process, it is classified as normal access paging; if the paging access is triggered by a page fault exception, it is classified as page fault access paging; and background process paging is identified and excluded through process scheduling information. Finally, the paging data within all time windows is integrated into a memory paging access sequence for subsequent memory access pattern analysis and optimization.

[0037] By recording the memory paging access situation and excluding the non-interactive paging activities of background processes, the analysis data can be made more accurate, avoiding misjudgment caused by non-user operations. Different types of paging data are divided by time windows, enabling the memory access pattern to be analyzed in a serialized manner, improving the traceability and temporal correlation of the data. By classifying and storing paging types, ordinary access, page fault access, and background process paging can be distinguished, providing a more fine-grained monitoring basis for anomaly detection, enabling the system to accurately identify abnormal memory access behaviors, and combining with remote environment analysis to improve the reliability of overall anomaly detection.

[0038] Please refer to Figure 5 , and the steps to obtain the remote environment analysis results are specifically as follows: Regarding each type of memory paging data in the memory paging access sequence as an independent data point, using the formula: ; Calculate the density of the th data point in the memory paging access sequence within the time window , and integrate to obtain the density distribution of various data points; Among them, is the number of accesses of the th data point in the memory paging access sequence within the corresponding time window, which can be obtained by sampling the memory access log, recording the paging access times of each process within each time window, and counting all access requests, is the time window of the memory paging access sequence, such as 10 seconds, 30 seconds, or 60 seconds, which determines the range of density calculation. The window length can be selected through experimental analysis, for example, by monitoring the memory paging trends of multiple time windows and selecting the window with the most obvious change as the analysis benchmark, represents all data points within the time window , indicating the total number of accesses within this time window, which represents the overall access volume within the time window and is obtained by accumulating all paging requests within the time window.

[0039] Assume that the set time window is 10 seconds, and multiple memory paging access data points are recorded within each window, The access count of the first data point , the access count of the second data point , the access count of the third data point , the access count of the fourth data point , the access count of the fifth data point .

[0040] Calculate the total access count: .

[0041] Calculate the density of each data point: , , , , .

[0042] The results show that the density of the first data point is 0.1951, indicating that its access frequency is at a medium level within this time window. The density of the second data point is 0.2439, which is the data point with the highest access frequency within this window, indicating that it has the most paging accesses. The density of the third data point is 0.2195, second only to the second data point, and still belongs to the data points with a relatively high access frequency. The density of the fourth data point is 0.1789, lower than the previous data points, indicating that its access frequency is relatively low. The density of the fifth data point is 0.1626, which is the data point with the lowest access frequency within this window, indicating that it has the fewest paging accesses.

[0043] Based on the density distribution of various data points, the Local Outlier Factor (LOF) algorithm is used to compare the density of each data point with that of the data points within its adjacent time window, using the formula: ; Calculate the local outlier degree of the th data point in the density distribution of various data points to obtain the local outlier factor of the corresponding data point relative to its adjacent data points; where is the density of the th data point within the time window , is the density of the th data point adjacent to the th data point within the time window , represents the total number of nearest neighbors selected for the th data point within the time window, usually screened based on the similarity of access behaviors within the time window. For example, if and , then

[0044] For example, the density of the first data point is 0.1951, indicating that its access frequency is at a medium level within this time window. The density of the second data point is 0.2439, which is the data point with the highest access frequency within this window, indicating that it has the most paging accesses. The density of the third data point is 0.2195, second only to the second data point, and still belongs to the data points with a relatively high access frequency. The density of the fourth data point is 0.1789, lower than the previous data points, indicating that its access frequency is relatively low. The density of the fifth data point is 0.1626. Assume that for each data point, its 4 nearest neighbor data points are selected for calculation.

[0045] Calculate the LOF of the first data point (neighbors: the second, third, fourth, and fifth data points): ; Calculate the LOF of the th data point (neighbors: the first, third, fourth, and fifth data points): ; Calculate the LOF of the th data point (neighbors: the first, second, fourth, and fifth data points): ; Calculate the LOF of the th data point (neighbors: the first, second, third, and fifth data points): ; Calculate the LOF of the th data point (neighbors: the first, second, third, and fourth data points): ; The results show that the local outlier factor of the first data point is 1.031, the local outlier factor of the second data point is 0.775, the local outlier factor of the third data point is 0.889, the local outlier factor of the fourth data point is 1.147, and the local outlier factor of the fifth data point is 1.288.

[0046] Compare the local outlier factor with the preset access density range to determine whether there are memory access abnormal behaviors, including process behavior abnormalities, sudden changes in memory access patterns, and remote control tool interventions, and generate remote environment analysis results; Perform threshold judgment on the calculated local outlier factor The preset access density range is set based on statistical analysis methods. The specific execution process is as follows: Statistically analyze the LOF values of normal data points within the historical time window, and calculate the mean and standard deviation. For example, the normal data points in the past 30 days The mean is 1.02 and the standard deviation is 0.03. These data are obtained by recording and statistically analyzing the paging access behavior during normal system operation, mainly analyzing the distribution of local anomaly factors in historical data. Set a threshold and use the mean + adjustment coefficient × standard deviation as the judgment basis. The selection of the adjustment coefficient also depends on the normal distribution principle in statistics and the fluctuation range of historical data: when taking 2 as the adjustment coefficient, 95% of the normal data is covered, that is: , if the data fluctuates greatly, the adjustment coefficient can be 3 to cover 99.7% of the normal data: , the specific selection of the adjustment coefficient is based on the statistical distribution of historical data and can be optimized by observing the false alarm rate and miss rate of anomaly detection.

[0047] Assume different levels of anomaly judgment thresholds: Normal range: (mean + 2 × standard deviation), Slight anomaly range: (mean + 2 - 3 × standard deviation), Obvious anomaly range: (above mean + 3 × standard deviation).

[0048] Compare the calculated LOF value with the threshold. The first data point: , Compare with the threshold: , It belongs to the normal range and no further processing is required.

[0049] The second data point: , Compare with the threshold: , It belongs to the normal range and no further processing is required.

[0050] The third data point: , Compare with the threshold: , It belongs to the normal range and no further processing is required.

[0051] The fourth data point: , Compare with the threshold: , It belongs to the slight anomaly range and the memory access behavior of this data point needs to be further analyzed to determine whether it is a short-term fluctuation or a persistent anomaly.

[0052] The fifth data point: , Compare with the threshold: , It belongs to the obvious anomaly range and the process to which this data point belongs needs to be monitored keyly. Combining the CPU load, memory occupancy and remote connection situation, determine whether it involves abnormal activities or remote intervention behaviors.

[0053] For For the data points, it is necessary to check the behavior patterns of the processes they belong to. For example, check for abnormal memory occupancy, abnormal CPU load, or high-frequency system calls, and further analyze whether malicious behavior or program errors have caused the anomalies. Mutation in memory access pattern: If the access behavior of a data point within a certain time window mutates compared to the historical data, such as a sharp increase or decrease in access density within a short period, it may be necessary to further monitor whether the process where the data point is located has an abnormal paging pattern, such as memory leakage or malicious code execution. Intervention by remote control tools: If the process to which the abnormal data point belongs has remote connection behaviors, such as remote desktop, remote management tools, etc., it is necessary to combine network traffic analysis to check for abnormal remote instruction execution or remote data transmission behaviors. Finally, generate the analysis results of the remote environment, output the list of abnormal data points according to the judgment results, record the corresponding processes, access patterns, and possible risk types, and conduct trend analysis in combination with historical data to evaluate the security of system operation.

[0054] By calculating the density distribution of memory paging access data and combining with the local outlier factor algorithm, it is possible to accurately identify abnormal access behaviors in the remote environment and improve the detection ability for process behavior mutations and interventions by remote control tools. The time window mechanism is adopted to enable the dynamic capture of memory access patterns, avoiding misjudgments caused by short-term fluctuations. At the same time, by statistically setting the outlier threshold based on historical LOF values, the stability and adaptability of anomaly detection are ensured. After classifying the anomaly levels, it is possible to monitor minor anomalies and conduct in-depth analysis of obvious anomalies. For example, by combining information such as CPU load and remote connections, further identify whether there is malicious behavior or remote intrusion, improving the accuracy of remote environment security analysis.

[0055] Please refer to Figure 6 , and the steps to obtain the user behavior analysis results are specifically as follows: Query the historical user input commands during remote login interactions, including the command execution order and the combination method of command parameters. Remove non-interactive commands such as duplicate inputs, auto-completion, and format adjustments to retain the actual user interaction operations, build a historical command sequence library, extract the current user input commands during remote login interactions, and build the current command sequence library; First, it is necessary to extract the complete sequence of user input commands from the system logs or terminal session records, filter out non-interactive input content caused by auto-completion, command format adjustment, etc., to ensure that the extracted commands only reflect the actual interactive operations of the user. Then, compare the list of commands entered by the user, remove duplicate commands, use the timestamp sorting method to ensure the consistency of the execution order of the command sequence, and store them separately according to different users or different remote login sessions to build a historical command sequence library. During the remote login interaction, the commands entered by the current user are extracted in real-time, and are also standardized according to the execution order and parameter combination method, removing irrelevant information, and stored in the current command sequence library to ensure that the historical command sequence library and the current command sequence library have the same structure.

[0056] Use the longest common subsequence algorithm to analyze the length of the longest common subsequence of the command sequences in the current command sequence library and the historical command sequence library and judge their similarity, compare the preset similarity range, judge whether there are abnormal command usage behaviors, and generate user behavior analysis results; Through the recurrence formula, analyze the current command sequence library and the historical command sequence library for the length of the longest common subsequence of the command sequences, where is the nth command in the current command sequence library , is the mth command in the historical command sequence library is the subsequence of the current command sequence library after removing the nth command, is the subsequence of the historical command sequence library after removing the and mean that if the current command and do not match, then calculate the length of the longest common subsequence after removing one element from the current sequence library , or the length of the longest common subsequence after removing one element from the historical sequence library , and take the larger value among them. is the maximum value function used to select the maximum value. represents that the nth command in the current command sequence is the same as the mth command in the historical command sequence, and at this time the length of the longest common subsequence increases by 1. When calculating the longest common subsequence, there are two possible cases: 1. Ignore the nth command in the current command sequence, that is, calculate , 2. Ignore the nth command in the historical command sequence, that is, calculate , and take the larger value of the two as the final result of the current .

[0057] Suppose the historical command sequence library is: ls -l, cd / var / log, cat auth.log, grep "error" auth.log, exit; Suppose the current command sequence library is: ls -l, cd / var / log, cat auth.log, grep "warning" auth.log, exit; Calculation of the longest common subsequence: Compare "exit" and "exit", they match ; Compare "grep warning auth.log" and "grep error auth.log", they do not match ; Compare "cat auth.log" and "cat auth.log", they match ; Compare "cd / var / log" and "cd / var / log", they match ; Compare "ls -l" and "ls -l", they match ; The final calculation result is .

[0058] After calculating the length of the longest common subsequence , the system compares it with the preset similarity range to determine whether the command usage is abnormal, using the formula: ; Calculate the similarity ; Among them, is the length of the current command sequence library, is the length of the historical command sequence library.

[0059] If the length of the current command sequence library: , the length of the historical command sequence library: , the length of the longest common subsequence: , then .

[0060] The setting of the similarity threshold includes statistical analysis based on historical data, analyzing the command input patterns of different users during remote login, calculating the command similarity distribution under the normal usage of each user, and finding the similarity interval of normal operations. It also includes the possible reasonable minor changes during the command input process, such as certain parameter changes, path adjustments, etc. Therefore, an appropriate range needs to be set to ensure that normal operations are not misjudged as abnormal. It also includes adjusting the similarity threshold according to the security policies set by the system administrator to ensure that abnormal behaviors can be identified and at the same time, too many false alarms are not triggered due to minor operation differences. It also includes using the past remote login command logs for regression analysis, testing the false alarm rate and missed alarm rate under different threshold settings, and selecting the best similarity threshold.

[0061] Assume that the similarity threshold is set as, normal behavior: (the similarity of most normal user behaviors is above 0.75), minor abnormality: (significant behavior changes, but still having a certain similarity with the historical command sequence), obvious abnormality: (great difference from the historical operation mode, may be an abnormal behavior). The current similarity is 0.8, which is greater than 0.75, so it is determined as normal behavior and no further processing is required.

[0062] Among them, normal behavior means that the current input command sequence has a high similarity with the historical command sequence, and the execution logic has not changed significantly. Slight anomalies occur when the similarity decreases but does not completely deviate from historical behavior. In this case, it is necessary to analyze whether it is a new task scenario or a short-term change in the user's behavior. For example, although a certain command parameter is adjusted, the overall command flow has not changed significantly. Obvious anomalies occur when the similarity is too low, indicating that the commands executed by the user are quite different from historical behavior, and there may be abnormal usage behaviors or potential security risks. For example, a completely different command set is entered, or there are suspicious remote control instructions. For slight anomalies: The system records this behavior and compares it with the user's past command input patterns. If this behavior persists in multiple time windows, it may indicate a change in the user's operation habits and further monitoring is required. For obvious anomalies: A security alert is triggered, analyze whether there are abnormal processes running, check for unknown remote access, and conduct further investigations in combination with other security logs (such as IP changes and geographical location anomalies). Finally, generate the user behavior analysis results, record the similarity of the current command sequence, the anomaly classification, and provide suggestions for further monitoring to ensure that the remote login behavior conforms to the historical operation mode and prevent abnormal behaviors or potential security threats.

[0063] By constructing the historical and current command sequence libraries and using the longest common subsequence algorithm to calculate the command similarity, the system can accurately identify abnormal situations in the command usage patterns during remote login. Compared with simply matching command strings, this method can consider factors such as command order and parameter combination methods, reduce false alarms, and at the same time can detect malicious tampering or abnormal command executions. By setting a similarity threshold based on historical data, the system can adapt to the normal behavior changes of users and effectively distinguish normal, slight anomalies, and obvious anomalies, improving the detection accuracy of remote login abnormal behaviors. In addition, for slight anomalies, the system can conduct long-term monitoring, while for obvious anomalies, it triggers a security alert and conducts further analysis in combination with other security logs, thereby enhancing the security protection ability of remote login.

[0064] Please refer to Figure 7 , the steps to obtain the remote login behavior recognition results are specifically as follows: Based on the packet behavior analysis results, remote environment analysis results, and user behavior analysis results, through a multi-factor anomaly analysis method, evaluate the overall anomaly degree of the remote login interaction to obtain the anomaly degree evaluation result; First, the packet behavior analysis results provide data such as packet size, interaction mode, and abnormal packet ratio. The remote environment analysis results provide characteristics of memory paging access, process behavior patterns, and the possibility of intervention by remote control tools. The user behavior analysis results provide characteristics such as command sequence similarity and command execution mode. These data are standardized to remove outliers and ensure the dimensional consistency of different types of data. Then, each analysis result is scored individually. For example, the abnormal score of packet behavior is calculated based on the deviation degree of packet size distribution, abnormal packet ratio, etc. The abnormal score of the remote environment is measured based on the running characteristics of the process, paging access mode, and historical behavior differences. The abnormal score of user behavior is analyzed based on factors such as command similarity, change in execution order, and abnormal ratio of key operations. Each score ranges from low to high, and the higher the value, the more serious the abnormal degree. After that, the scores are weighted and summed, and the weights are determined by the statistical results of historical data. For example, the weight increases when packet behavior contributes more in high-traffic attacks, and the weight is adjusted when user behavior contributes more in specific abnormal login events. Finally, the overall abnormal score is calculated comprehensively and used as the basis for evaluating whether the remote login interaction is abnormal. The combination of packet behavior, remote environment, and user behavior can improve the accuracy of anomaly detection. For example, relying solely on packet characteristics may lead to misjudgment because some legitimate high-traffic operations may also trigger alarms. Combining user behavior analysis can verify whether the operation conforms to user habits. At the same time, remote environment analysis can provide supporting information on underlying processes and memory activities to help distinguish normal and abnormal sessions. Therefore, multi-factor comprehensive analysis can reduce false alarms and enhance the robustness of detection.

[0065] Remote session identification is performed according to the abnormal degree evaluation results, including identifying normal remote login behavior and abnormal remote attack behavior, and obtaining the remote login behavior identification result; First, set the abnormal score threshold for remote session identification. Based on historical data analysis, determine the remote login behavior categories corresponding to different abnormal score ranges. For example, when the abnormal score is relatively low, it is determined as normal; when the abnormal score is in the medium range, it is determined as a suspicious behavior; when the abnormal score is relatively high, it is determined as an abnormal remote attack behavior. Compare the abnormal score of the current remote login interaction with the threshold. If the score is lower than the threshold, it is determined as a normal remote login behavior; otherwise, enter further analysis. Analyze the source of abnormal data to determine that the abnormality is mainly caused by packet behavior, remote environment, or user behavior, and verify it in combination with log data, system calls, process behavior, etc. For example, if the abnormal score of packet behavior is relatively high, it may mean large-scale data transmission or scanning behavior; if the abnormal score of the remote environment is relatively high, it may mean process abnormality or remote tool intervention; if the abnormal score of user behavior is relatively high, it may mean that the input command mode deviates from historical behavior. Finally, the system classifies the remote session as a normal session, a suspicious session, or an abnormal session according to the comprehensive analysis result, records the basis for determination, and outputs the recognition result of the remote login behavior.

[0066] Through the comprehensive analysis of packet, remote environment, and user behavior, the accuracy of remote login anomaly detection is improved, avoiding false positives or false negatives that may be caused by single-factor judgment. Through standardization processing and weighted scoring, different types of abnormal data can be evaluated under the same system, ensuring that the measurement of the overall abnormal degree is more objective and reasonable. Set the abnormal score threshold based on historical data statistics, enabling the system to dynamically adapt to different users' remote login behavior patterns and improving the ability to identify suspicious behaviors. Combine log data, system calls, and process behavior to further verify the source of the anomaly, ensure the reliability of the anomaly determination, and provide clear classification results, enabling remote sessions to be accurately classified as normal, suspicious, or abnormal, providing strong support for security policy adjustment and intrusion prevention.

[0067] The above is only a preferred embodiment of the present invention and does not limit the present invention in other forms. Any person skilled in the art may use the disclosed technical content to make changes or modifications into equivalent embodiments with equivalent changes and apply them to other fields. However, as long as it does not depart from the technical solution content of the present invention, any simple modification, equivalent change, and modification made to the above embodiments based on the technical essence of the present invention still fall within the protection scope of the technical solution of the present invention.

Claims

1. A computer remote login identification system, characterized in that: The system comprises: The traffic monitoring module obtains the packet data during the user's current remote login interaction, counts the occurrence probability of each packet size category in the packet data, and generates a probability distribution of the current packet size; The packet analysis module queries the historical occurrence probability of each packet size category, constructs the probability distribution of historical packet sizes, determines abnormal packet behavior based on the probability distribution of current packet sizes and the probability distribution of historical packet sizes, and generates packet behavior analysis results; The memory paging monitoring module obtains the memory paging data during the user's remote login interaction and divides it into time windows to generate a memory paging access sequence; The environment anomaly detection module regards each type of memory paging data in the memory paging access sequence as an independent data point, determines the abnormal behavior of memory access by calculating the density distribution of each type of data point and adjacent data points, and generates remote environment analysis results; The command sequence analysis module obtains the historical and current user input commands during the remote login interaction, builds the corresponding historical and current command sequence libraries, determines abnormal command usage behavior by calculating the similarity between the historical and current command sequence libraries, and generates user behavior analysis results; The security determination module performs remote session recognition based on the overall abnormality of the packet behavior analysis results, the remote environment analysis results and the user behavior analysis results to obtain a remote login behavior recognition result.

2. The computer remote login identification system according to claim 1, characterized in that: The steps for obtaining the probability distribution of the current packet size are specifically as follows: Collecting packet data during the user's current remote login interaction, parsing packet size information from the packet data and classifying them according to different byte ranges, and generating packet data classification results; Based on the packet data classification result, the occurrence probability of each packet size category after classification is statistically analyzed to generate a probability distribution of the current packet size.

3. The computer remote login identification system according to claim 1, characterized in that: The steps for obtaining the packet behavior analysis result are specifically as follows: Query each packet size category and the corresponding historical occurrence probability during the user's historical remote login interaction process, construct a probability distribution of historical packet sizes, and use KL divergence calculation to obtain a divergence value between the probability distribution of the current packet size and the probability distribution of the historical packet size; The divergence value is compared with a preset divergence threshold. If the divergence value exceeds the divergence threshold, the distribution of the corresponding packet size category is determined to be abnormal, and the abnormal packet size category and its packet behavior characteristics are extracted. The packet behavior characteristics include packet size change trend, abnormal packet ratio, and packet interaction mode, and the packet behavior analysis results are generated.

4. The computer remote login identification system according to claim 3, characterized in that: The divergence value between the probability distribution of the current packet size and the probability distribution of the historical packet size , the calculation formula is: ; in, is the probability distribution of the current packet size. The current probability of occurrence of the class packet size, is the probability distribution of historical packet sizes. The historical probability of occurrence of class packet size, is based on and Set the adjustment parameters.

5. The computer remote login identification system according to claim 1, characterized in that: The steps of obtaining the memory paging access sequence are specifically as follows: Obtain memory paging data during user remote login interaction, including the number of memory paging accesses, the number of paging faults, and the number of memory paging occupied by each process, and remove non-interactive paging activities generated by background processes. Non-interactive paging activities include background running processes, log writing processes, and cache management processes, so as to retain paging data related to user interaction and generate processed memory paging data; The processed memory paging data of different categories are divided by time windows, including common access paging, page-missing access paging and background process paging, to generate a memory paging access sequence.

6. The computer remote login identification system according to claim 1, characterized in that: The steps for obtaining the remote environment analysis results are specifically as follows: Treat each type of memory paging data in the memory paging access sequence as an independent data point, calculate the density distribution of each type of data point in different time windows, and integrate to obtain the density distribution of each type of data point; Based on the density distribution of the various data points, a local anomaly factor algorithm is used to compare the density of each data point with the data points in the adjacent time window, and the local anomaly factor of the corresponding data point relative to the adjacent data points is calculated; The local abnormality factor is compared with a preset access density range to determine whether there is abnormal memory access behavior, including abnormal process behavior, sudden change in memory access mode, and intervention by remote control tools, and a remote environment analysis result is generated.

7. The computer remote login identification system according to claim 1, characterized in that: The steps for obtaining the user behavior analysis results are specifically as follows: Query the historical user input commands during remote login interaction, including command execution order and command parameter combination mode, remove duplicate input, auto-complete, and format non-interactive commands to retain the user's actual interactive operations, build a historical command sequence library, extract the current user input commands during remote login interaction, and build a current command sequence library; The longest common subsequence algorithm is used to analyze the longest common subsequence length of the command sequences in the current command sequence library and the historical command sequence library and determine their similarity. By comparing the preset similarity range, it is determined whether there is abnormal command usage behavior and generate user behavior analysis results.

8. The computer remote login identification system according to claim 1, characterized in that: The steps for obtaining the remote login behavior recognition result are specifically as follows: Based on the packet behavior analysis results, the remote environment analysis results and the user behavior analysis results, the overall abnormality degree of the remote login interaction is evaluated by a multi-factor abnormality analysis method to obtain an abnormality degree evaluation result; Remote session identification is performed based on the abnormality degree assessment results, including identifying normal remote login behaviors and abnormal remote attack behaviors, to obtain remote login behavior identification results.

Citation Information

Patent Citations

  • Abnormal packet detection device and method

    CN114513323A

  • Internet data security protection method and system based on intelligent algorithm

    CN119272339A

  • Remote control method and system for multiple kitchen appliances based on context awareness

    CN119414722A

  • Computer network big data security protection method and system

    CN119892504A

  • Hybrid flow and packet anomaly detection system and method thereof

    TW202508258A