Method and system using decoy systems to identify malicious actions performed by attackers in an it infrastructure

WO2026177631A1PCT designated stage Publication Date: 2026-08-27PUBLICHNOE AKTSIONERNOE OBSHCHESTVO SBERBANK ROSSII (PAO SBERBANK)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/RU2025/000047
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-02-18
Filing Date
2025-02-21
Publication Date
2026-08-27

Smart Images

  • Figure RU2025000047_27082026_PF_FP_ABST
    Figure RU2025000047_27082026_PF_FP_ABST
Patent Text Reader

Abstract

The invention relates to a method and system using decoy systems to identify malicious actions performed by attackers in an IT infrastructure. Decoy systems are deployed in an IT infrastructure and subsequently configured. A machine learning model is used to generate baits for attackers, which are then distributed throughout the IT infrastructure. Data comprising information about all the attackers' activity when interacting with the decoy systems is gathered, and a machine learning model is then used to emulate additional services on the basis of the data obtained from the decoy systems. Data generated during the attackers' interaction with the emulated services is gathered using the decoy systems. Using a machine learning model, the gathered data is analyzed and malicious actions performed by the attackers are identified. The invention improves the security of an IT infrastructure.
Need to check novelty before this filing date? Find Prior Art

Description

METHOD AND SYSTEM FOR DETECTING MALICIOUS ACTIONS OF INVADERS IN IT INFRASTRUCTURE USING TRAPPING SYSTEMS. AREA OF TECHNOLOGY

[0001] The claimed technical solution generally relates to the field of computing technology, and in particular to an automated method and system for identifying malicious actions by intruders in IT infrastructure using honeypot systems. LEVEL OF TECHNOLOGY

[0002] With the development of information technology, IT solutions and information security systems have a significant impact on all areas of life, including the economy, healthcare, transportation, and industry. Currently, large IT companies and organizations are actively implementing and utilizing various information security solutions within their infrastructure to protect their assets, data, and reputation.

[0003] At the same time, cyberthreats are rapidly evolving. Ensuring the security of IT infrastructure requires new methods for rapid detection and the development of detection rules. One of the key challenges is the exploitation of zero-day vulnerabilities by attackers. These vulnerabilities can be actively exploited for a long time before being identified and patched, allowing attackers to gain access to a wide range of targets.

[0004] Attackers obtain information about accessible internet services using specialized search engines, such as Shodan, a search engine for internet-connected devices. These engines allow them to quickly find devices with vulnerable or misconfigured services. Furthermore, many services provide information about their version, name, and other characteristics when certain requests are submitted, making it easier for attackers to find targets. This creates a need to develop more sophisticated and flexible solutions to counter such approaches.

[0005] A drawback of known solutions in this area of ​​technology is the lack of the ability to automatically detect malicious activity by intruders in the IT infrastructure using honeypot systems. DISCLOSURE OF THE INVENTION

[0006] The proposed technical solution proposes a new approach to detecting malicious activity in IT infrastructure. This solution utilizes machine learning algorithms to automate the process of detecting malicious activity in IT infrastructure using honeypot systems.

[0007] This solves the technical problem of automated detection of malicious activity in IT infrastructure. t

[0008] The technical result achieved by solving this problem is an increase in the security of the IT infrastructure.

[0009] The specified technical result is achieved by implementing a computer-implemented method for identifying malicious actions of intruders in the IT infrastructure using honeypot systems, performed using at least one processor and containing the following stages: a) deploy decoy systems in the IT infrastructure to realistically interact with attackers; b) configure trap systems, as a result of which: • the type of trap system is selected; • the operating system (OS) is selected and installed; • network parameters and security settings are specified; c) generate attacker bait using a machine learning model containing data that matches the requirements of named entities within the IT infrastructure; d) distribute bait to attackers across the IT infrastructure containing emulated core services of interest to attackers; e) collect data containing information about all activity of attackers when interacting with honeypot systems; f) emulate, using a machine learning model, additional services of interest to attackers, based on data obtained from honeypot systems; g) collect data generated by interactions between attackers and emulated services using honeypot systems; h) analyze the collected data using a machine learning model; and i) detect malicious activities of intruders in the IT infrastructure using a machine learning model based on the analyzed data in step h).

[0010] In one particular implementation option of the method, decoy systems are additionally deployed on the Internet for realistic interaction with attackers.

[0011] In another particular embodiment of the method, language models are used as machine learning models.

[0012] In another particular embodiment of the method, lures for attackers are additionally distributed on the Internet.

[0013] In another particular implementation of the method, the bait for attackers is automatically updated.

[0014] In another particular embodiment of the method, data collection is carried out using specialized tools designed to collect system events and system calls.

[0015] In another particular embodiment of the method, a service is emulated that completely imitates the operation and functionality of a genuine service in the IT infrastructure.

[0016] In another particular embodiment of the method, the decoy systems dynamically expand and change, adapting to the actions of the attackers.

[0017] In another particular embodiment of the method, decoy systems transfer data from one session of interaction with the attacker to a new session to ensure uninterrupted interaction.

[0018] In another particular implementation of the method, a machine learning model interacts with the attacker in an emulated service. s

[0019] In another particular implementation option, a machine learning model for realistic interaction with attackers is assigned a role according to the properties of the emulated service.

[0020] In another particular embodiment of the method, at step i), a report is also prepared.

[0021] In another particular embodiment of the method, at stage i) I also develop recommendations for protecting the IT infrastructure.

[0022] In another particular embodiment of the method, at step i), rules for detecting illegitimate activity are also created.

[0023] In another particular embodiment of the method, at step i) indicators of compromise are also collected and used in detection rules.

[0024] In addition, the stated technical result is achieved through a system for detecting malicious actions by intruders in the IT infrastructure using honeypot systems containing at least one processor and memory storing machine-readable instructions. BRIEF DESCRIPTION OF DRAWINGS

[0025] Fig. 1 illustrates a block diagram of the claimed method.

[0026] Fig. 2 illustrates an example of the general view of a computing system that ensures the implementation of the claimed solution. IMPLEMENTATION OF THE INVENTION

[0027] Below we will describe the concepts and terms necessary for understanding this technical solution.

[0028] A model in machine learning (ML) is a set of artificial intelligence methods whose characteristic feature is not a direct solution to a problem, but learning in the process of applying solutions to many similar problems.

[0029] A software vulnerability is a flaw in a system that can be exploited to intentionally compromise its integrity and cause it to malfunction. Vulnerabilities can result from programming errors, design flaws, weak passwords, viruses and other malware, or script and SQL injection attacks. Vulnerabilities can be exploitable or unexploitable.

[0030] A honeypot instance is a deployed unit of a honeypot system that simulates a vulnerable system, service, or network to lure attackers, collect their actions, and analyze attacks. A honeypot instance can be a separate physical server, a virtual machine (VM), a container (e.g., Docker), or a process running on the host.

[0031] A honeypot is a cybersecurity tool that consists of a vulnerable or vulnerability-mimicking system designed to attract attackers. It is used to analyze their behavior, detect attacks, and improve defenses.

[0032] A honeytoken is a cybersecurity tool that consists of fake credentials or accounts created to detect and track unauthorized access or information leaks.

[0033] A management agent is a program running on a honeypot instance that:

[0034] Receives commands from the management console.

[0035] Reads JSON configuration files.

[0036] Emulates services based on configuration.

[0037] Sends requests to LLM and receives responses

[0038] Logs actions.

[0039] The management console (or management server) is the central component of the honeypot system, coordinating and managing honeypot instances through management agents. The console serves as the interface for security operators to interact with the honeypot network, allowing for centralized management, monitoring, and adaptation of honeypots in real time.

[0040] This technical solution can be implemented on a computer, in the form of an automated information system (AIS) or a machine-readable medium containing instructions for performing the above-mentioned method.

[0041] The technical solution can be implemented as a distributed computer system.

[0042] In this solution, the term “system” refers to a computer system, a computer (electronic computer), a numerical control (CNC), a PLC (programmable logic controller), computerized control systems, and any other devices capable of performing a given, clearly defined sequence of computing operations (actions, instructions).

[0043] A command processing unit is an electronic unit or integrated circuit (microprocessor) that executes machine instructions (programs) /

[0044] The command processing unit reads and executes machine instructions (programs) from one or more storage devices, such as random access memory (RAM) and / or read-only memory (ROM). ROM devices may include, but are not limited to, hard disk drives (HDD), flash memory, solid-state drives (SSD), optical storage media (CD, DVD, BD, MD, etc.), and others.

[0045] A program is a sequence of instructions intended for execution by a computer control unit or command processing device.

[0046] As shown in Fig. 1, the computer-implemented method for detecting malicious actions of an intruder in an IT infrastructure using honeypot systems (100) consists of several stages performed by at least one processor.

[0047] At stage (101), decoy systems are deployed in the IT infrastructure for realistic interaction with attackers.

[0048] In one embodiment of the invention, decoy systems are additionally deployed on the Internet to realistically interact with attackers.

[0049] In another embodiment of the invention, the decoy systems dynamically expand and change, adapting to the actions of the attackers.

[0050] In another embodiment of the invention, decoy systems transfer data from one session of interaction with an attacker to a new session to ensure uninterrupted interaction.

[0051] At stage (102), the trap systems are configured, as a result of which: • the type of trap system is selected; • the operating system (OS) is selected and installed; • network parameters and security settings are specified.

[0052] At step (103), attacker baits are generated using a machine learning model, containing data that matches the requirements of named entities within the IT infrastructure.

[0053] In one embodiment of the invention, language models are used as machine learning models.

[0054] At stage (104), bait for attackers is distributed across the IT infrastructure containing emulated basic services that are of interest to attackers.

[0055] In one embodiment of the invention, lures for attackers are additionally distributed on the Internet.

[0056] In another embodiment of the invention, the bait for attackers is automatically updated.

[0057] At stage (105), data is collected containing information about all the activity of intruders when interacting with honeypot systems.

[0058] In one embodiment of the invention, data collection is carried out using specialized tools designed to collect system events and system calls.

[0059] At step (106), additional services of interest to attackers are emulated using a machine learning model based on data obtained from honeypot systems;

[0060] In one embodiment of the invention, a service is emulated that completely imitates the operation and functionality of a genuine service in the IT infrastructure.

[0061] In another embodiment of the invention, a machine learning model interacts with the attacker in the emulated service.

[0062] In another embodiment of the invention, a machine learning model for realistic interaction with attackers is assigned a role according to the properties of the emulated service.

[0063] At stage (107), data generated during interactions between attackers and emulated services is collected using honeypot systems.

[0064] At step (108), the collected data is analyzed using a machine learning model.

[0065] At step (109), malicious actions of intruders in the IT infrastructure are identified using a machine learning model based on the data analyzed at step (108).

[0066] In one embodiment of the invention, at step (109), a report is also prepared.

[0067] In another embodiment of the invention, at step (109), I also generate recommendations for protecting the IT infrastructure.

[0068] In another embodiment of the invention, at step (109), rules for detecting illegitimate activity are also created.

[0069] In another embodiment of the invention, at step (109), indicators of compromise are also collected and used in detection rules.

[0070] Let us consider in more detail the implementation of the claimed invention.

[0071] During the initial stage, honeypot instances are deployed, containing management agents that run on hosts and facilitate interaction with a unified management console. These management agents retrieve configuration from JSON files (generated in subsequent stages), emulate specified services, and log attacker actions. They also process commands from the management console and dynamically adapt the instance configuration.

[0072] Instance infrastructure setup is automated through a single management console and may consist of the following steps: host selection—the host may be a physical server, virtual machine (VM), or container (Docker); setting a public or private IP address for the host; installing basic software—for example, installing Python and libraries for running the management agent; performing network configuration—checking that the required ports (e.g., TCP 22, TCP 80, TCP 3306) are open in the FW.

[0073] Next, a centralized honeypot instance management system is created using management agents that are located on each honeypot host.

[0074] These agents receive commands from the management console, start or stop emulated services, update configurations, transmit logs, and respond to dynamic changes in the infrastructure.

[0075] A monitoring mechanism is implemented to track instance statuses. The management console regularly receives information about the status of instances and their services. The console collects activity data on each instance, allowing it to identify which honeypots are most frequently attacked. The console alerts you to suspicious activity and automatically sends scheduled reports.

[0076] Next, instructions are generated for LLM to create honeytokens that comply with the internal requirements for naming entities within the IT infrastructure (simulating real ones, but not having actual value). The following parameters are set through a single system management console for LLM:

[0077] Temperature: To control the level of creativity of the generation. Low temperature (eg 0.2) - more accuracy and predictability. High temperature (0.8) - more creative and unique results.

[0078] Context: Adding additional data - task description and examples - so that LLM knows about the specifics of the infrastructure.

[0079] Example: "This data must comply with the organization's networking standards with 10.0.xx subnets. Computer FQDNs must begin with ARM, followed by a numeric value. User account names must match the pattern FirstName.LastName."

[0080] Next, honeytokens are prepared using requests to LLM, creating fake data (accounts, access credentials, and other bait) to lure attackers into the honeypot. Examples of honeytokens include: Credentials: Fake logins and passwords. Fake API keys. Network lures: IP addresses, DNS records. Counterfeit devices (e.g. IoT). System objects: Dummy databases. Names of services, files, folders.

[0081] To automate honeytoken generation, requests to LLM are automated via an API. Using a library (for example, the OpenAI API for ChatGPT), the generation process is initiated from a single management console.

[0082] After honeytokens are generated by LLM, the data is validated to ensure it meets established rules. For example, the username must match the pattern ^[az].[az]+$, duplicates must be excluded, and the data must be compared with real names to avoid obvious signs of forgery.

[0083] During the honeytoken preparation phase, LLM ensures high-quality and realistic data. This makes the honeypots more attractive to attackers and increases the effectiveness of the honeypot. Using automated queries to LLM allows for scalability and adaptability of the process to any cyberthreat scenario.

[0084] Next, honeytokens are distributed. Data generated to attract attackers' attention is integrated into the organization's ^ infrastructure or posted to public sources. This process involves selecting deployment methods, automation tools, and monitoring mechanisms.

[0085] Examples of placement methods include:

[0086] Deployment into local infrastructure. Honeytokens are deployed on systems used within the organization: Active Directory (AD): Create fake user accounts with fake privileges. Example: adding a test admin user with access to fake resources. Databases: Adding tables containing false data. Example: A dummy sensitive data table with data that appears valuable. File systems: Posting fake documents with enticing names. Example: file financial_report_2024.xlsx with fake content.

[0087] Placement in network systems DNS records: Adding fictitious domain names. Example: db-internal.example.com, which points to a honeypot. API keys: Injecting fake keys into configuration files. Example: a dummy key in the .env file that an attacker might try to use.

[0088] Posting in public sources

[0089] Honeytokens can be placed in open sources to attract external attackers: Public code repositories:

[0090] Placing fake secrets in files (e.g. AWS_SECRET_ACCESS_KEY on GitHub). Forums and websites: Creating fake ads that include fictitious information. Example: Posting a fake API key on a developer forum. E-mail: Sending fake data in phishing emails so that the attacker can try to use it.

[0091] Automating regular honeytoken refreshes is essential to ensure decoy tokens remain relevant, diverse, and consistent with current infrastructure and threats. This reduces the likelihood that attackers will detect honeytokens as counterfeit.

[0092] The objectives of the automation stage are: Regular data rotation – replacing outdated honeytokens with new ones. Ensuring uniqueness and freshness of data for each update. Reducing repetition - eliminating patterns that could reveal the bait as fake. Accounting for infrastructure changes – automating updates so that honeytokens adapt to the current IT infrastructure. Updating distribution locations - re-deploying honeytokens to systems that may attract attackers and eliminating locations where they remain undetected.

[0093] Mechanisms for automated honeytoken distribution can include orchestration tools (e.g., Ansible, Terraform) and task schedulers (e.g., cron or system utilities).

[0094] During the preparation of the required service configuration, which is a JSON file with a specified set of attributes (usernames, passwords, IP addresses, database or service names, etc.), a request is generated by the appropriate agent to the LLM. The generated JSON files serve as configuration files for emulating the required service on the target honeypot. The generated JSON files can be created and uploaded through a single management console.

[0095] An example of an agent-generated configuration file for an emulated service could be a file with the following content: { "service name": "postgresql service", f "version": "13.3", "users": [ {"username": "admin", "password": "AdminPassl23"}, {"username": "guest", "password": "GuestPass456"}], "db names": ["app_db", "test db", "logs db"], "ip address": "192.168.1.50", "port": 5432, "banner": "Welcome to PostgreSQL 13.3" }

[0096] After generating the JSON configuration files, one of the steps is to check the correctness of the generated file for: JSON syntax - using libraries to check the format; parameter validity: checking the compliance of IP addresses, ports, etc.; test run - trial deployment on a test honeypot instance.

[0097] The next step involves creating precise and effective instructions (prompts) to control LLM behavior so it can emulate the desired service or OS command shell. Prompts contain system settings, user requests, and additional parameters (e.g., temperature, context, tokens) that ensure realistic and predictable attacker interaction with the honeypot.

[0098] Next, a service banner is set. The banner helps create the illusion of a genuine system or service during scanning. The banner is used to respond to attacker requests and is designed to mimic real services, making them attractive for attacks. The goals of setting a banner are: emulating the service's reality - the banner creates a first impression of the system, so the attacker believes it is legitimate; increasing attacker interest - a fake banner can hint at the presence of vulnerabilities or valuable data; and doping and analysis - the response with the banner is recorded in the logs and can aid in analyzing attacker activity.

[0099] When specifying a banner, the following data is filled in: the name and version of the service, the name of the operating system, additional information (for example, license or configuration).

[0100] Banner generation can be performed using a script or by requesting the appropriate agent to LLM.

[0101] Next, based on the JSON configuration files, the emulated honeypot services are deployed to target instances. A single management console is used to send commands to the instances, where the JSON files are interpreted and the emulated services are launched.

[0102] The unified management console sends commands to honeypot instances to launch emulated services. This is accomplished through: selecting one or more instances – the target server or group of servers is selected; command transmission – the management console sends a command such as DEPLOY / path / to / config.json to the agent; and command delivery – the command is transmitted to the instance via a secure channel (e.g., SSH or API).

[0103] An agent is launched on the instance side, accepting commands from the console. This stage is achieved through the following steps: configuration loading – the JSON file is read by the agent on the instance; service initialization – service configuration scripts are run based on the JSON data; and the service emulation environment is deployed – the agent emulating the service interacts with the LLM to ensure connectivity and generate correct responses to attacker requests.

[0104] The honeypot infrastructure then automatically changes based on the attacker's activity. Using a management agent, the honeypot can adapt its behavior, add new services, change configurations, or scale to mimic a real IT infrastructure and collect additional data.

[0105] Dynamic adaptation to attacker actions involves creating new honeypots that match the attacker's interests and changing the honeypot's current parameters to maintain realism. Interaction escalation expands the set of emulated services to deeply simulate a real system.

[0106] Examples of triggers for expansion include: Interest in a specific service: If the attacker is actively interacting with MySQL, an additional SQL server can be launched. Scan attempt: Perform service emulation for network protocols such as FTP or SMB.

[0107] Next, continuous interaction between the attacker and the honeypot is ensured. If the attacker loses connection or their session expires, the system preserves the current context, including the state of the services and data with which the attacker interacted. Upon reconnection, the illusion is created that the interaction has continued from the same point. This approach achieves the goals and ensures comprehensive data for analyzing long-term attacks.

[0108] To ensure continuity of interaction, the following steps are taken:

[0109] Step 1: Save the current session state Session Identification: Honeypot generates a unique identifier (Session ID) for each attacker session. Context preservation: the system records session data - the attacker's IP address, the service used, active commands, as well as the state of services - current database data, files on the server, open connections. Session storage: session data is stored in the file system (e.g. JSON file), database (e.g. Redis for fast processing), cloud storage (if the honeypot is distributed).

[0110] Step 2: End the current session Activity Monitoring: The honeypot monitors attacker activity. If the user terminates the session or loses the connection, the attacker's activity is recorded. Examples of such attacker activity include inactivity for a specified period of time (timeout) or connection interruption. Saving state on termination: Before a session terminates, the current data is written to storage

[0111] Step 3: Restore session Reconnection detection: When a new connection is made, the honeypot checks whether there is an active or terminated session for the given IP address or user. Loading a saved context: If a saved session is found, its data is loaded into the current service. Creating the illusion of continuous interaction: the honeypot responds as if nothing has changed

[0112] Step 4: Updating the session in real time Dynamic state update: During sessions, data is constantly synchronized with the storage. Action doping: Each new attacker action is added to the current session.

[0113] Next, all actions performed by the attacker in the honeypot, with any emulated service, are recorded. Logging allows for collecting data on attack methods, tools used, and other aspects of the attacker's behavior. These logs become the basis for analysis, countermeasure development, and defense recommendations.

[0114] The doping mechanism has settings for selecting the type of doped events, which can be: Login attempts - successful or unsuccessful attempts to log into the system (SSH, FTP, etc.). Command execution - commands sent to the command shell. Requests to services - SQL queries, HTTP requests, read / write operations in the file system. Data transfer - downloading or uploading files. Network events - connections, port scanning, packet sending.

[0115] Logging of attacker activity can be performed both by the honeypot itself and by specialized log collection tools: syslog for recording system events; auditd for monitoring system calls and the file system; tepdump for analyzing network traffic.

[0116] Collected logs are transferred to the management console at specified intervals. An API can be used as the transfer mechanism.

[0117] The management console sends structured data to security event management (SIEM) systems (such as Splunk or Elastic Security) for analysis and correlation to identify attack patterns.

[0118] The next step involves using one of the agents in the multi-agent network to analyze logs using a large-scale language model (LLM). The multi-agent system distributes the data processing load and ensures coordination between agents for analysis, decision-making, and reporting of results.

[0119] Log analysis is optimized by using a single agent for deep analysis, with other tasks distributed among the remaining agents.

[0120] Coordination in the network is achieved by transmitting data from other agents to the selected analyst agent.

[0121] Each agent in the network collects data on the attacker's actions on its instance. The logs are then transmitted to the agent responsible for analysis (the analytics agent). This can be done via an API or a message queue (e.g., RabbitMQ, Kafka). The analytics agent is assigned by default, or the least loaded one is selected based on the overall workload of all agents.

[0122] The analytics agent preprocesses logs received from other agents: irrelevant events (such as system messages) are removed; the data is normalized for analysis.

[0123] The analyst agent is also responsible for generating requests to LLM for log analysis. An example of such a request: Analyze the following logs and identify suspicious activity: ( { "instance_id": "honeypot-1", "event_type": "ssh_login", "username": "root", "password": "12345", "status": "failed", "timestamp": "2024-11-21T14:32:01Z" ъ { "instance_id": "honeypot-2", "event_type": "sql_query", "query": " SELECT * FROM sensitive_data;", " timestamp": "2024-U-21T14:34:12Z" } 1 Identify the attacker's tactics and techniques and propose countermeasures.

[0124] LLM identifies anomalies and classifies actions. Sample answer: The following threats were detected: - Brute force SSH from IP 203.0.113.5. - SQL injection for data theft. Tactics: - Credential Access (T1110). - Data Exfiltration (T1041).

[0125] The next step involves generating a detailed report based on log analysis and the LLM response. The report describes the attacker's actions, techniques, and tactics used, and includes structured recommendations for protection. Report generation is fully automated and performed by an agent (agent reporter), eliminating the need for manual data processing.

[0126] The reporter agent systematizes the analysis results, converting LLM output into a structured and understandable report. It generates a detailed description of the attacker's activity, including techniques and tactics from the MITRE ATT&C database. The identified indicators of compromise are passed on to other security tools for further use.

[0127] The next stage involves analyzing logs and other attacker activity data to automatically generate recommendations for mitigating damage and improving the security of the system and IT infrastructure as a whole. Based on an analysis of the attacker's actions and known attack methods, the agent, using LLM, generates suggestions that may include updating security settings, blocking IP addresses, updating access policies, and other measures.

[0128] Examples of recommendations: 1. Block IP 203.0.113.5 at the firewall level. 2. Enable two-factor authentication for SSH. 3. Restrict access to the sensitive data table. 4. Close unused ports. 5. Analyze event logs to identify other suspicious IPs.

[0129] The agent also automatically generates commands for implementing generated recommendations. For example: "Iptables -A INPUT -s 203.0.113.5 -j DROP”, "systemctl enable ssh-2fa”, “mysql -e V'REVOKE ALL ON sensitive_data FROM ' guest V”, ”ufw deny 3306”

[0130] The final stage involves the automatic creation of threat detection rules based on log analysis, attacker activity, and identified techniques. These rules are used for integration into threat detection systems (IDS), SIEM, antivirus solutions, or EDR platforms.

[0131] This approach significantly accelerates the process of developing rules based on identified threats during attacker interactions with the emulated service. Rules are generated in a unified format (Sigma, Y ARA, or others), which is easily applied across various security tools.

[0132] To create Sigma or YARA rules, the agent generates a creation request to LLM. The generated rules are transmitted to other monitoring and security systems via APIs or configuration files.

[0133] Rules are saved in the management console for further use and management.

[0134] The implementation of this invention significantly increases the security of the IT infrastructure by identifying malicious actions of an intruder in the IT infrastructure.

[0135] Fig. 2 shows an example of a general view of a computing system (300) that ensures the implementation of the claimed method or is part of a computer system, for example, a server, a personal computer, part of a computing cluster, processing the necessary data for the implementation of the claimed technical solution.

[0136] In general, the system (300) comprises one or more processors (301), memory means such as RAM (302) and ROM (303), input / output interfaces (304), input / output devices (305), and a device for network interaction (306), united by a common information exchange bus.

[0137] The processor (301) (or several processors, a multi-core processor, etc.) can be selected from a range of devices that are widely used at present, for example, from such manufacturers as: Intel™, AMD™, Apple™, Samsung Exynos™, MediaTEK™, Qualcomm Snapdragon™, etc. Under the processor or one of the processors used in the system (300), it is also necessary to take into account a graphic processor, for example, NVIDIA GPU or Graphcore, the type of which is also suitable for the full or partial implementation of the method, and can also be used for training and applying machine learning models in various information systems.

[0138] RAM (302) is random access memory (RAM) and is designed to store machine-readable instructions executed by the processor (301) to perform the necessary logical data processing operations. RAM (302) typically contains executable instructions from the operating system and corresponding software components (applications, software modules, etc.). The available memory of a graphics card or graphics processor (GPU) can also serve as RAM (302).

[0139] ROM (303) is one or more permanent data storage devices, such as a hard disk drive (HDD), a solid-state drive (SSD), flash memory (EEPROM, NAND, etc.), optical storage media (CD-R / RW, DVD-R / RW, BlueRay Disc, MD), etc.

[0140] Various types of I / O interfaces (304) are used to organize the operation of system components (300) and to organize the operation of external connected devices. The choice of the appropriate interfaces depends on the specific design of the computing device, which may include, but are not limited to: PCI, AGP, PS / 2, IrDa, FireWire, LPT, COM, SATA, IDE, Lightning, USB (2.0, 3.0, 3.1, micro, mini, type C), TRS / Audio jack (2.5, 3.5, 6.35), HDMI, DVI, VGA, Display Port, RJ45, RS232, etc.

[0141] To ensure user interaction with the computing system (300), various means (305) of I / O information are used, for example, a keyboard, a display (monitor), a touch display, a touchpad, a joystick, a mouse, a light pen, a stylus, a touch panel, a trackball, speakers, a microphone, augmented reality means, optical sensors, a tablet, light indicators, a projector, a camera, biometric identification means (a retinal scanner, a fingerprint scanner, a voice recognition module), etc.

[0142] The network interaction means (306) ensures the transmission of data via an internal or external computer network, for example, an Intranet, the Internet, a LAN, etc. One or more means (306) may be, but are not limited to: an Ethernet card, a GSM modem, a GPRS modem, an LTE modem, a 5G modem, a satellite communication module, an NFC module, a Bluetooth and / or BLE module, a Wi-Fi module, etc.

[0143] The submitted application materials disclose preferred examples of the implementation of the technical solution and should not be interpreted as limiting other, particular examples of its implementation that do not go beyond the scope of the requested legal protection, which are obvious to specialists in the relevant field of technology.

Claims

FORMULA 1. A method for detecting malicious actions by intruders in an IT infrastructure using honeypot systems, performed using at least one processor and comprising the following steps: j) deploy decoy systems in the IT infrastructure to realistically interact with attackers; k) carry out the configuration of trap systems, as a result of which: • the type of trap system is selected; • the operating system (OS) is selected and installed; • network parameters and security settings are specified; l) generate attacker baits using a machine learning model containing data that matches the requirements of named entities within the IT infrastructure; m) distribute bait for attackers across the IT infrastructure containing emulated core services of interest to attackers; n) collect data containing information about all activity of attackers when interacting with honeypot systems; o) emulate additional services of interest to attackers using a machine learning model based on data obtained from honeypot systems; p) collect data generated during interactions between attackers and emulated services using honeypot systems; q) analyze the collected data using a machine learning model; and d) detect malicious actions of intruders in the IT infrastructure using a machine learning model based on the analyzed data in step h).

2. The method according to paragraph 1, characterized in that decoy systems are additionally deployed on the Internet for realistic interaction with intruders.

3. The method according to claim 1, characterized in that language models are used as machine learning models.

4. The method according to paragraph 1, characterized by the fact that additionally bait for intruders is distributed on the Internet.

5. The method according to paragraph 1, characterized by the fact that bait for intruders is automatically updated.

6. The method according to paragraph 1, characterized in that the data collection is carried out using specialized tools designed for collecting system events and system calls.

7. The method of claim 1, characterized by emulating a service that completely imitates the operation and functionality of a genuine service in the IT infrastructure.

8. The method of claim 1, characterized by the honeypot systems dynamically expanding and changing, adapting to the actions of the attackers.

9. The method according to claim 1, characterized in that the decoy systems transfer data from one session of interaction with the attacker to a new session to ensure uninterrupted interaction.

10. The method according to paragraph 1, characterized in that in the emulated service, a machine learning model interacts with the attacker. I. The method according to paragraph 10, characterized in that the machine learning model for realistic interaction with attackers is assigned a role according to the properties of the emulated service.

12. The method according to paragraph 1, characterized in that at stage i) a report is also prepared.

13. The method according to paragraph 1, characterized in that at stage i) I also carry out the development of recommendations for the protection of the IT infrastructure.

14. The method according to claim 1, characterized in that at step i) rules for detecting illegitimate activity are also created.

15. The method according to claim 1, characterized in that at step i) indicators of compromise are also collected and used in detection rules.

16. A system for detecting malicious actions by intruders in an IT infrastructure using honeypot systems, comprising at least one processor and memory storing machine-readable instructions that, when executed by the processor, implement the method according to any of paragraphs 1-15.