Generative sequence processing model for network security

By pre-training and fine-tuning the generated sequence processing model, the problem of heterogeneous data structures in cybersecurity tools is solved, enabling efficient cybersecurity data analysis and dynamic threat intelligence aggregation, and improving the professionalism and transparency of cybersecurity response.

CN121844314APending Publication Date: 2026-04-10GOOGLE LLC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-08-28
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing cybersecurity tools generate data with varying structures, making effective coordination and analysis difficult. This results in inefficiency, numerous vulnerabilities, and an inability to provide comprehensive protection.

Method used

A generative sequence processing model is used to fine-tune cybersecurity data, including pre-training and fine-tuning stages. It utilizes a variety of cybersecurity datasets and tasks to enhance the model's capabilities in specific cybersecurity domains, adapt to the needs of different roles, and enable interoperability within cybersecurity platforms.

Benefits of technology

It improves the efficiency of cybersecurity data analysis and processing, reduces false positives and false negatives, provides dynamic threat intelligence aggregation and professional cybersecurity response, reduces the overhead of managing multiple tools, and improves the efficiency and transparency of cybersecurity operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121844314A_ABST
    Figure CN121844314A_ABST
Patent Text Reader

Abstract

The invention provides a generation sequence processing model which is specially used for carrying out fine tuning on network security applications. This specialized generation sequence processing model can be trimmed over a wide variety of cyber-security data and associated trimming tasks, providing the model with rich capabilities to analyze, understand, describe real-time cyber-security data generated by one or more cyber-security operating tools, and take actions with respect to the real-time cyber-security data. Thus, the creation and use of the generative sequence processing model presents a solution to technical challenges to efficiently analyze and process a large amount of data generated by a plurality of different network security operating tools deployed by the organization.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Related applications

[0002] This application claims priority and benefit to U.S. Provisional Patent Application No. 63 / 579,251, filed August 28, 2023. U.S. Provisional Patent Application No. 63 / 579,251 is hereby incorporated by reference in its entirety. Technical Field

[0003] This disclosure generally relates to cloud-based cybersecurity platforms. More specifically, this disclosure relates to systems and methods that utilize generative sequence processing models that have been fine-tuned for cybersecurity applications to assist in the identification, analysis, and mitigation of cybersecurity threats. Background Technology

[0004] In today's digital age, organizations face an increasing number of complex cybersecurity threats. Cybersecurity refers to the practice of protecting systems, networks, and data from digital attacks, unauthorized access, and damage.

[0005] Traditional cybersecurity measures are often insufficient to provide comprehensive protection against such threats, leading to the emergence of a wide variety of cybersecurity operational tools, such as Security Orchestration, Automation and Response (SOAR) platforms, Security Information and Incident Management (SIEM) systems, Intrusion Detection Systems (IDS), Intrusion Prevention Systems (IPS), antivirus software, endpoint protection, vulnerability management tools, and so on.

[0006] Each of these tools generates massive amounts of cybersecurity data, often formatted according to different structures or formats that are not easily combined or harmonized. Analyzing and processing the staggering volume and diversity of data generated by the ever-growing number of cybersecurity operations tools is complex and tedious, leading to inefficiencies and vulnerabilities. Summary of the Invention

[0007] A system of one or more computers may be configured to perform specific operations or actions by installing software, firmware, hardware, or combinations thereof on the system, which, when operated, cause the system to perform these actions. One or more computer programs may be configured to perform specific operations or actions by means of instructions that, when executed by a data processing device, cause the device to perform the actions.

[0008] One general aspect includes a computer system for improving network security. The computer system includes one or more processors. The system also includes one or more non-transitory computer-readable media that collectively store a generative sequence processing model, wherein the generative sequence processing model has been fine-tuned on one or more fine-tuning tuples generated from one or more network security datasets, and wherein at least one of the one or more fine-tuning tuples may include fine-tuning input and fine-tuning labels, wherein the fine-tuning input may include network security data and a question about the network security data, and wherein the fine-tuning label may include a labeled answer to the question about the network security data.

[0009] Other embodiments of this aspect include corresponding computer systems, devices, and computer programs recorded on one or more computer storage devices, each of which is configured to perform the actions of these methods.

[0010] Example implementations may include various combinations of one or more of the following features. The computer system, wherein the cybersecurity dataset may include datasets from: Security Orchestration, Automation, and Response (SOAR) system data, Security Information and Incident Management (SIEM) system data, security blog information, analyst reports, signature-based detection files, malware scripts, vulnerability information, product documentation, secure code repositories, or cybersecurity and software development framework data. The generative sequence processing model may have been fine-tuned on the one or more fine-tuning tuples to perform one or more fine-tuning tasks, wherein the one or more fine-tuning tasks may include: classification tasks, summarization tasks, generation tasks, or extraction tasks. At least one second fine-tuning tuple in the one or more fine-tuning tuples may include a second fine-tuning input and a second fine-tuning label, wherein the second fine-tuning input may include a natural language query, and wherein the second fine-tuning label may include a query expressed in a domain-specific query language. At least one of the one or more fine-tuning tuples may have been manually generated. At least one of the one or more fine-tuning tuples may have been automatically generated using one or more templates. At least one of the one or more fine-tuning tuples may have been automatically generated using one or more machine learning models. The generated sequence processing model can communicate operationally with one or more cybersecurity operations tools. The computer system can be configured to provide an interface that allows users to query the generated sequence processing model in natural language. The generated sequence processing model can be fine-tuned to adapt to specific cybersecurity personas within a variety of different cybersecurity personas. These specific cybersecurity personas may include: a Security Operations Center (SOC) analyst persona; a threat intelligence analyst persona; a malware or code analyst persona; or a security architect persona. Implementation of the technology may include hardware, methods or processes, or computer software on a computer-accessible medium.

[0011] One general aspect includes a cybersecurity platform implemented by one or more computing devices. The cybersecurity platform includes multiple computer-implemented agents configured to interoperate to jointly receive and process cybersecurity data, thereby generating and executing cybersecurity actions in response to the cybersecurity data. The platform also includes each of the multiple agents comprising a machine learning-generated sequence processing model that has been fine-tuned to adopt a specific cybersecurity role image from a variety of different cybersecurity role images. Other embodiments of this aspect include corresponding computer systems, devices, and computer programs recorded on one or more computer storage devices, each configured to perform actions of these methods.

[0012] Example implementations may include various combinations of one or more of the following features. The cybersecurity platform, wherein the multiple agents correspond to multiple different cybersecurity role images, which may include at least: a Security Operations Center (SOC) analyst role image; a threat intelligence analyst role image; and a malware or code analyst role image. The multiple agents may operate according to a distributed operating architecture. Multiple generative sequence processing models, each associated with a different agent, may have been branched from a pre-trained model and then fine-tuned using appropriate parameter-efficient adapters. The multiple agents may operate according to a centralized planning architecture. The centralized planning architecture may include a planning agent configured to control other agents on the platform. The planning agent may be configured to invoke the other agents based on a tool-based framework. The planning agent may be configured to perform thought chain inference, and wherein the planning agent is configured to control the other agents on the platform based on the thought chain inference. Implementations of the technology may include hardware, methods or processes, or computer software on a computer-accessible medium. Attached Figure Description

[0013] Referring to the accompanying drawings, a detailed discussion of embodiments is set forth in this specification for those skilled in the art, in which: Figure 1 A block diagram depicting an example cybersecurity platform according to an example embodiment of the present disclosure.

[0014] Figure 2 An example process for fine-tuning a sequence processing model according to an example embodiment of the present disclosure is described.

[0015] Figure 3 An example process for fine-tuning a sequence processing model to adopt different character images, according to an example embodiment of the present disclosure.

[0016] Figure 4 An example process for performing inference using a finely tuned sequence processing model according to an example embodiment of this disclosure is described.

[0017] Figure 5 An example multi-agent environment is depicted within an example cybersecurity platform according to an example embodiment of this disclosure.

[0018] Figure 6 This is a flowchart illustrating an example method for training a machine learning model according to an example implementation of an aspect of this disclosure; Figure 7 This is a block diagram of an example processing flow for using a machine learning model to process input to generate output, based on an example implementation of aspects of this disclosure. Figure 8This is a block diagram of an example sequence processing model based on an example implementation of aspects of this disclosure; Figure 9 This is a block diagram of an example technique for filling an example input sequence for processing by a sequence processing model, based on an example implementation of aspects of this disclosure; Figure 10 This is a block diagram of an example model development platform based on an example implementation of aspects of this disclosure; Figure 11 This is a block diagram of an example training workflow for training a machine learning model, based on an example implementation of aspects of this disclosure. Figure 12 This is a block diagram of an inference system for performing inference by operating one or more machine learning models, based on an example implementation of aspects of this disclosure. Figure 13 This is a block diagram of an example networked computing system based on an example implementation of aspects of this disclosure; Figure 14 This is a block diagram of an example computing device implementing an aspect of this disclosure; and Figure 15 This is a block diagram of an example computing device that implements aspects of this disclosure.

[0019] The repeated reference numerals in multiple figures are intended to identify the same features in various implementations. Detailed Implementation

[0020] Some example aspects of this disclosure relate to cybersecurity systems and methods that utilize one or more generative sequence processing models specifically tuned for cybersecurity applications. Specialized generative sequence processing models can be fine-tuned on a wide variety of cybersecurity data and associated fine-tuning tasks, thereby providing the model with rich capabilities to analyze, understand, describe, and / or take actions based on real-time cybersecurity data generated by one or more cybersecurity operations tools. Therefore, the creation and use of generative sequence processing models presents a solution to the technical challenges of effectively analyzing and processing large amounts of diverse data generated by multiple different cybersecurity operations tools deployed by an organization.

[0021] More specifically, one aspect of this disclosure relates to a generative sequence processing model that has been fine-tuned for various forms of cybersecurity data and related tasks. The sequence processing model can be a model configured to analyze and generate sequence data. Specifically, the sequence processing model can include a sequence-to-sequence design that enables the model to receive and process input data sequences to produce corresponding output data sequences.

[0022] One example of a sequence processing model is the so-called “Large Language Model” (LLM). The term LLM can refer to a highly parameterized model configured to interpret and generate text with human-like patterns. Through extensive pre-training on large datasets, LLMs are capable of handling a wide range of natural language processing tasks, from text generation to complex question-answering mechanisms. Another example of a sequence processing model is the “Large Multimodal Model” (LMM). Similar to LLMs, LMMs can refer to highly parameterized models that can operate to generate outputs (e.g., sequential outputs, such as text, images, and / or audio words) based on inputs (e.g., sequential inputs). However, LMMs typically operate on two or more different modalities, including some combination of text, image, and / or audio inputs and / or outputs. Text inputs can include natural language text and / or structured language text, such as programming languages, machine code, and / or structured queries. Images can include still images or moving images (i.e., “movies”).

[0023] Specifically, sequence processing models can initially be pre-trained on a large amount of pre-training data. During this phase, the model is exposed to a large dataset, enabling it to develop a holistic understanding of data patterns, structure, and inherent relationships. The purpose of pre-training is to embed extensive knowledge into the model, thereby establishing a cognitive framework.

[0024] In some implementations, model pre-training can include training the model on various forms of cybersecurity data. Incorporating a wide range of cybersecurity data during the pre-training phase can help the model succeed subsequently on a diverse range of cybersecurity-related tasks. Specifically, cybersecurity data has historically been extremely rare in the typical pre-training data mixes of large models because it is not widely available online (e.g., due to its inherent sensitivity). Therefore, as suggested in this paper, continuing pre-training on the 'long tail' of cybersecurity data formats (e.g., especially those formats that are not well represented) is highly beneficial for enabling the trained model to generalize across security-focused tasks.

[0025] In some implementations, pre-training can be divided into multiple consecutive stages. In the initial stage, the model can be pre-trained on large-scale data across multiple different domains (e.g., including computer languages) to instill basic language knowledge and build the model's overall language capabilities. In subsequent or "continued" pre-training stages, the focus can shift to including a relatively larger amount or proportion of cybersecurity-specific data. Therefore, later stages can include more targeted pre-training that incorporates a higher proportion and / or more diverse range of cybersecurity-specific data, enabling the model to begin specifically learning cybersecurity information.

[0026] According to another aspect of this disclosure, after comprehensive learning provided by pre-training, the sequence processing model can undergo a fine-tuning process using a wide range of cybersecurity data and associated pre-training tasks. In some examples, the fine-tuning task may be more specific and / or more structured than the pre-training task. For instance, some example pre-training tasks may involve broadly attempting to predict the correct completion of a given prompt (e.g., predicting the remainder of a text set given some initial parts of the text), while some example fine-tuning tasks may seek to test the model's ability to apply inference to a given input, such as answering questions about the given input that require extrapolation, analysis, and / or application of logic to the input.

[0027] In some implementations, the fine-tuning phase focuses on enhancing and refining the model's capabilities within a specific cybersecurity domain or task. This is achieved by leveraging targeted cybersecurity datasets (which may emphasize certain cybersecurity-specific knowledge or patterns) to adjust the model's internal parameters, thereby optimizing the model's performance on the given cybersecurity task.

[0028] The scope of cybersecurity data and tasks that can be used during fine-tuning will be described in further detail with reference to the accompanying drawings. However, as an introduction, cybersecurity data may include data from Security Orchestration, Automation, and Response (SOAR) systems. SOAR data may include SOAR cases, queries, rules, operation manuals, alerts, or other SOAR data. As further examples, cybersecurity data may also include Security Information and Incident Management (SIEM) system data, security blog information, threat intelligence reports, signature-based detection files, malware scripts, vulnerability information, product documentation, secure code repositories, cybersecurity and software development framework data, and / or other forms of cybersecurity data.

[0029] As described above, some example fine-tuning tasks performed on cybersecurity data can include custom or structured tasks that go beyond basic language modeling approaches. As an example, fine-tuning tasks can include classification tasks such as file type classification, malware detection, vulnerability identification, actor attribution, file activity classification, or other classification tasks. As another example, fine-tuning tasks can include summarization tasks such as generating an execution summary of a threat intelligence report, interpreting programming or computer language content (e.g., malicious scripts) in natural language, and / or summarizing structured security operations tool data (e.g., cybersecurity incident and alert metadata) in natural language. As a further example, fine-tuning tasks can include generation tasks such as generating threat intelligence reports from structured programming language scripts, creating SOAR operations manuals, and / or translating natural language into specialized query languages, such as SOAR queries, which are typically expressed in tool-specific or system-specific query languages ​​or dictionaries. As a further example, fine-tuning tasks can include extraction tasks (such as extracting relevant vulnerability information or entities) and / or transformation tasks (such as translating cybersecurity threat intelligence into actionable security measures).

[0030] Furthermore, in some implementations, fine-tuning tasks may include having the model perform thought-chain-style inferences that combine the aforementioned tasks. For example, when analyzing structured cybersecurity incidents, an example approach might identify command-line invocation information, which the model would analyze to look for signs of suspicious behavior, and then use the output of the analysis (e.g., a URL or filename) to issue additional queries to the SIEM system to better understand the attack.

[0031] According to another aspect of this disclosure, data annotation can be performed on raw cybersecurity data to generate fine-tuning data for performing the aforementioned fine-tuning tasks. For example, a cybersecurity dataset can be preprocessed to generate fine-tuning label data. The fine-tuning label data can be metadata, attributes, or other information extracted from or otherwise generated from the cybersecurity dataset. As an example, if the cybersecurity dataset is a set of software code, the fine-tuning labels can include information such as: the number of network indicators associated with the software code; whether the code uses network interaction; whether the code edits the registry; and / or other information about the software code that helps train a sequence processing model to understand the software code and improve its ability to explain why the software code may be malicious or not. The data annotation process can be performed manually and / or automatically via the application of one or more templates and / or one or more machine learning models.

[0032] By training generative sequence processing models on these fine-tuning tasks and associated fine-tuning data, these models can be endowed with the potential to analyze and predict a range of cybersecurity threats and / or perform many other cybersecurity tasks. After fine-tuning, the sequence processing model can be included in or utilized by a cybersecurity platform to perform any of the aforementioned tasks.

[0033] Furthermore, because the generative sequence processing model has been fine-tuned on diverse data from different systems and sources, the resulting model is capable of operating on inherently heterogeneous data and systems, such as different security operation tools with their own unique internal syntax, patterns, query languages, and / or other data structures. For example, the sequence processing model may be able to interpret and / or generate queries and / or other structured data representations expressed in multiple different query languages ​​and / or data structures, each associated with multiple different security operation tools (e.g., multiple different SOAR systems).

[0034] By leveraging finely tuned sequence processing models, cybersecurity platforms can simplify and reduce the tools organizations need to protect their digital assets. As one example, the platform can autonomously generate cybersecurity designs, capabilities, and / or controls, thereby reducing the overhead associated with managing multiple cybersecurity environments. As another example, the platform can utilize generative sequence processing models to quickly summarize data from human users and / or respond to natural language questions provided by human users. Therefore, cybersecurity platforms can use generative sequence processing models to help users analyze cybersecurity incidents, develop potential cybersecurity protocols, and / or conduct in-depth threat assessments, thereby providing real-time response and adaptability.

[0035] Furthermore, because the model has a broad understanding of various tools and related data structures, it can be used to connect tools that were originally isolated. For example, the model can be used to interpret data from a combination of multiple different tools, enabling analysts or other users to gain a more comprehensive and holistic understanding of security operations.

[0036] According to another aspect of this disclosure, in some implementations, the generative sequence processing model described herein can be fine-tuned to adopt various role images tailored to specific roles within the cybersecurity field. The model can be fine-tuned to adopt specific role images using specific types of data and training tasks with corresponding role characteristics, thereby enabling the model to produce outputs that are context-consistent with the expectations and requirements of different cybersecurity professionals.

[0037] As an example, the model could adopt a role such as a Security Operations Center (SOC) analyst, focusing on rapid, action-based responses to threats; or a role such as a threat intelligence analyst, providing detailed, comprehensive analysis of cyber threats. Other examples include roles like malware or code analysts, focusing on technical analysis of code to identify malicious behavior; and security architects, providing a broad understanding of cybersecurity threats and defenses at the organizational level. Other roles include incident responders, compliance officers, cybersecurity administrators, cybersecurity researchers, chief information security officers (CISOs), and / or forensics analysts, each tailored to perform specific functions within their area of ​​expertise. This capability allows the model to simulate the actions these professionals typically perform, significantly enhancing its usability and applicability across a wide range of cybersecurity environments.

[0038] One benefit is that models can be specialized to focus on specific aspects of a task given the same input and adjust their output style accordingly. For example, a SOC analyst specialization model could be trained to broadly analyze security incidents based on specific strategies (e.g., jumping between events occurring on the same host), and its output might primarily focus on classifying specific sets of events or incidents. In contrast, a threat intelligence analyst role-playing model might be more interested in longitudinally observing similar types of attacks on multiple hosts, associating these attacks with known actor TTPs, and then providing a detailed report including appropriate confidence descriptions common in the field of intelligence analysis. Thus, specialization can be applied to and / or embodied in both the "input focus" (e.g., what is being analyzed) and the "output style." One benefit, particularly in agent-like systems, is that the system can specialize in the task and obtain various perspectives while analyzing something—for example, this might be similar to having an entire AI "security team" review data and then answer questions based on different analytical perspectives.

[0039] According to another aspect of this disclosure, in some implementations, one or more of the generative sequence processing models described herein can operate autonomously or semi-autonomously within a cybersecurity platform, performing typical tasks of various cybersecurity role models. For example, a model fine-tuned for a Security Operations Center (SOC) analyst role can independently generate rapid response strategies for real-time threats, while another model fine-tuned for a threat intelligence analyst can autonomously compile and analyze large amounts of threat data. This autonomous capability improves operational efficiency by allowing human analysts to focus on complex policy decision-making processes.

[0040] Specifically, some example implementations of this disclosure relate to cybersecurity platforms that include a multi-agent setup, where multiple generative sequence processing models, for example, partially or entirely employing different role images, can interoperate to process cybersecurity data and thereby perform cybersecurity actions. For example, each model can autonomously perform tasks consistent with its designated role image, collectively contributing to cybersecurity efforts. The interoperability of these models can be managed through a centralized system or a distributed framework, allowing the insights of one model to inform the actions of other models. This setup improves the efficiency and effectiveness of cybersecurity operations by allowing for division of labor among specialized models, thereby reducing the need for continuous human oversight and increasing interpretability.

[0041] Division of labor often improves results (e.g., true positive rate, etc.) because it does not require a single agent / model to exhaustively analyze all aspects of the problem (during the LLM decoding process, certain actions are effectively prevented regardless of the cueing policy).

[0042] The systems and methods disclosed herein provide several technical effects and benefits. As an example, the proposed techniques enhance digital network security. Specifically, the disclosed systems provide an advanced layer of protection against potential cybersecurity threats, thereby improving the integrity and functionality of digital environments such as cloud environments.

[0043] As another example of technical advantage, this disclosure provides a generative sequence processing model optimized for cybersecurity tasks. Fine-tuning the generative sequence processing model specifically for cybersecurity scenarios improves its ability to perform accurate threat detection and prediction. This reduces false positives and false negatives, thereby ensuring effective resource utilization in real-world threat scenarios.

[0044] As another example of technical advantage, this disclosure provides dynamic threat intelligence aggregation. The sequence processing model's ability to merge data and generate actionable information from various sources, such as SOAR systems or other cybersecurity operational tools or cybersecurity data sources, provides the capability to aggregate and process information from a comprehensive threat environment. This results in a more comprehensive understanding of potential threats, enabling proactive measures against emerging and evolving threats.

[0045] As another example of technological advantage, fine-tuning models for specific cybersecurity data and / or tasks can enhance model specialization. Each model can be fine-tuned to excel at specific tasks for a particular role, such as rapid threat response for a SOC analyst or in-depth code analysis for a malware analyst. This specialization improves the accuracy and efficiency of task execution because each model can leverage its customized training to more skillfully handle specific aspects of cybersecurity.

[0046] Another technological advantage stems from the interoperability between agents in a multi-agent system. By enabling agents to communicate and collaborate, the system ensures that insights and data generated by one agent can inform the actions and analyses of other agents. Furthermore, when each agent is embedded with thought chain analytics capabilities, they can operate with greater transparency and interpretability. By providing detailed explanations and rationales for their actions, agents allow human operators to understand and validate automated processes. This transparency enables verification that the system complies with organizational policies and ethical standards. Thought chains also significantly improve the performance of certain inference tasks, such as mathematical problems.

[0047] Exemplary embodiments of this disclosure will now be discussed in further detail with reference to the accompanying drawings.

[0048] Figure 1 A block diagram of an example computing system including a cybersecurity platform 112 according to an example embodiment of the present disclosure is provided. The cybersecurity platform 112 is configured to enhance security within a potential cloud-based environment. The cybersecurity platform 112 includes a sequence generation processing model 114. Model 114 can be fine-tuned, for example, to address potential security-specific scenarios.

[0049] Sequence processing model 114 can be a model configured to analyze and generate sequence data. Specifically, sequence processing model 114 can include a sequence-to-sequence design that enables the model to receive and process input data sequences to produce corresponding output data sequences.

[0050] Sequence processing model 114 may include or utilize different model architectures. While some examples may utilize the Transformer architecture, which is recognized for its dynamic self-attention mechanism, other examples may be implemented using different architectures, such as recurrent neural networks, convolutional neural networks, or even potential hybrid models.

[0051] One example of a sequence processing model is the so-called "Large Language Model" (LLM). The term LLM can refer to a highly parameterized model configured to interpret and generate text with human-like patterns. Through extensive pre-training on large datasets, LLMs are capable of handling a wide range of natural language processing tasks, from text generation to complex question-answering mechanisms. Another example of a sequence processing model is the "Large Multimodal Model" (LMM).

[0052] Specifically, the sequence processing model 114 can initially be pre-trained on a large amount of pre-training data. Pre-training serves as a crucial stage in developing an effective sequence processing model 114. During this stage, the model is exposed to a large dataset, enabling it to build a comprehensive understanding of data patterns, structure, and inherent relationships. The purpose of pre-training is to embed extensive knowledge into the model, thereby establishing a cognitive framework.

[0053] In some implementations, pre-training of Model 114 may include pre-training Model 114 on various forms of cybersecurity data. Incorporating a broad range of cybersecurity data during the pre-training phase can help Model 114 succeed subsequently on a range of different cybersecurity-related tasks. In some implementations, pre-training may be divided into multiple consecutive phases. In the initial phase, Model 114 may be pre-trained on a large scale on data across multiple different domains (e.g., including computer languages) to instill basic language knowledge and build Model 114's overall language capabilities. In subsequent or "continued" pre-training phases, the focus of pre-training may shift to including a relatively larger amount or proportion of cybersecurity domain-specific data. Therefore, subsequent phases may include more targeted pre-training that includes a higher proportion and greater diversity of cybersecurity-specific data, enabling Model 114 to begin specifically learning cybersecurity information.

[0054] According to this disclosure, after comprehensive learning provided by pre-training, the sequence processing model 114 can undergo a fine-tuning process using extensive cybersecurity data and associated pre-training tasks. This fine-tuning phase focuses on enhancing and refining the model's capabilities within a specific cybersecurity domain or task. By leveraging targeted cybersecurity datasets (which may emphasize certain cybersecurity-specific knowledge or patterns), the model's internal parameters are tuned to optimize its performance on the specified cybersecurity task. Reference Figure 2 Further details and examples of potential fine-tuning of the data and tasks are described.

[0055] Still referencing Figure 1 The cybersecurity platform 112 communicates with cybersecurity operation tools 116, 118, and 120. Cybersecurity operation tools 116, 118, and 120 can be various systems deployed by an organization to perform cybersecurity tasks such as threat detection and prevention. Each of the cybersecurity operation tools 116, 118, and 120 can generate cybersecurity data, which can be obtained by the cybersecurity platform 112 and processed by the sequence processing model 114. The cybersecurity platform 112 can be directly integrated with the cybersecurity operation tools 116, 118, and 120, and / or can communicate with them via an application programming interface (API).

[0056] For example, network security operation tools 116, 118, and 120 may be a security orchestration, automation, and response (SOAR) platform, a security information and incident management (SIEM) system, an intrusion detection system (IDS), an intrusion prevention system (IPS), antivirus software, endpoint protection, vulnerability management tools, and / or other tools. Each of these tools 116, 118, and 120 can create various forms of network security data, which can be transmitted to network security platform 112 (e.g., real-time transmission, on-demand transmission, and / or periodic transmission).

[0057] Therefore, as an example, cybersecurity platform 112 can be operationally integrated with a SOAR system. A SOAR system can include an integrated set of technologies, acting as a comprehensive solution for managing, analyzing, and responding to a range of cybersecurity incidents and threats. It can combine a wide variety of functions, including but not limited to threat and vulnerability management, incident response, and security automation. A SOAR system can be designed to integrate with various security tools and platforms, resulting in a more coherent and unified security infrastructure.

[0058] Key functionalities of a SOAR system may include: threat and vulnerability management; incident response; case management; summary forms and reports; and the creation and deployment of automated operation manuals. SOAR operation manuals may include automated workflows or structured task sequences within the SOAR system. These manuals can automate responses to specific cybersecurity incidents or scenarios, thereby integrating various tools, systems, and processes.

[0059] As another example, the cybersecurity platform 112 can be operationally integrated with a SIEM system. A SIEM system is a comprehensive solution in the cybersecurity field that enables real-time analysis of security alerts generated by various hardware and software infrastructures within an organization. Key functionalities of a SIEM system can include: data aggregation; log storage; event correlation; alerting; data analysis and visualization; forensics and analysis; and compliance reporting. SIEM focuses on detection and alerting, while the SOAR platform emphasizes automation, orchestration, and response. In many established Security Operations Centers (SOCs), SIEM and SOAR solutions work together to provide an end-to-end framework for security incident detection and response.

[0060] As another example, cybersecurity operations tools 116, 118, and 120 may include or provide cybersecurity blogs or information derived therefrom. Cybersecurity blogs may include digital publications or platforms that disseminate information, insights, and analysis related to various aspects of cybersecurity. While these blogs are primarily text-based, they may also include multimedia elements, code excerpts, or interactive demonstrations. They aim to enhance understanding of emerging threats, mitigation technologies, and advancements in the field of information security, although the content and approach may vary depending on the source.

[0061] As another example, cybersecurity operations tools 116, 118, and 120 may include or provide signature-based detection files. These signature-based detection files are used within the cybersecurity operations tools to identify and potentially mitigate known threats based on pre-determined patterns or “signatures.” These signatures, which may contain attributes such as file hashes, string patterns, or behavioral characteristics, can be derived from previous analysis of malicious activity or entities. For example, YARA detection rules utilize patterns to scan for and identify malware samples.

[0062] As another example, cybersecurity operations tools 116, 118, and 120 may include virus or malware detection systems. These systems are designed to identify, isolate, and eliminate malicious software (or “malware”), such as viruses, worms, Trojans, ransomware, etc. The system operates by analyzing system behavior, data patterns, and code to flag potential threats. These systems may also include the capability of automated response mechanisms that take immediate action upon detecting malware, thereby mitigating potential damage and minimizing system downtime. Furthermore, these systems may include components for generating comprehensive logs and reports for further analysis and proactive threat hunting, thereby enhancing the overall cybersecurity posture of entities using such systems.

[0063] As another example, cybersecurity operations tools 116, 118, and 120 may include open-source vulnerability analysis or information. This information may include data or information related to potential weaknesses or sensitivities in open-source software or systems. Such information is typically shared within the cybersecurity community and helps in the discovery, mitigation, and prevention of exploits that could compromise system integrity or security.

[0064] As another example, cybersecurity operations tools 116, 118, and 120 may include secure code repositories. Secure code repositories are dedicated storage spaces for code repositories with a strong emphasis on security. These repositories may include secure code excerpts, libraries, or frameworks. In some implementations, secure code repositories may include examples of malware scripts. Malware scripts can be scripts or sequences of code designed to perform malicious activities against a target system. These scripts may contain a range of functions, from data breaches and system corruption to privilege escalation.

[0065] As another example, cybersecurity operations tools 116, 118, and 120 can include cybersecurity and software development frameworks. These frameworks provide structured, standardized approaches to developing secure software and managing cybersecurity risks. For instance, the MITRE framework offers a comprehensive set of tools, technologies, and best practices for addressing cybersecurity challenges across various domains and industries.

[0066] As another example, cybersecurity operations tools 116, 118, and 120 may include or provide analyst reports. In the broadest sense, an analyst report is a comprehensive document prepared by a technically skilled professional who has conducted a detailed assessment of various aspects of a system, process, or entity. Analyst reports may include threat intelligence reports providing information on emerging cyber threats and vulnerabilities, malware reverse engineering reports dissecting malware to understand its origins, behavior, and mitigation strategies, and incident response reports detailing breach findings, clarifying causes, impacts, and recommending remediation measures.

[0067] Therefore, as an example, cybersecurity operations tools 116, 118, and 120 may include threat intelligence reports, data, or other information. Threat intelligence refers to the collection, analysis, and dissemination of information about potential or current cyber threats and attacks. Threat intelligence is the specialized field of cybersecurity that employs various data collection, processing, and analysis techniques to identify, track, and predict cybersecurity threats and vulnerabilities. Threat intelligence involves systematically collecting information from a variety of internal and external sources to provide comprehensive insights into the cyber threat landscape. This information may include, but is not limited to, detailed information about threat actors, their tools, tactics, and procedures (TTPs), threat indicators (IOCs), and the nature and impact of previous cyberattacks.

[0068] As another example, network security operation tools 116, 118, and 120 may include or provide product documentation. Product documentation may include a compilation of information related to a specific software or hardware product. This documentation is intended to provide guidance to users, administrators, or developers regarding the installation, configuration, use, and troubleshooting of the product.

[0069] The network security platform 112 can also communicate with one or more user devices 122. For example, user device 122 can be associated with a human user. In one example, the human user can be an analyst or administrator responsible for managing the information security of a computer network, for which some or all of the network security operation tools 116, 118, and 120 have been deployed. In other words, in some cases, some or all of the network security operation tools 116, 118, and 120 may be tools specifically deployed for a particular computer network or network group. Alternatively, some or all of the network security operation tools 116, 118, and 120 may be general-purpose tools that are not specific to or targeted at a particular computer network or network group, and are widely accessible regardless of whether the user belongs to a particular organization or network.

[0070] The cybersecurity platform 112 can communicate with one or more user devices 122 to receive questions or queries from user devices 122. The platform 112 can deploy a sequence processing model 114 to generate a response to the query based on the processing of cybersecurity data by the sequence processing model 114. For example, the sequence processing model 114 can retrieve (e.g., via an application programming interface) relevant information needed to answer queries from cybersecurity operations tools 116, 118, and 120. The sequence processing model 114 can then process the retrieved information to generate an appropriate response to the query, which may be, for example, a summary, extract, classification, suggested remedies, or other output. This information can then be supplied to user device 122 for display to the user.

[0071] Figure 2 A block diagram of an example method for fine-tuning a generative sequence processing model according to an example embodiment of this disclosure is provided. Figure 2 As shown, a cybersecurity dataset 202 can be obtained. The cybersecurity dataset 202 can be data generated by one or more cybersecurity operation tools or data associated with one or more cybersecurity operation tools, such as reference tools. Figure 1 Any cybersecurity operational tools described. Therefore, as an example, cybersecurity data 202 may include data from SOAR systems. SOAR data may include SOAR cases, queries, rules, operation manuals, alerts, or other SOAR data. As a further example, cybersecurity data may also include Security Information and Event Management (SIEM) system data, security blog information, threat intelligence reports, signature-based detection files, malware scripts, vulnerability information, product documentation, secure code repositories, cybersecurity and software development framework data, and / or other forms of cybersecurity data.

[0072] The computing system can perform data annotation 204 on the cybersecurity dataset 202. As a result of performing data annotation 204, a fine-tuning tuple 210 can be created. The fine-tuning tuple 210 can include fine-tuning input 206 and fine-tuning label 208.

[0073] The fine-tuning tuple 210 can be used to fine-tune the sequence processing model 114. For example, the fine-tuning input 206 can be provided to the model 114. In response, the model 114 can generate a model output 212. The loss function 214 can be evaluated. The loss function 214 can compare the model output 212 with the fine-tuning label 208. The loss function 214 can be used to update the parameter values ​​of the model 114. For example, the loss function 2114 can be backpropagated through the model 114 to update the parameter values ​​of the model 114.

[0074] More specifically, data annotation 204 may include generating and / or assigning relevant labels, classifications, or annotations (collectively referred to as labels 208) to each of several cybersecurity datasets 202, making them useful for supervised machine learning paradigms. Annotation tools, such as specialized annotation software, can facilitate efficient annotation processes, for example, by having capabilities for batch processing, collaborative annotation, and change monitoring. The annotation process can encompass manual, semi-automatic, or fully automatic methods, depending, for example, on the type of data being annotated.

[0075] Therefore, the fine-tuning label 208 can be metadata, attributes, or other information extracted from or otherwise generated about the cybersecurity dataset 202. In some cases, the fine-tuning input 206 may include some or all of the cybersecurity dataset 202. In some cases, in addition to or instead of the cybersecurity dataset 202, the fine-tuning input 206 may include additional input data, such as additional natural language prompts, questions, queries, or instructions.

[0076] In some implementations, the fine-tuning input 206 and fine-tuning label 208 can be structured according to a question-and-answer format that prompts model 114 to answer questions about the underlying cybersecurity dataset 202. Structured fine-tuning input 206 and fine-tuning label 208 in this way can challenge model 114 to answer using domain-specific terminology not explicitly included in the underlying cybersecurity dataset 202.

[0077] As an example, the cybersecurity dataset 202 could be a set of software code (e.g., potentially malicious code). The fine-tuning label 208 could include information such as: the number of network indicators associated with the software code; whether the code uses network interaction; whether the code edits the registry; and / or other information about the software code that helps train the sequence processing model 114 to understand the software code and improve its ability to explain why the software code might be malicious or not. For example, in this case, the fine-tuning input 206 could include the software code itself along with an additional natural language question or instruction prompting the sequence processing model 114 to generate the fine-tuning label 208. Thus, for example, the fine-tuning input 206 could include the software code and a natural language prompt asking, “How many network indicators are associated with this software code?” The fine-tuning label 208 could then be the labeled answer to that question: the correct or true number of network indicators associated with the software code that has been added as a label.

[0078] As another example, cybersecurity operations tools may have unique patterns and / or domain-specific query languages, enabling users to search for results or other information within the cybersecurity operations tool. In some implementations, cybersecurity dataset 202 may include queries expressed in a domain-specific query language and / or a result set structured according to patterns. For example, cybersecurity dataset 202 may be retrieved from the actual runtime logs of the cybersecurity operations tool, which displays past queries and responses. In this case, data annotation process 204 may include creating natural language queries corresponding to queries expressed in a domain-specific query language. Thus, in this case, fine-tuning input 206 may be a natural language query, and fine-tuning label 208 may be the original query expressed in a domain-specific query language. Creating fine-tuning tuples 210 using this structure enables sequence processing model 114 to learn to translate from natural language into a domain-specific query language and / or associated patterns. Alternatively or concurrently, the model may be trained to perform the opposite task, i.e., translating from structured queries (or similar queries) into natural language hints.

[0079] In some implementations, fine-tuning input 206 and / or fine-tuning label 208 may include examples that enable sequence processing model 114 to perform thought chain processing on the input. Unlike traditional hints where the model directly predicts the answer, the thought chain method prompts the model to generate intermediate inference steps, enabling the model to solve complex multi-step problems in arithmetic and common sense reasoning tasks by decomposing them into manageable components. Therefore, in some implementations, fine-tuning tuple 210 may provide examples that enable sequence processing model 114 to learn to perform thought chain response generation.

[0080] In some implementations, the fine-tuning input 206 and / or the fine-tuning label 208 may include examples that enable the sequence processing model 114 to use external tools. By providing the sequence processing model with access to external tools (e.g., via a text-to-text API), the model can learn to delegate tasks such as arithmetic operations, language translation, and / or accessing real-time information to external services. This allows the sequence processing model to generate more accurate and context-appropriate responses by incorporating information from these external tools into its output, ultimately making it more capable of solving a variety of tasks. Thus, in one example, to answer a question contained in the fine-tuning input 206, the sequence processing model may learn to initially formulate a query in a tool-specific query language, submit the query to the tool (e.g., a cybersecurity operations tool), and then receive and process the query results to generate output. For example, one or more fine-tuning labels 208 may provide illustrative demonstrations of this tool usage process.

[0081] The data annotation process 204 can be performed manually and / or automatically by applying one or more templates and / or one or more machine learning models. As an example, a human user can extract relevant information from a cybersecurity dataset and construct a natural language question associated with the extracted information. The natural language question can be included in the fine-tuning input 206, while the extracted information can be included in the fine-tuning labels 208.

[0082] As another example, one or more templates can be used in the data annotation process 204 to generate fine-tuned tuples 210 from the cybersecurity dataset 202. For example, the templates can apply an algorithm to generate fine-tuned tuples 210 from the cybersecurity dataset 202. In one example, the algorithm may include or apply a set of rules that encodes a “syntax” defining a given output based on received specific input. In some implementations, the algorithm may nondeterministically select the output to avoid overfitting to specific rules.

[0083] As another example, one or more machine learning models can be used in the data annotation process 204 to generate fine-tuned tuples 210 from the cybersecurity dataset 202. For example, the cybersecurity dataset 202 can be provided to the machine learning model, prompting it to generate useful fine-tuning inputs 206 and fine-tuning labels 208. As an example, the cybersecurity dataset 202 can be a query expressed in a domain-specific query language, and the machine learning model's task is to transform that query into a natural language query to generate fine-tuned tuples 210, which include the natural language query as fine-tuning input 206 and the query expressed in the domain-specific query language as fine-tuning labels 208. In some implementations, the one or more machine learning models used in the data annotation process 204 may differ from the sequence processing model 114. For example, the one or more machine learning models used in the data annotation process 204 can be a larger model, a teacher model, a model trained on different data, and / or other models. In some implementations, the one or more machine learning models used in the data annotation process 204 may be the sequence processing model 114 itself. For example, some of these implementations may be referred to as self-supervised training methods.

[0084] More generally, it can be executed Figure 2 The data annotation process 204 shown generates numerous fine-tuning tasks that go beyond basic language modeling tasks such as "next word prediction" or "masked language modeling." Instead, fine-tuning tasks can be tasks that require or trigger a deeper understanding of the concepts present in the cybersecurity data 202, thereby training the model to understand the cybersecurity data 202 more deeply and be able to perform more meaningful analysis on the cybersecurity data 202. As an example, fine-tuning tasks can include classification tasks, such as file type classification, malware detection, vulnerability identification, actor attribution, file activity classification, or other classification tasks. As another example, fine-tuning tasks can include summarization tasks, such as generating an execution summary of a threat intelligence report, interpreting programming or computer language content (e.g., malicious scripts) in natural language, and / or summarizing structured security operations tool data (e.g., cybersecurity incident and alert metadata) in natural language (programming language content). As a further example, fine-tuning tasks can include generation tasks, such as generating a threat intelligence report from a structured programming language script, creating a SOAR operations manual, and / or translating natural language into a specialized query language, such as a SOAR query. As a further example, fine-tuning tasks may include extraction tasks (such as extracting relevant vulnerability information or entities) and / or transformation tasks (such as transforming cybersecurity threat intelligence into actionable security measures).

[0085] Figure 3An example method is described in which the generative sequence processing model 114 is fine-tuned to adopt one or more different role models, each tailored to a different cybersecurity role. This capability enables model 114 to ultimately produce output that is context-aware and consistent with the expected requirements of a variety of cybersecurity professionals.

[0086] Various techniques can be used to fine-tune the generated sequence processing model 114 to adopt different role images. Example techniques may include differential privacy optimization (DPO), cue tuning (also known as “soft cue” tuning), and / or various types of parametrically efficient model adapters (e.g., low-rank adaptive (LORA) tuning). Each technique can be applied individually or in combination to enhance the model’s capabilities. Fine-tuning model 114 to adopt one or more different role images can enhance its adaptability and effectiveness across various cybersecurity roles.

[0087] This model can be fine-tuned to adopt or represent one or more different role models. One example is the Security Operations Center (SOC) Analyst model 310. The SOC Analyst role model 310 is characterized by an emphasis on rapid, tactical action-based responses. The fine-tuning process for this role model 310 can be designed to emphasize output generation, thereby facilitating immediate decision-making and enabling rapid responses to security threats.

[0088] The sequence processing model 114 can be fine-tuned to generate a SOC analyst role model 310 using specific types of data and / or training tasks. For example, fine-tuning data could include real-time threat alerts, event logs, and / or SOC-specific tactical response protocols. Example training (e.g., fine-tuning) tasks could focus on rapidly classifying threats, prioritizing event responses, and / or generating immediate action steps. Tasks can be structured to improve the model's ability to generate concise, actionable outputs. These outputs can facilitate rapid decision-making, which is crucial for the SOC analyst role. Training can leverage scenarios simulating environments requiring rapid response, enabling the model to operate effectively under conditions frequently encountered by SOC analysts.

[0089] Another example role model could be Threat Intelligence Analyst Model 312. Fine-tuning this role model 312 can emphasize the ability to provide comprehensive analysis, including detailed information and / or supporting material. For example, the output of this role model 312 could include detailed information such as context, policy implications, and / or the motivations behind cyber threats. In some cases, this role model 312 may exhibit a tendency to provide hedging and cautious interpretations while conducting lengthy analyses.

[0090] Model 114 can be fine-tuned to generate a threat intelligence analyst role model 312 using specific data types and / or training tasks. Fine-tuning data may include threat intelligence reports, policy documents, and / or analyses of historical cybersecurity incidents. Training tasks may include generating detailed summaries of materials, identifying and interpreting the impact of various cyber threats, and / or synthesizing information to outline potential policy implications. Furthermore, the model can be trained to manage tasks associated with generating cautious and hedging analyses that reflect the prudent considerations of a threat intelligence analyst. Training enables Model 312 to provide rich and detailed outputs suitable for strategy planning and decision-making in cybersecurity operations.

[0091] Another example role is like malware or code analyst model 314. This role-like model 314 can be tailored to focus on highly technical analysis of code behavior. For example, model 314 can be trained to identify specific behaviors and characteristics of code based on technical details.

[0092] The generated sequence processing model 114 can be fine-tuned to generate malware or code analyst persona model 314 using specific types of data and / or training tasks. Example fine-tuning data may include datasets of malware scripts, executable binaries, and / or code extracts. These elements may exhibit malicious behavior. Example training tasks can focus on classification challenges, such as distinguishing between benign and malicious code, identifying malware types, and / or detecting obfuscation techniques used by malware authors. Alternatively, model 314 can be trained on tasks involving annotating code with comments explaining the purpose and function of code segments. This can aid in understanding and identifying malicious intent within the code. Such focused training can enable model 314 to generate highly technical and detailed analyses, thus embodying the characteristics of a malware or code analyst.

[0093] It can also generate various other persona models 316. As an example, another example persona is the security architect persona. The security architect persona model can provide a conceptual understanding of cybersecurity threats and defenses. This persona model can generate outputs describing the technical nuances and analyses at the organizational level.

[0094] The generated sequence processing model 114 can be fine-tuned to generate a security analyst role model using specific types of data and / or training tasks. For example, training could include architecture diagrams, system configuration data, and / or corporate security policies. The model could access policy planning documents and / or high-level threat assessment reports. Further example training tasks could include generating an organizational security posture summary, proposing improvement recommendations based on emerging threats, and / or merging complex security requirements into actionable policies. Training can enable the model to produce outputs reflecting a broad understanding of cybersecurity. This understanding can be further customized to the specific organizational context.

[0095] Another example role model is the incident responder role model. This role model can focus on making a rapid and effective response to security vulnerabilities or incidents. The model can be fine-tuned to prioritize rapid assessment and mitigation strategies. It can provide step-by-step guidance for the containment and remediation process.

[0096] Another example role model is the compliance officer role model. This role model can be configured to address regulatory and compliance requirements. The model can generate outputs that identify potential compliance issues, recommend corrective actions, and / or ensure compliance with applicable laws and standards. Examples of these standards include GDPR, HIPAA, or PCI-DSS.

[0097] Another example role model is that of a network security administrator. This role model can focus on the security aspects of network infrastructure. The model can be trained to analyze network traffic and configurations, thereby providing insights and recommendations to enhance network security and prevent unauthorized access.

[0098] Another example role model is that of a cybersecurity researcher. This role model can be designed to explore and analyze new and emerging cybersecurity threats and technologies. Models with this role model can be used to simulate cyberattacks, analyze threat patterns, and / or help develop new defense mechanisms.

[0099] Another example role model is the Chief Information Security Officer (CISO) role model. This executive role model can be designed to provide strategic insights, assess the organization's overall security posture, and / or provide risk assessments. Models with this role model can be used to provide policy recommendations and / or develop strategic plans to enhance the organization's cybersecurity resilience.

[0100] Therefore, the generative sequence processing model can be fine-tuned to simulate different role models. This allows the resulting model to produce outputs tailored to meet the needs of various cybersecurity professionals and / or to simulate the actions performed by various cybersecurity professionals. The model's ability to adopt different role models enhances its practicality and applicability in various contexts within the cybersecurity field.

[0101] Figure 4 An example process for performing inference using a finely tuned sequence processing model 114 according to an exemplary embodiment of this disclosure is described. More specifically, once the sequence processing model 114 has been trained, it can be deployed to perform many different inference tasks. Figure 4 As shown, in order to perform inference tasks, sequence processing model 114 can receive and process model input 402 to generate model output 404. Because model 114 has been fine-tuned for a wide range of tasks, it can also perform a large number of different inference tasks. Optionally, model 114 may or may not undergo further fine-tuning to adopt a specific role image.

[0102] In some implementations, the inference task can be a classification task. As an example, model 114 can perform classification on file types. File type classification can include automatically classifying input files into one or more different file formats based on unique attributes or patterns discovered within the data. As an example, model input 402 can include one or more files with unrecognized formats, including documents, images, audio files, and / or videos. Example model output 404 can classify and label each file, such as "Document-PDF", "Image-JPEG", "Audio-MP3", and "Video-MP4".

[0103] As another example, model 114 can perform malware detection. Malware detection can be performed using model 114 to scan, analyze, and determine whether software or code snippets contain malicious intent that could compromise the security of digital systems. Example model input 402 can include one or more executable files, some of which may contain hidden malware or spyware components. After analysis, example model output 404 may label certain files as "malicious-Trojan" or "safe-standard executable," providing an immediate risk assessment for each file. In some implementations, model output 404 may include a natural language explanation of the reasons for a particular file classification. For example, model output 404 may describe indicators or other information from or about the file to explain the reasons for giving a certain malware classification.

[0104] As another example, Model 114 can perform vulnerability identification in code. Model 114 can evaluate the code structure to identify potential vulnerabilities that could be targets for hackers. Example model input 402 can include one or more pieces of software code that may or may not contain embedded flaws. Example model output 404 can highlight specific sections of code, labeling them as "Vulnerable - SQL Injection Risk" or "Safe - No Vulnerability Detected." This helps developers take preventative measures before deployment. Again, in some implementations, model output 404 can include a natural language explanation of the reasons for the specific classification of a file or detected vulnerability.

[0105] In some implementations, the reasoning task can be a summarizing task. As another example, model 114 can generate an execution summary of a report generated by threat intelligence or other analysts. Generating execution summaries from threat intelligence can help decision-makers by distilling complex cybersecurity intelligence into easily understandable reports. Example model input 402 can include a detailed and verbose report on recent cybersecurity vulnerabilities, detailing attack pathways, affected systems, and potential consequences. Example model output 404 can be a concise execution summary highlighting the primary threat source, the most affected business units, the critical vulnerabilities exploited, and / or recommended corporate preventative measures.

[0106] In some implementations, Model 114 may perform certain tasks (e.g., summarizing tasks) using specific organizational data to personalize the model's output. For example, in some implementations, one or more personalized threat profiles targeting a specific entity, network, or organization may be incorporated along with the original data to be summarized. In this case, Model 114 can specialize the summary based on the specific needs defined or inferred from the personalized threat profiles. For instance, the summary generated by Model 114 might ignore threats irrelevant to a particular organization, focusing only on those that are highly probable or have a significant impact. As an example, if a personalized threat profile indicates that the organization does not use a certain software program, the summary generated by Model 114 may not discuss a vulnerability in that software program. On another front, the summary generated by Model 114 may ensure that new attack tools used by threat actors against other similar organizations are prominently highlighted.

[0107] In some implementations, model 114 can perform summarization using an iterative, hierarchical approach, where model 114 summarizes large, individual sets of data elements (e.g., alerts) separately to generate individual alert summaries. The model can then reprocess all the individual alert summaries together to generate a single, more comprehensive alert summary.

[0108] As another example, the model 114 can interpret malicious scripts. For example, the model can translate technical malicious code behaviors into more understandable content. This summary helps to understand the true intent and impact of malware. The example model input 402 can include a malicious script (e.g., a Python script that launches a keylogger and sends the recorded keystrokes to a remote server). The example model output 404 can include a short summary statement: "This script is designed to secretly record a user’s keystrokes and send the data to an external server, potentially compromising private information."

[0109] As another example, the model 114 can summarize structured JSON content in natural language. Converting structured JSON content into a natural language summary can help bridge the gap between machine-readable data and human understanding, enabling more direct data interpretation. The example model input 402 can include a JSON object such as ‘{"name": "Erica", "age": 28, "profession": "Engineer", "city": "New York"} ({"姓名": "Erica", "年龄": 28, "职业": "工程师", "城市": "纽约"})’. The example output can include a more understandable summary statement: "Erica is a 28-year-old Engineer from New York."

[0110] In some implementations, the inference task can be a generation task. As an example, model 114 could generate a threat intelligence report from structured programming language content, such as a structured JSON file. This task could include automatically generating a comprehensive cybersecurity narrative based on structured data such as a JSON file. Example model input 402 could include a structured JSON file containing attributes such as {“threatActor”: “BT468”, “vector”: “Phishing”, “targetedSector”: “Finance”, “malwareUsed”: “DancingBear”}. Example model output 404 could include a generated report stating, “The threat actor BT468 has recently launched a phishing attack targeting the finance sector, deploying the malware known as DancingBear.”

[0111] As another example, Model 114 can automatically generate a SOAR operations manual. The synthesis of the SOAR operations manual by Model 114 can utilize specific parameters to formulate a systematic action plan for the SOAR system. Example model input 402 can include parameters such as "ransomware detection on endpoints," "isolate compromised machines," "scan the network for similar threats," and "notify the IT manager." Example model output 404 can include a SOAR operations manual detailing the following steps, such as "1. Once ransomware is detected on any endpoint, immediately isolate the compromised machine. 2. Initiate a full network scan to identify similar threats. 4. Automatically send a notification to the IT manager detailing the incident."

[0112] As another example, Model 114 can translate natural language into a domain-specific query language. This inference task can include translating human-language queries into specialized query syntax, enabling users with minimal technical knowledge to extract data from a database or perform specialized tasks. Example model input 402 can include user input stating: “Find all employees in the sales department who started after January 2020.” Example model output 404 can include a translated SQL query, such as “SELECT…” FROM employees WHERE department = 'sales' ANDstartDate > '2020-01-01'. "

[0113] As another example, Model 114 can automatically generate computer code to perform various cybersecurity-related tasks. For instance, Model 114 can automatically generate computer code that, when executed, "obfuscates" an existing codebase (e.g., open-source code). This computer code, which may be referred to as a "obfuscator," can iteratively test the existing codebase or "obfuscate" it to test for vulnerabilities. Example model input 402 may include an existing codebase and multiple pre-existing obfuscation frameworks or libraries. Example model output 404 may include obfuscator code that, when executed, obfuscates the existing codebase.

[0114] In some implementations, the inference task can be a content extraction task. As an example, model 114 could extract relevant vulnerabilities from a blog into a structured representation such as a JSON file. The act of extracting vulnerability details from a blog narrative and presenting those details in a structured format (e.g., a JSON file) can streamline the cybersecurity analysis process. This is helpful for security professionals who need to quickly absorb information from multiple sources. Example model input 402 could include a blog post detailing, “The latest vulnerability discovered in the Doors Operating System allows unauthorized access to user data. Dubbed ‘OpenDoor,’ this flaw primarily affects version 16 and can be mitigated using the patch released on August 5, 2023.” The example model output 404 can be a structured JSON object, such as {“vulnerabilityName”: “OpenDoor”, “affectedSystem”: “Doors OS”, “version”: “16”, “mitigation”: “Patch released on August 5, 2023”}.

[0115] As another example, Model 114 can identify important / relevant entities in unstructured natural language content. Extracting relevant entities from unstructured natural language material enables a more focused and contextual understanding of the data. In this task, Model 114 can accurately identify specific keywords, names, dates, or other valuable information fragments within large amounts of text. Example model input 402 could include an article stating, "Moonshot's Resolution rover successfully landed on Phobos on February 18, 2026, marking a historic moment in space exploration." The example model output 404 can include extracted entities such as {“organization”: “Moonshot”, “object”: “Resolution rover”, “event”: “landed”, “location”: “Phobos”, “date”: “February 18, 2026”} ({“organization”: “Moonshot”, “object”: “Resolution rover”, “event”: “landed”, “location”: “Phobos”, “date”: “February 18, 2026”}).

[0116] As another example, model 114 can extract relevant information from file metadata. For instance, some malware or other malicious files may be password-protected, but the password may be embedded in the filename in an obscure way. Therefore, example model input 402 could include the filename. Example model output 404 could include the password already extracted from the filename. Extracting passwords in this way allows network security systems to scan files to detect malware and take appropriate countermeasures. For example, model 114 could suggest one or more appropriate remediation measures based on detected malware. For example, model 114 could select one or more remediation measures from a set of possible remediation measures.

[0117] As another example, model 114 can summarize or extract relevant information from the attack graph. The attack graph can include structured information describing a series of steps an attacker will take to achieve a specific goal. Summarizing or extracting information from the attack graph can help users understand and avoid potential threats. Example model input 402 can include a structured attack graph. Example model output 404 can include a natural language summary of the attack graph and / or a list of extracted information (such as entities).

[0118] Figure 5 This diagram illustrates the architecture of a role-based multi-agent cybersecurity system within a cybersecurity platform 550. This configuration is characterized by a planning agent 500, which acts as a central hub coordinating operations among various role-based models. These models include a SOC analyst role-based model 502, a threat intelligence analyst role-based model 504, a malware or code analyst role-based model 506, and possibly other role-based models 508. Each model is fine-tuned according to its assigned role-based character to perform specific cybersecurity tasks, as shown in the reference diagram. Figure 3 As described.

[0119] The planning agent 500 is responsible for managing and guiding the flow of cybersecurity data 552 to appropriate role-image models, and ensuring that the outputs from these models are effectively used to generate cybersecurity actions 554. This agent 500 can handle task assignment and prioritize actions based on real-time threat assessment.

[0120] In some implementations, each role model operates autonomously or semi-autonomously. For example, the SOC analyst role model 502 can be fine-tuned to quickly respond to detected cybersecurity threats by processing real-time data streams and executing predefined response strategies. Meanwhile, the threat intelligence analyst role model 504 can autonomously collect and synthesize large amounts of cybersecurity data to generate detailed reports that aid in strategic decision-making. The malware or code analyst role model 506 can focus on technically inspecting potential malicious code, providing insights that help understand and mitigate malware contamination.

[0121] Interactions between these models can be managed through a centralized management system facilitated by Planning Agent 500, and / or via a distributed framework where models can communicate and collaborate directly. This multi-agent setup can improve the efficiency and effectiveness of cybersecurity operations by dividing tasks among specialized models.

[0122] In operation, the planning agent 500 can receive input including cybersecurity data 552, such as alerts to suspicious network activity, and assign tasks among role models based on their specialization. For example, upon receiving detailed information about suspicious activity, the planning agent 500 can instruct the SOC analyst role model 502 to immediately take containment action, while instructing the malware or code analyst role model 506 to conduct a more in-depth technical analysis of any associated payloads. The threat intelligence analyst role model 504 can then use insights from these analyses to assess the broader impact of the attack and develop policy adjustment recommendations by consulting with the security architect model.

[0123] In some implementations, the planning agent 500 can synthesize the outputs of various models (which are then received back as input to the agent 500). These inputs can come from SOC analyst model 502, malware / code analyst model 506, threat intelligence analyst model 504, and security architect model 508. The planning agent 500 can generate a set of output actions 555. These actions can include immediate steps to further contain and eliminate threats. These actions can also include long-term strategic measures aimed at enhancing the organization's overall cybersecurity posture.

[0124] For example, by Figure 5 The cybersecurity actions 554 generated by the multi-agent system in the context can include a variety of specific measures tailored to effectively address identified threats. For example, these actions may include automatically isolating affected network segments to prevent malware propagation, deploying patches or updates to vulnerable systems, initiating password resets or blocking access to compromised user accounts, and / or configuring firewalls or intrusion detection systems to enhance defenses against detected threats. As a further example, cybersecurity actions 554 may include generating alerts about ongoing security incidents and disseminating them to relevant stakeholders, creating detailed incident reports for forensic analysis, and / or recommending policy changes to cybersecurity policies or practices based on insights gathered from role-playing models.

[0125] Following an event, the planning agent 500 can collect and analyze feedback regarding the effectiveness of deployed actions. This data can be used to refine the model. This refinement can improve the model's accuracy and effectiveness in predicting future events.

[0126] In some implementations, each model or agent within the cybersecurity platform is fine-tuned to employ a thought-chain style of analysis and / or prompted to follow such analysis. This enhancement allows models such as SOC analyst role-like model 502, threat intelligence analyst role-like model 504, and malware or code analyst role-like model 506 to effectively handle complex multi-step cybersecurity issues by outlining their inference process in a step-by-step manner. This approach simulates human-like problem-solving, enabling each agent to break down complex cybersecurity scenarios into easily understandable pieces and provide detailed rationales for their decisions. This capability increases the transparency and trustworthiness of these systems.

[0127] In some implementations, each model in the system is fine-tuned to meet performance standards based on criteria developed by subject matter experts. These criteria define specific standards and benchmarks, outlining the essential skills and knowledge required for each role. Adhering the fine-tuning process to these expert-defined criteria ensures that each model not only performs tasks autonomously but also meets the high-performance standards expected in real-world operations.

[0128] In some implementations, the planning agent 500 may include a machine learning model trained to assess the situation and efficiently manage resource allocation among the agents. In other implementations, the agent's behavior may be modulated through a framework of prompts, constraints, criteria, and / or guidance to ensure that each agent operates within predefined operational boundaries and conforms to organizational policies.

[0129] In some implementations, the planning agent 500 can leverage a tool usage framework to effectively manage and guide the activities of other agents, such as SOC analyst role-like model 502, threat intelligence analyst role-like model 504, and malware or code analyst role-like model 506. This framework is analogous to how tools work with large language models (LLMs), where tools are essentially functions or APIs that the LLM can call to perform specific tasks. For example, the planning agent 500 can use APIs to call other models, such as SOC analyst role-like model 502 or threat intelligence analyst role-like model 504. These API calls can be dynamically generated as outputs of the planning agent 500, ensuring that the appropriate model is activated based on real-time analysis and contextual needs.

[0130] In some implementations, various generative sequence processing models associated with different agents, such as SOC analyst role model 502, threat intelligence analyst role model 504, and malware or code analyst role model 506, are initially forked from a common pre-trained model. This base model provides a broad foundation of general functionality, which is then specialized using parametrically efficient adapters. These adapters selectively fine-tune each forked model to enhance its specific capabilities associated with its corresponding role within the cybersecurity platform.

[0131] therefore, Figure 5 A role-based multi-agent cybersecurity system is demonstrated. This system can exhibit collaborative capabilities. Each component or model within the system can specialize to perform different tasks. These tasks can correspond to various roles commonly found in cybersecurity operations. Assembling these models under the supervision of 500 planning agents ensures that cybersecurity threats are managed in a unified and comprehensive manner.

[0132] Although Figure 5 An example platform utilizing a centralized planning approach is shown, but other example implementations can leverage distributed planning methods. In such a distributed environment, each agent within the cybersecurity platform can autonomously interact with other agents without the need for a centralized control entity. This allows for dynamic problem-solving, where tasks can be subdivided and assigned among agents based on real-time needs and agent expertise.

[0133] In this distributed model, each role-based agent can operate independently as a planning agent, capable of initiating communications with other agents to request information or actions beyond its own area of ​​expertise. For example, a malware or code analyst role might detect anomalies in a code sequence and can then interact directly with a threat intelligence analyst role to see if similar patterns have been observed in recent cybersecurity threats. This direct inter-agent interaction facilitates a more flexible and responsive approach to threat detection and response, as agents can dynamically collaborate based on contextual information.

[0134] Furthermore, in some implementations, these agents can operate recursively, meaning that an agent responsible for a complex problem can break it down into smaller, more manageable tasks and delegate these tasks to other agents, or even to itself at varying capacities. This recursive subdivision enables the platform to dynamically scale its problem-solving capabilities to adapt to the complexity and urgency of the threats it encounters.

[0135] Figure 6 A flowchart depicts a method 600 for training one or more machine learning models according to various aspects of this disclosure. For example, an example machine learning model may include a generative sequence processing model.

[0136] One or more portions of Example Method 600 may be implemented by a computing system (such as the computing system described, for example, with reference to other figures) including one or more computing devices. Each corresponding portion of Example Method 600 may be performed by any one (or any combination of) of the one or more computing devices. Furthermore, one or more portions of Example Method 600 may be implemented on the hardware components of the apparatus described herein, for example, to train one or more systems or models. Figure 6 For illustrative and discussion purposes, elements executed in a specific order are depicted. Those skilled in the art will understand using the disclosure provided herein that elements of any of the methods discussed herein can be adapted, rearranged, extended, omitted, combined, or modified in various ways without departing from the scope of this disclosure. Figure 6 The description is for illustrative purposes only and with reference to elements / terms described with respect to other systems and figures, and is not intended to be limiting. One or more portions of example method 600 may be performed additionally or alternatively by other systems.

[0137] At 602, example method 600 may include obtaining training instances. The training dataset may include multiple training instances partitioned across multiple datasets (e.g., training datasets, validation datasets, or test datasets). Training instances may be labeled or unlabeled. Although referred to as “training” instances in example method 600, it should be understood that runtime inference may form training instances when a model is trained using an evaluation of the model’s performance on that runtime instance (e.g., online training / learning). Example data types for training instances and various tasks associated with them are described throughout this disclosure.

[0138] At 604, example method 600 may include using one or more machine learning models to process training instances to generate output. The output may be obtained directly from the one or more machine learning models, or it may be a downstream result of a chain of processing operations that includes the output of the one or more machine learning models.

[0139] At 606, example method 600 may include receiving an evaluation signal associated with the output. The evaluation signal can be obtained using a loss function. Various losses can be used, such as mean squared error, likelihood loss, cross-entropy loss, hinge loss, contrastive loss, or various other loss functions. The evaluation signal can be computed using known baseline true labels (e.g., supervised learning), predicted or estimated labels (e.g., semi-supervised or self-supervised learning), or without labels (e.g., unsupervised learning). The evaluation signal can be a reward (e.g., for reinforcement learning). The reward can be computed using a machine learning reward model configured to generate a reward based on the received output. The reward can also be computed using feedback data describing human feedback to the output.

[0140] At 608, example method 600 may include using an evaluation signal to update a machine learning model. For example, in some embodiments, various training or learning techniques, such as backpropagation, may be used to learn the values ​​of the parameters of the machine learning model. For example, the evaluation signal may be backpropagated from the output (or another source of the evaluation signal) through the machine learning model to update one or more parameters of the model (e.g., based on the gradient of the evaluation signal relative to the parameter values). For example, a system containing one or more machine learning models may be trained in an end-to-end manner. Gradient descent techniques may be used to iteratively update the parameters over multiple training iterations. In some implementations, performing backpropagation of error may include performing truncated backpropagation over time. Example method 600 may include implementing various generalization techniques (e.g., weight decay, dropout, etc.) to improve the generalization ability of the model being trained.

[0141] In some implementations, example method 600 can be implemented to train a machine learning model from an initial state to a fully trained state (e.g., when the model exhibits a desired performance profile, such as based on accuracy, precision, recall, etc.).

[0142] In some implementations, example method 600 can be implemented for a specific stage of the training process. For example, in some implementations, example method 600 can be implemented for pre-training a machine learning model. Pre-training can include, for example, large-scale training on potentially noisy data to achieve a broad performance level base across multiple tasks / data types.

[0143] In some implementations, Example Method 600 can be implemented for fine-tuning a machine learning model. Fine-tuning may include, for example, smaller-scale training on higher-quality (e.g., labeled, curated, etc.) data. Fine-tuning can affect all or part of the parameters of the machine learning model. For example, parts of the machine learning model may be “frozen” for certain training phases. For example, parameters associated with the embedding space may be “frozen” during fine-tuning (e.g., to preserve information learned from a broader domain than that present in the fine-tuned dataset). In some implementations, Example Method 600 uses an adapter module. The adapter may be a small trainable layer inserted between pre-existing layers of the pre-trained model. During the fine-tuning process, the original parameters of the pre-trained model are typically frozen, and only the parameters of the adapter are updated.

[0144] In some implementations, example method 600 can be implemented to perform parameter-efficient fine-tuning methods, such as Residual Hierarchical Optimization (LoRA). LoRA can refine a pre-trained model with minimal adjustments to the original parameters. This can be achieved by introducing trainable low-rank matrices that modify the behavior of the pre-trained weights without directly altering them. In some implementations, only these auxiliary matrices are updated during fine-tuning, which significantly reduces the number of parameters being trained.

[0145] Example fine-tuning methods include reinforcement learning. Reinforcement learning can be based on user feedback on model performance during use.

[0146] Figure 7 This is a block diagram of an example processing flow for using machine learning model 1 to process input 2 to generate output 3.

[0147] Machine learning model 1 can be or includes one or more machine learning models or model components. Example machine learning models can include neural networks (e.g., deep neural networks). Example machine learning models can include non-linear or linear models. Instead of or in addition to neural networks, example machine learning models can use other architectures. Example machine learning models can include decision tree-based models, support vector machines, hidden Markov models, Bayesian networks, linear regression models, k-means clustering models, etc.

[0148] Machine learning model 1 can be, include, or otherwise represent any one or more of the machine learning models described above with respect to the foregoing figures. For example, machine learning model 1 can be, include, or otherwise represent any one or more of the generative sequence processing models, etc. Although the various features, variations, and implementations described below are described with respect to machine learning model 1, it should be understood that such features, variations, and implementations should be understood to be described with respect to each of the generative sequence processing models, etc., and any other machine learning components described herein.

[0149] Example neural networks can include feedforward neural networks, recurrent neural networks (RNNs) (including recurrent neural networks based on long short-term memory (LSTM), convolutional neural networks (CNNs), diffusion models, generative adversarial networks, or other forms of neural networks. Example neural networks can also be deep neural networks. Some example machine learning models may utilize attention mechanisms, such as self-attention. For example, some example machine learning models may include multi-head self-attention models.

[0150] Machine learning model 1 may include one or more instances of the same model configured to operate on data from input 2. Machine learning model 1 may also include multiple different models or multiple different model parts configured to operate on data from input 2.

[0151] Machine learning model 1 can include an ensemble of different models that can collaborate and interact to process data from input 2. For example, the model ensemble can include multiple models with different properties (e.g., different architectures, trained with different formulas, etc.). The ensemble can output an overall output based on the individual outputs of the constituent models. In this way, for example, different constituent models can work together to provide system-level robustness by effectively aggregating the individual strengths and weaknesses of any given model. The corresponding individual outputs can be combined using weighted combinations, voting or routing mechanisms, or learned output layers (e.g., one or more feedforward or fully connected layers).

[0152] Machine learning model 1 can employ a hybrid expert structure. For example, see Zhou et al. Mixture of Experts with Expert Choice Routing , arXiv:2202.09368v2 (October 14, 2022). For example, different parts of the model can learn (explicitly or implicitly) different expert domains, where the path through the model is selected by a learned routing mechanism that uses the appropriate expert for a given input (e.g., a given part of the input, such as on a per-word basis). For example, the feedforward network can be sparsely activated for that part of the input based on the output of the routing mechanism that processes that given part of the input. For example, in this way, the reorganized activated weights can form the “expert” selected by the router. In each forward pass, only a subset of the total model weights can be used, thus reducing the amount of operations performed to process a given input compared to a densely activated model. In this way, for example, the expressiveness and interpretability of high parameter counting models can be achieved through computationally more efficient forward passes.

[0153] Input 2 can typically include or otherwise represent various types of data. Input 2 can include one type or many different types of data. Output 3 can be data of the same type or a different type compared to Input 2. Output 3 can include one type or many different types of data.

[0154] Example data types for input 2 or output 3 include natural language text data, software code data (e.g., source code, object code, machine code, or any other form of computer-readable instruction or programming language), machine code data (e.g., binary code, assembly code, or other forms of machine-readable instruction that can be executed directly by a computer's central processing unit), assembly code data (e.g., a low-level programming language that uses a symbolic representation of machine code instructions to program the processing unit), genetic data or other chemical or biochemical data, image data, audio data, audiovisual data, tactile data, biometric data, medical data, financial data, statistical data, geographic data, astronomical data, historical data, and generally sensor data (e.g., digital or analog values, such as voltage or other absolute or relative level measurements from real or artificial inputs, such as from audio sensors, light sensors, displacement sensors, etc.). Data can be raw or processed and can be in any format or mode.

[0155] In multimodal input 2 or output 3, example combinations of data types include image data and audio data, image data and natural language data, natural language data and software code data, image data and biometric data, sensor data and medical data, etc. It should be understood that any combination of data types in input 2 or output 3 can exist.

[0156] Example input 2 may include one or more data types, such as the example data types noted above. Example output 3 may include one or more data types, such as the example data types noted above. The data type of input 2 may be the same as or different from the data type of output 3. It should be understood that the example data types noted above are provided for illustrative purposes only. The data types contemplated within the scope of this disclosure are not limited to those examples noted above.

[0157] Figure 8 This is a block diagram illustrating an example implementation of an example machine learning model configured to process information sequences. For example, an example implementation of machine learning model 1 could include machine learning sequence processing model 4. The example system could pass input 2 to sequence processing model 4. Sequence processing model 4 could include one or more machine learning components. Sequence processing model 4 could process the data from input 2 to obtain input sequence 5. Input sequence 5 could include one or more input elements 5-1, 5-2, ..., 5- obtained from input 2. M The sequence processing model 4 can use the prediction layer 6 to process the input sequence 5 to generate the output sequence 7. The output sequence 7 may include one or more output elements 7-1, 7-2, ..., 7- generated based on the input sequence 5. N The system can generate output 3 based on output sequence 7.

[0158] Sequence processing models 4 may include one or more machine learning model components configured to ingest, generate, or otherwise infer sequences of information. For example, some example sequence processing models in the text domain are referred to as “large language models” or LLMs. See, for example, the PaLM 2 technology report, Google, https: / / ai.google / static / documents / palm2techreport.pdf (nd). Other example sequence processing models may operate in other domains, such as the image domain, for example, see Dosovitskiy et al. An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale (An image is equivalent to 16x16 characters: large-scale image) Like the Transformer for recognition) arXiv:2010.11929v2 (June 3, 2021); audio domain, for example, see Agostinelli et al. MusicLM: Generating Music From Text happy)arXiv:2301.11325v1 (January 26, 2023); Biochemical domain, for example, see Jumper et al., Highlyaccurate protein structure prediction with AlphaFold, 596 Nature 583 (August 26, 2021). Sequence processing model 4 can process one or more types of data simultaneously. Sequence processing model 4 can include relatively large models (e.g., more parameters, computationally intensive, etc.), relatively small models (e.g., fewer parameters, computationally lightweight, etc.), or both.

[0159] Generally, sequence processing model 4 can use data from input 2 to obtain input sequence 5. For example, input sequence 5 may include a representation of the data from input 2 in a format understood by sequence processing model 4. One or more machine learning components of sequence processing model 4 may ingest data from input 2, parse the data into fragments compatible with the processing architecture of sequence processing model 4 (e.g., via “word segmentation”), and project the fragments into the input space associated with prediction layer 6 (e.g., via “embedding”).

[0160] Sequence processing model 4 can ingest data from input 2 and parse the data into a sequence of elements to obtain input sequence 5. For example, a portion of the input data from input 2 can be decomposed into segments, which collectively represent the content of said portion of the input data. The segments can provide the elements of the sequence.

[0161] In some cases, elements 5-1, 5-2, ..., 5- M Elements can represent building blocks used to capture or express meaningful information within a specific data domain. For example, an element can describe "atomic units" across one or more domains. For instance, for a text input source, an element can correspond to a group of one or more words or sub-word components (such as a set of one or more characters).

[0162] For example, elements 5-1, 5-2, ..., 5- M This can represent the tokens obtained using a tokenizer. For example, a tokenizer can process a given part of the input source and output a series of tokens representing that part of the input source (e.g., with input elements 5-1, 5-2, ..., 5-). M Correspondingly). Various word segmentation methods can be used. For example, byte-pair encoding (BPE) technology can be used to segment the text input source. See, for example, Kudo et al. SentencePiece: A simple and language independent subword tokenizer and detokenizer for Neural Text Processing (SentencePiece: A simple and language-independent word segmenter and desegmentation tool for neural text processing) (Word segmenter)Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing (SystemDemonstrations), pp. 66–71 (October 31–November 4, 2018), https: / / aclanthology.org / D18-2012.pdf. Tokenization of image-based input sources can be performed by extracting pieces from images and serializing these pieces.

[0163] Generally, any data type can be serialized and processed into the input sequence 5. It should be understood that... Figure 8 The elements depicted in the text are 5-1, 5-2, ..., 5- M It can be a word element or its embedded representation.

[0164] Prediction layer 6 can predict one or more output elements 7-1, 7-2, ..., 7- based on the input elements. N Prediction layer 6 may include one or more machine learning model architectures, such as one or more learning parameter layers, which manipulate and transform the input to obtain the input elements 5-1, 5-2, ..., 5- M The higher-order meanings and relationships between the input elements are extracted. In this way, for example, example prediction layer 6 can predict new output elements based on the context provided by input sequence 5.

[0165] Prediction layer 6 can evaluate the associations between portions of the input sequence 5 and specific output elements. These associations can inform predictions about the likelihood that a particular output follows the input context. For example, consider the text snippet “The carpenter's toolbox was small and heavy. It was full of ___. ” Example prediction layer 6 can identify that “It” refers back to “toolbox” by determining the relationships between the corresponding embeddings. Example prediction layer 6 can also associate “it” with attributes of the toolbox, such as “small” and “heavy”. Based on these associations, for example, prediction layer 6 can assign a higher probability to the word “nails” than to the word “sawdust”.

[0166] Transformer is an example architecture that can be used in prediction layer 4. For example, see Vaswani et al. Attention is all you need. arXiv:1706.03762v7 (August 2, 2023). A transformer is an example of a machine learning model architecture that uses an attention mechanism to compute associations between items within a context window. The context window can include a sequence containing an input sequence 5 and potentially one or more output elements 7-1, 7-2, ..., 7-N. A transformer block can include one or more attention layers and one or more post-attention layers (e.g., feedforward layers, such as a multilayer perceptron).

[0167] In addition to or instead of transformer-based architectures, prediction layer 6 can include other machine learning model architectures. For example, recurrent neural networks (RNNs) and long short-term memory (LSTM) models, as well as convolutional neural networks (CNNs) can be used. Generally, prediction layer 6 can utilize various artificial neural networks that can understand or generate sequences of information.

[0168] The output sequence 7 may include or otherwise represent the same or different data type as the input sequence 5. For example, the input sequence 5 may represent text data, and the output sequence 7 may represent text data. The input sequence 5 may represent image, audio, or audiovisual data, and the output sequence 7 may represent text data (e.g., describing image, audio, or audiovisual data). It should be understood that the prediction layer 6 and any other gap model components of the sequence processing model 4 can be configured to receive multiple data types from the input sequence 5 and output multiple data types in the output sequence 7.

[0169] Output sequence 7 can have various relationships with input sequence 5. Output sequence 7 can be a continuation of input sequence 5. Output sequence 7 can be a supplement to input sequence 5. Output sequence 7 can translate, transform, expand, or otherwise modify input sequence 5. Output sequence 7 can answer, evaluate, acknowledge, or otherwise respond to input sequence 5. Output sequence 7 can implement instructions provided via input sequence 5 (or describe instructions for implementing said instructions).

[0170] Output sequence 7 can be generated autoregressively. For example, for some applications, the output of one or more prediction layers 6 can be passed through one or more output layers (e.g., softmax layers) to obtain a probability distribution of an output vocabulary (e.g., a text or symbol vocabulary) conditioned on the set of input elements in a context window. In this way, output sequence 7 can be generated autoregressively, for example, by sampling possible next output elements, adding that element to the context window, regenerating the probability distribution based on the updated context window, and sampling possible next output elements, etc.

[0171] Output sequence 7 can also be generated non-autoregressively. For example, multiple output elements of output sequence 7 can be predicted together without explicit order conditions. See, for example, Saharia et al., Non-Autoregressive Machine Translation with Latent Alignments, arXiv:2004.07437v3 (November 16, 2020).

[0172] Output sequence 7 may include one or more parts or elements. In the example content generation configuration, output sequence 7 may include multiple elements corresponding to multiple parts of the generated output sequence (e.g., text sentences, values ​​of discrete waveforms, computer code, etc.). In the example classification configuration, output sequence 7 may include a single element associated with the classification output. For example, the output "vocabulary" may include the set of classes to which the input sequence will be classified. For example, a visual transformer block may pass latent state information to a multilayer perceptron, which outputs possible class values ​​associated with the input image.

[0173] Figure 9 This is a block diagram of an example technique for populating example input sequence 8. Input sequence 8 may include various functional elements that form part of the model infrastructure, such as element 8-0 obtained from task indicator 9, which signals to any model processing input sequence 8 that a specific task is in progress (e.g., helps adapt the model's performance to that specific task). Input sequence 8 may include various data elements from different data modalities. For example, input modality 10-1 may include one data modality. Data-to-sequence model 11-1 may process the data from input modality 10-1 to project the data into a format compatible with input sequence 8 (e.g., determining one or more vectors of dimensions based on the dimensions of input sequence 8) to obtain elements 8-1, 8-2, 8-3. Another input modality 10-2 may include a different data modality. Data-to-sequence model 11-2 may project the data from input modality 10-2 into a format compatible with input sequence 8 to obtain elements 8-4, 8-5, 8-6. Another input modality 10-3 may include yet another different data modality. The data to sequence model 11-3 can project the data from the input mode 10-3 into a format compatible with the input sequence 8 to obtain elements 8-7, 8-8, and 8-9.

[0174] Input sequence 8 may be the same as or different from input sequence 5. Input sequence 8 may be a multimodal input sequence containing elements representing data from different modalities using a common dimension representation. For example, the embedding space may have PThe input sequence 8 can be configured to contain dimensions. P Multiple elements in a dimension. In this way, for example, the example implementation can facilitate information extraction and inference across different data modalities by projecting data onto elements in the same embedding space to compare, combine or otherwise compute between them.

[0175] For example, elements 8-0, ..., 8-9 can indicate specific locations within a multidimensional embedding space. Some elements can map to a discrete set of locations in the embedding space. For example, elements corresponding to discrete members in a predetermined lexicon can map to discrete locations in the embedding space associated with those lexicons. Other elements can be continuously distributed across the embedding space. For example, some data types can be decomposed into continuously defined parts (e.g., image patches), which can be described using continuously distributed locations within the embedding space.

[0176] In some implementations, the expressive power of an embedding space may not be limited to meanings associated with any particular set of lexical units or other building blocks. For example, a continuous embedding space can encode a series of higher-order information. Individual segments of information (e.g., lexical units) can be mapped to specific points in the space: for example, lexical units of the word "dog" can be projected to embedding values ​​that point to specific locations in the embedding space associated with dog-related information. Similarly, image patches of a dog on grass can be projected into the embedding space. In some implementations, the projection of the dog image can be similar to, and also similar to, the projection of the word "grass," but potentially different from both. In some implementations, the projection of the image patch may not be perfectly aligned with any single projection of a single word. In some implementations, the projection of the image patch can be aligned with a combination of the projections of the words "dog" and "grass." In this way, for example, a higher-order embedding space can encode information that is independent of the data modality expressing the information.

[0177] Task indicator 9 may include a model or model component configured to identify the ongoing task and inject the input value represented by element 8-0 into input sequence 8, which signals which task is in progress. For example, the input value may be provided as a data type associated with an input modality and projected along with that modality (e.g., the input value may be a text task label embedded along with other text data in the input; the input value may be a pixel-based representation of the task embedded along with other image data in the input; etc.). The input value may be provided as a data type different from or at least independent of other inputs. For example, the input value represented by element 8-0 may be learned within a continuous embedding space.

[0178] Input modes 10⁻¹, 10⁻², and 10⁻³ can be associated with a variety of different data types (e.g., as described above with respect to input 2 and output 3).

[0179] Data-to-sequence models 11-1, 11-2, and 11-3 may be the same as or different from each other. Data-to-sequence models 11-1, 11-2, and 11-3 can be adapted to each corresponding input modality 10-1, 10-2, and 10-3. For example, a text data-to-sequence model can subdivide a portion of the input text and project the subdivisions onto elements in input sequence 8 (e.g., elements 8-1, 8-2, 8-3, etc.). An image data-to-sequence model can subdivide the input image and project the subdivisions onto elements in input sequence 8 (e.g., elements 8-4, 8-5, 8-6, etc.). An arbitrary data type data-to-sequence model can subdivide the input of that arbitrary data type and project the subdivisions onto elements in input sequence 8 (e.g., elements 8-7, 8-8, 8-9, etc.).

[0180] Data-to-sequence models 11-1, 11-2, and 11-3 can form part of machine learning sequence processing model 4. Data-to-sequence models 11-1, 11-2, and 11-3 can be trained jointly with machine learning sequence processing model 4 or independently of it. Data-to-sequence models 11-1, 11-2, and 11-3 can be trained end-to-end with machine learning sequence processing model 4.

[0181] Figure 10 This is a block diagram of an example model development platform 12, which facilitates the creation, adaptation, and refinement of example machine learning models (e.g., machine learning model 1, sequence processing model 4, etc.). The model development platform 12 can provide several different toolkits that developer systems can use to develop new or adapted machine learning models.

[0182] The model development platform 12 can provide one or more model libraries 13 containing building blocks for new models. Model library 13 can include one or more pre-trained base models 13-1, which can provide a backbone of processing power across a variety of tasks. Model library 13 can include one or more pre-trained expert models 13-2, which can focus on performance in a specific domain of expertise. Model library 13 can include various model primitives 13-3, which can provide low-level architectures or components (optionally pre-trained) that can be assembled in various arrangements as needed. Model primitives 13-3 can include a library of pre-trained adapters or LoRA modules that can adjust the baseline base model to align its output with a desired performance profile, enhance model capabilities (e.g., to adapt to different input modalities, etc.), etc.

[0183] The model development platform 12 can receive selections of various model components 14. The model development platform 12 can transfer the selected model components 14 to the workbench 15, which combines the selected model components 14 into the development model 16.

[0184] Workbench 15 can facilitate further refinement and adaptation of the development model 16 by utilizing multiple different toolkits integrated with the model development platform 12. For example, workbench 15 can facilitate the use of model alignment toolkit 17 to align the development model 16 with expected performance profiles for various tasks.

[0185] The model alignment toolkit 17 can provide a variety of tools for enabling the development model 16 to generate outputs aligned with desired behavioral characteristics. Alignment can include increasing the accuracy, precision, recall, etc., of the model output. Alignment can include enforcing output styles, patterns, or other preferred characteristics of the model output. Alignment can be general or domain-specific. For example, the pre-trained base model 13-1 can start from an initial performance level across multiple domains. Alignment of the pre-trained base model 13-1 can include improving performance in a specific information or task domain (e.g., even at the expense of performance in another information or task domain).

[0186] The model alignment toolkit 17 can integrate one or more datasets 17-1 used to align development models 16. Selected datasets 17-1 may include labeled or unlabeled training data. Datasets 17-1 can be obtained from public domain datasets. Datasets 17-1 can also be obtained from private datasets associated with one or more developer systems used to align custom machine learning models tailored for private use cases.

[0187] The pre-training pipeline 17-2 may include a machine learning model training workflow configured to update the development model 16 on a large-scale, potentially noisy dataset. For example, pre-training may utilize unsupervised learning techniques (e.g., denoising, etc.) to process a large number of training instances to update model parameters from an initial state and achieve the desired baseline performance. The pre-training pipeline 17-2 may utilize the unlabeled dataset from dataset 17-1 for pre-training. Workbench 15 may implement the pre-training pipeline 17-2 to pre-train the development model 16.

[0188] The fine-tuning pipeline 17-3 may include a machine learning model training workflow configured to refine the model parameters of the development model 16 using higher-quality data. The fine-tuning pipeline 17-3 can update the development model 16 through supervised training using a labeled dataset from dataset 17-1. The fine-tuning pipeline 17-3 can also update the development model 16 through reinforcement learning using reward signals from user feedback. Workbench 15 enables the fine-tuning pipeline 17-3 to fine-tune the development model 16.

[0189] Hint library 17-4 may include a set of inputs configured to induce behavior aligned with desired performance criteria. Hint library 17-4 may include few-shot hints (e.g., providing examples of the desired model output for appending a header to the desired runtime query), thought chain hints (e.g., providing stepwise inference within an example to facilitate comprehensive inference by the model), and so on.

[0190] Sample hints can be retrieved from the available repository of hint library 17-4. Sample hints can be facilitated by one or more developer systems using workbench 15.

[0191] In some implementations, pre-trained or fine-tuned models can achieve satisfactory performance even when there are no paradigms in the input. For example, zero-shot hints can include inputs lacking paradigms. Zero-shot hints can be within the domain of the training dataset or outside the training domain.

[0192] Hint library 17-4 may include one or more hint engineering tools. Hint engineering tools can provide a workflow for retrieving or learning optimized hint values. Hint engineering tools can facilitate the direct learning of hint values ​​(e.g., input element values) based on one or more training iterations. Workbench 15 can implement the hint engineering tools within development model 16.

[0193] Hint library 17-4 may include a pipeline for hint generation. For example, input can be generated using development model 16 itself or other machine learning models. In this way, for example, a first model can process information about the task and output input for a second model to process in order to perform the steps of the task. The second model may be the same as or different from the first model. Workbench 15 can implement the hint generation pipeline within development model 16.

[0194] Hint library 17-4 may include a pipeline for context injection. For example, if additional context is provided for performing a specific task, the performance of development model 16 on that task can be improved. Hint library 17-4 may include software components configured to identify desired context, retrieve context from external sources (e.g., databases, sensors, etc.), and add the context to input hints. Workbench 15 can implement the context injection pipeline in development model 16.

[0195] Although the various training examples described in this document regarding model development platform 12 refer to "pre-training" and "fine-tuning," it should be understood that the model alignment toolkit 17 generally supports a wide variety of training techniques adapted to training a wide range of machine learning models. Example training techniques may correspond to the example training method 600 described above.

[0196] The model development platform 12 may include a model plug-in toolkit 18. The model plug-in toolkit 18 may include a variety of tools configured to enhance the functionality of the machine learning model by integrating it with other systems, devices, and software components. For example, the machine learning model can use tools to improve performance quality where appropriate. For instance, deterministic tasks can be offloaded to dedicated tools instead of performing tasks probabilistically when the risk of error increases. For example, instead of autoregressively predicting solutions to a system of equations, the machine learning model can identify the tools invoked to obtain solutions and pass the system of equations to the appropriate tool. This tool can be a conventional equation solver that operates deterministically to solve the system of equations. The tool's output can be returned in response to the original query. In this way, tool usage can allow some example models to focus on the strengths of the machine learning model—e.g., understanding the intent in unstructured requests for a task—while enhancing model performance by offloading certain tasks to more focused tools to mechanically apply deterministic algorithms to well-defined problems.

[0197] The model plugin toolkit 18 may include a validation tool 18-1. The validation tool 18-1 may include tools capable of parsing and verifying the output of a machine learning model. The validation tool 18-1 may include engineered heuristics that establish certain thresholds applied to the model output. For example, the validation tool 18-1 may base the output of the machine learning model on a structured data source (e.g., to mitigate "illusion").

[0198] The model plugin toolkit 18 may include a toolkit 18-2 for implementing one or more tools, which may include scripts or other executable code that can be executed with the development model 16. The toolkit 18-2 may include one or more inputs configured to enable a machine learning model to implement the tools (e.g., few-shot hints that induce the model to output tool calls with correct syntax). For example, the toolkit 18-2 may include fine-tuned training data for training the model to use the tools.

[0199] The model plugin toolkit 18 may include interfaces for calling external application programming interfaces (APIs) 18-3. For example, attached to or replacing the direct implementation of tool calls or tool code using development model 16, development model 16 may be aligned with output instructions that initiate API calls to send or retrieve data via external systems.

[0200] The model plugin toolkit 18 can be integrated with the hint library 17-4 to create a catalog of available tools for use with the development model 16. For example, the model can receive a catalog of available tools in its input, and the model can generate output that selects a tool from the available tools and initiates a tool call for using that tool.

[0201] Model development platform 12 may include a suite of computational optimization tools 19 for optimizing the computational performance of development model 16. For example, tools for model compression 19-1 may allow development model 16 to be reduced in size while maintaining the desired performance level. For example, model compression 19-1 may include quantization workflows, weight pruning, and sparsification techniques. Tools for hardware acceleration 19-2 may facilitate the configuration of model storage and execution formats for optimal operation on different hardware resources. For example, hardware acceleration 19-2 may include tools for optimally sharding the model for distributed processing across multiple processing units to increase bandwidth, reduce uniform memory requirements, etc. Tools for refinement 19-3 may provide tools for training a lighter model based on knowledge encoded in development model 16. For example, development model 16 may be a large, high-performance machine learning model optimized using model development platform 12. To obtain a lightweight model for operation in resource-constrained environments, the smaller model may be a “student model” that learns from and imitates development model 16 as the “teacher model.” In this way, for example, the investment in learning and developing the parameters and configuration of model 16 can be efficiently transferred to a smaller model for more efficient inference.

[0202] Workbench 15 may implement one or more of the toolkits implemented in model development platform 12, or may not implement any toolkits. Workbench 15 may output output model 20 based on development model 16. Output model 20 may be a deployment version of development model 16. Output model 20 may be a development or training checkpoint of development model 16. Output model 20 may be a refined, compressed, or otherwise optimized version of development model 16.

[0203] Figure 11 This is a block diagram of an example training process for training a machine learning development model 16. One or more parts of the example training process can be implemented by a computing system (such as the computing system described, for example, with reference to other figures) that includes one or more computing devices. Each corresponding part of the example training process can be performed by any one (or any combination of) of the one or more computing devices. Furthermore, one or more parts of the example training process can be implemented on the hardware components of the apparatus described herein, for example, to train one or more systems or models. Figure 11 For illustrative and discussion purposes, the elements are depicted in a specific order. Those skilled in the art will understand using the disclosure provided herein that elements of any of the methods discussed herein can be adapted, rearranged, extended, omitted, combined, or modified in various ways without departing from the scope of this disclosure. Figure 11The descriptions are for illustrative purposes only and refer to elements / terms described with reference to other systems and diagrams, and are not intended to be limiting. One or more parts of the example training process may be performed additionally or alternatively by other systems.

[0204] Initially, development model 16 can be kept in its initial state as initialization model 21. Development model 16 can be initialized using weight values. The initial weight values ​​can be random or based on an initialization pattern. The initial weight values ​​can be based on previous pre-training for the same or different models.

[0205] The initialization model 21 can undergo pre-training in the pre-training phase 22. The pre-training phase 22 can be implemented using one or more pre-training pipelines 17-2 on data from dataset 17-1. For example, if the initialization model 21 has already been pre-trained (e.g., the development model 16 contains, is, or is based on a pre-trained base model or expert model), pre-training can be omitted.

[0206] The pre-trained model 23 can then be a new version of the development model 16, which can remain as the development model 16 or be a new development model. If the development model 16 has already been pre-trained, the pre-trained model 23 can be in its initial state. The pre-trained model 23 can undergo fine-tuning in the fine-tuning phase 24. The fine-tuning phase 24 can be implemented using one or more fine-tuning pipelines 17-3 on data from dataset 17-1. For example, fine-tuning can be omitted if the pre-trained model has satisfactory performance, if the model has already been fine-tuned, or if other tuning methods are preferred.

[0207] Then, the fine-tuned model 29 can be a new version of the development model 16, which can remain as the development model 16 or a new development model. If the development model 16 has already been fine-tuned, then the fine-tuned model 29 can be in its initial state. The fine-tuned model 29 can undergo refinement 26 using user feedback. For example, refinement 26 using user feedback can optionally include reinforcement learning based on human feedback from human users of the fine-tuned model 25. Since reinforcement learning can take the form of fine-tuning, it should be understood that the fine-tuning phase 24 can include a phase for refinement 26 using user feedback. Refinement 26 using user feedback can produce a refined model 27. The refined model 27 can be output to the downstream system 28 for deployment or further development.

[0208] In some implementations, computational optimization operations can be applied before, during, or after each stage. For example, initializing model 21 may undergo computational optimization 29-1 (e.g., using computational optimization toolkit 19) before pre-training stage 22. Pre-trained model 23 may undergo computational optimization 29-2 (e.g., using computational optimization toolkit 19) before fine-tuning stage 24. Fine-tuned model 25 may undergo computational optimization 29-3 (e.g., using computational optimization toolkit 19) before refinement 26 utilizing user feedback. Refined model 27 may undergo computational optimization 29-4 (e.g., using computational optimization toolkit 19) before outputting to downstream system 28. Computational optimizations 29-1, ..., 29-4 may be all the same, all different, or include at least some different optimization techniques.

[0209] Figure 12 This is a block diagram of an inference system used to operate one or more machine learning models 1 for inference (e.g., for training, for deployment, etc.). A model host 31 can receive machine learning models 1. Model host 31 can host one or more model instances 31-1, which can be one or more instances of one or more models. Model host 31 can use available computing resources 31-2 associated with model host 31 to host model instances 31-1.

[0210] Model host 31 can perform inference on behalf of one or more clients 32. Client 32 can transmit input request 33 to model host 31. Using input request 33, model host 31 can obtain input 2 to feed into machine learning model 1. Machine learning model 1 can process input 2 to generate output 3. Using output 3, model host 31 can return output payload 34 in response to input request 33 from client 32. Output payload 34 can include or be based on output 3.

[0211] Model host 31 can utilize various other resources and tools to enhance the inference task. For example, model host 31 can communicate with tool interface 35 to facilitate the use of tools by model instance 31-1. Tool interface 35 may include local or remote APIs. Tool interface 35 may include integrated scripts or other software functions. Model host 31 can use online learning interface 36 to facilitate continuous improvement of machine learning model 1. For example, online learning interface 36 can be used in reinforcement learning loops to retrieve user feedback on inference served by model host 31. Model host 31 can access runtime data source 37 for enhancing input 2 with additional contextual information. For example, runtime data source 37 may include knowledge graph 37-1 that facilitates structured information retrieval for information associated with input request 33 (e.g., search engine service). Runtime data source 37 may include public or private, external or local database 37-2 that can store information associated with input request 33 for enhancing input 2. The runtime data source 37 may include account data 37-3, which can be retrieved in association with the user account corresponding to the client 32 to customize the behavior of the model host 31 accordingly.

[0212] The model host 31 may be implemented by one or more computing devices or systems. The client 2 may be implemented by one or more computing devices or systems, which may include computing devices or systems shared with the model host 31.

[0213] For example, model host 31 can operate on a server system that provides machine learning services (e.g., via a local area network or wide area network) to client devices operating client 32. The client device can be an end-user device used by an individual. The client device can also be a server system that operates client 32 to provide various functionalities as services to downstream end-user devices.

[0214] In some implementations, model host 31 may operate on the same device or system as client 32. Model host 31 may be a machine learning service that runs on the device to provide machine learning capabilities to one or more applications operating on the client device, which may include the application implementing client 32. Model host 31 and client 32 may be part of the same application. For example, model host 31 may be a subroutine or method implemented as part of the application, and client 32 may be another subroutine or method that uses model host 31 to perform inference functionality within the application. It should be understood that model host 31 and client 32 may have various different configurations.

[0215] Model instance 31-1 may include one or more machine learning models that can be used to perform inference. Model instance 31-1 may include weights or other model components stored in persistent storage, temporary caches, or loaded into memory. Model instance 31-1 may include multiple instances of the same model (e.g., for parallel execution of more requests on the same model). Model instance 31-1 may include instances of different models. Model instance 31-1 may include intermediate states of cached active or inactive models, which are used to accelerate inference for those models. For example, an inference session with a particular model can generate a significant amount of computational results that can be reused for future inference runs (e.g., using a KV cache for a transformer-based model). These computational results can be stored in association with the inference session, allowing for more efficient execution when the session resumes.

[0216] Computing resource 31-2 may include one or more processors (central processing unit, graphics processing unit, tensor processing unit, machine learning accelerator, etc.) connected to one or more memory devices. Computing resource 31-2 may include a dynamic pool of available resources shared with other processes. Computing resource 31-2 may include a memory device large enough to fit an entire model instance into a single memory instance. Computing resource 31-2 may also share model instances across multiple memory devices (e.g., using data parallelization or tensor parallelization). Doing so can increase parallelization or execute large models using multiple memory devices that, individually, may not be able to fit the entire model into memory.

[0217] Input request 33 may include data for input 2. Model host 31 can process input request 33 to obtain input 2. Input 2 can be obtained directly from input request 33 or retrieved using input request 33. Input request 33 can be submitted to model host 31 via API.

[0218] Model host 31 can perform inference on multiple batch input requests 33 in parallel. For example, model instance 31-1 can be configured with an input structure having a batch dimension. Individual inputs 2 can be distributed across batch dimensions (e.g., rows of an array). Individual inputs 2 can include completely different contexts. Individual inputs 2 can be multiple inference steps for the same task. Individual inputs 2 can be interleaved in the input structure, such that any given inference loop can operate on different parts of the corresponding inputs 2. In this way, for example, model host 31 can perform inference on batches in parallel, such that output 3 can also contain batch dimensions and return the inference results of batch inputs 2 in parallel. In this way, for example, multiple batch input requests 33 can be processed in parallel to achieve higher throughput of output payload 34.

[0219] The output payload 34 may include or be based on the output 3 from the machine learning model 1. The model host 31 may process the output 3 to obtain the output payload 34. This may include linking multiple rounds of inference (e.g., iteratively, recursively, across the same model or different models) to obtain the final output of the task so that it can be returned in the output payload 34. The output payload 34 may be transferred to the client 32 via an API.

[0220] Online learning interface 36 can facilitate reinforcement learning for machine learning model 1. Online learning interface 36 can facilitate reinforcement learning with human feedback (RLHF). Online learning interface 36 can facilitate federated learning for machine learning model 1.

[0221] The model host 31 can access libraries of pre-trained adapters or LoRA modules that can tune the baseline model to align its output with the desired performance profile, enhance the model's capabilities (e.g., to adapt to different input modalities), etc. For example, the model host 31 can receive input requests to load a custom model, and it can retrieve one or more components to adapt the baseline model to the custom profile. The model host 31 can determine which specific functionality is required for a particular task (e.g., based on the output of the model with pre-processed input) and retrieve pre-trained components accordingly.

[0222] Model host 31 can execute machine learning model 1 to perform inference for various tasks using various types of data. For example, various different inputs 2 and outputs 3 can be used for various different tasks. In some implementations, input 2 may be or otherwise represent image data. Machine learning model 1 can process image data to generate outputs. As an example, machine learning model 1 can process image data to generate image recognition outputs (e.g., image data identification, latent embedding of image data, encoded representation of image data, hashing of image data, etc.). As another example, machine learning model 1 can process image data to generate image segmentation outputs. As another example, machine learning model 1 can process image data to generate image classification outputs. As another example, machine learning model 1 can process image data to generate image data modification outputs (e.g., image data alterations, etc.). As another example, machine learning model 1 can process image data to generate encoded image data outputs (e.g., encoded and / or compressed representations of image data, etc.). As another example, machine learning model 1 can process image data to generate enlarged image data outputs. As another example, machine learning model 1 can process image data to generate predictive outputs.

[0223] In some implementations, the task is a computer vision task. In some cases, the input 2 includes pixel data from one or more images, and the task is an image processing task. For example, an image processing task could be image classification, where the output is a set of scores, each score corresponding to a different object class and representing the probability that one or more images depict an object belonging to that object class. An image processing task could be object detection, where the image processing output identifies one or more regions in one or more images, and for each region, identifies the probability that that region depicts an object of interest. As another example, an image processing task could be image segmentation, where the image processing output defines a corresponding probability for each class in a predetermined set of categories for each pixel in one or more images. For example, the set of categories could be foreground and background. As another example, the set of categories could be object classes. As another example, an image processing task could be depth estimation, where the image processing output defines a corresponding depth value for each pixel in one or more images. As another example, an image processing task could be motion estimation, where the network input includes multiple images, and the image processing output defines the motion of the scene depicted at each pixel in one of the input images between the images in the network input.

[0224] In some implementations, input 2 may be or otherwise represent natural language data. Machine learning model 1 can process the natural language data to generate output. As an example, machine learning model 1 can process natural language data to generate language-encoded output. As another example, machine learning model 1 can process natural language data to generate latent text embedding output. As another example, machine learning model 1 can process natural language data to generate translation output. As another example, machine learning model 1 can process natural language data to generate classification output. As another example, machine learning model 1 can process natural language data to generate text segmentation output. As another example, machine learning model 1 can process natural language data to generate semantic intent output. As another example, machine learning model 1 can process natural language data to generate expanded text or natural language output (e.g., text or natural language data of higher quality than the input text or natural language, etc.). As another example, machine learning model 1 can process natural language data to generate predictive output (e.g., the next part of one or more predictions of natural language content).

[0225] In some implementations, input 2 can be or otherwise represent speech data (e.g., data describing spoken natural language, such as audio data, text data, etc.). Machine learning model 1 can process the speech data to generate output. As an example, machine learning model 1 can process speech data to generate speech recognition output. As another example, machine learning model 1 can process speech data to generate speech translation output. As another example, machine learning model 1 can process speech data to generate latent embedding output. As another example, machine learning model 1 can process speech data to generate encoded speech output (e.g., encoded and / or compressed representations of speech data, etc.). As another example, machine learning model 1 can process speech data to generate amplified speech output (e.g., speech data of higher quality than the input speech data, etc.). As another example, machine learning model 1 can process speech data to generate text representation output (e.g., text representations of the input speech data, etc.). As another example, machine learning model 1 can process speech data to generate predictive output.

[0226] In some implementations, input 2 can be or otherwise represent latent encoded data (e.g., a latent space representation of the input). Machine learning model 1 can process the latent encoded data to generate an output. As an example, machine learning model 1 can process the latent encoded data to generate an identification output. As another example, machine learning model 1 can process the latent encoded data to generate a reconstruction output. As another example, machine learning model 1 can process the latent encoded data to generate a search output. As another example, machine learning model 1 can process the latent encoded data to generate a re-clustering output. As yet another example, machine learning model 1 can process the latent encoded data to generate a prediction output.

[0227] In some implementations, input 2 may be or otherwise represent statistical data. Statistical data may be, represent, or otherwise include data calculated and / or computed from another data source. Machine learning model 1 can process the statistical data to generate output. As an example, machine learning model 1 can process the statistical data to generate an identification output. As another example, machine learning model 1 can process the statistical data to generate a prediction output. As another example, machine learning model 1 can process the statistical data to generate a classification output. As another example, machine learning model 1 can process the statistical data to generate a segmentation output. As another example, machine learning model 1 can process the statistical data to generate a visualization output. As another example, machine learning model 1 can process the statistical data to generate a diagnostic output.

[0228] In some implementations, input 2 can be or otherwise represent sensor data. Machine learning model 1 can process the sensor data to generate output. As an example, machine learning model 1 can process sensor data to generate identification output. As another example, machine learning model 1 can process sensor data to generate prediction output. As another example, machine learning model 1 can process sensor data to generate classification output. As another example, machine learning model 1 can process sensor data to generate segmentation output. As another example, machine learning model 1 can process sensor data to generate visualization output. As another example, machine learning model 1 can process sensor data to generate diagnostic output. As another example, machine learning model 1 can process sensor data to generate detection output.

[0229] In some implementations, machine learning model 1 can be configured to perform a task that includes encoding input data to achieve reliable and / or efficient transmission or storage (and / or corresponding decoding). For example, the task could be an audio compression task. The input could include audio data, and the output could include compressed audio data. In another example, the input includes visual data (e.g., one or more images or videos), the output includes compressed visual data, and the task is a visual data compression task. In another example, the task could include generating embeddings for input data (e.g., input audio or visual data). In some cases, the input includes audio data representing spoken utterances, and the task is a speech recognition task. The output could include text output mapped to spoken utterances. In some cases, the task includes encrypting or decrypting the input data. In some cases, the task includes microprocessor performance tasks such as branch prediction or memory address translation.

[0230] In some implementations, the task is a generation task, and machine learning model 1 can be configured to output content generated based on input 2. For example, input 2 can be, or otherwise represent, data of one or more modalities, which encodes the context used to generate additional content.

[0231] In some implementations, the task can be a text completion task. Machine learning model 1 can be configured to process input 2 representing text data and generate output 3, whereby the output represents additional text data following the text sequence of input 2. For example, machine learning model 1 can be configured to generate output 3 to complete a sentence, paragraph, or section of text following a portion of the text represented by input 2.

[0232] In some implementations, the task can be an instruction-following task. Machine learning model 1 can be configured to process input 2 representing instructions for performing a function and generate output 3, which advances to satisfy the objective of the instruction function (e.g., at least one step of a multi-step process for performing the function). Output 3 can represent data of the same or different modality as input 2. For example, input 2 can represent text data (e.g., natural language instructions for a task to be performed), and machine learning model 1 can process input 2 to generate output 3, which represents text data in response to the instructions (e.g., a natural language response, a programming language response, a machine language response, etc.). Input 2 can represent image data (e.g., image-based instructions for a task to be performed, optionally accompanied by text instructions), and machine learning model 1 can process input 2 to generate output 3, which represents text data in response to the instructions (e.g., a natural language response, a programming language response, a machine language response, etc.). One or more outputs 3 can be generated iteratively or recursively to sequentially process and complete steps toward completing the requested function. For example, the initial output can be executed by an external system or processed by machine learning model 1 to complete the initial steps of performing the function. Multiple steps can be performed to obtain the final output in response to the initial instruction.

[0233] In some implementations, the task can be a question-answering task. Machine learning model 1 can be configured to process input 2 representing a question to be answered and generate output 3 that advances towards the goal of returning an answer to the question (e.g., at least one step in a multi-step process for performing the function). Output 3 can represent data of the same or different modality as input 2. For example, input 2 can represent text data (e.g., natural language instructions for a task to be performed), and machine learning model 1 can process input 2 to generate output 3 representing text data in response to the question (e.g., a natural language response, a programming language response, a machine language response, etc.). Input 2 can represent image data (e.g., image-based instructions for a task to be performed, optionally accompanied by text instructions), and machine learning model 1 can process input 2 to generate output 3 representing text data in response to the question (e.g., a natural language response, a programming language response, a machine language response, etc.). One or more outputs 3 can be generated iteratively or recursively to process and complete the steps toward answering the question sequentially. For example, the initial output can be executed by an external system or processed by a machine learning model 1 to complete the initial steps to obtain an answer to the question (e.g., querying a database, performing calculations, executing scripts, etc.). Multiple steps can be performed to obtain the final output in response to the question.

[0234] In some implementations, the task can be an image generation task. Machine learning model 1 can be configured to process input 2, which represents context regarding a desired portion of the image content. Context can include text data, image data, audio data, etc. Machine learning model 1 can be configured to generate output 3, which represents image data depicting the image in context. For example, machine learning model 1 can be configured to generate pixel data of an image. The values ​​of the channels associated with pixels in the pixel data can be selected based on context (e.g., based on probabilities determined according to the context).

[0235] In some implementations, the task can be an audio generation task. Machine learning model 1 can be configured to process input 2, which represents context regarding a desired portion of the audio content. Context can include text data, image data, audio data, etc. Machine learning model 1 can be configured to generate output 3, which represents audio data relevant to the context. For example, machine learning model 1 can be configured to generate waveform data in the form of an image (e.g., a spectrogram). The values ​​of channels associated with pixels in the image can be selected based on the context. Machine learning model 1 can be configured to generate waveform data in the form of a sequence of discrete samples of a continuous waveform. The values ​​of the sequence can be selected based on the context (e.g., based on probabilities determined according to the context).

[0236] In some implementations, the task can be a data generation task. Machine learning model 1 can be configured to process input 2, which represents context regarding a desired portion of the data (e.g., data from various data domains, such as sensor data, image data, multimodal data, statistical data, etc.). For example, the desired data can be synthetic data used to train other machine learning models. The context can include any data type. Machine learning model 1 can be configured to generate output 3, which represents data aligned with the desired data. For example, machine learning model 1 can be configured to generate data values ​​to populate a dataset. The values ​​of data objects can be selected based on context (e.g., based on probabilities determined according to the context).

[0237] Figure 13This is a block diagram of an example networked computing system that can perform aspects of the exemplary implementations of this disclosure. The system may include multiple computing devices and systems communicatively coupled via network 49. Example computing device 50 is described as an example of a computing device capable of performing any aspect of this disclosure (e.g., implementing model host 31, client 32, or both). Example server computing system 60 is described as an example of a server computing system capable of performing any aspect of this disclosure (e.g., implementing model host 31, client 32, or both). Computing device 50 and server computing system 60 can interact collaboratively (e.g., via network 49) to perform any aspect of this disclosure (e.g., implementing model host 31, client 32, or both). Model development platform system 70 is an example system that can host or provide a model development platform 12 for developing machine learning models. Third-party system 80 is an example system that can interact with any of computing device 50, server computing system 60, or model development platform system 70 when performing various aspects of this disclosure (e.g., using third-party tools, accessing third-party databases or other resources, etc.).

[0238] Network 49 can be any type of communication network, such as a local area network (e.g., intranet), a wide area network (e.g., the Internet), or some combination thereof, and can include any number of wired or wireless links. Generally, communication via network 49 can be conducted using a wide variety of communication protocols (e.g., TCP / IP, HTTP, SMTP, FTP), encodings or formats (e.g., HTML, XML), or protection schemes (e.g., VPN, Secure HTTP, SSL) via any type of wired or wireless connection. Network 49 can also be implemented via a system bus. For example, Figure 13 One or more devices or systems may be located in the same place as, contained in, or otherwise integrated into one or more other devices or systems.

[0239] Computing device 50 can be any type of computing device, such as, for example, a personal computing device (e.g., a laptop or desktop computer), a mobile computing device (e.g., a smartphone or tablet), a game console or controller, a wearable computing device, an embedded computing device, a server computing device, a virtual machine operating on a host device, or any other type of computing device. Computing device 50 can be a client computing device. Computing device 50 can be an end-user computing device. Computing device 50 can be a computing device that provides services to an end user (who may use another computing device to interact with computing device 50).

[0240] Computing device 50 may include one or more processors 51 and memory 52. ​​Processor 51 may be any suitable processing device (e.g., processor core, microprocessor, ASIC, FPGA, controller, microcontroller, etc.) and may be a single processor or multiple processors operatively connected. Memory 52 may include one or more non-transitory computer-readable storage media, such as HBM, RAM, ROM, EEPROM, EPROM, flash memory devices, disks, etc., and combinations thereof. Memory 52 may store data 53 and instructions 54, which may be executed by processor 51 to cause computing device 50 to perform operations. These operations may implement any or more of the features described herein. The operations may implement the example methods and techniques described herein.

[0241] The computing device 50 may also include one or more input components for receiving user input. For example, the user input component may be a touch-sensitive component (e.g., a touch-sensitive display or touchpad) that is sensitive to the touch of a user input object (e.g., a finger or stylus). The touch-sensitive component can be used to implement a virtual keyboard. Other example user input components include a microphone, camera, LiDAR, physical keyboard or other buttons, or other components through which the user can provide input.

[0242] Computing device 50 may store or include one or more machine learning models 55. Machine learning model 55 may include one or more machine learning models 1, such as sequence processing model 4. Machine learning model 55 may include one or more model instances 31-1. Machine learning model 55 may be received from server computing system 60, model development platform system 70, third-party system 80 (e.g., application distribution platform), or developed locally on computing device 50. Machine learning model 55 may be loaded into memory 52 and used by processor 51 or otherwise implemented. Computing device 50 may implement multiple parallel instances of machine learning model 55.

[0243] Server computing system 60 may include one or more processors 61 and memory 62. Processor 61 may be any suitable processing device (e.g., processor core, microprocessor, ASIC, FPGA, controller, microcontroller, etc.) and may be a single processor or multiple processors operatively connected. Memory 62 may include one or more non-transitory computer-readable storage media, such as HBM, RAM, ROM, EEPROM, EPROM, flash memory devices, disks, etc., and combinations thereof. Memory 62 may store data 63 and instructions 64, which may be executed by processor 61 to cause server computing system 60 to perform operations. These operations may implement any or more features described herein. The operations may implement the exemplary methods and techniques described herein.

[0244] In some implementations, the server computing system 60 includes one or more server computing devices or is otherwise implemented by one or more server computing devices. In instances where the server computing system 60 includes multiple server computing devices, such server computing devices may operate according to a sequential computing architecture, a parallel computing architecture, or some combination thereof.

[0245] Server computing system 60 may store or otherwise include one or more machine learning models 65. Machine learning model 65 may be the same as or different from machine learning model 55. Machine learning model 65 may include one or more machine learning models 1, such as sequence processing model 4. Machine learning model 65 may include one or more model instances 31-1. Machine learning model 65 may be received from computing device 50, model development platform system 70, third-party system 80, or developed locally on server computing system 60. Machine learning model 65 may be loaded into memory 62 and used by processor 61 or otherwise implemented. Server computing system 60 may implement multiple parallel instances of machine learning model 65.

[0246] In the example configuration, machine learning model 65 may be included in or otherwise stored and implemented by server computing system 60 to establish a client-server relationship with computing device 50 for service model inference. For example, server computing system 60 may implement model host 31 on behalf of client 32 on computing device 50. For example, machine learning model 65 may be implemented by server computing system 60 as part of a web service (e.g., a remote machine learning model hosting service, such as an online interface for performing machine learning model operations on server computing system 60 over a network). For example, server computing system 60 may communicate with computing device 50 via a local intranet or internet connection. For example, computing device 50 may be a workstation or endpoint communicating with server computing system 60, where the implementation of machine learning model 65 is managed by server computing system 60 to remotely perform inference (e.g., for runtime or training operations), where output is returned (e.g., projected, streamed, etc.) to computing device 50. Machine learning model 65 may work collaboratively or interoperably with machine learning model 55 on computing device 50 to perform various tasks.

[0247] The model development platform system 70 may include one or more processors 71 and memory 72. Processor 71 may be any suitable processing device (e.g., processor core, microprocessor, ASIC, FPGA, controller, microcontroller, etc.) and may be a single processor or multiple processors operatively connected. Memory 72 may include one or more non-transitory computer-readable storage media, such as HBM, RAM, ROM, EEPROM, EPROM, flash memory devices, disks, etc., and combinations thereof. Memory 72 may store data 73 and instructions 74, which may be executed by processor 71 to cause the model development platform system 70 to perform operations. These operations may implement any one or more features described herein. These operations may implement the example methods and techniques described herein. Example operations include the functionality described herein with respect to model development platform 12. This functionality, and other functionalities, may be implemented by developer tools 75.

[0248] The third-party system 80 may include one or more processors 81 and memory 82. Processor 81 may be any suitable processing device (e.g., processor core, microprocessor, ASIC, FPGA, controller, microcontroller, etc.) and may be a single processor or multiple processors operatively connected. Memory 82 may include one or more non-transitory computer-readable storage media, such as HBM, RAM, ROM, EEPROM, EPROM, flash memory devices, disks, etc., and combinations thereof. Memory 82 may store data 83 and instructions 84, which may be executed by processor 81 to cause the third-party system 80 to perform operations. These operations may implement any one or more features described herein. These operations may implement the example methods and techniques described herein. Example operations include the functionality described herein regarding tools and other external resources (e.g., third-party resource 85) invoked when training machine learning models 1, 4, 16, 20, 55, 65, etc., or when performing inference using said machine learning models.

[0249] Figure 13An example arrangement of a computing system that can be used to implement the present disclosure is shown. Other computing system configurations may also be used. For example, in some implementations, one or both of computing system 50 or server computing system 60 may implement all or part of the operation of model development platform system 70. For example, computing system 50 or server computing system 60 may implement developer tool 75 (or extensions thereof) to develop, update / train, or refine machine learning models 1, 4, 16, 20, 55, 65, etc., using one or more techniques described herein with respect to model alignment toolkit 17. In this way, for example, computing system 50 or server computing system 60 may develop, update / train, or refine machine learning models based on local datasets (e.g., for model personalization / customization, as permitted by user data preference selection).

[0250] Figure 14 This is a block diagram of an example computing device 98 implemented according to an example embodiment of the present disclosure. The computing device 98 may be a user computing device or a server computing device (e.g., computing device 50, server computing system 60, etc.). The computing device 98 may implement model host 31. For example, the computing device 98 may include multiple applications (e.g., application 1 to N). Each application may contain its own machine learning library and machine learning model. For example, each application may include a machine learning model. Example applications include text messaging applications, email applications, dictation applications, virtual keyboard applications, browser applications, etc. Figure 14 As shown, each application can communicate with multiple other components of the computing device, such as, for example, one or more sensors, a context manager, a device state component, or additional components. In some implementations, each application can communicate with each device component using an API (e.g., a public API). In some implementations, the API used by each application is application-specific.

[0251] Figure 15 This is a block diagram of an example computing device 99 implemented according to an example embodiment of the present disclosure. Computing device 99 may be the same as or different from computing device 98. Computing device 99 may be a user computing device or a server computing device (e.g., computing device 50, server computing system 60, etc.). Computing device 98 may implement model host 31. For example, computing device 99 may include multiple applications (e.g., applications 1 to N). Each application may communicate with a central intelligence layer. Example applications include text messaging applications, email applications, dictation applications, virtual keyboard applications, browser applications, etc. In some implementations, each application may use an API (e.g., a common API across all applications) to communicate with the central intelligence layer (and the model stored therein).

[0252] The central intelligence layer can include multiple machine learning models. For example, such as Figure 15 As shown, a corresponding machine learning model can be provided for each application, and this corresponding machine learning model is managed by a central intelligent layer. In other implementations, two or more applications can share a single machine learning model. For example, in some implementations, the central intelligent layer can provide a single model for all applications. In some implementations, the central intelligent layer is included within the operating system of the computing device 99 or otherwise implemented by the operating system.

[0253] The central intelligence layer can communicate with the central device data layer. The central device data layer can be a centralized data storage repository for computing device 99. For example... Figure 15 As shown, the central device data layer can communicate with multiple other components of the computing device, such as, for example, one or more sensors, a context manager, a device state component, or additional components. In some implementations, the central device data layer can communicate with each device component using an API (e.g., a private API).

[0254] This paper discusses technologies related to servers, databases, software applications, and other computer-based systems, as well as the actions taken and the information sent to and from such systems. The inherent flexibility of computer-based systems allows for a wide range of possible configurations, combinations, and partitions of tasks and functions between and within components. For example, the processes discussed herein can be implemented using a single device or component, or multiple devices or components working in combination. Databases and applications can be implemented on a single system or distributed across multiple systems. Distributed components can operate sequentially or in parallel.

[0255] While the subject matter has been described in detail with respect to various specific example embodiments, each example is provided by way of explanation and not limitation. Modifications, alterations, and equivalents to such embodiments will be readily apparent to those skilled in the art upon understanding the foregoing. Therefore, this disclosure does not exclude such modifications, alterations, or additions to the subject matter that will be readily understood by those of ordinary skill in the art. For example, features shown or described as part of one embodiment may be used with another embodiment to produce yet another embodiment. Therefore, it is intended that this disclosure cover such modifications, alterations, and equivalents.

[0256] The aspects of this disclosure have been described with reference to their illustrative embodiments. Any and all features of the appended claims can be combined or rearranged in any possible manner, including combinations of claims not expressly listed together, because the illustrative claims dependencies listed herein should not be construed as limiting the scope of possible combinations of the features disclosed herein. Therefore, the scope of this disclosure is illustrative rather than limiting, and this disclosure does not exclude such modifications, alterations, or additions to the subject matter that will be readily understood by one of ordinary skill in the art. Furthermore, terms are described herein using a list of illustrative elements connected by conjunctions such as “and,” “or,” “but,” etc. It should be understood that such conjunctions are provided for illustrative purposes only. For example, a sequence of terms and other items connected by a specific conjunction such as “or” may refer to “and / or,” “at least one of,” “any combination,” etc., of the illustrative elements listed therein. Terms such as “based on” should be understood as “at least partially based on.”

[0257] The term "capable" should be understood as referring to the possibility of a feature in various implementations, rather than specifying a capability that must exist in every implementation. For example, the phrase "X is capable of Y" should be understood as indicating that in various implementations, X may be configured to perform Y, rather than indicating that X must always be capable of performing Y in every instance. It should be understood that in various implementations, X may not be able to perform Y and is still within the scope of this disclosure.

[0258] The term "may" should be understood as referring to the possibility of a feature in various implementations, rather than specifying a capability that must exist in every implementation. For example, the phrase "X may perform Y" should be understood as indicating that in various implementations, X may be configured to perform Y, rather than indicating that X must always be able to perform Y in every instance. It should be understood that in various implementations, X may not be able to perform Y and is still within the scope of this disclosure.

Claims

1. A computer system for improving cybersecurity, the computer system comprising: one or more processors; and one or more non-transitory computer-readable media collectively storing a generated sequence processing model, wherein the generated sequence processing model has been fine-tuned on one or more fine-tuning tuples generated from one or more cybersecurity datasets, and wherein at least one fine-tuning tuple of the one or more fine-tuning tuples comprises a fine-tuning input and a fine-tuning label, wherein the fine-tuning input comprises cybersecurity data and a question about the cybersecurity data, and wherein the fine-tuning label comprises an annotated answer to the question about the cybersecurity data.

2. The computer system of claim 1, wherein the cybersecurity dataset comprises a dataset from security orchestration, automation, and response (SOAR) system data, security information and event management (SIEM) system data, security blog information, analyst reports, signature-based detection files, malware scripts, vulnerability information, product documentation, security code repositories, or cybersecurity and software development framework data. a classification task, a summarization task, a generation task, or an extraction task.

3. The computer system of any preceding claim, wherein the generation sequence processing model has been fine-tuned on the one or more fine-tuning tuples to perform one or more fine-tuning tasks, wherein the one or more fine-tuning tasks comprise:

4. The computer system of any preceding claim, wherein at least a second fine-tuning tuple of the one or more fine-tuning tuples comprises a second fine-tuning input and a second fine-tuning label, wherein the second fine-tuning input comprises a natural language query, and wherein the second fine-tuning label comprises a query expressed in a domain-specific query language.

5. The computer system of any preceding claim, wherein at least one fine-tuning tuple of the one or more fine-tuning tuples has been generated manually.

6. The computer system of any preceding claim, wherein at least one fine-tuning tuple of the one or more fine-tuning tuples has been generated automatically using one or more templates.

7. The computer system of any preceding claim, wherein at least one fine-tuning tuple of the one or more fine-tuning tuples has been generated automatically using one or more machine learning models.

8. The computer system of any preceding claim, wherein the generated sequence processing model is in operative communication with one or more cybersecurity operations tools.

9. The computer system of any preceding claim, wherein the computer system is configured to provide an interface that enables a user to query the generated sequence processing model in natural language.

10. The computer system of any preceding claim, wherein the generated sequence processing model has been fine-tuned to assume a particular cybersecurity persona among a plurality of different cybersecurity personas.

11. The computer system of claim 10, wherein the particular cybersecurity persona comprises: a security operations center (SOC) analyst persona; a threat intelligence analyst persona; a malware or code analyst persona; or a security architect persona.

12. A cybersecurity platform implemented by one or more computing devices, wherein the cybersecurity platform comprises: ​ ​ a plurality of computer-implemented agents configured to interoperate to collectively receive and process cyber-security data to generate and perform cyber-security actions responsive to the cyber-security data; wherein each of the plurality of agents includes a machine-learned generative sequence processing model that has been fine-tuned to assume a particular cyber-security role persona from among a plurality of different cyber-security role personas.

13. The cyber-security platform of claim 12, wherein the plurality of agents correspond to the plurality of different cyber-security role personas, the plurality of different cyber-security role personas including at least: a security operations center (SOC) analyst role persona; a threat intelligence analyst role persona; and a malware or code analyst role persona.

14. The cyber-security platform of claim 12 or 13, wherein the plurality of agents operate according to a distributed operating architecture.

15. The cyber-security platform of claim 12 or 13, wherein the plurality of agents operate according to a centralized planning architecture.

16. The cyber-security platform of claim 15, wherein the centralized planning architecture includes a planning agent configured to control other agents of the platform.

17. The cyber-security platform of claim 16, wherein the planning agent is configured to invoke the other agents according to a tool usage framework.

18. The cyber-security platform of claim 16 or 17, wherein the planning agent is configured to perform a chain-of-thought inference, and wherein the planning agent is configured to control the other agents of the platform based on the chain-of-thought inference.

19. The cyber-security platform of any one of claims 12 to 18, wherein the plurality of generative sequence processing models respectively associated with the plurality of different agents have been forked from a pre-trained model and then fine-tuned using respective parameter efficient adapters.

20. One or more non-transitory computer-readable media collectively storing: a generative sequence processing model, wherein the generative sequence processing model has been fine-tuned on one or more fine-tuning tuples generated from one or more cyber-security datasets, and wherein at least one of the one or more fine-tuning tuples includes a fine-tuning input and a fine-tuning label, wherein the fine-tuning input includes cyber-security data and a question about the cyber-security data, and wherein the fine-tuning label includes an annotated answer to the question, and instructions for running the generative sequence processing model to process a model input to generate a model output. ​