Automated generation of cybersecurity risks and threats-related datasets for machine learning

WO2026202853A1PCT designated stage Publication Date: 2026-10-01NITTOOR SOFTWARE RESEARCH CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/IB2026/053061
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-27
Filing Date
2026-03-27
Publication Date
2026-10-01

Smart Images

  • Figure IB2026053061_01102026_PF_FP_ABST
    Figure IB2026053061_01102026_PF_FP_ABST
Patent Text Reader

Abstract

The system and method envisaged by the present disclosure collect information related to cybersecurity-related vulnerabilities corresponding to a data management environment, and generate datasets for training predetermined machine learning models to identify the possibility of occurrence of corresponding cybersecurity attacks on the data management environment. An artificial intelligence (AI) model stores information indicative of cybersecurity-related vulnerabilities embodied in the data management environment. The AI model identifies cybersecurity-related factors that influence each of the cybersecurity-related vulnerabilities, applies a pre-defined set of statistics to the cybersecurity-related factors, and determines the severity of the cybersecurity-related factors based on the set of statistics. Subsequently, the AI model generates training vectors whose elements include an indication of whether the data management environment is vulnerable to cybersecurity attacks, considering the cybersecurity-related factors and their severity. A training dataset comprising the training vectors is also created to train the machine learning models.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] AUTOMATED GENERATION OF CYBERSECURITY RISKS AND THREATS- RELATED DATASETS FOR MACHINE LEARNING TECHNICAL FIELD

[0002] The present disclosure relates to machine learning (ML) and cybersecurity-related threat assessment. The present disclosure more particularly relates to an automated system that utilizes artificial intelligence-based functionalities and models to create training datasets suitable for training machine learning models that undertake cybersecurity-related risk assessments.

[0003] BACKGROUND

[0004] Cyber insurance is an essential component of risk management for businesses that handle voluminous amounts of confidential and critical data, and thus face increasing threats from data breaches and various types of cyberattacks. Cyberattacks pose a threat to the security, especially the data security, of enterprises. Cyberattacks are typically aimed at entire enterprise networks that store large volumes of sensitive, confidential enterprise data. Cyberattacks can sabotage the entire enterprise network or a few computing devices that are connected to the enterprise network. Cyberattacks typically create cyber threats to the confidential and sensitive enterprise data stored on or accessible through the enterprise network, by way of compromising the integrity of such data and by envisaging misuse of such confidential and sensitive data for malicious and fraudulent purposes. Cyberattacks are typically manifested as malware attacks or ransomware attacks triggered by the injection of malicious software into a computing device connected to the enterprise network. Cyberattacks are manifested either through an external endpoint or a computing entity internal to the enterprise network. Cybersecurity-related vulnerabilities present in the enterprise network or in one or more computing devices connected to the enterprise network facilitate the cyberattacks, enabling malicious and unauthorized users to clandestinely access the enterprise network and misuse the sensitive and confidential data stored therein for nefarious activities such as fraudulent fund transfers, theft of user credentials and confidential useridentification information, rendering critical data inaccessible to genuine users, sabotaging the integrity of the confidential and sensitive data, and the like.

[0005] The risk assessment models in cyber insurance typically rely on training datasets for identifying various cybersecurity-related threats and risks. Training datasets are collections of data used to train machine learning models. Training datasets contain amapping between input data and the corresponding appropriate output data, using which the machine learning models are trained to a) identify the correlation between the input data and the output data and b) appropriately infer the output data based on the input data. The machine learning models are programmed to deduce patterns from the training dataset and adjust their parameters, based on the training dataset, to make accurate predictions or classifications on new, unprocessed data.

[0006] Conventional machine learning models used in cyber insurance often require large, labeled training datasets, whose conceptualization and creation are often an arduous process. Moreover, conventional machine learning models are often static by configuration, thus requiring periodic updates that may not reflect the rapidly evolving landscape of cyber threats. This limitation, on the part of the machine learning models, hampers the ability to generate timely and precise cyber risk assessments.

[0007] Hence, there is a need to automate, using the principles of artificial intelligence, the creation of training datasets usable for training machine learning models to identify cybersecurity-related threats and risks, thereby rendering the process of generating training datasets comparatively efficient, effective, and dynamic. Additionally, there is a need for self-updating unsupervised machine learning models and self-updating deep learning models that can continuously refine cyber risk predictions, automatically based on an evolving training dataset that contains information regarding various types of cybersecurity risks and threats, without requiring extensive human intervention.

[0008] OBJECTS

[0009] An object of the present disclosure is to envisage a computer-implemented system and method for creating training datasets containing cybersecurity risks and threats-related data.

[0010] Yet another object of the present disclosure is to envisage a computer-implemented system and method that automatically enlarges and shrinks the training datasets based on the probability of the training datasets contained therein.

[0011] Still a further object of the present disclosure is to envisage a computer-implemented system and method that automatically creates training datasets indicative of various types of cybersecurity-related vulnerabilities and attacks.

[0012] Another object of the present disclosure is to envisage a computer-implemented system and method that accurately represents real-world cybersecurity risks and threat scenariosin terms of training data, and thus facilitates effective training of machine learning models in terms of identifying such real-world cybersecurity risks and threat scenarios.

[0013] Yet another object of the present disclosure is to envisage a computer-implemented system and method that trains the machine learning models to accurately predict the onset of cybersecurity-related vulnerabilities and attacks.

[0014] Yet another object of the present disclosure is to envisage a computer-implemented system and method that makes use of up-to-date cybersecurity risks and threat-related data to train the machine learning models to identify real-world cybersecurity risk and threat scenarios.

[0015] Still a further object of the present disclosure is to envisage a computer-implemented system and method that is well aware of a variety of cybersecurity-related vulnerabilities plaguing a data management environment, and is enabled to accurately identify a specific set of cybersecurity threats and risks that such a data management environment would attract, given the presence of the said cybersecurity-related vulnerabilities therein.

[0016] Another object of the present disclosure is to envisage accurately trained machine learning models that connect specific cybersecurity-related vulnerabilities to specific cybersecurity-related attacks.

[0017] One more object of the present disclosure is to envisage machine learning models that can accurately predict the probability of occurrence of specific cybersecurity attacks, based on training sets that are indicative of specific cybersecurity-related vulnerabilities. Yet another object of the present disclosure is to envisage a computer-implemented system and method that accurately correlates the cybersecurity-related vulnerabilities to specific types of cybersecurity attacks, while generating the training data for training the machine learning models to identify the onset of specific cybersecurity-related attacks.

[0018] SUMMARY

[0019] The various embodiments of the present disclosure envisage a computer-implemented system configured to automate the creation of training datasets usable for training machine learning models to accurately identify cybersecurity-related risks and threats corresponding to a data management environment, for instance, a business institution that handles large volumes of customer data, a banking institution that handles large volumes of financial data, and an information technology service provider that handles large volumes of client-specific data. The present disclosure envisages the use of agenerative artificial intelligence (Al) platform or a generative Al model to create the training datasets. The system and method envisaged by the present disclosure enhance the efficiency of cybersecurity-related risk and threat evaluation by leveraging generative Al to generate high-quality training datasets, thus leading to improved predictions of the onset of ransomware threats, business-to-business (B2B) risks, and cyberattacks, inter alia, on the data management environment.

[0020] In accordance with the present disclosure, automation of the process of creation of training datasets provides for an improvement in terms of the accuracy with which a cybersecurity score - that signifies at least the cybersecurity risks and threats faced by the data management environment and the current set of cybersecurity measures deployed within the data management environment - is assigned to the data management environment, thereby facilitating informed and considered decision making in terms of cyber risks and threats assessment. The present disclosure also envisages a self-updating, unsupervised machine learning model that is capable of operating with a reduced training dataset, ensuring efficiency and adaptability in accommodating evolving cyber threat landscapes.

[0021] In accordance with the present disclosure, the generative Al model is pre-programmed to contain information directed to cybersecurity vulnerabilities corresponding to the data management environment. In accordance with the present disclosure, the system and method dynamically generate and update the training datasets using the generative Al model. In accordance with the present disclosure, the generative Al model is also used to automate the creation of training datasets for machine learning models that forecast the probability of the data management environment being exposed to cybersecurity attacks, based on the cybersecurity-related vulnerabilities embodied in the said data management environment.

[0022] The system and method envisaged by the present disclosure are also useful in addressing the challenges posed by a rapidly changing and evolving cyber threat environment, and in enhancing the performance of unsupervised machine learning models and reinforcement learning models in terms of forecasting the likelihood of cyberattacks, based on an analysis of the underlying cybersecurity-related vulnerabilities. By obtaining cybersecurity vulnerabilities and the corresponding cybersecurity attacks-related data, at least in part from publicly available information resources, the generative Al model, envisaged by the present disclosure, automatically creates training datasets that are up-to-date and embody the most recent and relevantcybersecurity threats and risk-related information. The generative Al model, by providing the hitherto described training datasets, also enables the machine learning models to retain their relevancy in the face of continually evolving cyber threats and risks landscape, and perform effectively in a multitude of cybersecurity attack-based scenarios. In addition, the system and method envisaged by the present disclosure facilitate regular and automated updating of the training dataset, thereby ensuring the corresponding machine learning model stays aligned with any changes to the underlying training data that is based on the cybersecurity-related vulnerabilities and the corresponding cybersecurity attacks.

[0023] In accordance with the present disclosure, the generative Al model is employed to enable dynamic updating of unsupervised machine learning model parameters, thereby allowing unsupervised machine learning models to rapidly adapt to evolving cyber threats and risk-related data distributions. The rapid adaptation of the unsupervised machine learning models ensures that the model accuracy is retained even in those cyber threat environments where underlying cyber threat and risk patterns frequently shift. By leveraging the generative Al model, the system and method envisaged by the present disclosure efficiently collect and synthesize the latest statistical data corresponding to the cybersecurity-related vulnerabilities, thereby recalibrating and dynamically adjusting the training datasets for training the underlying machine learning model. This generative Al-based approach, envisaged by the present disclosure, enhances the robustness and accuracy of the machine learning model by rendering the machine learning model more capable of operating effectively and efficiently in a dynamic and unpredictable cybersecurity environment.

[0024] In accordance with the present disclosure, a cyber profile is created for the data management environment based on an assessment of a multitude of cybersecurity-related vulnerabilities that are embodied in the said data management environment. Such cybersecurity-related vulnerabilities are considered risks to the digital and cyber infrastructure operated by the data management environment. Preferably, the cyber profile illustrates a range of critical cybersecurity-related vulnerabilities, including but not restricted to Wi-Fi vulnerabilities, phishing email vulnerabilities, risk of business email compromise, version, and browser-related vulnerabilities. Preferably, the other key cybersecurity-related vulnerabilities include the operating system-related vulnerabilities that are determined based on whether updates and security patches have been regularly installed on all the computing devices operating under the data managementenvironment. Further, multi-factor authentication vulnerabilities are analyzed to evaluate the effectiveness of data access control mechanisms. The cyber profile also identifies risks arising from third-party applications installed on the computing systems operating under the data management environment. Further, social engineering attack vulnerabilities are evaluated to determine the susceptibility of employees to manipulation tactics, and the ensuing analytical information is codified into the training datasets by the generative Al model. Further, user authentication vulnerabilities are analyzed to determine the robustness of user identity and role verification mechanisms.

[0025] In accordance with the present disclosure, the machine learning models trained using the training datasets conceptualized by the generative Al model are programmed to predict the onset of ransomware attacks, inter alia, on the data management environment. Preferably, the ransomware risks are forecasted by correlating a multitude of ransomware families and their underlying breach and attack patterns to the cyber profile corresponding to the data management environment.

[0026] The machine learning models, envisaged by the present disclosure, continually refine their forecasts by learning from real-time cybersecurity attack-related statistics that are infused into the underlying training datasets by the generative Al model. The automated adjustments to the training datasets ensure accurate and up-to-date cybersecurity threat and risk assessments.

[0027] According to the present disclosure, the machine learning algorithms forecast the cybersecurity-related threats and risks corresponding to the data management environment, based on the cyber profile of the data management environment. Preferably, the major cybersecurity-related vulnerabilities embodied in the data management environment are analyzed to determine the extent and seventy of cybersecurity attacks the data management environment could encounter. In accordance with the present disclosure, the machine learning models continually analyze a multitude of cybersecurity-related vulnerabilities and identify whether such cybersecurity-related vulnerabilities would facilitate one or more cybersecurity-related attacks on the data management environment.

[0028] BRIEF DESCRIPTION OF THE DRAWINGS FIG. 1 is a block diagram of a computer-implemented system that generates training datasets based on cybersecurity-related vulnerabilities to train a machine learning model to detect cybersecurity attacks; andFIG.2 is a flow diagram illustrating the steps involved in a method for generating training datasets based on cybersecurity-related vulnerabilities to train a machine learning model to detect cybersecurity attacks.

[0029] DETAILED DESCRIPTION

[0030] The present disclosure envisages a computer-implemented system and method that facilitates detection of cybersecurity-related vulnerabilities in an enterprise environment and the ensuing cybersecurity attacks against the enterprise. The computer-implemented system and method also envisage training a machine learning model to detect such cybersecurity-related vulnerabilities and the ensuing cybersecurity attacks. The present disclosure envisages a computer-implemented system that utilizes the principles of generative Artificial Intelligence (Al) to automate the creation of datasets used for training predetermined machine learning models, viz., unsupervised machine learning models and deep learning models. In accordance with the present disclosure, the predetermined machine learning models are trained to determine whether a data handling environment, for example, a business entity that handles large volumes of confidential client-related data, a banking institution that handles large volumes of financial data, an insurance provider company that handles large volumes of customer insurance-related data as well as personal data, and a hospital that handles large volumes of patient data, is vulnerable to any of the cybersecurity attacks, including business mailbox compromise, theft or wrongful obtainment of confidential, personal information, unauthorized and illegal fund diversions, email spoofing attacks, ransomware attacks, malware attacks, misuse of authorization credentials, hijacking of valid computer sessions, wrongful use of user privileges, wrongful use of confidential internal object identifiers (for example, database keys and file paths in publicly accessible Universal Resource Locators), attacks on RDP (Remote Desktop Protocol) and SSH (Secure Shell) ports, injection of malicious program codes, buffer overflows, broken object level authorizations, Wi-Fi eavesdropping and packet sniffing, denial of service attack, data exfiltration and cloud security attacks, man-in-the-middle attack, loT vulnerability exploitation, exploitation of system vulnerabilities, supply chain attacks, and data breaches, inter alia.

[0031] The system and method envisaged by the present disclosure automate the creation of training datasets that, in turn, include training vectors that correlate the cybersecurity vulnerabilities to the possibility of the onset of cybersecurity attacks. The trainingdatasets generated by the system and method (of the present disclosure) are used to train predetermined machine learning models to identify the onset of the above-mentioned cybersecurity attacks.

[0032] The system and method envisaged by the present disclosure enhance the efficiency and effectiveness with which cybersecurity risks and threats are analyzed and evaluated. The system and method envisage the use of a generative Artificial Intelligence (Al) model that performs a real-time search for cybersecurity-related vulnerabilities and the cybersecurity attacks that originated because of such cybersecurity-related vulnerabilities. The generative Al model analyzes time-varying information, including real-time cyberattack events and cybersecurity-related vulnerabilities that facilitated such cyberattacks, most recent cybersecurity attack events, and the latest cybersecurity-related vulnerabilities, inter alia, and creates an output that includes an up-to-date aggregation of cybersecurity-related vulnerabilities and the resultant cybersecurity attacks.

[0033] The generative Al model, in accordance with the present disclosure, typically performs a search for cybersecurity-related vulnerabilities and the ensuing cybersecurity attacks on the Internet using web-indexing. Alternatively, or in addition, the generative Al model performs a search for the cybersecurity-related vulnerabilities and the ensuing cybersecurity attacks on predetermined exhaustive data sources. Further, the generative Al model may parse the contents of weblinks that were unearthed during a search, and use the parsed contents as a part of the input to identify the cybersecurity-related vulnerabilities and the ensuing cybersecurity attacks. In addition, the generative Al model postulates cybersecurity attacks based on the underlying cybersecurity-related vulnerabilities and one or more cybersecurity-related incidents and events. The generative Al model derives cybersecurity attacks from documented real-world behavior of various computing and data handling environments, and from the cybersecurity architecture and cybersecurity policies implemented in real-world computer networks. The computer-implemented system and method envisage the use of the generative Al model to create high-quality training datasets, which, when used to train a predetermined machine learning model (for example, an unsupervised learning model), enhance the efficiency of the machine learning model in terms of correlating the cybersecurity-related vulnerabilities and the cybersecurity-related attacks, correlating previously-occurred cyber incidents and cyber risks to the onset of cybersecurity-related attacks, and forecasting the onset of cybersecurity-related attacks based on the analysis of thecorresponding cybersecurity-related vulnerabilities.

[0034] The present disclosure also envisages a self-updating machine learning model (for example, an unsupervised machine learning model) that operates with both reduced training datasets and enhanced training datasets, created by the generative Al model. In this manner, the machine learning model exhibits improved efficiency and adaptability in terms of analyzing evolving cyberthreat landscapes.

[0035] In accordance with the present disclosure, the generative Al model reduces the size of the training dataset while preserving the essential cybersecurity vulnerability-related data. Essentially, the reduced training dataset is created based on dimensionality reduction, which, in turn, involves considering a lesser number of cybersecurity-related vulnerability features while creating the training dataset. Preferably, the reduced training datasets created by the generative Al model do not contain redundant and noisy cybersecurity-related vulnerability features, and, instead, emphasize on essential cybersecurity risk, threat, and attack patterns. Such reduced training datasets are also characterized by improved generalization of new cybersecurity vulnerability-related data. Further, reduced training datasets embodying reduced dimensionality provide for improved visualization of clusters and patterns of cybersecurity vulnerability-related data. In accordance with the present disclosure, the generative Al model, when reducing the training dataset, ensures that the ensuing smaller dataset embodies the same probability distribution as the originally created training dataset. Further, such reduced training datasets also reduce the learning time taken by the machine learning model, thereby facilitating faster and quicker learning iterations and comparatively easier deployment.

[0036] In accordance with the present disclosure, the generative Al model also enhances the training datasets, i.e., increases the size of the training datasets. Preferably, the generative Al model generates new data points that follow the underlying distribution of the original training data, thereby ensuring that the expanded dataset remains realistic and informative, and is devoid of biased data samples and noise. In accordance with the present disclosure, enhanced datasets contain diversified cybersecurity vulnerability-related data, which, in turn, enables the machine learning model to identify subtle patterns, leading to improved accuracy in terms of predicting the ensuing cybersecurity attacks. Furthermore, enhanced datasets contain variation in terms of the data elements embodied therein, and prevent the machine learning model from relying solely on a smaller, limited number of samples. Furthermore, enhanced datasets incorporate abroader range of cybersecurity vulnerability-related data that covers all the possible cybersecurity-related vulnerabilities, including the cybersecurity-related vulnerabilities that have been newly identified and / or occur sporadically. Preferably, enhanced datasets embodying a broader range of cybersecurity vulnerability-related data perform reliably and efficiently in terms of forecasting the occurrence of specific cybersecurity-related attacks, based on the underlying cybersecurity-related vulnerabilities. Furthermore, such enhanced datasets provide the corresponding machine learning model with increased exposure to a wide variety of cybersecurity-related vulnerabilities and consequently a wide variety of cybersecurity attack-related scenarios, and render the machine learning model resilient to variations and anomalies in the cybersecurity vulnerabilities-related data.

[0037] Furthermore, the enhanced training datasets created by the generative Al model also include distinct data points corresponding to the cybersecurity vulnerabilities-related data, thereby enhancing the efficacy of the machine learning model (in terms of the training) and reducing the risk of the machine learning model failing to anticipate / forecast the onset of scarcer cybersecurity-related attacks. Furthermore, in cases wherein the information on a specific type of cybersecurity-related vulnerability is sparse (due to the fact that the occurrence of the corresponding cybersecurity-related vulnerability is rare), the generative Al model, by expanding the datasets, increases the data points (training vectors) belonging to such a rare cybersecurity-related vulnerability and thus prevents the machine learning model from being biased towards other types of cybersecurity-related vulnerability which have been repeatedly recorded. In accordance with the present disclosure, the generative Al model makes use of the relative frequency of the data points (i.e., training vectors) as a benchmark to enhance the training datasets. In this manner, the generative Al model creates high-quality synthetic data points (training vectors) from the originally available data points (training vectors). Further, the enhanced datasets created by the generative Al model also expedite the process of training the machine learning model, all the while enabling the machine learning model to learn more comprehensive and robust cybersecurity-related vulnerabilities and cybersecurity-related attack patterns in comparatively fewer training iterations.

[0038] In accordance with the present disclosure, the computer-implemented system and method dynamically generate and update the training datasets, using the generative Al model. The generative Al model automates the creation of training datasets usable for training the machine learning model. The automated creation of the training datasetsaddresses data gathering and data analysis-related challenges associated with the continually evolving cybersecurity landscape, and enhances the performance of the machine learning model in terms of efficiently correlating the cybersecurity vulnerability-related data to the onset of cybersecurity attacks. Since the generative Al model procures the cybersecurity vulnerability-related data, necessary for creating the training datasets, from a plurality of data sources accessible via the Internet, the training datasets thus created incorporate up-to-date information regarding the cybersecurity-related vulnerabilities. In this manner, the generative Al model ensures that the machine learning model remains relevant, in terms of its efficiency and effectiveness in correlating the cybersecurity vulnerability-related data to the cybersecurity attacks and forecasting the occurrence of the cybersecurity attacks, even in a rapidly evolving cybersecurity landscape where new cybersecurity risks and threats are being regularly unearthed. Further, by continually updating the training datasets, the generative Al model ensures that the machine learning model is always well-aligned with any changes to the cybersecurity landscape, and, in particular, any changes to the cybersecurity vulnerability-related data.

[0039] In accordance with the present disclosure, the computer-implemented system and method leverage the capabilities of the generative Al model to update the parameters of the machine learning model, including the network data flow-related characteristics, user behavior-related features, host behavior-related features, and cybersecurity attack-related features, inter alia. In accordance with the present disclosure, by updating and adjusting the machine learning model parameters, the generative Al model ensures that the machine learning model rapidly adapts to the changes in the cybersecurity vulnerability-related data distributions, thereby maintaining the accuracy of the machine learning model (in terms of correlating the cybersecurity vulnerability-related data and the ensuing cybersecurity attacks), regardless of the changes in the cybersecurity vulnerability-related data patterns.

[0040] In accordance with the present disclosure, the generative Al model overcomes the disadvantages associated with traditional static machine learning models, which are typically agnostic to any changes / modifications to the underlying cybersecurity vulnerability-related data distributions and cybersecurity vulnerability-related data patterns. The generative Al model creates training datasets based on the most relevant and recent statistical data, such that the machine learning model being trained using such training datasets can be effortlessly recalibrated and adjusted as per the changesto the underlying cybersecurity vulnerability-related data distributions and cybersecurity vulnerability-related data patterns. The use of the training datasets created by the generative Al model also enhances the robustness and accuracy of the machine learning model, in terms of correlating the cybersecurity vulnerability-related data to the onset of cybersecurity attacks, and renders the machine learning model immune to dynamic and unpredictable changes to the cybersecurity vulnerability-related data. In accordance with the present disclosure, the computer-implemented system and method create a cyber risk profile corresponding to the data management environment based on the cybersecurity-related vulnerabilities embodied in the computing systems, computing networks, and data servers, inter alia, that form a part of the computing infrastructure of the data management environment. Furthermore, the cyber profile also highlights the key cybersecurity-related drawbacks or shortcomings that foster the cybersecurity-related vulnerabilities within the computing infrastructure of the data management environment. The key cybersecurity-related aspects covered by the cyber profile typically include business mailbox compromise, theft or wrongful obtainment of confidential, personal information, unauthorized and illegal fund diversions, email spoofing attacks, ransomware attacks, malware attacks, misuse of authorization credentials, hijacking of valid computer sessions, wrongful use of user privileges, wrongful use of confidential internal object identifiers (for example, database keys and file paths in publicly accessible Universal Resource Locators), attacks on RDP (Remote Desktop Protocol) and SSH (Secure Shell) ports, injection of malicious program codes, buffer overflows, broken object level authorizations, Wi-Fi eavesdropping and packet sniffing, denial of service attack, data exfiltration and cloud security attacks, man-in-the-middle attack, loT vulnerability exploitation, exploitation of system vulnerabilities, supply chain attacks, and data breaches, inter alia.

[0041] In accordance with the present disclosure, the machine learning model forecasts the onset of cybersecurity attacks, also based on the cyber risk profile corresponding to the data management environment, by analyzing the cybersecurity-related vulnerabilities specified in the cyber risk profile. Additionally, the computer-implemented system and method also generate a cyber compliance score that rates the data management environment in terms of at least the cybersecurity-related vulnerabilities embodied in the computing systems, computing networks, and data servers, inter alia, of the data management environment, the cybersecurity attacks stemming from the cybersecurity-related vulnerabilities, and the countermeasures undertaken at the data managementenvironment to thwart or at least mitigate the cybersecurity attacks.

[0042] Referring to FIG.1 , there is shown a computer-implemented system 115 that includes a processor 101 , a memory 103, a generative Al model 105, and a machine learning model 107. Preferably, the processor 101 is one of a Central Processing Unit (CPU), an x86-based processor, an x64-based processor, a Reduced Instruction Set Computing (RISC) processor, an Application-Specific Integrated Circuit (ASIC) processor, and a Complex Instruction Set Computing (CISC) processor. In accordance with the present disclosure, the memory 103 is a non-transitory computer-readable storage medium that stores a set of computer-executable instructions executable by the processor 101 for implementing the generative Al model 105 and the machine learning model 107. Preferably, the processor 101 is communicably coupled to the generative Al model 105 (step 200 in FIG.2), and triggers the generative Al model 105 to execute the functionalities described in detail below.

[0043] In accordance with the present disclosure, the computer-implemented system 115 is communicably coupled to a data management environment 109. In accordance with the present disclosure, the term “data management environment” refers to any entity that stores, processes, analyzes, and manages large volumes of critical and confidential information. In accordance with the present disclosure, the term “data management environment” can also refer to an enterprise network that stores, manages, and transmits large volumes of critical and confidential information. Typically, the data management environment 109 is prone to cybersecurity attacks given that it stores, processes, analyzes, and manages large volumes of critical and confidential information, which is otherwise inaccessible to unauthorized users, such as data hackers and fraudsters who intend to misuse such critical and confidential information for fraudulent purposes, including extortion, financial cyber fraud, and the like.

[0044] In accordance with the present disclosure, the system 115 further includes a generative Al model 105 that connects to the internet to collect historical, documented information pertaining to the various types of cybersecurity-related vulnerabilities located within data management environments 109 and / or enterprise networks. Preferably, the generative Al model 105 obtains documented historical information pertaining to the presence of a plurality of predetermined cybersecurity-related vulnerabilities across the plurality of data management environments and / or enterprise networks, from publicly accessible resources (for example, web resources accessible from the Internet). Likewise, the generative Al model 105 also collects, from publicly accessible resources, thedocumented historical statistical information corresponding to various types of cybersecurity attacks that were orchestrated on those data management environments and / or enterprise networks either directly or indirectly due to the presence of the aboveidentified cybersecurity-related vulnerabilities.

[0045] Subsequently, the generative Al model 105 determines the likelihood of occurrence of specific types of cyberattacks on any particular data management environment / enterprise network, given the presence of any of the above-identified cybersecurity-related vulnerabilities within the said particular data management environment / enterprise network.

[0046] In accordance with the present disclosure, the generative Al model 105 collects information pertaining to the following cybersecurity-related vulnerabilities: manipulative and fraudulent email communications or web notifications, improperly configured computer systems, email misconfigurations, presence of outdated versions of software programs, operating systems, and anti-virus systems, improper password management, failure to acquire data backups at regular time intervals, continued use of legacy systems, lack of safeguards against spam emails (for example, lack of spam blockers), weak access control mechanisms, lack of browser updates and browser protection, misconfigured operating system services, downloading data from unknown third-party resources, use of weak passwords, lack of strong, multi-factor authentication, failure to restrict data access on the basis of user roles, leaving Remote Desktop Protocol (RDP) and secure shell (SSH) ports directly exposed to the Internet, failure to validate, filter, or sanitize user inputs that are processed by computing systems, use of outdated memory drivers, use of corrupted memory control software, writing more data than the memory storage (buffer) can hold, Application Program Interfaces (APIs) lacking strong authentication, inadequate or weak Wi-Fi encryption, inadequate network architecture, misconfigured cloud services, inadequate data encryption, connecting insecure devices to the enterprise network, transmitting non-encrypted sensitive, confidential data over insecure networks, transmitting sensitive and confidential data without appropriate encryption, failure to apply security patches on enterprise computing resources that are accessible through the Internet, presence of less-secure parties in an enterprise network (for example, less-secure third party software service provider resources, less secure third party vendor websites, and less-secure managed service providers), and inadequate control over employee activities performed on the enterprise network.

[0047] In accordance with the present disclosure, the generative Al model 105 performs a rootcause analysis by using the information pertaining to the above-identified cybersecurity-related vulnerabilities, and correlates specific cybersecurity-related vulnerabilities to the onset of corresponding specific cybersecurity attacks. For instance, the generative Al model 105 correlates the presence of manipulative and fraudulent email communications or web notifications in a data management environment 109 to a cyberattack involving business email compromise (BEC). Likewise, the generative Al model 105 correlates the presence of improperly configured computer systems within the data management environment 109 to a cyberattack that involves fraudulent fund transfers. Further, the generative Al model 105 correlates the presence of misconfigured email servers or misconfigured enterprise mail clients to a cyberattack that involves email spoofing and domain impersonation.

[0048] Likewise, the generative Al model 105 correlates the presence of outdated versions of software programs, operating systems, and anti-virus systems, improper password management, failure to acquire data backups at regular time intervals, continued use of legacy systems, lack of safeguards against spam emails (for example, lack of spam blockers), weak access control mechanisms, lack of browser updates and browser protection, misconfigured operating system services, downloading data from unknown third-party resources, and the use of weak passwords, to a ransomware-based cyberattack. In accordance with the present disclosure, the generative Al model 105 correlates each of the cybersecurity-related vulnerabilities that promote a ransomware attack to corresponding specific ransomware families.

[0049] Further, the generative Al model 105 correlates the use of weak passwords and lack of strong, multi-factor authentication (in a data management environment 109) to a specific type of cyberattack that involves session hijacking and credential stuffing. Likewise, the generative Al model 105 correlates the failure to restrict data access on the basis of user roles to a specific cyberattack that involves privilege escalation and insecure direct object reference. Likewise, the generative Al model 105 correlates the direct exposure of the Remote Desktop Protocol (RDP) and secure shell (SSH) ports directly to the Internet to a specific cyberattack that involves initial access brokerage.

[0050] Likewise, the generative Al model 105 correlates the failure to validate, filter, or sanitize the user inputs processed by computing systems to a specific cyberattack that involves SQL injection and cross-site scripting. Likewise, the generative Al model 105 correlates the use of outdated memory drivers, the use of corrupted memory control software, and the writing of more data than the memory storage (buffer) could store, to a specificcyberattack that involves buffer overflow. Likewise, the generative Al model 105 correlates the presence of Application Programming Interfaces (APIs) that lack strong authentication to a specific cyberattack that involves broken object-level authorization (BOLA). Furthermore, the generative Al model 105 correlates inadequate or weak Wi-Fi encryption and inadequate network architecture to a specific cyberattack that involves Wi-Fi eavesdropping and packet sniffing.

[0051] Further, the generative Al model 105 correlates the presence of inadequate network architecture to a distributed denial of service cyberattack. Furthermore, the generative Al model 105 correlates the presence of misconfigured cloud services and inadequate data encryption to cloud security attacks and data exfiltration. Furthermore, the generative Al model 105 correlates the aspects of connecting insecure devices to the enterprise network, transmitting non-encrypted sensitive, confidential data over insecure networks, and transmitting sensitive and confidential data without appropriate encryption to man-in-the-middle attacks. Furthermore, the generative Al model 105 correlates the failure to apply security patches on enterprise computing resources that are accessible through the Internet to a specific type of cyberattack that involves vulnerability exploitation. Furthermore, the generative Al model 105 correlates the presence of less-secure parties in an enterprise network or in a data management environment 109 (for example, less-secure third-party software service provider resources, less-secure third-party vendor websites, and less-secure managed service providers) to a supply chain attack. Further, the generative Al model 105 correlates inadequate control over employee activities performed on the enterprise network or in the data management environment to a data breach.

[0052] In accordance with the present disclosure, the generative Al model 105 generates a descriptive analysis corresponding to each of the above-identified cybersecurity-related vulnerabilities embodied in the data management environment 109 (step 202 in FIG.2). Preferably, the descriptive analysis created by the generative Al model 105 includes an identification of the cybersecurity-related factors that influence or contribute significantly to the presence of each of the cybersecurity-related vulnerabilities, within the data management environment 109.

[0053] For instance, the generative Al model 105 determines that the factors, such as a) the use of unsecured computer networks, b) the use of weaker encryption protocols, c) the use of default router configurations, and d) the lack of network segmentation, largely contribute to the occurrence of Wi-Fi eavesdropping and packet sniffing in a datamanagement environment 109.

[0054] Furthermore, the generative Al model 105 determines that the cybersecurity-related factors, such as a) the presence of phishing emails in an enterprise mailbox, b) use of weak, default, and predictable passwords, c) use of unpatched, non-updated, outdated software programs and operating systems, d) misconfigured data security settings and cloud security settings, and e) insecure connections between an enterprise network and a third party outsider network, largely contribute to data breach-based cyberattack in a data management environment 109.

[0055] Likewise, the generative Al model 105 determines that the cybersecurity-related factors, such as a) higher levels of dependency on third-party vendor-controlled systems and software programs, b) extensive use of unverified open-source software, and c) interconnected digital ecosystems, largely contribute to a supply chain attack on a data management environment 109.

[0056] Likewise, the generative Al model 105 determines that the cybersecurity-related factors, such as a) use of insecure and / or public Wi-Fi networks, b) lack of end-to-end encryption, c) improper security certificate validation, d) the use of outdated firmware, e) lack of pinning specific servers with corresponding security certificates, and f) forceful downgrade of a browser connection to HTTP, largely contribute to a man-in-the-middle attack in a data management environment 109.

[0057] Likewise, the generative Al model 105 determines that the cybersecurity-related factors, such as a) cloud misconfigurations, b) failure to use multi-factor authentication, c) insecure APIs and interfaces, d) continued use of compromised credentials, e) limited visibility into cloud environments and the safety thereof, largely contribute to a cloud security attack on a data management environment 109.

[0058] Likewise, the generative Al model 105 determines that the cybersecurity-related factors, such as a) lack of a limitation on the number of requests a user can make, b) presence of unpatched software vulnerabilities, c) use of insecure loT devices, d) use of open DNS resolvers, and e) buffer overflow, largely contribute to a distributed denial of service attack on a data management environment 109.

[0059] Further, the generative Al model 105 determines that the cybersecurity-related factors, such as a) direct conversion of user input into SQL queries, b) insufficient validation and sanitization of user input, c) use of raw SQL queries instead of pre-compiled SQLstatements with parameterized inputs, d) use of over-privileged database accounts, and e) use of improperly configured object-relational mappers, largely contribute to an SQL injection-based cyberattack on a data management environment 109.

[0060] Further, the generative Al model 105 determines that the cybersecurity-related factors, such as a) reuse of the same username-password pair on different devices and b) lack of multi-factor authentication, largely contribute to a credential sniffing-based cyberattack on a data management environment 109.

[0061] Further, the generative Al model 105 determines that the cybersecurity-related factors, such as a) the use of an insecure Wi-Fi network, b) lack of full-site HTTPS, c) use of predictable session IDs, d) storing session IDs in URLs, and e) unsafe browser habits, largely contribute to a credential sniffing-based cyberattack on a data management environment 109.

[0062] Likewise, the generative Al model 105 determines that the cybersecurity-related factors such as a) the use of excessive permissions, b) repeated use of default credentials, c) misconfiguration of user roles and privileges, d) the use of unsecured software services, e) the use of unpatched software and operating systems, f) buffer overflows, and g) the use of insecure APIs, largely contribute to a privilege hijacking-based cyberattack on a data management environment 109.

[0063] Likewise, the generative Al model 105 determines that the cybersecurity-related factors, such as a) leaving Remote Desktop Protocol port open to the Internet, b) use of unpatched software and operating systems, c) use of weak user credentials, and d) lack of multi-factor authentication, largely contribute to an initial access brokerage-based cyberattack on a data management environment 109.

[0064] Likewise, the generative Al model 105 determines that the cybersecurity-related factors, such as a) accessing and replying to phishing emails, b) presence of weaker password management strategies, c) the use of unpatched software and operating systems, and d) lack of backup for confidential and sensitive data, largely contribute to a ransomware attack on a data management environment 109.

[0065] Likewise, the generative Al model 105 determines that the cybersecurity-related factors, such as a) lack of output encoding, b) lack of user input validation, c) unsafe handling of user-generated inputs, d) use of weaker content security policies, and e) use of vulnerable application frameworks, largely contribute to cross site scripting cyberattackon a data management environment 109.

[0066] Likewise, the generative Al model 105 determines that the cybersecurity-related factors such as a) failure to validate the user’s ownership of data objects, b) use of sequential and thus predictable object IDs, c) exposure of internal database keys, d) insecure API design, e) lack of tenant isolation in shared database architectures, and f) inconsistent authorization enforcement, largely contribute to broken object level authorization (BOLA)-based cyberattack on a data management environment 109.

[0067] In accordance with the present disclosure, the generative Al model 105 applies a set of predetermined statistics to each of the cybersecurity-related factors that influence and largely contribute to the presence of the corresponding cybersecurity-related vulnerabilities (step 204 in FIG.2). The statistics applied by the generative Al model 105 include but are not restricted to a total number of cybersecurity-related factors present in the data management environment 109, the duration for which the cybersecurity-related factors were present in the data management environment 109, the effect that each of the cybersecurity-related factors had on the packet data flowing through the data management environment 109, the frequency of occurrence of each of the cybersecurity-related factors within the data management environment 109, the patterns underlying the occurrence of each of the cybersecurity-related factors within the data management environment 109, and the like.

[0068] Subsequently, the generative Al model 105 determines the severity of each of the cybersecurity-related factors based on the above-identified set of statistics applied to each of the cybersecurity-related factors. Preferably, each of the cybersecurity-related factors is categorized either as having a low severity, a medium severity, a high severity, or an extremely high severity. Subsequently, the generative Al model 105 generates a training dataset containing a plurality of training vectors (step 206 in FIG.2). The training vectors contained within the training dataset are indicative of a) the cybersecurity-related vulnerabilities found within the data management environment 109, b) the cybersecurity-related factors that contributed to the presence of each of the cybersecurity-related vulnerabilities, and c) the severity of each of the cybersecurity-related factors.

[0069] Subsequently, the generative Al model 105 trains a predetermined machine learning model 107, by using the training datasets, to determine whether the data management environment 109 is vulnerable to one or more of the cybersecurity attacks, and to determine the likelihood of occurrence of one or more cybersecurity attacks on the datamanagement environment 109.

[0070] Preferably, the machine learning model 107, based on the training datasets created by the generative Al model 105, identifies a) the cybersecurity-related vulnerabilities present within the data management environment 109, b) the cybersecurity-related factors that contributed to the presence of each of the cybersecurity-related vulnerabilities, within the data management environment 109, and c) the seventy of each of the cybersecurity-related factors.

[0071] The machine learning model 107 subsequently correlates the presence of the cybersecurity-related vulnerabilities within the data management environment 109, the number of the cybersecurity-related factors that contributed to the presence of the said cybersecurity-related vulnerabilities, and the severity of each of the cybersecurity-related factors, to the possibility of occurrence of one or more cyberattacks on the data management environment 109, and determines the likelihood of onset of one or more cyberattacks on the data management environment 109.

[0072] In accordance with the present disclosure, the generative Al model 105 expands or truncates the training dataset prior to training the machine learning model 107. In an exemplary embodiment of the present disclosure, the training dataset created by the generative Al model 105 contains the following exemplary training vectors: {a, b, a, a, b, a, a, a, a, a}. The generative Al model 105 determines the probability of training vectors ‘a’ and ‘b’ appearing in the training dataset.

[0073] In this case, the probability of training vector ‘a’ appearing in the training dataset is 0.8, and the probability of training vector ‘b’ appearing in the training dataset is 0.2. Preferably, the probability of training vector ‘a’ appearing in the training dataset is determined based on a relative frequency of occurrence of training vector ‘a’ in relation to the relative frequency of occurrence of the remaining training vectors (in this case, the training vector ‘b’), and the total number of training vectors present in the training dataset. The same analogy is applied by the generative Al model 105 to determine the probability of training vector ‘b’ appearing in the training dataset.

[0074] Preferably, the frequency-based probability of training vector ‘a’ being included in the training dataset is 0.8, and the frequency-based probability of training vector ‘b’ being included in the training dataset is 0.2. Further, the generative Al model 105 converts the training dataset into a reduced training dataset or a shrunken training dataset, based on the probability of occurrence, i.e., the frequency-based probability of occurrence, of eachof the training vectors ‘a’ and ‘b’. In this case, a reduced dataset created by the generative Al model 105 will have the following training vectors: {a, a, a, a, b}.

[0075] In accordance with the present disclosure, the generative Al model 105 also expands the above-mentioned reduced training dataset based on the frequency-based probability of occurrence of training vectors ‘a’ and ‘b’.

[0076] The generative Al model 105 expresses the frequency-based probability of occurrence of training vector ‘a’ as a function of the total number of instances of training vector ‘a’ divided by the total number of training vectors (i.e. , in this case, 4 divided by 5). Likewise, the generative Al model 105 expresses the frequency-based probability of occurrence of training vector ‘b’ as “1 divided by 5”.

[0077] For instance, the probability of training vector ‘a’ appearing in the reduced training dataset is 0.8, and the probability of training vector ‘b’ appearing in the reduced training dataset is 0.2. Therefore, the generative Al model 105 expands the training dataset based on the frequency-based probability of occurrence of training vectors ‘a’ and ‘b’, i.e., {(0.8, a), (0.2, b)}.

[0078] Furthermore, the generative Al model 105 represents {(0.8, a), (0.2, b)} as {(4 / 5, a), (1 / 5, b)}, based on the frequency-based probability of occurrence of training vectors ‘a’ and ‘b’, which are “4 / 5” and “1 / 5”, respectively. Subsequently, the generative Al model 105 expands the reduced training dataset to include “four” instances of training vector ‘a’ and “one” instance of training vector ‘b’, followed again by “four” instances of training vector ‘a’ and “one” instance of training vector ‘b’.

[0079] TECHNICAL ADVANTAGES

[0080] The present disclosure addresses the need for a computer-implemented system and method capable of automatically generating cybersecurity threats and risk-related training datasets. The present disclosure describes a computer-implemented system and method that utilizes artificial intelligence-based computing principles to dynamically generate and update training datasets embodying cybersecurity threats and risk-related information. Such dynamically generated training datasets are accurate and adapt rapidly to the ever-changing cybersecurity threats and risks landscape. By leveraging cybersecurity threats and risk-related information, the system and method envisaged by the present disclosure continuously refine the training sets, thereby enhancing the effectiveness of the underlying machine learning models. The system and methodenvisaged by the present disclosure also enable automatic adjustment of machine learning model parameters, thus allowing the machine learning models to adapt rapidly to any shift in the patterns corresponding to the cybersecurity threats and risks-related information, and to thus remain relevant and accurate as far as the detection of cybersecurity-related threats and risks is concerned. Further, the computer-implemented system and method envisaged by the present disclosure also creates cyber profiles representative of the cyber posture of data handling environments by analyzing their cybersecurity vulnerabilities across various cyber-related operational aspects, including Wi-Fi-related vulnerabilities, phishing email-related vulnerabilities, business email-related vulnerabilities, software version-related vulnerabilities, systems maintenance-related vulnerabilities, authentication-related vulnerabilities, third-party applications-related vulnerabilities, and social engineering-related vulnerabilities. Furthermore, the system and method envisaged by the present disclosure predict the cyber risks likely to be faced by various data handling environments, such as business entities, financial institutions, and the like, by evaluating their specific cyber profiles. The system and method envisaged by the present disclosure are pre-programmed to assess the cybersecurity-related vulnerabilities, determine potential cybersecurity-related threats, and identify potential cyber-attack targets. The machine learning model envisaged by the present disclosure also generates a cyber compliance security score to measure the overall security effectiveness of data handling environments, based on the cybersecurity-related vulnerabilities embodied in the data handling environments and the cybersecurity-related countermeasures adapted by the data handling environments.

Claims

CLAIMSWhat is claimed is:

1. A computer-implemented system for collecting data related to cybersecurity-related vulnerabilities and generating at least one dataset based on said cybersecurity- related vulnerabilities to train at least one predetermined machine learning model to identify a possibility of occurrence of corresponding cybersecurity attacks in a data management environment, said computer-implemented system comprising:a processor;a memory communicably coupled to said processor, said memory storing computer-executable instructions, which when executed by said processor, cause said processor to:establish a communicable coupling with a generative artificial intelligence (Al) model storing thereon information indicative of at least a plurality of cybersecurity-related vulnerabilities embodied in said data management environment, and trigger said generative Al model to:generate a descriptive analysis corresponding to said plurality of cybersecurity-related vulnerabilities embodied in said data management environment, said descriptive analysis comprising an identification of respective cybersecurity-related factors that influenced each of said cybersecurity-related vulnerabilities;apply a pre-defined set of statistics to each of said cybersecurity-related factors and determine a seventy of each of said cybersecurity-related factors, based on said set of statistics;generate a plurality of training vectors, and wherein elements within each of said training vectors include at least an indication of whether said data management environment is vulnerable to one or more of said cybersecurity attacks, in consideration of said cybersecurity-related factors and said severity thereof;create at least one training dataset comprising said training vectors, and expand said training dataset based at least on a relative frequency of occurrence of a specific training vector within said training dataset to eachof remaining training vectors within said training dataset and a total number of said training vectors embodied in said training dataset.

2. The system as claimed in claim 1, wherein said generative Al model is further configured to represent said relative frequency of occurrence of said specific training vector to each of said remaining training vectors, as a frequency-based probability, said generative Al model further configured to create a truncated version of said training dataset by incorporating with said truncated version a limited number of said training vectors, determined based on said frequency-based probability.

3. The system as claimed in claim 1, wherein said generative Al model is further configured to create a cause-effect mapping by linking each of said cybersecurity- related vulnerabilities to corresponding cybersecurity attack family fingerprints, based on a root cause analysis of each of said cybersecurity attack family fingerprints.

4. The system as claimed in claim 1, wherein said generative Al model is further configured to determine a possibility of occurrence of any of said cybersecurity attacks on said data management environment, by identifying unique cybersecurity-related vulnerabilities embodied within said data management environment.

5. A computer-implemented method for collecting data related to cybersecurity-related vulnerabilities and generating at least one dataset based on said cybersecurity-related vulnerabilities to train at least one predetermined machine learning model to identify a possibility of occurrence of corresponding cybersecurity attacks in a data management environment, said computer-implemented method comprising the following steps: establishing, by a processor, a communicable coupling with a generative artificial intelligence (Al) model storing thereon information indicative of at least a plurality of cybersecurity-related vulnerabilities embodied in said data management environment; andtriggering, by said processor, said generative Al model to:generate a descriptive analysis corresponding to said plurality of cybersecurity-related vulnerabilities embodied in said data management environment, said descriptive analysis comprising an identification of respective cybersecurity-related factors that influenced each of said cybersecurity-related vulnerabilities;apply a pre-defined set of statistics to each of said cybersecurity-related factors and determine a seventy of each of said cybersecurity-related factors, based on said set of statistics;generate a plurality of training vectors, and wherein elements within each of said training vectors include at least an indication of whether said data management environment is vulnerable to one or more of said cybersecurity attacks, in consideration of said cybersecurity-related factors and said severity thereof;create at least one training dataset comprising said training vectors, and expand said training dataset based at least on a relative frequency of occurrence of a specific training vector within said training dataset to each of remaining training vectors within said training dataset and a total number of said training vectors embodied in said training dataset.

6. The method as claimed in claim 5, wherein the method further includes the following steps:representing, by said generative Al model, said relative frequency of occurrence of said specific training vector to each of said remaining training vectors, as a frequency-based probability;creating, by said generative Al model, a truncated version of said training dataset by incorporating with said truncated version a limited number of said training vectors, determined based on said frequency-based probability.

7. The method as claimed in claim 5, wherein the method further includes a step of creating, by said generative Al model, a cause-effect mapping by linking each of said cybersecurity-related vulnerabilities to corresponding cybersecurity attack family fingerprints, based on a root cause analysis of each of said cybersecurity attack family fingerprints.

8. The method as claimed in claim 5, wherein the method further includes a step of determining, by said generative Al model, a possibility of occurrence of any of said cybersecurity attacks on said data management environment, by identifying unique cybersecurity-related vulnerabilities embodied within said data management environment.

9. A non-transitory computer-readable storage medium having computer-executable instructions embodied thereon, said computer-executable instructions, when executed by a processor, cause said processor to:establish a communicable coupling with a generative artificial intelligence (Al) model storing thereon information indicative of at least a plurality of cybersecurity- related vulnerabilities embodied in a data management environment, and trigger said generative Al model to:generate a descriptive analysis corresponding to said plurality of cybersecurity-related vulnerabilities embodied in said data management environment, said descriptive analysis comprising an identification of respective cybersecurity-related factors that influenced each of said cybersecurity-related vulnerabilities;apply a pre-defined set of statistics to each of said cybersecurity-related factors and determine a severity of each of said cybersecurity-related factors, based on said set of statistics;generate a plurality of training vectors, and wherein elements within each of said training vectors include at least an indication of whether said data management environment is vulnerable to one or more of said cybersecurity attacks, in consideration of said cybersecurity-related factors and said severity thereof;create at least one training dataset comprising said training vectors, and expand said training dataset based at least on a relative frequency of occurrence of a specific training vector within said training dataset to each of remaining training vectors within said training dataset and a total number of said training vectors embodied in said training dataset.

10. The non-transitory computer-readable storage medium as claimed in claim 9, wherein said computer-executable instructions, when executed by said processor, further cause said processor to trigger said generative Al model to:represent said relative frequency of occurrence of said specific training vector to each of said remaining training vectors, as a frequency-based probability, said generative Al model further configured to create a truncated version of said trainingdataset by incorporating with said truncated version a limited number of said training vectors, determined based on said frequency-based probability; create a cause-effect mapping by linking each of said cybersecurity-related vulnerabilities to corresponding cybersecurity attack family fingerprints, based on a root cause analysis of each of said cybersecurity attack family fingerprints; and determine a possibility of occurrence of any of said cybersecurity attacks on said data management environment, by identifying unique cybersecurity-related vulnerabilities embodied within said data management environment.