Method and system for generating a response to an internet of things (IOT) security-related query
The method and system leverage a language model to generate tailored IoT security responses by processing diverse datasets, addressing the challenge of providing accessible and actionable intelligence for various user groups.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-09-03
- Publication Date
- 2026-03-12
AI Technical Summary
Existing solutions for Internet of Things (IoT) security fail to provide timely, actionable, and easily accessible threat intelligence to both technical and non-technical users, as they often offer raw data that is difficult to understand and do not effectively leverage insights from diverse datasets.
A method and system using a language model to generate responses to IoT security-related queries by retrieving and processing document chunks from multiple IoT security-related datasets, adapting selection based on query relevance and user context, and converting diverse data formats into retrievable documents.
Provides reliable, relevant, and user-friendly IoT security insights tailored to different user expertise levels, ensuring up-to-date and actionable responses.
Smart Images

Figure SG2025050582_12032026_PF_FP_ABST
Abstract
Description
METHOD AND SYSTEM FOR GENERATING A RESPONSE TO AN INTERNET OF THINGS (loT) SECURITY-RELATED QUERYCROSS-REFERENCE TO RELATED APPLICATION
[0001] This application claims the benefit of priority of Singapore Patent Application No. 10202402752Q filed on 4 September 2024, the content of which being hereby incorporated by reference in its entirety for all purposes.TECHNICAL FIELD
[0002] The present invention generally relates to a method and a system for generating a response to an Internet of Things (loT) security-related query using a language model, as well as an loT security-related query system comprising the above-mentioned system for generating a response to an loT security-related query, and an loT security-related dataset processing system configured to process contents of the loT security-related datasets into document chunks.BACKGROUND
[0003] Internet of Things (loT) generally refers to a vast network of physical and virtual entities (e.g., inter-networking of physical devices, such as smart devices), which may be characterized by their sensing and / or actuation capabilities, programmability features, and unique identifiers This interconnected infrastructure has rapidly expanded, and loT connections surpassed non-IoT connections in 2020. It is estimated that over 40 billion loT devices will be integrated into homes and workplaces via sensors, processors, and software by 2030. With the rapid developments of loT, increased connectivity and complexity of loT ecosystems also introduce various vulnerabilities, making them attractive targets for attacks. Consequently, loT security has become a critical issue for individuals, organizations, and governments worldwide.
[0004] Over the decades, loT security has garnered significant attention from researchers, covering both defensive and offensive strategies. As the large language model (LLM) has made significant strides in recent years, it has been explored in the context of loT security as well, such as threat / vulnerability identification, perceive loT sensor data, device management and labelling, and loT trust semantics enhancements. This surge in interest has led to many domainspecific datasets (loT security-related datasets of different security domains) that offerextensive insights covering various aspects of loT security, such as vulnerabilities and exploits, tactics, techniques, and procedures (TTPs), and industry-standard guidelines. However, most existing works focus on discovering new vulnerabilities / threats or designing novel defensive / offensive techniques but pay less emphasis on leveraging the insights contained in datasets to assist or guide users in enhancing their security practices. Although certain sources offer information search services, they typically only provide raw data (e.g., vulnerability descriptions) which is often challenging for non-technical users to understand. Even experienced security analysts may find it difficult to extract actionable insights from massive unprocessed information. Consequently, there is an urgent need for solutions that deliver timely, actionable, and easily accessible loT security and threat intelligence to a wide range of users, including both technical and non-technical ones.
[0005] A need therefore exists to provide a method of generating a response to an loT security-related query, as well as a system thereof, that is reliable and relevant (or with superior or enhanced reliability and relevance). It is against this background that the present invention has been developed.SUMMARY
[0006] According to a first aspect of the present invention, there is provided a method of generating a response to an loT security-related query using a language model, the method comprising: receiving an loT security -related query regarding security of one or more loT devices; retrieving document chunks relating to the loT security-related query from selected one or more loT security-related document databases of a plurality of loT security-related document databases based on the loT security -related query; generating a response prompt based on the loT security-related query and the document chunks retrieved; and generating a response to the loT security-related query using the language model based on the response prompt, wherein the plurality of loT security-related document databases is generated from a plurality of loT security -related datasets, respectively.
[0007] According to a second aspect of the present invention, there is provided a system for generating a response to an loT security-related query using a language model, the system comprising:at least one memory; and at least one processor communicatively coupled to the at least one memory and configured to: receive an loT security-related query regarding security of one or more loT devices; retrieve document chunks relating to the ToT security-related query from selected one or more loT security-related document databases of a plurality of loT security-related document databases based on the loT security -related query; generate a response prompt based on the loT security-related query and the document chunks retrieved; and generate a response to the loT security-related query using the language model based on the response prompt, wherein the plurality of ToT security-related document databases is generated from a plurality of ToT security -related datasets, respectively.
[0008] According to a third aspect of the present invention, there is provided an loT security-related query system comprising: a system for generating a response to an ToT security-related query using a language model according to the above-mentioned second aspect of the present invention; and an ToT security -related dataset processing system comprising: at least one memory; and at least one processor communicatively coupled to the at least one memory and configured to process, for each of the plurality of ToT security -related datasets, contents of the ToT security-related dataset into document chunks for a corresponding ToT security-related document database of the plurality of ToT security-related document databases.BRIEF DESCRIPTION OF THE DRAWINGS
[0009] Embodiments of the present invention will be better understood and readily apparent to one of ordinary skill in the art from the following written description, by way of example only, and in conjunction with the drawings, in which:FIG. 1 depicts a schematic diagram of a method of generating a response to an ToT security-related query using a language model, according to various embodiments of the present invention;FIG. 2 depicts a schematic block diagram of a system for generating a response to an loT security-related query using a language model, according to various embodiments of the present invention;FIG. 3 depicts a schematic block diagram of an loT security-related query system, according to various embodiments of the present invention;FIG. 4 depicts a schematic diagram of a method of processing contents (or data) of the loT security-related datasets, according to various embodiments of the present invention;FIG. 5 depicts a schematic block diagram of an example loT security-related query system, herein referred to as CHATIoT, according to various example embodiments of the present invention;FIG. 6 shows Table 1 presenting example selector configurations for selecting selfquerying retrievers (and thus, selects the corresponding loT security-related document database), according to various example embodiments of the present invention;FIG. 7 depicts a schematic block diagram of an example self-querying retriever, according to various example embodiments of the present invention;FIGs. 8A to 8C show the metadata fields information and examples for the self-querying retriever corresponding to various loT security-related datasets, according to various example embodiments of the present invention;FIG. 9 shows the generated queries and filters for VARIoT (Vulnerability and Attack Repository for loT) vulnerabilities and CLS (Cybersecurity Labelling Scheme (in Singapore)) list for an example submitted query by a Security Analyst, according to various example embodiments of the present invention;FIG. 10 depicts a schematic block diagram of an example data processing system, herein referred to as DataKit, according to various example embodiments of the present invention;FIG. 11 presents the example detailed specifications for four example properties of the use case, according to various example embodiments of the present invention;FIG. 12 shows Table 2 presenting key actions for each of the five example user roles, according to various example embodiments of the present invention;FIG. 13 shows Table 3 presenting the LLMs-based field selection for page content and metadata, according to various example embodiments of the present invention;FIG. 14 shows Table 4 presenting the specific fields selected for page content and metadata of each of the five example loT security-related datasets used in the CHATIoT, according to various example embodiments of the present invention;FIG. 15 shows Table 5 presenting context (precision, recall) of different chunking configurations for VARIoT vulnerabilities, exploits, MITRE, ATT&CK and threat reports, according to various example embodiments of the present invention;FIG. 16 shows Table 6 presenting optimized chunking strategy for VARIoT vulnerability, exploits, ICS, and threat reports, according to various example embodiments of the present invention;FIG. 17 shows a prompt template for LMM-based evaluation of outputs, according to various example embodiments of the present invention;FIG. 18A shows Table 7 presenting the experimental results for comparison of CHATIoT with LLM-A for moderate LLMs LLaMA3:8B, LLaMA3.1:8B, and GPT-4o-mini;FIG. 18B shows Table 8 presenting the experiment results for comparison of CHATIoT 500 with LLM-A for more advanced LLMs LLaMA3. 1 :70B and GPT-4o; andFIG. 19 presents the results of human evaluations comparing CHATIoT with GPT-4o (LLM-A).DETAILED DESCRIPTION
[0010] Various embodiments of the present invention provide a method of generating a response to an Internet of Things (loT) security-related query using a language model, and a system thereof In addition, various embodiments of the present invention provide an loT security-related query system comprising the above-mentioned system for generating a response to an loT security-related query, and an loT security -related dataset processing system configured to process contents of the loT security-related datasets into document chunks.
[0011] As discussed in the background, with the rapid developments of loT, increased connectivity and complexity of loT ecosystems also introduce various vulnerabilities, making them attractive targets for attacks. Consequently, loT security has become a critical issue for individuals, organizations, and governments worldwide. Therefore, various embodiments of the present invention seek to provide a method of generating a response to an loT security-related query, as well as a system thereof, that is reliable and relevant (or with superior or enhanced reliability and relevance).
[0012] FIG. 1 depicts a schematic diagram of a method 100 of generating a response to an loT security-related query using a language model, according to various embodiments of the present invention. The method 100 comprises: receiving (at 106) an loT security-related query regarding security of one or more loT devices; retrieving (at 108) document chunks relating tothe loT security-related query from selected one or more loT security-related document databases of a plurality of loT security-related document databases based on the loT security- related query; generating (at 110) a response prompt based on the loT security-related query and the document chunks retrieved; and generating (at 112) a response to the loT security- related query using the language model based on the response prompt. Furthermore, the plurality of loT security-related document databases is generated from a plurality of loT security-related datasets, respectively. Accordingly, the method 100 of generating a response is based on retrieval -augmented generation (RAG).
[0013] The method 100 of generating a response to an loT security-related query is advantageously reliable and relevant (or with superior or enhanced reliability and relevance). Firstly, multiple loT security-related document databases are generated (or built) from multiple loT security-related datasets (e.g., comprising loT security-related datasets of different security domains), respectively. It will be understood by a person skilled in the art that the present invention is not limited to any parti cular / specific number and / or types of loT security-related datasets and any number and / or types of loT security-related datasets (e.g., any loT security- related dataset known in the art) may be utilized as desired or as appropriate without going beyond the scope of the present invention. Therefore, the method 100 is advantageously able to leverage or utilize loT security-related information / contents contained in multiple loT security- related datasets for generating the response to the loT security-related query. Tn addition, by having the multiple loT security-related document databases generated from the multiple loT security-related datasets, respectively, each loT security-related document database is able to be kept up-to-date with latest document chunks derived from the latest loT security-related information / contents contained in the corresponding loT security-related dataset, such as new vulnerabilities or threats. Furthermore, the method 100 retrieves document chunks relating to the loT security-related query from selected one or more loT security-related document databases based on the loT security-related query. Therefore, the method 100 advantageously adaptively selects loT security -related document database(s) that are determined to be relevant to the loT security-related query for retrieving document chunks, rather than simply attempting to retrieve document chunks from all available loT security-related document databases. This adaptive selection from the available loT security-related document databases not only avoids negative effects to the quality of the generated response (answer) by document chunks from irrelevant loT security-related document databases, but also avoids unnecessary resource costs (e.g., computational overhead) Therefore, the method 100 of generating a response to an loTsecurity-related query is advantageously reliable and relevant (or with superior or enhanced reliability and relevance). These advantages or technical effects, and / or other advantages or technical effects, will become more apparent to a person skilled in the art as the method 100 of generating a response to an loT security-related query, as well as the corresponding system for generating a response to an loT security -related query, is described in more detail according to various embodiments and example embodiments of the present invention.|0014| In various embodiments, each of the plurality of loT security-related document databases is generated from a corresponding loT security-related dataset of the plurality of loT security-related datasets. Accordingly, each loT security-related document database is dedicated to store document chunks generated from a corresponding or specific loT security- related dataset. Therefore, each loT security-related document database is able to be kept up- to-date with latest document chunks derived from the latest loT security-related information / contents contained in the corresponding loT security-related dataset, such as new vulnerabilities, exploits or threats. Furthermore, since each loT security-related document database is dedicated to store document chunks generated from a corresponding or specific loT security-related dataset, particular or specific one or more loT security-related document database(s) that are determined to be relevant to the loT security-related query can be advantageously selected for retrieving document chunks, rather than simply attempting to retrieve document chunks from all available loT security-related document databases.
[0015] In various embodiments, each of the plurality of loT security-related document databases is a vector database comprising document chunks generated from the corresponding loT security-related dataset. In this regard, each document chunk comprises chunked texts, embedding vectors (of the chunked texts), and metadata (for facilitating filtering of the document chunk).
[0016] In various embodiments, for each of the plurality of loT security-related document databases, the document chunks therein each has one or more content fields and one or more metadata fields defined for the document chunks generated from the corresponding loT security-related dataset. For example, one or more content fields and one or more metadata fields are specifically defined for chunked texts and metadata for all document chunks (or all documents) generated from the loT security-related dataset for the corresponding loT security- related document database In various embodiments, one or more content fields and one or more metadata fields may be selected from a set of content fields and a set of metadata fields using a language model (e g., LLM)
[0017] In various embodiments, for each of one or more loT security-related document databases of the plurality of loT security-related document databases, non-textual contents in the corresponding loT security-related dataset are converted to textual contents in the loT security-related document database. In various embodiments, the plurality of loT security- related datasets comprises loT security-related datasets of different security domains. In this regard, loT security-related datasets may include loT security-related information / contents of various data formats, such as but not limited to, text, tables, figures and codes. Therefore, according to various embodiments, the method 100 advantageously enables the conversion of diverse data formats into retrievable documents (or document chunks) for the loT security- related document databases, thereby enabling loT security-related datasets of diverse data formats (e.g., multi-modal contents) to be utilized and maximizing the utilization of loT security-related information / contents (e g , multi-modal contents) contained in the loT security- related datasets.
[0018] In various embodiments, the method 100 further comprises selecting one or more loT security-related document databases from the plurality of loT security-related document databases based on the loT security-related query and descriptions of the plurality of loT security-related document databases to obtain the selected one or more loT security-related document databases. As described hereinbefore, the method 100 advantageously adaptively selects loT security-related document database(s) that are determined to be relevant to the loT security-related query for retrieving document chunks, rather than simply attempting to retrieve document chunks from all available loT security-related document databases. In various embodiments, such a selection of loT security-related document database(s) is performed further based on the descriptions of the plurality of loT security-related document databases.
[0019] In various embodiments, the above-mentioned selecting the one or more loT security-related document databases comprises: generating a selector prompt based on the loT security-related query and the descriptions of the plurality of loT security-related document databases (e g., the selector prompt comprises the loT security-related query and the descriptions of the plurality of loT security-related document databases); generating loT security-related document database selection information using the language model based on the selector prompt; and selecting the one or more loT security-related document databases based on the loT security-related document database selection information.
[0020] In various embodiments, the selector prompt is generated further based on user information corresponding to a user associated with the loT security-related query (e.g., theselector prompt further comprises the user information). In various embodiments, the user information comprises user type information and associated background information. In this regard, various embodiments note that users in loT ecosystems have diverse user requirements, for example, ranging from consumers to security analysts, each user type with distinct expertise levels and specific security concerns. Various embodiments found that naively combining LLM with loT security information would struggle to effectively cater to these varied requirements, limiting its ability to provide meaningful and contextually appropriate responses for different user groups. In contrast, by utilizing the user’s background (such as, but not limited to, knowledge, goals and requirements) according to various embodiments of the present invention, the method 100 is advantageously able to generate a response to an loT security- related query that is not only reliable, relevant and technical, but also user-friendly. For example, the method 100 is able to generate the latest reliable, relevant, and technical loT security answers tailored to both query contexts and users’ requirements. Accordingly, by tailoring for different user types, the method 100 guides its generated answers based on the users’ roles (consumer, security analyst, and so on) and their corresponding backgrounds. This allows each kind of user to get insights or solutions that are relevant, understandable, and actionable based on their specific requirements and expertise level.
[0021] In various embodiments, the user information of the user is selected from a predefined set of user information for different user types. For example, in various embodiments, a set of common use case specifications is defined to guide the LLM in generating answers aligned with users’ specific needs and expertise levels.
[0022] In various embodiments, the above-mentioned retrieving document chunks relating to the loT security-related query comprises: generating, for each of the selected one or more loT security-related document databases, one or more structured queries for the selected loT security-related document database based on the loT security -related query; and retrieving, from each of the selected one or more loT security-related document databases, one or more document chunks relating to the loT security-related query from the selected loT security- related document database based on the one or more structured queries and search parameters for the selected loT security-related document database. For example, for each selected security-related document database, when the loT security-related query is passed to a query constructor associated with the selected security-related document database, a LLM associated with the selected security-related document database may generate internal query language elements based on pre-defined metadata field information and metadata examples. A querytranslator associated with the selected security-related document database may convert these elements into a structured query with appropriate filters. The structured query and search parameters may then be applied to the selected loT security-related document database (vector store) to retrieve documents (or document chunks).
[0023] In various embodiments, for each of the plurality of loT security-related document databases, the loT security-related document database is updated based on new loT security- related data in the corresponding loT security-related datasets. Accordingly, each loT security- related document database is kept up-to-date with latest document chunks derived from the latest loT security-related information / contents contained in the corresponding loT security- related dataset, such as new vulnerabilities, exploits or threats. For example, in various embodiments, the loT security-related document databases may each obtain or download latest document chunks derived from the latest ToT security-related information / contents contained in the corresponding loT security-related dataset periodically, when latest loT security-related information / contents are available, or in real-time to keep up-to-date.
[0024] In various embodiments, the language model is a large language model (LLM).
[0025] FIG. 2 depicts a schematic block diagram of a system 200 for generating a response to an loT security-related query using a language model, according to various embodiments of the present invention, corresponding to the above-mentioned method 100 of generating a response to an loT security-related query as described hereinbefore according with reference to FIG. 1 according to various embodiments of the present invention. The system 200 comprises: at least one memory 202; and at least one processor 204 communicatively coupled to the at least one memory 202 and configured to perform the method 100 of generating a response to an loT security -related query according to various embodiments of the present invention. Accordingly, the at least one processor 204 is configured to: receive an loT security-related query regarding security of one or more loT devices; retrieve document chunks relating to the loT security-related query from selected one or more loT security-related document databases of a plurality of loT security-related document databases based on the loT security-related query; generate a response prompt based on the loT security-related query and the document chunks retrieved; and generate a response to the loT security-related query using the language model based on the response prompt. Furthermore, the plurality of loT security-related document databases is generated from a plurality of ToT security-related datasets, respectively.
[0026] It will be appreciated by a person skilled in the art that the at least one processor 204 may be configured to perform various functions or operations through set(s) of instructions(e g., software modules) executable by the at least one processor 204 to perform various functions or operations. Accordingly, as shown in FIG. 2, the system 200 may comprise: an loT security-related query receiving module (or an loT security-related query receiving circuit) 206 configured to receive an loT security-related query regarding security of one or more loT devices; a document chunk retrieving module (or a document chunk retrieving circuit) 208 configured to retrieve document chunks relating to the loT security-related query from selected one or more loT security-related document databases of a plurality of loT security-related document databases based on the loT security-related query; a response prompt generating module (or a response prompt generating circuit) 210 configured to generate a response prompt based on the loT security-related query and the document chunks retrieved; and a response generating module (or a response generating circuit) 212 configured to generate a response to the loT security-related query using the language model based on the response prompt.
[0027] It will be appreciated by a person skilled in the art that the above-mentioned modules of the system 200 are not necessarily separate modules, and two or more modules may be realized by or implemented as one functional module (e g., a circuit or a software program) as desired or as appropriate without deviating from the scope of the present invention. For example, two or more of the loT security-related query receiving module 206, the document chunk retrieving module 208, the response prompt generating module 210, and the response generating module 212 may be realized (e g., compiled together) as one executable software program (e g., software application), which for example may be stored in the at least one memory 202 and executable by the at least one processor 204 to perform the corresponding functions or operations as described herein according to various embodiments of the present invention.
[0028] In various embodiments, the system 200 for generating a response to an loT security -related query corresponds to the method 100 of generating a response to an loT security-related query as described hereinbefore with reference to FIG. 1, therefore, various operations, functions or steps configured to be performed by the at least one processor 204 may correspond to various operations, functions or steps of the method 100 of generating a response to an loT security-related query described hereinbefore according to various embodiments, and thus need not be repeated with respect to the system 200 for generating a response to an loT security-related query for clarity and conciseness. In other words, various embodiments described herein in context of methods (e.g., the method 100 of generating a response to an loTsecurity-related query) are analogously valid for the corresponding systems or devices (e.g., the system 200 for generating a response to an loT security-related query), and vice versa.
[0029] A computing system, a controller, a microcontroller or any other system providing a processing capability may be provided according to various embodiments in the present invention. Such a system may be taken to include one or more processors and one or more computer-readable storage mediums. For example, the system 200 for generating a response to an loT security-related query described hereinbefore includes at least one processor 204 and at least one computer-readable storage medium (or memory) 202 which are for example used in various processing carried out therein as described herein. A memory or computer-readable storage medium used in various embodiments may be a volatile memory, for example a DRAM (Dynamic Random Access Memory) or a non-volatile memory, for example a PROM (Programmable Read Only Memory), an EPROM (Erasable PROM), EEPROM (Electrically Erasable PROM), or a flash memory, e g., a floating gate memory, a charge trapping memory, an MRAM (Magnetoresistive Random Access Memory) or a PCRAM (Phase Change Random Access Memory).
[0030] In various embodiments, a “circuit” may be understood as any kind of a logic implementing entity, which may be special purpose circuitry or a processor executing software stored in a memory, firmware, or any combination thereof. Thus, in an embodiment, a “circuit” may be a hard-wired logic circuit or a programmable logic circuit such as a programmable processor, e.g., a microprocessor (e.g., a Complex Instruction Set Computer (CISC) processor or a Reduced Instruction Set Computer (RISC) processor). A “circuit” may also be a processor executing software, e.g., any kind of computer program, e.g., a computer program using a virtual machine code, e.g., Java. Any other kind of implementation of various functions or operations may also be understood as a “circuit” in accordance with various other embodiments. Similarly, a “module” may be a portion of a system according to various embodiments in the present invention and may encompass a “circuit” as above, or may be understood to be any kind of a logic-implementing entity therefrom.
[0031] Some portions of the present disclosure may be explicitly or implicitly presented in terms of algorithms and functional or symbolic representations of operations on data within a computer memory. These algorithmic descriptions and functional or symbolic representations are the means used by those skilled in the data processing arts to convey most effectively the substance of their work to others skilled in the art. An algorithm may be, and generally, conceived to be a self-consistent sequence of steps leading to a desired result.
[0032] The present specification also discloses a system (e.g., which may also be embodied as devices or apparatuses), such as the system 200, for performing various operations, functions or steps of various methods described herein. Such a system may each be specially constructed for the required purposes or may comprise a general purpose computer system selectively activated or reconfigured by a computer program stored in the computer system. Tn general, various algorithms that may be presented herein are not limited to being implemented or executed by any particular computer system. Alternatively, the construction of more specialized computer system to perform various operations, functions or steps of various methods described herein may be provided as desired or as appropriate without going beyond the scope of the present invention.
[0033] In addition, the present specification also at least implicitly discloses computer program(s) or software / functional module(s), in that it would be apparent to a person skilled in the art that various operations, functions or steps of various methods described herein may be put into effect by computer code. The computer program(s) is not intended to be limited to any particular programming language and implementation thereof, and it will be appreciated by a person skilled in the art that a variety of programming languages and coding thereof may be used to implement the computer program(s). Moreover, the computer program(s) is not intended to be limited to any particular control flow as there are a variety of programming languages which can use different control flows. It will be appreciated by a person skilled in the art that a computer program may be stored on any computer-readable storage medium (non- transitory computer-readable storage medium), such as but not limited to, a magnetic disk, an optical disk or a memory chip. For example, a computer program stored on a computer-readable storage medium may be loaded and executed on a computer system to implement various operations, functions or steps of various methods described herein according to various embodiments of the present invention.
[0034] Accordingly, in various embodiments, there is provided a computer program product, embodied in one or more computer-readable storage mediums (non-transitory computer-readable storage medium), comprising instructions (e g., the ToT security-related query receiving module 206, the document chunk retrieving module 208, the response prompt generating module 210, and / or the response generating module 212) executable by one or more computer processors to perform the method 100 of generating a response to an ToT security- related query as described hereinbefore with reference to FIG. 1 according to various embodiments of the present invention Accordingly, various computer programs or softwaremodules described herein may be stored in a computer program product receivable by a system therein, such as the system 200 for generating a response to an loT security-related query as shown in FIG. 2, for execution by at least one processor 204 of the system 200 to perform various operations, functions or steps of various methods described herein according to various embodiments of the present invention.
[0035] It will be appreciated by a person skilled in the art that various modules of systems described herein (e.g., the loT security-related query receiving module 206, the document chunk retrieving module 208, the response prompt generating module 210, and / or the response generating module 212) may be software module(s) realized by computer program(s) or set(s) of instructions executable by a computer processor to perform various functions or operations. Various modules described herein (e.g., the loT security-related query receiving module 206, the document chunk retrieving module 208, the response prompt generating module 210, and / or the response generating module 212) may also be implemented as hardware module(s) being functional hardware unit(s) designed to perform various functions or operations. More particularly, in the hardware sense, a module is a functional hardware unit designed for use with other components or modules. For example, a module may be implemented using discrete electronic components, or it can form a portion of an entire electronic circuit such as an Application Specific Integrated Circuit (ASIC) or a Field Programmable Gate Array (FPGA). Numerous other possibilities exist It will also be appreciated by a person skilled in the art that a combination of hardware and software modules may be implemented. Furthermore, various operations, functions or steps of various methods described herein may be performed in parallel rather than sequentially as desired or as appropriate (e.g., as long as it does not render the method(s) inoperable or unsatisfactory for its intended purpose).
[0036] FIG. 3 depicts a schematic block diagram of an loT security-related query system 300 comprising: the system 200 for generating a response to an loT security-related query according to various embodiments of the present invention; and an loT security-related dataset processing system 350 comprising: at least one memory 352; and at least one processor 354 communicatively coupled to the at least one memory 352 and configured to process, for each of the plurality of loT security-related datasets, contents of the loT security -related dataset into document chunks for a corresponding loT security-related document database of the plurality of loT security -related document databases.
[0037] FIG. 4 depicts a schematic diagram of a method 400 of processing contents (or data) of the loT security-related datasets comprising processing (at 406) contents of the loT security- related datasets into document chunks for the plurality of loT security-related document databases according to various embodiments of the present invention, and more particularly, processing, for each of the plurality of ToT security-related datasets, contents of the loT security-related dataset into document chunks for a corresponding loT security-related document database of the plurality of loT security-related document databases.
[0038] In various embodiments, the above-mentioned processing contents of the loT security-related dataset into document chunks for the corresponding loT security-related document database comprises: parsing raw data (or raw contents) of the loT security-related dataset into content elements comprising textual content elements and non-textual content elements (such as, but not limited to, tables, figures and codes); converting non-textual content elements of the parsed content elements into textual content elements; and generating the document chunks for the corresponding loT security-related document database based on the textual content elements (i.e., both the textual content elements obtained directly from parsing the raw data and the textual content elements obtained from converting the non-textual content elements). Therefore, loT security-related datasets of diverse data formats (e.g., multi-modal contents) can be utilized and the utilization of loT security-related information / contents (e.g., multi-modal contents) contained in the ToT security-related datasets can be maximized In various embodiments, the plurality of loT security-related datasets comprises loT security- related datasets of different security domains.
[0039] In various embodiments, the method 400 further comprises defining one or more content fields and one or more metadata fields for the document chunks generated for the corresponding loT security-related document database In this regard, the corresponding loT security-related document database is a vector database for storing the document chunks generated from the corresponding loT security-related dataset, and each document chunk comprises chunked texts, embedding vectors, and metadata. For example, one or more content fields and one or more metadata fields are specifically defined for chunked texts and metadata for all document chunks (or all documents) generated from the loT security-related dataset for the corresponding loT security-related document database.
[0040] In various embodiments, the one or more content fields and the one or more metadata fields for the document chunks are defined based on a selection from a set of content fields and a set of metadata fields using a language model. In this regard, various embodimentsnote that contents may be stored across multiple fields. Instead of simply using all fields, according to various embodiments of the present invention, one or more content fields and one or more metadata fields determined to be most relevant to or best represent the contents and metadata of the loT security-related document database are determined or selected for the documents (or document chunks) generated for the corresponding ToT security-related document database.100411 It will be appreciated by a person skilled in the art that the at least one processor 354 may be configured to perform various functions or operations through set(s) of instructions (e.g., software modules) executable by the at least one processor 354 to perform various functions or operations. Accordingly, as shown in FIG. 3, the loT security-related dataset processing system 350 may comprise: a content processing module (or a content processing circuit) 356 configured to process, for each of the plurality of ToT security-related datasets, contents (or data) of the loT security-related dataset into document chunks for a corresponding loT security-related document database of the plurality of loT security-related document databases. Various operations, functions or steps configured to be performed by the least one processor 354 may correspond to various operations, functions or steps of the method 400 described hereinbefore according to various embodiments, and thus need not be repeated with respect to the loT security -related dataset processing system 350 for clarity and conciseness.
[0042] It will be appreciated by a person skilled in the art that the terminology used herein is for the purpose of describing various embodiments only and is not intended to be limiting of the present invention. As used herein, the singular forms “a”, “an” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms “comprises” and / or “comprising,” when used in this specification, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0043] Any reference to an element or a feature herein using a designation such as “first”, “second” and so forth does not limit the quantity or order of such elements or features, unless stated or the context requires otherwise. For example, such designations may be used herein as a convenient way of distinguishing between two or more elements or instances of an element. Thus, a reference to first and second elements does not necessarily mean that only two elements can be employed, or that the first element must precede the second element, unless stated or thecontext requires otherwise. In addition, a phrase referring to “at least one of’ a list of items refers to any single item therein or any combination of two or more items therein.
[0044] In order that the present invention may be readily understood and put into practical effect, various example embodiments of the present invention will be described hereinafter by way of examples only and not limitations. It will be appreciated by a person skilled in the art that the present invention may, however, be embodied in various different forms or configurations and should not be construed as limited to the example embodiments set forth hereinafter. Rather, these example embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the present invention to those skilled in the art
[0045] loT has gained widespread popularity, revolutionizing industries and daily life. However, it has also emerged as a prime target for attacks. Numerous efforts have been made to improve loT security, and substantial loT security (e.g., vulnerabilities or threats) information, such as datasets and reports, have been developed. However, existing research often falls short in leveraging these insights to assist or guide users in harnessing loT security practices in a clear and actionable way. To address this technical problem or gap, various embodiments of the present invention seek to provide a method of generating a response to an loT security-related query, as well as a system thereof, that is reliable and relevant (or with superior or enhanced reliability and relevance). As an illustrative example according to various example embodiments of the present invention, a large language model (LLM)-based loT security assistant, herein referred to as CHATIoT (e.g., corresponding to the loT security- related query system 300 described hereinbefore with reference to FIG. 3 according to various embodiments of the present invention), is provided, which is designed to disseminate loT security intelligence (e g., vulnerabilities or threats). By leveraging the versatile property of retrieval-augmented generation (RAG), CHATIoT successfully integrates the advanced language understanding and reasoning capabilities of LLM with fast-evolving loT security information. Moreover, according to various example embodiments of the present invention, an end-to-end data processing toolkit (e.g., corresponding to the loT security-related dataset processing system 350 described hereinbefore with reference to FIG. 3 according to various embodiments of the present invention) is developed to handle heterogeneous datasets. The data processing toolkit converts datasets of various formats into retrievable documents (or document chunks) and optimizes chunking strategies for efficient retrieval. Additionally, according to various example embodiments of the present invention, a set of common use case specifications(e g., corresponding to the predefined set of user information for different user types described hereinbefore according to various embodiments of the present invention) is defined to guide the LLM in generating answers aligned with users’ specific needs and expertise levels. To demonstrate the effectiveness of the CHATIoT (and thus, the method 100 and system 200 for generating a response to an loT security-related query, as well as the loT security-related query system 300, described hereinbefore according to various embodiments of the present invention), an example prototype of CHATIoT was implemented and extensive experiments were conducted with different LLMs, such as LLaMA3, LLaMA3.1, and GPT-4o. Experimental evaluations demonstrate that CHATIoT is able to generate more reliable, relevant, and technical in-depth answers for most use cases. For example, when evaluating the answers with LLaMA3:70B, CHATIoT was found to improve the above metrics by over 10% on average, particularly in relevance and technicality, compared to using LLMs alone.
[0046] Based on LLM’s advanced capability, there may be two different approaches to generate a response to an loT security-related query: i) fine-tuning the LLM specifically on loT security information and ii) RAG, which retrieves loT security information to augment the LLM’ s generation. However, naive application of these approaches are still facing one or more challenges, such as one or more of the following.
[0047] Challenge 1 : loT Security Evolves Rapidly. The field of loT security is continuously developing, with new vulnerabilities, exploits, and security protocols emerging regularly. For example, in August 2024, the VARIoT vulnerability dataset received 79 updates, and cybersecurity websites almost report new developments daily. As a consequence, various example embodiments found that fine-tuning an LLM on static loT data would quickly become outdated, failing to provide the latest threats or innovations in loT security.
[0048] Challenge 2: Heterogeneous Dataset Formats loT security-related datasets come in a variety of formats, such as structured data (e.g., vulnerability databases), unstructured text reports, and product compliance lists (such as Cybersecurity Labelling Schemes (CLS) in Singapore, Finland, Germany, and the United States). There may also be a wide range of loT security-related datasets, such as but not limited to, loT security vulnerabilities, exploits, common attacking tactics, techniques, and procedures (TTPs), and assessments of loT device security levels. Various example embodiments found that fine-tuning / augmenting an LLM on such diverse data types without a robust data processing mechanism would significantly limit the utilization of information contained in these datasets.
[0049] Challenge 3: Diverse User Requirements. Users in loT ecosystems range from consumers to security analysts, each with distinct expertise levels and specific security concerns. Various example embodiments found that naively combining LLM with loT security information may struggle to effectively cater to these varied requirements, limiting its ability to provide meaningful and contextually appropriate responses for different user groups
[0050] In light of the knowledge gap and technical challenges analyzed above, various example embodiments seek to enable the utilization of large language models to effectively disseminate loT security assistance to various key users of the loT ecosystems in an understandable and actionable manner to provide better loT security guarantees. In this regard, as described above, various example embodiments provide CHATIoT - a large language model (LLM)-based security assistant for loT with RAG, for example, an LLM-based loT security assistant augmented with external information retrieval, to answer this question affirmatively. On the paradigm of RAG, according to various example embodiments of the present invention, CHATIoT is designed by combining the advanced language understanding and reasoning capabilities of LLM with fast-evolving loT security-specific information to generate reliable, relevant, and technical answers. According to various example embodiments of the present invention, to handle heterogeneous datasets, an end-to-end data processing toolkit is developed that integrates a range of existing technologies, enabling the conversion of diverse data formats into retrievable documents (or document chunks) and optimizing chunking strategies to improve retrieval performance. Additionally, according to various example embodiments of the present invention, several use-case specifications may be defined (e.g., corresponding to the predefined set of user information for different user types described hereinbefore according to various embodiments of the present invention) and the user’s background may be provided to ensure that the LLM generated answers are also user-friendly, that is, aligned with each user’s specific needs and expertise levels. Accordingly, CHATIoT offers the following contributions.
[0051] loT security assistant driven by LLM and threat intelligence. The loT security assistant, CHATIoT, is designed to leverage advanced LLM understanding and reasoning capabilities and threat intelligence at the same time. According to various example embodiments of the present invention, self-querying retrievers are developed for diverse loT threat datasets and dynamically activate selected retrievers (and thus, dynamically select the corresponding loT security-related document databases) based on the user’s background and the query contexts. In various example embodiments, each self-querying retriever is dedicated to retrieving document chunks generated from a corresponding loT security-rel ted dataset, andmay comprise an loT security-related document database (or vector database, which may also be referred to as a vector store) generated or built from the corresponding loT security-related dataset and a database retrieving module configured to retrieve documents (or document chunks) from the loT security-related document database. In this regard, the database retrieving modules of the self-querying retrievers may collectively correspond to the document chunk retrieving module 208 described hereinbefore according to various embodiments of the present invention. In this way, LLM extracts loT threat intelligence from the relevant retrieved documents (or document chunks). This integration allows the loT security assistant to generate the latest reliable, relevant, and technical loT security answers tailored to both query contexts and users’ requirements. To the best of the inventors’ knowledge, it is the first time to make use of LLM and loT threat intelligence to offer loT security assistance.
[0052] Data processing toolkit According to various example embodiments, a toolkit (herein may also be referred to as DataKit) is designed to handle datasets in diverse formats. The toolkit first parses the raw data and converts the parsed contents into text. For example, a LLM may be utilized to i) select the appropriate content for page content and metadata of documents, and ii) process the multimodal contents, such as figures, tables, and codes, into text summarization. Additionally, according to various example embodiments, the Ragas library may be integrated or utilized to optimize the chunking strategy, including splitter methods and chunk sizes. Accordingly, various example embodiments advantageously provide an end-to- end data processing toolkit configured to process contents of the loT security-related datasets into document chunks.
[0053] Implementation and Evaluation. To implement the effectiveness of the CHATIoT, an example prototype of CHATIoT was developed implemented and five example use cases were defined. For each user case, the user’s background is specified in terms of knowledge, goals, and requirements to guide the LLM in generating answers that are not only reliable, relevant, and technical but also user-friendly. Extensive evaluations show the improvements achieved by CHATIoT. For example, CHATIoT’ s answers were compared with those generated directly from LLaMA3, LLaMA3.1, and GPT-4o in terms of reliability, relevance, technical depth, and user-friendliness. When evaluated with LLaMA3:70B, CHATIoT was found to improve all metrics by over 10% on average, particularly in relevance and technicality. Human evaluations also confirm that CHATIoT provides better answersloT Threat Intelligence
[0054] loT threat intelligence has emerged as a critical element in addressing security challenges, it refers to the collected and analyzed data related to threats, vulnerabilities, and attack patterns. In the context of loT, threat intelligence helps to monitor and analyze threats specific to ToT ecosystems, including device vulnerabilities, communication protocols, malware, tactics, techniques, and procedures (TTPs) used by attackers, and others. Recent research has highlighted the importance of loT threat intelligence in identifying zero-day vulnerabilities and mitigating threats using collaborative defence mechanisms. The dynamic and fast-evolving nature of loT systems, coupled with their diverse deployments, including both consumer and industrial settings, highlight the need for up-to-date and automated threat intelligence solutions capable of addressing its unique security challenges.Large Language Model (LLM)
[0055] LLM is pre-trained on billions of available datasets, enabling it to capture vast amounts of linguistic, factual, and contextual knowledge. Due to the extensive training data, LLM can be directly employed for many downstream tasks, ranging from language translation to complex reasoning tasks. However, LLM has certain limitations, particularly when it comes to handling highly specialized or constantly evolving domains. Since LLM relies solely on the data they were trained on, which may not always reflect the latest information, it may generate incomplete or outdated responses for specific queries. This is where Retrieval-Augmented Generation (RAG) comes into play.
[0056] RAG is a powerful approach that combines the generative capabilities of LLM with the latest external information retrieval. Instead of solely relying on the LLM’s internal knowledge, RAG augments the LLM by retrieving relevant, up-to-date information from external databases or documents at the time of generation. The process involves querying an external knowledge source, retrieving the most relevant documents, and using the retrieved information to guide or enhance the LLM’s response generation. By bridging the gap between static pre-trained knowledge and dynamic, real-world data, RAG enhances the performance of LLM on tasks that require both advanced language understanding, reasoning, and up-to-date information. This makes it particularly effective for specialized domains where access to the latest, domain-specific information is crucial.Design Considerations
[0057] According to various example embodiments, CHATIoT is designed to include one or more of the following features.
[0058] Coping with the rapid evolution of loT threat intelligence . In various example embodiments, CHATIoT is designed to stay up-to-date with the rapidly evolving loT threat intelligence. By combining the advanced capabilities of LLM with external information retrieval, it is enhanced by the latest loT threat information, such as new vulnerabilities, to deliver timely and technical insights.
[0059] Filtering relevant information. In various example embodiments, to prevent information overload, CHATIoT uses advanced retrieval and filtering techniques to prioritize important, highly relevant information while discarding irrelevant data. This ensures that CHATIoT generates actionable intelligence while avoiding overlooking critical information and filtering out unnecessary data.
[0060] Tailored for different user types. In various example embodiments, CHATIoT guides its generated answers based on the users’ roles (such as, but not limited to, Consumer, Security Analyst, and so on) and their corresponding backgrounds. This allows each kind of user to get insights or solutions that are relevant, understandable, and actionable based on their specific requirements and expertise level.System Architecture
[0061] FIG. 5 depicts a schematic block diagram of an example CHATIoT 500 according to various example embodiments of the present invention. The CHATIoT 500 comprises a data processing toolkit 550 (herein referred to as DataKit) (e.g., corresponding to the loT security- related dataset processing system 350 described hereinbefore according to various embodiments of the present invention) and a retrieval and generation system 510 (abbreviated as loT-RG) (e.g., corresponding to the system 200 for generating a response to an loT security- related query described hereinbefore according to various embodiments of the present invention).
[0062] Data Processing. A variety of technologies is integrated to build the end-to-end data processing toolkit (DataKit) 550. It is configured to convert the collected loT security-related datasets, which may be provided in various formats, into documents (or document chunks), which are suitable for retrieval and LLM processing. In various example embodiments, for each dataset, it extracts specific content for page content and metadata for documents (or documentchunks), and a corresponding retriever, as well as the corresponding loT security-related document databases (or vector database, which may also be referred to as a vector store) (e.g., included as part of the retriever), is subsequently constructed. DataKit 550 helps CHATIoT 500 keep up with the rapid evolution of loT threat intelligence.
[0063] Adaptive Retrieval According to various example embodiments, adaptive retrieval may be achieved through Selector 512 and Self-Querying mechanisms. When a user submits a query 0, the Selector 512 may invoke the equipped LLM 516 to generate a configuration (selection information) that determines which retrieversto activate (and thus, which loT security-related document database to selected). Once the configuration and query are passed to the retrieversloT-RG 510 executes the activated self-querying retrievers (e.g.,and 3?5) to get highly relevant documents from the corresponding selected loT security-related document databases while filtering out irrelevant ones. This approach prevents overload, ensuring that only relevant and important documents (or document chunks) will be utilized.
[0064] Guided Generation. According to various example embodiments, loT-RG 510 synthesizes the user’s background, the query 0 and the retrieved documents (or document chunks) in a response prompt (generated by a prompt template (UPT)) to instruct the LLM 516 to generate answers or response to the query Q. In this way, the answers are generated not only based on the advanced language understanding capabilities of the LLM 516 and the retrieved loT security -related information, but also aligned with the user’s background.
[0065] Accordingly, the CHATIoT 500 may operate in the following manner. Firstly, the DataKit 550 processes loT security-related datasets of various formats to documents (or document chunks) for building retrievers, including the corresponding loT security-related document databases, respectively. Then, the loT-RG 510 provides an interface for the user to submit the query (e.g., via a live user interface) and retrieves relevant documents (or document chunks) adaptively: i) Selector 512 first determines which retrievers should be activated, and ii) the activated self-querying retrievers retrieve similar documents (or document chunks) from the corresponding selected loT security-related document databases, respectively, and filter out unsatisfied documents (or document chunks) based on metadata. 3) LLM 516 synthesizes all information (the query Q, the user’s background and the retrieved documents) to generate the answer to the query Q. In this manner according to various example embodiments, CHATIoT 500 effectively provides reliable, relevant, and technical insights on loT security tailored to different users.
[0066] An example construction or architecture of the CHATIoT 500 will now be described in detail according to various example embodiments of the present invention, including the loT- RG 510 and the data processing toolkit 550.Construction of lo T-RG 510
[0067] As shown in FIG. 5, the loT-RG 510 is configured to perform adaptive retrieval 520 and guided generation 530.
[0068] Adaptive Retrieval 520. As described above, the loT-RG 510 comprises multiple retrievers, each retriever dedicated to retrieving documents (or document chunks) from a corresponding loT security-related document database (e.g., included in the retriever) generated from a corresponding and specific loT security-related dataset (accordingly, each of the plurality of loT security-related document databases is generated from a corresponding loT security-related dataset of the plurality of loT security-related datasets). In various example embodiments, the adaptive retrieval mechanism may be achieved in two aspects: i) selecting which retrievers should be activated and ii) trying to retrieve only relevant documents (or document chunks) from activated retrievers (and thus, from selected loT security-related document databases) while discarding irrelevant ones, even for the activated retrievers (e.g., corresponding to retrieving document chunks relating to the loT security-related query from selected one or more loT security-related document databases of a plurality of loT security- related document databases based on the loT security-related query).
[0069] Design of Selector 512. When user role submits a query Q a straightforward and static approach is to Q to all retrievers and gather retrieved documents from them. However, this method has the following drawbacks: i) Irrelevant Retrievers - Documents of some retrievers may not be relevant to the query. Retrieving them not only fails to improve the quality of the generated answer but may even negatively affect it; ii) Resource Costs - Retrieving unnecessary retrievers leads to increased resource costs, such as computational overhead, during both the retrieval and generation processes.
[0070] To address these issues, according to various example embodiments, an adaptive Selector 512 (e.g., LLM-based) is provided. As shown in FIG. 5, according to various example embodiments, the user role with the background, the query Q, and the descriptions of all retrievers are included in a Selector Prompt (SPT) (accordingly, the Selector Prompt is generated based on the user role with the background (or user information), the query 0, and the descriptions of all retrievers) and the SPT is inputted to LLM 516, which generates{SJp=1(e.g., corresponding to the loT security-related document database selection information described hereinbefore according to various embodiments of the present invention), where SL= True indicates that the retrievershould be activated (and thus, indicates that the corresponding loT security-related document database is selected), and St= False indicates that the retriever J 'it should not be activated (and thus, indicates that the corresponding loT security-related document database is not selected). By way of an illustrative example and without limitation, Table 1 shown in FIG. 6 present example selector configurations (loT security-related document database selection information)generated by LLaMA3:8B for a number of example queries submitted by different types or roles of users, where V indicates True (activated / selected) and X indicates False (activated / selected). Accordingly, in various example embodiments, the strong reasoning capability of the LLM 516 is advantageously utilized to generate the selector configurations for selecting suitable or relevant retrievers (and thus, selecting suitable or relevant loT security- related document databases) with respect to the query Q based on the user role with the background (or user information), the query O, and the descriptions of all retrievers.|0071| Self-Querying Retrievers. Various example embodiments found that the activated retrievers may return information (e.g., document chunks) that does not meet requirements, e.g., mismatch id and products. To address this problem, various example embodiments provide a self-querying technique to filter documents by metadata
[0072] FIG. 7 depicts a schematic block diagram of an example self-querying retriever700 according to various example embodiments of the present invention. For example, the retriever700 may be based on LangChain and Elastic. When Sj = True for a retriever tRj 700, the retriever700 retrieve documents (or document chunks) that are semantically similar to query (? and filtered by metadata.
[0073] As illustrated in FIG. 7, in various example embodiments, the retriever 700 comprises a database retrieving module comprising a query constructor 702 and a query translator 704. The retriever 700 may further comprises a vector store 708 formed or built from a corresponding loT security-related dataset. For example, when the user’s query Q is passed to the query constructor 702, an LLM of the retriever 700 may generate internal query language elements based on pre-defmed metadata field information and metadata examples (e g., the metadata examples show how to use metadata to generate filters (e.g., the vulns example shown in FIG. 8A)). The query translator 704 may then convert these elements into a structured query with appropriate filters. Accordingly, for each of the selected one or more loT security-related document databases, one or more structured queries for the selected loT security-related document database is generated based on the loT security-related query. Finally, the structured query and search parameters (e.g., the search parameters may be configured by developers, whereby the langchain provides many options for each parameter) are applied to the vector store 708 to retrieve documents (or document chunks). Accordingly, from each of the selected one or more loT security-related document databases, one or more document chunks relating to the loT security-related query are retrieved from the selected loT security-related document database based on the one or more structured queries and search parameters for the selected loT security-related document database. Accordingly, the vector store 708 stores documents in manageable document chunks to provide efficient processing and retrieval. The vector store 708 further stores embedding vector for each document chunk. For example, the embedding vector may be determined by applying an embedding model to the chunked text. The embedding vectors are useful for the retrieval process to facilitate in identifying the relevant document chunks in response to a query.
[0074] In various example embodiments, to specialize the self-querying retrieverfor loT security, the following steps are taken.
[0075] Metadata & Examples. Various example embodiments provide the metadata field information and examples from relevant loT security-related datasets. For instance, in the VARIoT vulnerabilities dataset, fields id and products are utilized as metadata. By way of examples for illustration purposes only, the corresponding metadata field info and examples are illustrated in FIG. 8A, and details for other example loT security-related datasets are illustrated in FIGs. 8B and 8C. In particular, FIG. 8A shows the metadata fields information and examples for the self-querying retriever corresponding to the VARIoT vulnerabilities dataset. FIGs. 8B and 8C show the metadata information and examples for the self-querying retrievers respectively corresponding to the VATIoT exploits, MITRE ATT&CK ICS, and CLS. This ensures the LLM of the retriever can gain the necessary loT security-specific knowledge to generate effective internal query language elements from the user’s query.
[0076] C 'reale Structured Queries. Based on the above customized internal query language elements, the query translator 704 can create specific structured queries for the retriever. By way of an example for illustration purposes only, FIG. 9 shows how to enable the retrieval of VARIoT vulnerabilities and CLS lists that are both semantically similar to the query and appropriately filtered by their respective metadata. In particular, FIG. 9 shows the generatedqueries and filters for VARIoT vulnerabilities and CLS list for an example submitted query “What are the security issues with DLink DCS-942 camera?” by a Security Analyst.
[0077] Guided Generation. After retrieval, one simple approach is feeding the retrieved documents and query to LLM to generate the answer. However, this simple approach may result in user-unfriendly outputs. For example, consumers often lack the expertise needed to fully comprehend highly technical content, making such answers unhelpful and not actionable for them.
[0078] To address this issue, various example embodiments incorporate user-specific backgrounds (or background specifications or information), including knowledge, goals, and requirements, into each user type’s user-friendly prompt template UPT (e g., corresponding to the response prompt described hereinbefore according to various embodiments of the present invention). This adjustment guides the LLM 516 in generating answers tailored to the user’s needs. By way of an example, an example background of the general consumer utilized to guide the CHATloT 500 to generate consumer-friendly outputs may be provided as follows: for knowledge, the consumer may not have formal technical training but are familiar with using loT devices for daily convenience such as smart home systems. The consumer has a basic understanding of device operation but may not be aware of the intricate security risks that exist; for goals, for example, the primary aim is to understand whether a device is secure and how to maintain or improve its security, ensure safety, security, and reliability of ToT devices within their homes or personal environments; for requirements, for example, the answers should be practical, easy to follow, and focused on actionable steps the general consumer can take. Therefore, the example background for the general consumer demonstrates how content can be simplified for easy understanding. As further examples, example background specifications of other user types, namely, Security Analyst, Technical Officer, Developer, and Trainer, utilized to guide the CHATloT 500 to generate answers, are provided below. It will be appreciated by a person skilled in the art that the present invention is not limited to the above-mentioned example user types, and background specifications for additional or other types of users may be defined as desired or as appropriate.
[0079] Example Background of Security; Analyst. Knowledge: Security Analyst is an expert in identifying vulnerabilities, analyzing threats, and ensuring loT devices are secure from cyber threats. Security Analyst possesses in-depth technical knowledge of security protocols, vulnerabilities, and exploits, and are proficient in interpreting complex security data; Goals: the primary aim is to conduct in-depth analyses of ToT security threats, vulnerabilities, and exploits,contributing to the development of secure loT systems, and provide deep insights into potential attack vectors, technical analysis, and mitigation strategies; Requirements: Security Analyst requires detailed information about the vulnerabilities, exploits, and technical configurations of loT devices.
[0080] Example Background of Technical Officer. Knowledge: Technical Officer is familiar with security patch management, ensuring devices adhere to organizational security standards, and handling technical troubleshooting; Goals: Technical Officer is responsible for overseeing the implementation and maintenance of secure loT systems within an organization, applying security patches, enforcing security policies, and troubleshooting security issues; and Requirements: Technical Officers focus is on implementing security measures within the organization’ s infrastructure. Practical steps are needed to deploy security updates and verify compliance with security standards.
[0081] Example Background of Developer. Knowledge: Developer works on the technical design and architecture of loT devices, with a focus on incorporating security into product design. Developer has a deep understanding of device security, encryption protocols, and compliance with security regulations; Goals: Developer is responsible for ensuring loT products meet industry security standards and are resilient against known threats, and designing and developing secure loT devices, and Requirements: Developer needs insights into current vulnerabilities, designs best practices, and how to avoid common security pitfalls in future product iterations.
[0082] Example Background of Trainer. Knowledge: Trainer creates educational material or conducts training sessions to teach loT security to a broader audience, including technical and non-technical participants. Trainer understands both technical and pedagogical aspects of loT security and can explain complex concepts in a simplified manner; Goals: Trainer aims to guide others in the best practices for loT security, helping to raise awareness and improve security practices across different user groups; and Requirements: Trainer needs information that can be used in a training environment, with clear examples, case studies, and simplified explanations for different levels of learners.
[0083] As an example implementation, an example workflow or process of loT-RG 510 is implemented in Algorithm 1 below.Data Processing Toolkit 550
[0084] The raw data of the collected loT security-related datasets may be of different formats, for example, PDF, Word, XLS, JSON and XML However, these data or file formats are not suitable for RAG processing directly. To address this and enable the CHATIoT 500 to handle loT security-related datasets of various or diverse data formats, various example embodiments develop an loT security-related dataset processing system 550, herein referred to as an end-to-end data processing toolkit or DataKit. FIG. 10 depicts a schematic block diagram of an example DataKit 550 according to various example embodiments of the present invention.As shown in FIG. 10, DataKit may be configured to perform the following operations.
[0085] 1) Parse Raw Data. An initial step involves parsing raw data (e g., in various formats) into distinct content elements such as text, tables, figures, and code. Leveraging existing tools like the unstructured library helps in extracting textual content from threat reports, while the JSON library is useful for handling structured data from sources like VARIoT and MITRE ATT&CK.
[0086] 2) Convert Multi-Modal Elements to Text. Once parsed, any multi-modal elements(e.g., tables, figures, code) are converted into text descriptions for further processing. For example, LLMs like LLaVA (for images), LLaMA3 :8B (for tables), and CodeLlama (for code)may be employed to generate these descriptions. Therefore, all parsed multi-model elements are represented in text form. Accordingly, for each of multiple security-related document databases, non-textual contents in the corresponding loT security-related dataset are converted to textual contents in the loT security-related document database. This modular approach allows for easy integration of other LLMs to handle different types of multi-modal content.
[0087] 3) Field Selection for Page Content & Metadata. In structured formats like JSON, content is often stored across multiple fields. Instead of using all fields, according to various example embodiments, relevant fields (e g., most relevant fields) for retrievers are identified or selected and utilized. In various example embodiments, the relevant fields for retrievers are identified or selected by sampling example items from each field and using an LLM (e.g., LLaMA3:8B) to intelligently select fields that best represent the document’s page content and metadata, and an example prompt (field selection prompt) is shown in FIG. 10. Accordingly, in various example embodiments, for each of the plurality of loT security-related document databases, the document chunks therein each has one or more content fields and one or more metadata fields defined for the document chunks generated from the corresponding loT security-related dataset. Example selected fields for page content and metadata of various example datasets will be described later below according to various example embodiments of the present invention.
[0088] 4) Optimized Chunking Strategy. Subsequently, an appropriate chunking strategy is selected, including the chunking size, overlap, and splitting method. For example, the Ragas library may be used to optimize this process. As an example implementation, an example process for optimized chunking is implemented in Algorithm 2 below.
[0089] Accordingly, as shown in FIG. 10, after collecting various loT security-related datasets, the DataKit 550 may first parse the raw data to get elements of multi-modal (step 1 shown in FIG. 10), and then converts the multi-modal elements into text by utilizing LLM (step 2 shown in FIG. 10). Subsequently, the DataKit 550 may use a LLM to select fields for the page content and metadata of documents (step 3 shown in FIG. 10), and optimizes the chunking strategy (step 4 shown in FIG. 10).
[0090] After obtaining the optimized chunking strategy, the documents are split into small chunks (document chunks) and a language model (e g., all-MiniLM model) is used for chunked text embedding. The document chunks comprise chunked text, embedding (or embedding vectors), and metadata. For example, an embedding model is applied to chunked texts to generate embedding vectors, which are used in the retrieval process, and metadata is produced for each document chunk to enhance the retrieval process by providing additional context and filtering capabilities. This approach ensures that the loT security-related (e g., vulnerability and threat) datasets are processed efficiently and ready for further analysis or use in the LLM 516. Accordingly, various example embodiments advantageously develop an end-to-end data processing toolkit for allowing loT security-related datasets of diverse data formats (e.g., multimodal contents) to be utilized and maximizing the utilization of loT security-related information / contents (e g., multi-modal contents) contained in the loT security-related datasets, to enhance usefulness in practical applications.
[0091] In various example embodiments, the loT security-related document databases may each obtain or download latest document chunks derived from the latest loT security-related information / contents contained in the corresponding loT security-related dataset periodically, when latest loT security-related information / contents are available, or in real-time to keep up- to-date.Use Cases Specialization
[0092] By way of examples only for illustration purposes, five example specialized use cases of CHATIoT 500 are defined. Each use case may be defined by four example fundamental properties: user role or type, background or goal, actions, and example query, within the loT security domain FIG. 1 1 presents the example detailed specifications for the four example properties of the use case, whereby for each property, its description and type are provided. For example, the user roles or types include Consumer, Security Analyst, Technical Officer, Developer, and Trainer. The background property has been described hereinbefore. The actions and example query properties will now be described below in further detail.
[0093] FIG. 12 shows Table 2 presenting key actions for each of the five example user roles, such as assessing the security of loT devices, deploying security patches, or developing training programs on loT security. Additionally, it provides example queries for each user role, demonstrating how CHATIoT 500 can be utilized to address the unique needs of various users. This structured approach ensures that CHATIoT 500 caters to a diverse range of users, offering tailored assistance and enhancing loT security management across different scenarios. It will be understood by a person skilled in the art that the present invention is not limited to the above- mentioned five example use cases or user roles / types, and that their definitions / specifications may be defined or varied as desired or as appropriate. For example, the use cases can be easily changed or extended by defining new user roles, specifying the background (including knowledge, goals, and requirements), and outlining actions. Example queries can also be added as desired or as appropriate to further clarify the context and functionality of each user role.Experimental Evaluation
[0094] To demonstrate the effectiveness of the CHATIoT 500, an example prototype of CHATIoT 500 was implemented and extensive experiments and evaluations were conducted. For example, various example embodiments seek to address the following questions.• How does DataKit 550 extract appropriate field selection from loT security-related datasets and convert them into well-structured documents for retrieval and LLM analysis? What are the optimal chunking strategies for each kind of document?• What are certain advantages of CHATIoT 500? For example, can CHATIoT 500 effectively generalize and improve the capabilities of the most advanced LLMs available in processing loT security issues?• Can CHATIoT 500 be a practical useful loT assistant over using LLM alone? Howabout the feedback from real-world human evaluation?Experimental Setup
[0095] Testbed. CHATIoT 500 as implemented in Python 3.10.13, utilizing large language models LLaMA3:8B & 70B and LLaMA3.1 :8B & 70B provided by Groq, GPT-4o-mini and 4o provided by OpenAI. All these LLMs are utilized by calling their APIs. For building the vector store, Elasticsearch 8.13.2 was employed, running on Docker Desktop 4.29.0. All components were integrated using the LangChain library (version 0.2.5). The WebApp was developed using Streamlit (version 1.33.0). Experiments were conducted on a MacBook Pro equipped with an Apple M3 Pro CPU (11 cores) and 18 GB of RAM, running macOS 14 6.1 with the Darwin 23.6.0 kernel.
[0096] loT Security-Related Data Sources. Five example kinds or types of loT security- related datasets (of different security domains) were collected from the public Internet: i. VARloT vulnerabilities: This dataset catalogs known vulnerabilities in various loT devices, offering detailed information about the potential risks associated with each vulnerability. ii. VARloT exploits: This dataset contains exploits targeting loT devices, providing insights into the techniques and methods attackers use to compromise these systems. iii. MITRE ATT&CK ICS TTPs: This dataset outlines the tactics, techniques, and procedures (TTPs) employed by adversaries specifically in industrial control systems (ICS), which often include loT-related TTPs as well. iv. Threat reports: 17 public threat reports were collected from VXUG about emerging threats and vulnerabilities, offering analysis and recommendations for mitigating potential risks. v. Cybersecurity labelling schemes: These schemes provide information on the security posture of various loT products, helping consumers and organizations assess the security standards and certifications achieved by specific devices.
[0097] These loT security-related datasets provide comprehensive insights into the current landscape of loT threats, which facilitate the enhancement of the CHATIoT’ s capabilities.Fields Selection A Chunking Strategy
[0098] The experimental evaluation for field selection and chunking strategy optimization will now be discussed.
[0099] FIG. 13 shows Table 3 presenting the Large Language Models (LLMs)-based field selection for page content and metadata, where • / denotes for page content Q denotes metadata, and X indicates unused fields.
[0100] Fields for Page Content <C Metadata. As described hereinbefore, according to various example embodiments, for each loT security-related dataset, the fields for page content and metadata for documents (or document chunks) are determined before building selfquerying retrievers. For each loT security-related dataset, for example, 3 items may be sampled, the fields’ names may then be listed, and the LLM may be instructed to select the suitable fields. The results are discussed below.
[0101] Three LLMs: LLaMA3:8B, LLaMA3.1 :70B, and GPT4-O, were used to select fields for VARIoT vulnerabilities, exploits, and MITRE ATT&CK ICS. The experimental results are summarized in Table 3 in FIG. 13. While there are slight variations in their selections, it can be seen that there is consensus on the key decisions: For instance, in the case of the VARIoT vulnerabilities, all LLMs select title and description for page content, and id and products for metadata. Similar selections can be observed for the VARIoT exploits and MITRE ATT&CK ICS datasets.
[0102] For threat reports, which are typically unstructured, the report’s content was used as page content and the title was used as metadata (Note that self-querying retrieval is not used (i.e., not using filters) for threat reports. Instead, because threat report is unstructured pdf / text, in various example embodiments, a similarity search (e.g., mar) may be utilized to query the texts in the vector store and select the top-k document chunks). The CLS schemes consist solely of metadata with no descriptive content, so the page content may be left blank and the metadata may be utilized for self-querying retrievers.
[0103] By way of examples only for illustration purposes, the specific fields selected for page content and metadata of each of the five example loT security-related datasets used in the CHATIoT 500 are presented in Table 4 shown in FIG. 14.
[0104] Chunking Evaluation. To optimize the chunking strategy for documents’ page_content, the Ragas library was utilized in conjunction with all-MiniLM14 (for embedding) and LLaMA3:8B (for evaluation) to search the most suitable chunking size, overlap, and splitting method for each loT security -related dataset. Content precision and recall were used as the key metrics, whereby:• Precision measures whether all relevant items retrieved by the model are ranked higher than the irrelevant items;• Recall measures how much of the relevant content is retrieved based on the annotated answers and the retrieved context.
[0105] Both precision and recall are evaluated within the range [0, 1], where a higher score indicates better performance. As the ToT security-related datasets contain a huge number of samples, for practical efficiency, a subset of 1,000 samples was selected from each dataset except threat reports, a testset of 50 items was generated, and conducted evaluations based on the subset and testset. This may not result in an optimal chunking size, overlap, and splitter method, but is sufficient to obtain a reasonable and useful chunking strategy for practical applications. FIG. 15 shows Table 5 presenting context (precision, recall) of different chunking configurations for VARIoT vulnerabilities, exploits, MITRE, ATT&CK and threat reports. The best (precision, recall) in the experimental settings are marked in bold. As shown in Table 5 in FIG. 15, the following commonly used configurations were tested:• chunk sizes: {500, 1000, 1500, 2000};• overlaps: {50, 100, 150, 200};• splitters: Character, RecursiveCharacter, and TokenText
[0106] The objective is to achieve high precision and recall simultaneously, ensuring that the system retrieves as many relevant documents as possible while minimizing irrelevant content. From the experimental results in Table 5 shown in FIG. 15, it is easy to see that using the RecursiveCharacter splitter with a chunk size of 500 and an overlap of 100 is the most effective strategy for the VARIoT vulnerabilities, offering the best or an optimal trade-off between precision and recall. Similarly, suitable chunking strategies for VARIoT exploits, MITRE ATT&CK ICS, and threat reports can be chosen. FIG. 16 shows Table 6 presenting optimized chunking strategy for VARIoT vulnerability, exploits, ICS, and threat reports, whereby RecurChar is an abbreviation for RecursiveCharacter.LLMs-based Evaluation of Outputs
[0107] As there is no public Question-Answer dataset about loT security and threat intelligence, 50 common loT security-related questions (10 questions for each kind of user) were synthesized. To evaluate improvements in the CHATIoT 500 according to various example embodiments of the present invention, CHATIoT’ s outputs are compared with the answers generated by underlying LLM alone (denoted as LLM-A), which is not equipped with the above-mentioned loT security -related datasets. Another LLM was utilized as the evaluatorand the quality of outputs was measured by four metrics: Reliability, Relevance, Technical, and Friendliness as follows:• Reliability: the trustworthiness and reliability of each answer, ensuring it is plausible and aligns with known ToT best practices and standards.• Relevance: Assess how well the answer addresses the specific question and meets the user’s needs, considering their role and context in the loT ecosystem.• Technical: Judge the appropriateness and precision of technical language, including loT research, standards, protocols, and relevant technical aspects. Ensure that the answer demonstrates a solid understanding of ToT technologies.• Friendliness: Determine how easy the answer is to comprehend, focusing on clarity for the user’s role, and how well the answer provides actionable steps or solutions tailored to the user’s ToT security needs.
[0108] All scores are in [0, 5], where 5 is the highest. For each question, the answers generated by CHATIoT 500 and corresponding LLM-A are inputted into the evaluator simultaneously. This approach ensures that both answers for each question are evaluated within the same internal state of the evaluator. Doing so aims to reduce the impact of LLM randomness as much as possible and enable a fair comparison between CHATIoT 500 and LLM-A. The prompt template for LMM-based evaluation of outputs is shown in FIG. 17.
[0109] Five versions of CHATIoT 500 were developed using LLaMA3 :8B, LLaMA3.1 :8B, LLaMA3.1 :70B, GPT-4omini, and GPT-4o, and LLaMA3:70B was employed to evaluate them. FIG. 18A shows Table 7 presenting the experimental results for comparison of CHATIoT 500 with LLM-A for moderate LLMs LLaMA3:8B, LLaMA3.1 :8B, and GPT-4o-mini; and FIG. 18B shows Table 8 presenting the experiment results for comparison of CHATIoT 500 with LLM-A for more advanced LLMs LLaMA3.1 :70B and GPT-4o. LLaMA3:70B was used as the evaluator for all experiments. The improved scores over LLM-A are also presented in the Tables 7 and 8. From these results, several key observations can be made as described below.
[0110] CHATIoT 500 significantly enhances the moderate LLM’s performance in the loT security domain. As shown in Table 7 in FIG. 18A, CHATIoT 500 achieves higher scores across most metrics for example use cases: Consumer, Security Analyst, Technical Officer, and Developer. This is expected, as CHATIoT 500 integrates domain-specific loT security knowledge, e.g., vulnerabilities and exploits, and tailors responses to be more user-friendly andrelevant. As illustrated in Table 8 in FIG. 18B, CHATIoT 500 can also improve the advanced LLMs’ capabilities in loT problems.
[0111] However, CHATIoT 500 may not always outperform the baseline LLMs. Taking the use case Trainer, when using LLaMA series models, CHATIoT 500 even performs slightly worse; when using GPT-4o-mini and GPT-4o, the improvements achieved by CHATIoT 500 are much less than that for the other use cases. This is likely due to the external data introduced in CHATIoT 500 focusing mainly on vulnerabilities, exploits, and TTPs, while lacking sufficient information on course training materials. As a result, CHATIoT 500 excels at producing technical and security-centric content but may overlook broader aspects like training programs.
[0112] The above analysis also highlights the importance of incorporating external knowledge to bolster LLMs in specialized domains. Fortunately, additional information, such as training materials, can easily be integrated into CHATIoT 500 using the DataKit toolkit 550 according to various example embodiments of the present invention.Analysis of Human Evaluation
[0113] FIG. 19 presents the results of human evaluations comparing CHATIoT 500 with GPT-4o (LLM-A), noting that CHATIoT 500 is built on top of GPT-4o in this experiment. Four Q&A pairs were selected for each user type and displayed as a survey. Human evaluators were asked to select the answer they find more suitable. To aid in their decisions, evaluation metrics were provided, but did not require them to score each metric. This approach was intended to streamline the evaluation process, making it more akin to real-world scenarios where users prioritize ease of decision-making.
[0114] From the experimental results, it is clear that CHATIoT 500 consistently outperforms GPT-4o across all use cases. This aligns with expectations, as CHATIoT 500 integrates additional loT-specific intelligence into the LLM 516. Notably, CHATIoT 500 demonstrates the greatest improvement for Security Analyst and the least for Technical Officer. The former result aligns with Table 8 shown in FIG. 18B, where the comparison for Security Analyst shows the most significant difference. Additionally, a use case comparison for Security Analyst was given between GPT-4o and GPT-4o-based CHATIoT 500. The formal analysis will be provided below. However, while Table 8 suggests that the least improvement is for Trainer, the human evaluation indicates Technical Officer experiences the smallest gains. This discrepancy can be explained by: i) The improvements for Technical Officer, though better thanthose for Trainer, particularly in top metric scores, may not be as easily discernible to humans as other use cases, making it harder for them to identify notable differences; ii) The Q&A tasks for Technical Officer are generally more complex and technical than those for Trainer, making it easier for users to select a better answer in Trainer case even when improvements are less significant Accordingly, CHATIoT 500 was found to consistently provide better results over GPT-4o alone in the loT security-related domain, demonstrating its enhanced capability to address loT-specific challenges.
[0115] For the above user case study for Security Analyst, the generated outputs of CHATIoT 500 were compared to that of GPT-4o to show the benefits from external loT security and threat information and knowledge. The generated answers of GPT-4o for the inputs (User Role = Security Analyst, Query= “Investigate the vulnerabilities and exploits associated with the TP-Link Routers) are presented belowTP-Link routers, like many other networking devices, have been subject to various vulnerabilities and exploits over the years. These vulnerabilities can range from simple misconfigurations to more complex software flaws that can be exploited by attackers. Here are some notable vulnerabilities and exploits associated with TP-Link routers:1. Default Credentials and Weak PasswordsMany TP-Link routers come with default administrative credentials (e.g., admin / admin). If these credentials are not changed, attackers can easily gain access to the router’s administrative interface. Weak passwords can also be brute-forced.2. Remote Code Execution (RCE)Several TP-Link router models have been found to have vulnerabilities that allow remote code execution. For example:• CVE-2019-7406: A vulnerability in the TP-Link WR940N allows remote attackers to execute arbitrary code via a crafted HTTP request.• CVE-2020-12109: A buffer overflow vulnerability in the TP-Link Archer C5 allows remote attackers to execute arbitrary code.3. Cross-Site Scripting (XSS)Cross-Site Scripting vulnerabilities can allow attackers to inject malicious scripts into the web interface of the router. For example:• CVE-2019-7405: An XSS vulnerability in the TP-Link WR940N allows remote attackers to inject arbitrary web script or HTML.4. Cross-Site Request Forgery (CSRF)CSRF vulnerabilities can allow attackers to perform actions on behalf of the authenticated user without their consent. For example:• CVE-2018-11714: A CSRF vulnerability in the TP-Link Archer C1200 allows remote attackers to change the router’s settings.5. Command InjectionCommand injection vulnerabilities allow attackers to execute arbitrary commands on the router’s operating system. For example:• CVE-2017-13772: A command injection vulnerability in the TP-Link WR1 043ND allows remote attackers to execute arbitrary commands via the ping feature.6. Information DisclosureInformation disclosure vulnerabilities can expose sensitive information about the router or the network. For example:• CVE-2019-7404: An information disclosure vulnerability in the TP-Link WR940N allows remote attackers to obtain sensitive information via a crafted HTTP request.7. Firmware VulnerabilitiesOutdated firmware can contain numerous vulnerabilities that have been patched in later versions. It’s crucial to keep the router’s firmware up to date to mitigate these risks.8. MisconfigurationsImproper configurations, such as enabling remote management without proper security measures, can expose the router to external attacks.Mitigation StrategiesTo protect against these vulnerabilities, consider the following best practices:• Change Default Credentials: Always change the default username and password to something strong and unique.• Update Firmware: Regularly check for and apply firmware updates from TP- Link.• Disable Unnecessary Services: Turn off services like remote management if they are not needed.• Use Strong Encryption: Ensure that Wi-Fi networks use strong encryption methods like WPA3.• Network Segmentation: Segment your network to limit the impact of a compromised device.• Monitor Network Traffic: Use network monitoring tools to detect unusual activity.ConclusionTP-Link routers, like any other network devices, can be vulnerable to various types of attacks if not properly secured. Regular updates, strong passwords, and proper configuration can significantly reduce the risk of exploitation. Always stay informed about the latest vulnerabilities and apply necessary patches and mitigations promptly.
[0116] The generated answers of GPT-4o-based CHATIoT 500 for the inputs (User Role = Security Analyst, Query = “Investigate the vulnerabilities and exploits associated with the TP- Link Routers) are presented below.Technical Analysis of Vulnerabilities and Exploits in TP-Link RoutersOverviewTP-Link routers have been identified with multiple security vulnerabilities that pose significant risks to users. This analysis focuses on the vulnerabilities, associated risks, and actionable security measures to mitigate these threats.Identified Vulnerabilities1) Denial of Service (DoS) Vulnerability• Affected Models: TP-Link Archer C2, Archer C20i• Description: This vulnerability allows attackers to disrupt the normal functioning of the router, rendering it unavailable to legitimate users.• Risk: High. A successful DoS attack can lead to network downtime, affecting all connected devices.• Mitigation: Implement rate limiting and anomaly detection to identify and block malicious traffic. Regularly update firmware to patch known vulnerabilities.2) Security Bypass Vulnerability• Affected Models: TP-Link Archer C2, Archer C20i• Description: This vulnerability allows attackers to bypass security mechanisms, gaining unauthorized access to the router.• Risk: High. Unauthorized access can lead to data breaches and further exploitation.• Mitigation: Enable strong authentication mechanisms, such as multi-factor authentication (MFA). Regularly review and update access control lists (ACLs).3) Cross-Site Request Forgery (CSRF)• Affected Model s : TP-Link TL-WR 1043N• Description: This vulnerability allows attackers to perform unauthorized actions on behalf of authenticated users by tricking them into clicking malicious links.• Risk: Medium. CSRF can lead to unauthorized configuration changes, such as altering admin passwords.• Mitigation: Implement anti-CSRF tokens in web forms. Educate users about the risks of clicking on unknown links.4) Path Traversal Vulnerability• Affected Model s : TP-Link TL-R600 VPN• Description: This vulnerability allows attackers to access sensitive system files by manipulating URL paths.• Risk: High. Unauthorized access to system files can lead to information disclosure and further exploitation• Mitigation: Validate and sanitize user inputs to prevent path traversal attacks. Regularly update firmware to patch known vulnerabilities.5) Command Injection Vulnerability• Affected Models: TP-Link Archer C2, Archer C20i• Description: This vulnerability allows attackers to execute arbitrary commands on the router.• Risk: Critical. Command injection can lead to complete system compromise.• Mitigation: Implement input validation and sanitization. Use security mechanisms like Web Application Firewalls (WAF) to detect and block malicious inputs.Exploits Analysis1) Proof-of-Concept (PoC) Exploits• Description: Researchers have developed PoCs to demonstrate the feasibility of exploiting these vulnerabilities.• Risk: High. PoCs can be used by malicious actors to develop real-world attacks.• Mitigation: Monitor security advisories and apply patches promptly. Conduct regular security assessments to identify and mitigate potential vulnerabilities.2) Unauthenticated Reflected XSS• Affected Models: TP-Link TD-W895 IND• Description: This vulnerability allows attackers to inject malicious scripts into the browser of an unauthenticated user.• Risk: Medium. XSS can lead to session hijacking and unauthorized actions.• Mitigation: Implement Content Security Policy (CSP) and sanitize user inputs. Regularly update firmware to patch known vulnerabilities.3) Authenticated Reflected XSS• Affected Models: TP-Link TD-W895 IND• Description: This vulnerability allows attackers to inject malicious scripts into the browser of an authenticated user.• Risk: Medium. XSS can lead to session hijacking and unauthorized actions.• Mitigation: Implement CSP and sanitize user inputs Educate users about the risks of clicking on unknown links.4) Reset Admin Password CSRF• Affected Models: TP-Link TD-W895 IND• Description: This vulnerability allows attackers to reset the admin password by tricking the administrator into clicking a malicious link.• Risk: High. Unauthorized access to the admin account can lead to complete system compromise.• Mitigation: Implement anti-CSRF tokens and educate users about the risks of clicking on unknown links.Actionable Security Measures1) Firmware Updates Regularly update router firmware to patch known vulnerabilities.Enable automatic updates if available.2) Strong Authentication Use strong, unique passwords for admin accounts. Enable multi-factor authentication (MFA) where possible.3) Network Segmentation Segment the network to limit the impact of a compromised device. Use VLANs to isolate sensitive devices.4) Monitoring and Logging Implement network monitoring to detect and respond to suspicious activities. Enable logging to maintain an audit trail of administrative actions.5) User Education Educate users about the risks of phishing and social engineering attacks. Encourage users to report suspicious activities.DisclaimerThe recommendations provided in this report are based on the information available at the time of writing. We are not legally responsible for any actions taken based on these recommendations. Always consult with a professional security expert before implementing any security measures.
[0117] Key advantages of CHATIoT 500 over GPT-4o are summarized below.1001181 Reliability: GPT-4o provides general vulnerabilities and examples but lacks specific details, such as CVE references for many vulnerabilities. It misses certain critical vulnerabilities. On the other hand, CHATIoT 500 offers precise information about vulnerabilities, including affected models and specific risks. It references real-world CVEs, enhancing factual reliability.
[0119] Relevance: GPT-4o covers abroad range of vulnerabilities but does not specifically target the needs of a security analyst, making it less relevant for professionals. CHATIOT’ s answer is tailored for a security analyst, focusing on vulnerabilities impacting network security. It provides detailed exploit analysis and mitigation strategies, making it highly relevant.
[0120] Technical: GPT-4o is a general LLM, so it lacks comprehensive technical details about each vulnerability and does not explain them in depth or provide mitigation strategies. CHATIoT 500 offers detailed technical specifics for each vulnerability, including descriptions, risks, and mitigation strategies The structured analysis of exploits is suitable for technical audiences.
[0121] Friendliness: GPT-4o uses straightforward language, making it easy for a general audience to understand, but lacks engagement for detailed insights. Our approach maintains a professional tone while being informative. The structured layout enhances readability, making it user-friendly for professionals seeking specific information
[0122] In a nutshell, CHATIoT 500 outperforms using only GPT-4o to process loT security-related queries / questions in reliability, relevance, and technical depth, making it more suitable for a security analyst’s needs. While the GPT-4o’s answer is friendly and accessible, it lacks the detailed information and focus required for professionals in loT security.
[0123] Internet of Things (loT) has seen rapid advancements in recent years, becoming an integral part of various domains, such as smart industries and homes, and serving as a key enabler in modern society However, despite its growth, loT continues to face numerous security challenges, prompting significant research efforts aimed at improving loT security. With the rise of artificial intelligence (Al), machine learning (ML) and deep learning (DL)-based approaches have become increasingly popular in designing defence mechanisms for loT devices, including malicious traffic classification, malware detection, vulnerability discovery, and others. On the other hand, various example embodiments integrate loT threat intelligence of various data sources (loT security-related datasets) into CHATIoT 500, which is able to assist multiple kinds of users. Furthermore, in various example embodiments, an end-to-end toolkit 550 to process data in various formats, not limited to PDF. By integrating LLMs with loT- specific threat intelligence, these models can be guided to meet the unique challenges posed by the loT ecosystem. Moreover, the continuous advancements in the LLM community, combined with increasingly accessible loT datasets, are likely to further drive the adoption of LLMs in loT -related research and practical applications.
[0124] Accordingly, CHATIoT 500, an LLM-based loT security assistant, is provided according to various example embodiments of the present invention, and extensive evaluations were conducted on several common use cases. In particular, the advanced language understanding and reasoning capabilities of LLM and loT security and threat information were leveraged to provide loT security assistance, and an easy-to-use data processing toolkit 550 was developed. With the LLM-based generation system 510 and developed toolkit 550, CHATIoT 500 is easily scalable to integrate different kinds of loT threat intelligence from multiple loT security -related datasets or data sources. Accordingly, CHATIoT 500 incorporates LLM with datasets with loT security -related datasets, such as security vulnerabilities, exploits, threat reports, and cybersecurity labels associated with devices connected with the Internet to provide a knowledgebase and up-to-date threat intelligence to end users. In various example embodiments, an automated library may be integrated within CHATIoT 500 that may continuously update and ingest the latest or real-time loT threat intelligence, significantly reducing the manual efforts involved to gather loT security-related datasets from various sources. In various example embodiments, the LLM 516 may be re-trained or fine-tuned specifically in the loT security domain. Although a general -purpose LLM 516 may be utilized and enhanced by external data sources, re-training or fine-tuning allows the LLM 516 to more deeply understand the nuances and technical challenges unique to loT security. Such an approach may be integrated into CHATIoT 500 to provide even more reliable, technical, and actionable insights, pushing the boundaries of current loT security solutions. Together, these enhancements can make CHATIoT 500 not only even more effective in processing and presenting loT security and threat intelligence but also more capable of adapting to the fastevolving landscape of cybersecurity threats.
[0125] While embodiments of the invention have been particularly shown and described with reference to specific embodiments, it should be understood by those skilled in the art that various changes in form and detail may be made therein without departing from the scope of the invention as defined by the appended claims. The scope of the invention is thus indicated by the appended claims and all changes which come within the meaning and range of equivalency of the claims are therefore intended to be embraced.
Claims
CLAIMS1. A method of generating a response to an Internet of Things (loT) security -related query using a language model, the method comprising: receiving an ToT security-related query regarding security of one or more loT devices; retrieving document chunks relating to the loT security-related query from selected one or more loT security-related document databases of a plurality of loT security-related document databases based on the loT security -related query; generating a response prompt based on the loT security-related query and the document chunks retrieved; and generating a response to the loT security-related query using the language model based on the response prompt, wherein the plurality of loT security-related document databases is generated from a plurality of loT security-related datasets, respectively.
2. The method according to claim 1, wherein each of the plurality of loT security-related document databases is generated from a corresponding loT security-related dataset of the plurality of loT security -related datasets.
3. The method according to claim 2, wherein each of the plurality of loT security-related document databases is a vector database comprising document chunks generated from the corresponding loT security-related dataset, each document chunk comprising chunked texts, embedding vectors, and metadata.
4. The method according to claim 3, wherein for each of the plurality of loT security- related document databases, the document chunks therein each has one or more content fields and one or more metadata fields defined for the document chunks generated from the corresponding loT security-related dataset.
5. The method according to any one of claims 2 to 4, wherein for each of one or more loT security-related document databases of the plurality of ToT security -related document databases, non-textual contents in the corresponding ToT security-related dataset are converted to textual contents in the ToT security-related document database6. The method according to any one of claims 1 to 5, wherein the plurality of loT security- related datasets comprises loT security-related datasets of different security domains.
7. The method according to any one of claims 1 to 6, further comprising selecting one or more loT security-related document databases from the plurality of loT security-related document databases based on the loT security-related query and descriptions of the plurality of loT security-related document databases to obtain the selected one or more loT security-related document databases.
8. The method according to claim 7, wherein said selecting the one or more loT security- related document databases comprises: generating a selector prompt based on the loT security-related query and the descriptions of the plurality of loT security-related document databases; generating loT security-related document database selection information using the language model based on the selector prompt; and selecting the one or more loT security-related document databases based on the loT security-related document database selection information.
9. The method according to claim 8, wherein the selector prompt is generated further based on user information corresponding to a user associated with the loT security-related query.
10. The method according to claim 9, wherein the user information comprises user type information and associated background information.
11. The method according to claim 9 or 10, wherein the user information of the user is selected from a predefined set of user information for different user types.
12. The method according to any one of claims 1 to 11, wherein said retrieving document chunks relating to the loT security-related query comprises: generating, for each of the selected one or more loT security-related document databases, one or more structured queries for the selected loT security-related document database based on the loT security-related query; andretrieving, from each of the selected one or more loT security-related document databases, one or more document chunks relating to the loT security-related query from the selected loT security-related document database based on the one or more structured queries and search parameters for the selected loT security -related document database.
13. The method according to any one of claims 1 to 12, wherein for each of the plurality of loT security-related document databases, the loT security-related document database is updated based on new loT security-related data in the corresponding loT security-related dataset.
14. The method according to any one of claims 1 to 13, wherein the language model is a large language model.
15. A system for generating a response to an Internet of Things (loT) security-related query using a language model, the system comprising: at least one memory; and at least one processor communicatively coupled to the at least one memory and configured to: receive an loT security-related query regarding security of one or more loT devices; retrieve document chunks relating to the loT security-related query from selected one or more loT security-related document databases of a plurality of loT security-related document databases based on the loT security -related query; generate a response prompt based on the loT security-related query and the document chunks retrieved; and generate a response to the loT security-related query using the language model based on the response prompt, wherein the plurality of loT security-related document databases is generated from a plurality of loT security -related datasets, respectively.
16. The system according to claim 15, wherein each of the plurality of loT security-related document databases is generated from a corresponding loT security-related dataset of the plurality of loT security -related datasets.
17. The system according to claim 16, wherein each of the plurality of loT security-related document databases is a vector database comprising document chunks generated from the corresponding loT security-related dataset, each document chunk comprising chunked texts, embedding vectors, and metadata.
18. The system according to claim 17, wherein for each of the plurality of loT security- related document databases, the document chunks therein each has one or more content fields and one or more metadata fields defined for the document chunks generated from the corresponding loT security-related dataset.
19. The system according to any one of claims 16 to 18, wherein for each of one or more loT security-related document databases of the plurality of loT security-related document databases, non-textual contents in the corresponding loT security-related dataset are converted to textual contents in the loT security-related document database.
20. The system according to any one of claims 15 to 19, wherein the plurality of loT security-related datasets comprises loT security-related datasets of different security domains.21 . The system according to any one of claims 15 to 20, wherein the at least one processor is further configured to select one or more loT security-related document databases from the plurality of loT security-related document databases based on the loT security-related query and descriptions of the plurality of loT security-related document databases to obtain the selected one or more loT security-related document databases.
22. The system according to claim 21, wherein said select the one or more loT security- related document databases comprises: generating a selector prompt based on the loT security-related query and the descriptions of the plurality of loT security-related document databases; generating loT security-related document database selection information using the language model based on the selector prompt; and selecting the one or more loT security-related document databases based on the loT security-related document database selection information.
23. The system according to claim 22, wherein the selector prompt is generated further based on user information corresponding to a user associated with the loT security-related query.
24. The system according to claim 23, wherein the user information comprises user type information and associated background information.
25. The system according to claim 23 or 24, wherein the user information of the user is selected from a predefined set of user information for different user types.
26. The system according to any one of claims 15 to 25, wherein said retrieve document chunks relating to the loT security-related query comprises: generating, for each of the selected one or more loT security-related document databases, one or more structured queries for the selected loT security-related document database based on loT security-related query; and retrieving, from each of the selected one or more loT security-related document databases, one or more document chunks relating to the loT security-related query from the selected loT security-related document database based on the one or more structured queries and search parameters for the selected loT security -related document database.
27. The system according to any one of claims 15 to 26, wherein for each of the plurality of loT security-related document databases, the loT security-related document database is updated based on new loT security-related data in the corresponding loT security-related dataset.
28. The system according to any one of claims 15 to 27, wherein the language model is a large language model.
29. The system according to any one of claims 15 to 28, further comprising the plurality of loT security-related document databases.
30. An Internet of Things (loT) security-related query system comprising: a system for generating a response to an loT security-related query using a language model according to any one of claims 15 to 29; andan loT security-related dataset processing system comprising: at least one memory; and at least one processor communicatively coupled to the at least one memory and configured to process, for each of the plurality of loT security-related datasets, contents of the loT security-related dataset into document chunks for a corresponding ToT security-related document database of the plurality of loT security-related document databases.
31. The loT security-related query system according to claim 30, wherein said process contents of the loT security-related dataset into document chunks for the corresponding loT security-related document database comprises: parsing raw data of the loT security-related dataset into content elements comprising textual content elements and non-textual content elements; converting non-textual content elements of the parsed content elements into textual content elements; and generating the document chunks for the corresponding loT security-related document database based on the textual content elements.
32. The loT security-related query system according to claim 31, further comprising defining one or more content fields and one or more metadata fields for the document chunks generated for the corresponding loT security -related document database, wherein the corresponding loT security-related document database is a vector database for storing the document chunks generated from the corresponding loT security -related dataset, and each document chunk comprises chunked texts, embedding vectors, and metadata.
33. The loT security-related query system according to claim 32, wherein the one or more content fields and the one or more metadata fields for the document chunks are defined based on a selection from a set of content fields and a set of metadata fields using a language model.
34. The loT security-related query system according to any one of claims 31 to 33, wherein the plurality of loT security-related datasets comprises loT security-related datasets of different security domains.
Citation Information
Patent Citations
Question and answer model training method and intelligent question and answer method and device in network security field
CN116933075A
Safe knowledge question-answering method and device
CN117609446A
Interactive cyber security user interface
US20240045990A1
Generating security reports
US20240256780A1