Bridge network device manager
By scanning and classifying the bridge network through Device Manager 100, and utilizing physical domain-specific knowledge and a learning database, the problem of existing technologies being unable to identify operational technical devices is solved. This enables accurate identification and safe management of operational technical devices in the bridge network, meeting the regulatory needs of the maritime industry.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-28
- Publication Date
- 2026-03-27
AI Technical Summary
Existing IT-based network security systems are unable to identify and manage operational technology devices in bridge networks, resulting in insufficient security, especially in the maritime sector, where these devices interact with the physical world and use different protocols, making them difficult to identify and monitor over the network.
The bridge network is scanned by Device Manager 100. Using physical domain-specific knowledge and queries of expected behavior, combined with classifiers and learning databases, the operating technology devices are progressively identified and classified. This includes identifying the specific characteristics and types of the devices and dynamically updating the database to improve identification accuracy.
It enables accurate identification and classification of operating technical equipment in the bridge network, improves network security, and can generate detailed equipment lists and security assessment reports to meet the regulatory requirements of the maritime industry.
Smart Images

Figure CN121753306A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to networked operating technology devices, and more particularly, to a bridge network device manager for networked operating technology devices. Background Technology
[0002] Cybersecurity is a global issue. People are aware of vulnerabilities within connected computer systems and seek to improve their security. Attempts are made to address the security of information technology (IT)-based systems communicating within networks. Tools used to test IT systems identify computers and other devices by detecting the operating system and other known network information. Unfortunately, these attempts have not addressed the use of bridged networks with specialized equipment that operates in the physical world.
[0003] A bridge can be a bridge for a ship, aircraft, or operating facility, or any other bridge used to monitor and direct the operation of physical equipment. For example, a bridge can be the conning tower of a ship, the cockpit or flight deck of an aircraft or space shuttle, the operations control center of a plant, the control room of a water treatment plant, or any other location that includes equipment for directing the operation of physical equipment.
[0004] In such a domain, many different devices communicate with each other. These devices include IT devices that interact with other IT devices via a network. They also include operational technology devices that interact both with devices via a network and with the physical world. Existing IT-based systems may be able to identify the presence of IT devices but fail to identify domain-specific operational technology devices, thus producing inaccurate outputs for domain-specific networks that include such devices. For example, an IT-based system might indicate the presence of 32 devices, four of which are personal computers and 28 are unknown. This not only fails to identify most of the devices, but the four personal computers may not actually be personal computers at all, but rather operational technology devices that happen to respond to a specific aspect of the personal computer identification process.
[0005] One example is the maritime sector, which comprises ships and ports. The maritime sector differs from computer systems connected via networks because it uses bridges and specialized equipment in control rooms similar to those on ships. This equipment includes operational technology equipment, which is distinctly different from IT equipment because it interacts with the physical world. For example, maritime operational technology equipment includes mapping systems, radar systems, steering control systems, satellite positioning system sensors, navigation systems, and other ship and port operational equipment. This equipment is difficult to identify on a network because it uses different protocols than traditional network equipment. It is even difficult to identify physically because it is often hidden in unseen locations, such as enclosed server racks.
[0006] As mentioned above, the goal of a network security system is to address problems with known IT devices (such as personal computers and routers), but not with specialized operational technology devices. IT systems can identify the presence of multiple nodes, but cannot identify nodes other than IT devices like computers and routers. Therefore, network security systems cannot provide sufficient security for specialized operational technology devices.
[0007] Seafarers need to comply with evolving regulations that require the maintenance of onboard asset inventories. However, the diverse range of equipment in maritime systems, such as radar and mapping systems, uses different protocols, making them difficult to identify. Furthermore, malicious devices, such as USB drives, can easily be attached to concealed locations on board, infecting systems. This presents significant challenges for penetration testers, ethical hackers, and seafarers in properly securing vessels and adhering to regulations.
[0008] The solutions described in this section are feasible solutions, but not necessarily solutions that have been previously conceived or practiced. Therefore, unless otherwise specified, no solution described in this section should be assumed to be prior art simply because it is included in this section. Furthermore, no solution described in this section should be assumed to be well-known, conventional, or routine simply because it is included in this section. Attached Figure Description
[0009] To describe how the advantages and features of this disclosure are obtained, the description of this disclosure will be given by reference to specific embodiments illustrated in the accompanying drawings. These drawings depict only exemplary embodiments of this disclosure and should not be considered as limiting its scope. For clarity, the drawings may have been simplified and are not necessarily drawn to scale.
[0010] Figure 1 This is an example illustration of a device manager for a bridge network according to one possible embodiment; Figure 2 This is a flowchart illustrating a method according to one possible embodiment; Figure 3 This is an example illustration of a ship area network according to one possible embodiment; Figure 4 This is an example illustration of a network packet flow topology according to one possible embodiment; Figure 5 This is an example illustration of an asset count histogram according to one possible embodiment; Figure 6 This is an example illustration of a port count distribution diagram drawn according to one possible embodiment; Figure 7 This is an example illustration of an open port heatmap according to one possible embodiment; Figure 8 This is a block diagram illustrating a computer system according to one possible embodiment; and Figure 9 It is a block diagram of a basic software system according to one possible embodiment. Detailed Implementation
[0011] In the following description, numerous specific details are set forth for illustrative purposes in order to provide a thorough understanding of the invention. However, it will be clear that the invention can be practiced without these specific details. In other instances, well-known structures and devices are shown in block diagram form to avoid unnecessarily obscuring the invention.
[0012] General Overview Figure 1 This is an example illustration of a device manager 100 according to one possible embodiment of a bridge network 102. In one possible implementation, the device manager 100 can be considered a profiler that scans the bridge network 102 and identifies devices on the network based on a first query 102. The device manager 100 then searches for an additional layer of information based on physical domain-specific knowledge, such as the use of a specific protocol, the use of a specific port, a specific communication protocol, and other physical domain-specific knowledge. The device manager 100 builds a second-order query based on this additional layer of information. The second-order query extracts information based on the expected behavior of a specific device. The device manager 100 categorizes devices using expected behavior based on knowledge about how devices of a specific category operate. The device manager 100 can then search for another layer of additional information to identify as many specific characteristics of the actual devices as possible. This information becomes part of a knowledge set, thereby increasing the information in the identification database and improving the identification database. Thus, this process can progress from general identification to specific identification. For example, over time, the process could progress from knowing that a device is a certain type of marine equipment to identifying that the device is a specific radar, model 35, firmware level 6.
[0013] For example, Device Manager 100 executes queries to extract device characteristics. It can send domain-specific queries to ports already identified as open to construct feature sets that are mapped to a database to identify the device. Upon receiving a response to a query for a specific device, Device Manager 100 can create an instantiation of a device record, such as a feature identifier for the device. The device record indicates which features of the feature set are present in the device. This feature information is fed into a classifier, which produces a prediction. For example, the classifier predicts the type of device based on the features extracted from the device record.
[0014] Device prediction provides information about how well the query is performing. Device Manager 100 uses this information to improve the query and iterate until the desired level of identification is achieved. For example, subsequent queries can be selected based on device prediction to refine the predictions. Each subsequent query is selected based on additional information obtained from previous queries and predictions, and the subsequent queries further narrow the prediction range. Thus, after identifying the category / type of the devices, Device Manager 100 can perform an additional scanning layer to identify the specific details of the devices, thereby generating a detailed list of devices present on network 102.
[0015] The database can include information about a collection of devices. Some devices that might look like personal computers could be communication convergents, specific control systems, or other domain-specific devices. A list of ports can be used to identify specific devices. For example, to identify whether a device is part of a specific control system, Device Manager 100 can determine that port 25 is open and can send one or more queries to that port, allowing Device Manager 100 to identify (e.g., classify) the device as a domain-specific device, such as a specific control system. After performing this second-level identification, additional queries can be sent to identify the device's brand, model, and other information.
[0016] This can be viewed as a learned branching tree that narrows the range of downstream nodes. As the tree is traversed further, more details are gained. The database is no longer static but a dynamic learning environment. Thus, through chained identification queries, more specific identifications are obtained, such as manufacturer, brand, model, and so on. This ever-increasing knowledge improves the database, allowing for faster and more accurate identification of devices previously identified in other environments, which are then used to update the database.
[0017] This process can be repeated until the classifier has a sufficiently good recognition of the device. Thus, a series of queries allows the classifier to become accurate enough to identify the device.
[0018] In a specific maritime example, Device Manager 100 can couple to a bridge on a vessel and identify the operational technology equipment communicating with the bridge. Using a query and classification process, Device Manager 100 can indicate that it has found two mapping devices belonging to a certain category and model. This classification detail is provided using a learning database that Device Manager 100 is using. This differs from existing tools, which can only identify the presence of nodes on a network but cannot identify the type of operational technology equipment, such as a specific mapping device at a node.
[0019] At least some embodiments can improve security within the bridging department for monitoring and directing the operation of physical equipment. At least some embodiments can also allow prediction of device types among various operational technology devices. At least some embodiments can further identify devices on the bridging and create a topology of which devices are communicating with each other, such as an electronic charting system that collects data from sensors. At least some embodiments can further provide a dataset of maritime system hardware for use with a classifier for predicting devices on the network.
[0020] The embodiments may provide different implementations, each unique to a different aspect of this disclosure. One unique implementation may be a feature set of operating technology devices, which has sufficient features to classify operating technology devices operating in a particular environment. Another implementation may be a query set constructed for querying the feature set. Another implementation may be selecting queries from the query set to extract feature information from the operating technology devices. The obtained information may be extracted and encoded in a manner that allows a classifier to appropriately classify (e.g., predict) the operating technology devices. Another implementation may be identifying intermediate devices to generate a query, thereby identifying devices connected to the intermediate devices. Other unique aspects of the implementations will be described below.
[0021] Network catcher Device Manager 100 may include Network Detector 104. Network Detector 104 is a data collection module that begins the process of growing a database by initiating network reconnaissance and collecting information from bridge network 102. It determines which devices are connected on network 102, their Internet Protocol (IP) addresses, their Medium Access Control (MAC) addresses, which device ports are open, which open ports are vulnerable, and other information. It can also identify device-specific and department-specific information, such as specific communication protocols. For example, Network Detector 104 can determine whether a device is communicating using the National Marine Electronics Association (NMEA) protocol, the Automatic Identification System (AIS) protocol, and / or other specific protocols. After receiving network configuration in the form of network domain addresses from network 102, Network Detector 104 captures information such as network traces as Packet Capture (PCAP) files, data packets, and / or other data formats.
[0022] Topology Builder Device Manager 100 may include Topology Builder 106. Topology Builder 106 creates a network communication flow graph based on information collected by Network Capturer 104. For example, a PCAP file is processed by Topology Builder 106, which creates a directed topology graph 108 depicting network connections and system communication patterns.
[0023] Feature extractor Topology builder 106 can pass network information to feature extractor 110, which collects feature data to construct and / or update dataset 112. To this end, feature extractor 110 scans network information and / or network 102, collects information about open ports, manufacturers, protocols, operating systems, and extracts these features of devices on network 102.
[0024] Feature extractor 110 collects feature information to generate a set of records for identifying devices on network 102. Specifically, feature extractor 110 selects a query and executes the query on a device to collect information about the device's features, and generates an instantiation of a device record that includes information about that feature in the feature set of the device record. According to a possible example, each query can provide information about a specific feature in the record, and the instantiation of the record can indicate whether that specific feature is present or absent in the feature set. The collected feature information can be used to select subsequent queries to refine the information in the feature set, thereby correctly identifying the device.
[0025] For example, for each feature of interest, there may be one or more queries that will produce the information under different circumstances. Feature extractor 110 identifies the desired feature to be extracted and selects from a set of queries that are likely to produce the desired information. Depending on the result of the query, feature extractor 110 may then select other queries from a known list to extract the desired information. When the information is extracted, the knowledge of which query successfully extracted the information is itself encoded as information about the feature.
[0026] For example, feature extractor 110 sends a query to extract desired information from a device (e.g., a node on network 102 and / or topology 108). This information may or may not be obtained. The fact that this information was obtained is a small additional piece of information about the characteristics of that node. Feature extractor 110 may select another query that may or may not obtain the desired information. Feature extractor 110 may continue to select and send queries until it extracts the complete set of desired features applicable to the identified node types. If a node is found to be an intermediate node, feature extractor 110 can send queries from an additional set of queries to connected nodes through the intermediate node to determine the device type of the connected node.
[0027] Knowing which queries identify a node as a piece of information itself, this information is added to dataset 112 to help select queries for each node. The set of queries in the database can be selected based on current knowledge about the nodes being queried. The acquired knowledge about successful query combinations provides a more refined query selection, which feature extractor 110 further enhances using these queries.
[0028] As an example, the manufacturer ID could be the first few digits of a MAC address. The feature extractor now knows that the device comes from a specific manufacturer. Some queries in the dataset are biased towards devices from specific manufacturers. The feature extractor 110 can select one of these queries as the next query. The feature extractor 110 can also know that the device is a specific type of device from that manufacturer, such as a radar display. The feature extractor can then select queries that are more likely to receive a positive response for radar displays.
[0029] Therefore, based on the knowledge that the device comes from a specific manufacturer, the feature extractor 110 can know that the next desired feature is likely to be selected by one of the queries. The feature extractor 110 constructs knowledge of the query set to extract the feature set. As the query set evolves, the feature extractor 110 becomes increasingly better at acquiring this information, and the redundancy in its selection of query methods decreases.
[0030] For example, topology builder 106 can identify a node on network 102. If topology builder 106 does not know the type of device, it can identify the node as an unknown device. The query builder in feature extractor 110 can add queries to determine how to identify the device, thereby expanding and improving the query set used to identify unknown devices.
[0031] Dataset The features of the devices collected by feature extractor 110 are incorporated into dataset 112, such as device records in a database. Dataset 112 is populated with the identifying features from the device records. Each device record may have a feature set, such as fields for certain features. As described above, feature extractor 110 selects a query and executes the query on the device to collect information about the device's features, and generates an instantiation of the device record that includes information about that feature from the feature set. Each query can provide information about a specific feature of a record in dataset 112, and the instantiation of the record can indicate whether that specific feature is present in the feature set. Thus, dataset 112 may include a data structure that records a feature set of a particular device. This data structure may be an instance of a device record having a feature set of features populated from the query results of feature extractor 110.
[0032] For example, a device record may have a data structure that includes a certain number of features. Feature extractor 110 uses queries to generate instantiations of the data structure. Instantiation can indicate whether each feature in the feature set exists, does not exist, or is unknown.
[0033] Feature set Queries are used to generate instantiations of data structures. For example, a feature set is a collection of features recorded in a record. A query for feature extraction produces an instance of that record. Each query can provide one piece of information from a specific record for a specific device. This record can indicate whether each feature exists, does not exist, is unknown, or otherwise identifies the presence and / or absence of that feature. More specifically, an instance of a device record can be a feature record that provides values from fields in the feature set. A feature set comprises a list of features used to identify a specific device. Each query can provide one piece of information, i.e., a feature, within that record, and any random query used for feature extraction will produce an instance of that record.
[0034] A feature set may include a sufficient number of features to uniquely identify the category of equipment, uniquely identify the model of equipment within the category, and perform other equipment identification. The number of features may depend on the specific field in which the scanned equipment operates. For example, the feature set may include a sufficient number of features to identify operational technology equipment operating in a specific physical field. In the maritime field, this number could be 21 features, 26 features, or any number of features sufficient to identify maritime operational technology equipment. Other numbers of features are also possible for maritime operational technology equipment, aviation operational technology equipment, factory operational technology equipment, and operational technology equipment operating in other fields.
[0035] The feature set may include characteristics such as device type, manufacturer, model, version number, operating system, communication protocol used by the device, open port numbers of the device, and other features that can identify the device. Certain features may be particularly helpful in classifying operational technology devices, or at least narrow the scope of the query process for classifying operational technology devices. Such features may include the protocol the device responds to, the open ports of the device, and other features that, individually or in combination, can classify operational technology devices into specific types of devices. Once the device type is determined, manufacturer identifier features can be obtained to identify specific devices of that type.
[0036] In one possible embodiment, a query set is constructed to identify certain features that have been identified as sufficient to classify (e.g., identify or predict) maritime equipment. Example feature sets may include port numbers such as: 21 (File Transfer Protocol or FTP), 22 (Secure Shell or SSH), 23 (Telnet), 25 (Simple Mail Transfer Protocol or SMTP), 445 (Microsoft Server Message Block (SMB)), 53 (Domain), 80 (Hypertext Transfer Protocol or HTTP), 81 (tiny / turbo / throttling HTTP server (THTTPD) – device specific), 137 (Netbios), 443 (HTTPS), 3389 (Windows Remote Desktop Protocol (RDP)), 5800 (VNC), 5900 (VNC), 5901 (VNC), 502 (MODBUS protocol – device specific), 4000 / 4001 (zmtp, remoteanything – device specific), 4800 (MOXA). UDP (device-specific), 10010 (rxapi – device-specific), 139 (Netbios-SS), 445 (Microsoft-DS), 161 (Simple Network Management Protocol or SNMP), and 123 (Network Time Protocol or NTP). Ports 139 (Netbios-SS), 445 (Microsoft-DS), 161 (Simple Network Management Protocol or SNMP), and 123 (Network Time Protocol or NTP) can be UDP ports. In addition to the protocols mentioned in the open port list, the example feature set may also include other device characteristics such as: operating system, device model / version, device manufacturer, and device protocols such as NMEA and AIS. Other protocols may include Supervisory Control and Data Acquisition (SCADA) protocols for controlling, monitoring, and analyzing industrial equipment and processes, and Controller Area Network (CAN) protocols that allow devices to communicate with each other in a vehicle without a host computer and other control applications. Other protocols may also be used for devices operating in other physical domains.
[0037] The characteristics of protocols allow for different methods of device identification. In the NMEA example, Device Manager 100 can look for latitude and longitude fields within packets because standard TCP / IP packet protocols typically do not include this information, but the NMEA protocol does. For SCADA, Device Manager 100 can view fields indicating commands such as on, off, left, and right—all based on protocol information. Thus, Device Manager 100 can use the characteristics of known operating technology protocols to identify device features. Other fields, such as aviation, manufacturing, rail, and others, can utilize additional feature sets specific to their operating technologies and devices.
[0038] encoder Encoder 114 transforms and prepares the data in dataset 112, generating encoded data based on features in instances recorded by the device for use by classifier 116. When a new classification test is performed, feature information is extracted from dataset 112 and fed into encoder 114, which converts the information into encoded data usable by classifier 116. For example, encoder 114 can provide yes / no indications for different features, such as by using 1 and 0. As a more specific example, port 421 can be a feature where a value of one indicates that port 421 is open, while a value of zero indicates that port 421 is not open.
[0039] Classifier Encoded feature data is provided from encoder 114 to classifier 116. Classifier 116 can be any classifier, such as a random forest classifier, artificial neural network classifier, support vector machine classifier, Bayesian classifier, or any other classifier. Classifier 116 predicts the type of device based on features identified in instances of device records. In a general sense, classifier 116 feeds the encoded data into various buckets, with continuously improving resolution for these buckets. Classifier 116 can predict the type of device by comparing features with stored device feature data. Classifier 116 uses bridging domain-specific knowledge, such as device profiles for maritime, aerospace, or other equipment, to predict the device.
[0040] For example, classifier 116 can receive all existing features of the device in an instance of device recorded information, determine that these features indicate the device has sensors and is communicating on a certain port, and predict that the device is radar. To this end, classifier 116 can compare the recorded features with a known set of feature combinations to identify the device. For example, classifier 116 can map features in an instance of device recorded information to an existing dataset to predict the device.
[0041] Equipment Profile Classifier 116 creates a profile 118 for each device discovered in network 102, and the output is fed back to dataset 112. For example, the feature set of a particular device record in dataset 112 is updated to include information about the predicted device type, which could be one of the features in the feature set.
[0042] Model validator Model Validator 120 validates the model and calculates the classification accuracy score. The accuracy score serves as an indicator of model performance; a higher accuracy score indicates that the model is more accurately able to identify devices and distinguish between different devices.
[0043] recorder The visualization and recorder 122 can generate graphs, heatmaps, and other images to visually depict information about the assets, which will be included in reports generated after the entire testing process. For example, the recorder 122 can use information from each device profile 118 and topology diagram 108 to provide visual information about network 102 in an easily understandable way to different categories of personnel involved in bridge operation, diagnostics, and other processes.
[0044] In detail, the recorder 122 can output information from the device profile 118 in various formats. These formats may include files, PDF documents, hard copies, displayed information, text, tables, graphs, port heatmaps, and other formats. Charts, graphs, and heatmaps can help provide information in a form that is meaningful to domain users unfamiliar with the deep technical details of different devices. For example, the domain user might be a seaman who is not an IT professional, and the report would indicate devices found on the ship's network. Some may be obvious, such as radar systems, while others may be less obvious, such as a rudder control unit using a certain brand of controller. The device profile can be mapped to regulatory requirements to ensure that specific domains meet those requirements.
[0045] In one example, a heatmap can identify which ports are open to different devices and how frequently they are used. It can also identify which devices are similar in terms of vulnerabilities. Furthermore, it can identify known vulnerabilities for specific devices. For example, the Voyage Data Recorder (VDR) is a black box on ships used to store evidence and data from accident investigations. Device documentation has indicated that VDRs have vulnerabilities similar to those found on Windows personal computers, allowing for security measures to be taken against these vulnerabilities.
[0046] Additional queries In one possible embodiment, the feature extractor 110 can initiate multiple queries without further refinement to obtain a set of feature data in the device records, and the classifier 116 can generate predictions based on the obtained records.
[0047] In another possible embodiment, the query can be refined based on the predictions made by classifier 116. In this embodiment, feature extractor 110 executes the query to extract features, classifier 116 makes predictions based on the results of the query, and feature extractor 110 selects a subsequent query based on an updated feature dataset that includes information from the predictions. Feature extractor 110 performs this subsequent query on the device to gather additional feature information about the device. Using information from previous predictions to select the subsequent query improves the accuracy of the query.
[0048] Then, feature extractor 110 updates the device records to include additional features in the feature set and stores the records in a database. Encoder 114 generates encoded data of the features in the updated records. Classifier 116 predicts the updated type of the device by comparing the updated features with the stored device feature data and updates the device profile 118. The device records in dataset 112 can then be updated based on the updated device profile 118. Feature extractor 110 can select additional queries based on each update of the device records, thereby improving the accuracy of the queries, the accuracy of the features, the accuracy of each prediction, the accuracy of the device profile 118, and the accuracy of the device records in dataset 112. Thus, classifier 116 performs better classification work as it updates the records with each query based on each subsequent prediction.
[0049] Initially, there may be no prior prediction information, so random queries can be used to determine initial device features for the initial prediction. As the prediction scope narrows, the randomness of query selection decreases because the queries are based on refined features from the predictions. As the predictions become more refined, better queries are selected to extract additional feature information. Query refinement can continue, for example, until no more queries can add information, thus achieving accurate device predictions.
[0050] intermediate equipment Many network topologies that include operational technology devices operating in the physical domain are unusual compared to classic network topologies. For example, in a classic network with IT devices operating in a network domain, everything uses Transmission Control Protocol (TCP) / IP over Ethernet. In many topologies with operational technology devices operating in the physical domain, some devices are connected via wires to intermediate converters, such as protocol converters. Specifically, the intermediate converter acts as a gateway to another device, which is often a more important device connected to the converter. Device Manager 100 can utilize information about the physical domain and treat the converter and its connected devices as a single unit by recognizing the converter as an intermediate node. For example, when Device Manager 100 recognizes a serial-to-TCP / IP converter, it might not treat the converter as an end node, whereas a classic topology scanner would. Device Manager 100 then submits a query to identify what device is behind the converter. Device Manager 100 can submit these queries based on the protocol being used. For example, the converter can be identified as an NMEA or RS422 to Ethernet converter that communicates with devices using the NMEA or RS422 protocol, and the device manager 100 can submit queries based on that protocol.
[0051] Therefore, in some cases, device manager 100 (e.g., feature extractor 110) can perform a two-stage scan when identifying a converter. For example, device manager 100 can perform a first scan and identify an intermediate device, such as a converter. Device manager 100 can then perform an additional scan by submitting additional queries to identify devices connected to that intermediate device using received information (e.g., knowledge of the protocol used). Device manager 100 can infer the protocol from contextual information, such as from identified features. Device manager 100 can also infer the type of connected device based on protocol information.
[0052] For example, Device Manager 100 can identify the converter in use and detect packets coming out of the converter to identify the underlying protocol, such as IP. Device Manager 100 can combine this with other identified features to determine that the primary reason for using a converter from an underlying protocol (e.g., IP) is that there is a connected NMEA serial device at the back end. For example, the protocol from the connected device to the converter might be serial NMEA, while the outgoing protocol might be NMEA over TCP / IP. Device Manager 100 can deconstruct IP packets to extract the NMEA protocol, construct queries using NMEA over TCP / IP packets, and send the constructed packets to the converter in subsequent queries. The converter can convert the constructed packets to NMEA on the serial interface, allowing Device Manager 100 to send queries to the device to obtain characteristics of the connected device. Thus, Device Manager 100 can send packets utilizing the NMEA serial protocol to query the connected device for additional characteristics. After determining the connected device, Topology Builder 106 can add the converter and connected device to the topology graph 108.
[0053] Operating method Figure 2 Example flowchart 200 illustrates a method of operation of a bridge network device manager 100 according to one possible embodiment. At 202, the method may include performing a first query on devices of an operational technology device network. Operational technology devices are hardware and software system devices that monitor and control the physical world. For example, operational technology devices operate in conjunction with physical processes by monitoring or controlling physical processes, control physical objects, receive information from the physical world, monitor and control the execution of physical devices, detect or directly alter physical properties, receive sensor data, measure physical characteristics, and / or may otherwise operate physical processes. As a particular example, an operational technology device may be a maritime mapping system, a computing device with a display, that receives information from other devices, such as position and velocity information, then maps and displays the information. Operational technology devices may also be radar, rudder control systems, Global Positioning System (GPS) sensors, speed sensors, and / or any other operational technology devices. The network may also include information technology devices. Thus, the network may include physical domain-specific operational technology devices and information technology devices.
[0054] At 204, the method may include collecting first characteristic information of the device based on a first query. At 206, the method may include generating a first instance of a device record that includes the first characteristic information. A specific device record is an instantiation of a feature set corresponding to a device with a specific operating technology.
[0055] At 208, the method may include selecting a second query based on a first instance of the device record. The second query is selected from a query set of characteristics of operating technology devices operating in a specific physical domain. The specific physical domain may be a ship at sea, an aircraft, a factory, or any other physical domain.
[0056] At 210, the method may include performing a second query on the device. At 212, the method may include collecting second characteristic information of the device based on the second query. At 214, the method may include updating the device record based on the second characteristic information to generate an updated device record. At 216, the method may include predicting the type of the device based on the updated device record. The method can be repeated by performing additional queries and updating the prediction based on information received in each query.
[0057] In one possible embodiment of selecting a second query, the method includes predicting a first type of device based on a first instance of a device record. The method also includes updating the first instance of the device record based on the predicted first type of device. Selecting a second query at point 208 includes selecting a second query based on the updated first instance of the device record.
[0058] One possible example of selecting a second query could involve selecting the second query based on identifying whether the device is an intermediate device. In this example, predicting the first type of the device includes predicting that the first type of the device is an intermediate device that converts a first communication protocol of the connected devices to a second communication protocol. Selecting the second query at 208 includes selecting the second query based on the first communication protocol and the second communication protocol. Collecting second feature information at 212 includes collecting second feature information of the connected device based on the second query. Updating the device record at 214 includes storing the connected device record based on the second feature information. Predicting the type of the device at 216 includes predicting the type of the connected device based on the connected device record.
[0059] One possible embodiment involves a device record comprising a feature set. The method may include selecting a first query from a query set of features of an operating technology device operating in a particular physical domain. Each instance of the device record includes a predetermined feature set of features of the operating technology device operating in the particular physical domain. Each instance of the device record indicates the presence or absence of a feature in the predetermined feature set. The feature set of the device record may include fields for each feature in the set. Fields may indicate the presence or absence of the corresponding feature, for example, by using one or zero. Fields may also include alphanumeric text of the feature, such as the protocol used by the device, the port used by the device (e.g., at least one port number), the operating system of the device, the manufacturer of the device, the model of the device, the IP address of the device, the predicted type of the device, and other information. These fields may also indicate whether the information for the corresponding field is unknown.
[0060] In one possible embodiment, the query set is constructed to inquire about the characteristics of operational technology equipment operating in a specific physical domain. In another possible embodiment, prediction is performed by a classifier trained on data collected from a set of cyber-physical testbeds, which is based on maritime hardware equipment configured as a vessel bridge.
[0061] In one possible embodiment, the second query may be a query about the communication protocol of the operating technology device. The operating technology device communication protocol can be considered an industrial communication protocol. The second query may also include queries about a specific port number, an operating system, a model / version, a manufacturer, a communication protocol, and any other queries about the characteristics of the operating technology device.
[0062] In one possible embodiment, the specific physical domain is a ship, the device is a ship operation technology device coupled to a bridge with the ship, and the second query is a query of the ship operation technology device's communication protocol. For example, the ship operation technology device communication protocol could be the National Marine Electronics Association (NMEA) protocol or the Automatic Identification System (AIS) protocol.
[0063] In one possible embodiment, the method may include creating a network communication flow graph of an operating technology device network. The network communication flow graph includes a first network connection between a first device and an intermediate device using a first communication protocol, and a second network connection between the intermediate device and a second device using a second communication protocol, wherein the second communication protocol differs from the first communication protocol.
[0064] At least some embodiments may provide one or more non-transitory computer-readable storage media storing instructions that, when executed by one or more processors, cause the operation of the methods in the disclosed embodiments.
[0065] Maritime Asset Manager In some embodiments, the network device manager 100 operates within the context of a marine vessel bridge, such as an asset profiler for a marine vessel bridge. Features of these embodiments can be incorporated into the embodiments identified above. In these embodiments, the device manager 100 may include a feature extractor 110 that performs a first query on devices coupled to the operational technology device network of the marine vessel bridge. The operational technology devices operate in conjunction with the physical processes of the marine vessel. The feature extractor 110 also performs a process of collecting feature information of the devices based on at least the first query and generates a device record that includes features from the feature set.
[0066] Device manager 100 may include a database storing device records, such as dataset 112. Device manager 100 may include encoder 114 that generates first encoded data based on a feature set. Device manager 100 may include classifier 116 that predicts a device type based on the first encoded data by comparing it with stored device feature data. Device manager 100 updates the feature set in the device records with information about the predicted device type.
[0067] Feature extractor 110 selects a second query based on information about the predicted equipment type. The second query is selected from a query set of features of marine vessel operation technology equipment. The second query is executed on the equipment, and additional feature information of the equipment is collected based on the second query.
[0068] Device manager 100 updates the feature set in the device record with information about additional features. Encoder 116 generates second encoded data based on the feature set including the additional features. Classifier 116 predicts the updated type of the device based on the second encoded data by comparing it with stored device feature data.
[0069] While many embodiments are not limited to the maritime bridging environment of marine vessel bridging devices, such a field is provided as an example for illustrative purposes. The maritime bridging environment is a heterogeneous ecosystem comprised of complex systems from various maritime operations. As part of new requirements from the International Association of Classification Societies (IACS), vessel operators are now required to maintain an inventory of assets on board vessels specifically to improve their cybersecurity. Vessel-specific versions of Device Manager 100, such as the Maritime Asset Profiler, not only automatically identify and record the devices present but also provide in-depth analysis of the attributes and characteristics of these devices in an intelligent and user-friendly manner. With the increasing prevalence of cyberattacks in the maritime industry, proper testing of vessel systems is crucial to ensuring vessel safety and minimizing the risk of cyberattacks. Device Manager 100 for bridging environments acts as a tool for profiling devices, helping personnel make faster, more informed decisions and serving as part of a broader audit framework.
[0070] This Device Manager 100 can also be referred to as a vessel bridging asset profiler or an asset profiler for penetration testing in heterogeneous marine bridging environments. Device Manager 100 is used to automatically identify all devices on a vessel's bridging system. Furthermore, Device Manager 100 provides information about the devices, such as through generated PDF reports or displays including graphs and charts. In one implementation, Device Manager 100 uses a classifier algorithm, and the information it provides enables auditors or penetration testers to perform tests and automate audits, while also providing engineers and seafarers with comprehensive information for regulatory compliance.
[0071] Introduction to Maritime Asset Analyzer The maritime industry is a complex multi-billion dollar sector and a vital component of the global economy. Countries like the United States import approximately 90% of their goods by sea, while China is a major importer of resources such as oil and iron. As technology advances, the complex systems on board ships are adapting to new functionalities to make operation easier and more efficient. With the advent of network connectivity, emerging topics such as Artificial Intelligence (AI) and Machine Learning (ML) have appeared in traditional operating environments. While these often provide improved security, usability, and comfort, they also introduce new challenges, such as cyber vulnerabilities or flaws in critical shipboard systems that could be exploited by cybercriminals.
[0072] A cybersecurity audit is a process that helps identify digital threats within a defined scope. This involves a comprehensive review of system vulnerabilities, compliance with policies and regulations, and an assessment of cyber risks. One of the first steps in a cybersecurity audit is gathering information to identify scope and assets. The IASME Maritime Network Baseline, developed in November 2021 by the Information Assurance for Small and Medium Enterprises (IASME) and supported by the Royal Institution of Naval Architects (RINA), is an audit process that uses a checklist that allows ship owners and operators to demonstrate the compliance of security controls and processes. Within the scope of the audit, the checklist requires a register of all information and operational technology (IT / OT) assets, along with their brand, model, and other characteristics. The assessment also requires listing all networks on board, their functions, how they are segmented, routers, firewalls, and gateways.
[0073] Identifying systems and compiling equipment inventories are also necessary for compliance with certain requirements, standards, and policies, such as the Unified Requirements (URs) of the International Association of Classification Societies (IACS). IACS is the organization of classification societies that establishes technical standards for the shipping and maritime industry. IACS produces Unified Requirements (URs), which are adopted resolutions regarding the minimum requirements for matters covered by the classification societies. To ensure network resilience on ships, IACS has produced two new URs, URE26 (regarding network resilience of vessels) and UR E27 (regarding network resilience of onboard systems and equipment), effective January 1, 2024. Both URs apply to ships built on or after January 1, 2024. UR E26 addresses the minimum requirements for establishing network resilience on vessels, while UR E27 addresses establishing network resilience for onboard systems, not the vessel itself.
[0074] The first objective in UR E26, “Identification,” mentions identifying all computer-based systems (CBSs) on board, their interrelationships, dependencies, and the resources involved. This includes creating and maintaining a list of all CBSs and related networks on board throughout the vessel's service life. UR also requires detailed information on the systems, such as manufacturer, brand, model, and their logical connections on the network. As part of Section 3.1 of the UR E27 document, information on equipment, hardware, operating systems, configuration files, and network processes, as well as plans and policies, must be submitted to the classification society for review and approval. This is followed by maintaining a list of equipment names, manufacturers, models, and software versions, as well as a software list that includes at least the installation date, version number, maintenance, and access control policies.
[0075] Given the increased requirements and guidelines introduced by maritime authorities to improve onboard network security, a proper asset management process is essential. Therefore, Device Manager 100 identifies onboard assets / equipment and provides relevant information to testers / auditors monitoring this process. Furthermore, Device Manager 100 helps audit / identify any unused and unwanted devices connected to the network, which could potentially become vulnerabilities in the overall environment. Device Manager 100 generates a concise, user-friendly PDF report containing all discovered and analyzed asset and network information, which can be used in conjunction with maintaining the asset register. Device Manager 100 can also be integrated into automated penetration testing systems to aid in testing for system vulnerabilities.
[0076] Asset Identification Figure 3This is an example illustration of a ship area network 300 according to one possible embodiment. Many types of networks exist, especially in complex heterogeneous environments. Different systems communicate using different protocols, creating separate subnetworks (often IT or operational technology (OT) specific) or system clusters. Assets in a ship environment include equipment, communication interfaces, and networks critical to the smooth operation of the vessel. Each organization may define the term "asset" differently, but in this disclosure, the term refers to any equipment networked on a ship's bridge for bridge operation (e.g., navigation, emergency communications) and having an assigned IP address for communication.
[0077] Ship bridging systems typically include a variety of equipment, including Electronic Chart Displaying and Information Systems (ECDIS), Vehicle Data Receipts (VDRs), Automatic Identification Systems (AIS), radar, Very High Frequency (VHF) equipment, Global Maritime Distress and Safety Systems (GMDSS), compasses, gyroscopes, and more. Ship safety and security standards mandate certain equipment, but the types of equipment may vary depending on the ship's class, size, and type. For example, Chapter 5 of the Safety of Life at Sea (SOLAS)—Navigation Safety Requirements—requires that ships built on or after July 1, 2002, or roll-on / roll-off passenger ships built before July 1, 2002, or vessels with a gross tonnage of 3,000 tons or more (excluding passenger ships) built on or after July 1, 2002, be equipped with VDRs. However, considering the size and type of vessel, the regulation states that vessels can be equipped with S-VDR (Simplified VDR), which captures less data than a standard VDR. This difference in equipment type alters the scope and characteristics of management and testing. Therefore, using Equipment Manager 100 to identify and analyze equipment is a useful reconnaissance tool for engineers / seafarers to maintain their checklists in compliance with regulations. More use cases will be discussed below.
[0078] Asset maintenance and inventory list An asset inventory provides useful situational awareness for maintenance and incident response. An asset inventory, including information such as device type, IP address, MAC address, open ports, manufacturer information, and version number, makes managing these devices much easier for those responsible. For example, it allows them to identify any obsolete / unused devices connected to the network. Removing such devices reduces the network's threat surface without impacting operations. According to the 2020 Global Network Insights Report, based on an assessment of over 800,000 IT network devices, 47.9% of enterprise network assets are obsolete, with an average of twice as many vulnerabilities per device (42.2) compared to aging devices (26.8) and current devices (19.4). Therefore, Device Manager 100 allows operators or asset owners to better understand their systems and maintain asset inventories to comply with regulatory requirements. The table below describes the IACS URs that Device Manager 100 can resolve.
[0079] Table 1: IACS UR Information gathering by penetration testers / auditors Maritime cybersecurity for ships is a relatively new discipline concerning the protection of onboard systems and surrounding infrastructure. Understanding gaps and vulnerabilities is a crucial step in this discipline. Penetration testing, or penetration testing, is the process by which authorized personnel attack a system within a defined area to discover exploitable vulnerabilities and threats. When conducting penetration testing in a real-time, complex environment with department-specific equipment, a lack of system knowledge can be challenging and disruptive to operations. Traditionally, penetration testing has been used to test IT systems; however, today, OT penetration testing, as well as IT systems used to monitor or control OT, are becoming increasingly common.
[0080] One of the first steps in penetration testing is gathering information to identify the scope and assets. In an IT environment, this is fairly straightforward, as most devices are computers, networked devices, or small IoT devices. However, in a ship's bridging environment, networked devices are more customized and have diverse uses, making it more difficult, yet still necessary, to understand the existing systems and networks. Currently, this is done manually by penetration testers or auditors boarding the ship. This is very time-consuming for penetration testers / auditors and requires appropriate technical qualifications and certifications. With automated tools like Device Manager 100, testers not only don't need to be familiar with all the devices and protocols in the department being tested, but the tool also provides guidance to non-specialists, making it faster than experts manually inspecting assets. On a ship's bridging equipment, the equipment may also be hidden out of sight, making using Device Manager 100 a less intrusive and disruptive process.
[0081] Network topology generators for IT systems identify, list, and visualize networks. Auditors use network topology generator tools to view the status of their networks and monitor them. The tool takes the IP addresses of seed devices (typically the main switch in the network), scans for devices, and draws a map with the IP addresses. Network topology generators are used in IT environments where common devices include routers, switches, servers, firewalls, VMware hosts, and wireless access points. These tools are used by network administrators in IT and office environments where devices are mostly PCs, routers, firewalls, and network gateways. They do not include domain-specific devices or those without internet connectivity. They also face additional challenges: identifying devices outside the system's scope, unidentified devices from the same manufacturer, dynamically growing datasets and the ability to learn new devices, and robustness of features.
[0082] Device Manager Device Manager 100 is responsible for monitoring ship systems to bridge the gap between IoT and IT systems.
[0083] Components and tools Device Manager 100, such as Asset Profiler, may include the following components, tools, and terms.
[0084] Topology Builder 106: This module creates a network communication flow graph. After receiving network configuration in the form of network domain addresses, Device Manager 100 captures network traces as Packet Capture Approval Files (PCAP files). These files are processed by Topology Builder 106, which creates a directed topology graph 108 depicting network connections and system communication patterns.
[0085] Feature Extractor 110: Network information from Topology Builder 106 is passed to Feature Extractor 110, which scans the network, collects information about open ports, manufacturers, and operating systems, and extracts their features. The features are then incorporated into Dataset 112.
[0086] Dataset 112: This feature dataset can be fed into the Machine Learning (ML) module for data analysis. ML and Artificial Intelligence (AI) can be used for various purposes, including predicting vessel categories and using millions of AIS data records to predict flow density, reduce fuel emissions, and avoid collisions. Dataset 112 was created for the classification, profiling, and testing of onboard hardware and electronic equipment.
[0087] Encoder 114: Encoder 114 transforms and prepares the data in the dataset for use by classifier 116. When a new test is performed, information is extracted from dataset 112 and fed into encoder 114, which transforms it into information that classifier 116 can use.
[0088] Classifier 116: Classifier 116 is used to classify all the discovered devices. The Random Forest classifier has been found to be a useful classifier that produces accurate results and works robustly on limited data, but other classifiers can also be easily used. Classifier 116 creates profiles 118 for all devices discovered in the network, and the output is fed back into dataset 112.
[0089] Scikit-learn (Sklearn): The Sklearn library is used to create training and test sets.
[0090] Model Validator 120: The model validator module 120 validates the model and calculates the classification accuracy score. The accuracy score serves as an indicator of model performance; a higher accuracy score indicates that the model is more accurately able to identify devices and distinguish between different devices.
[0091] Recorder 122: Device Manager 100 will also utilize visualization and Recorder 122 to automatically generate graphics, heatmaps and other images to visually depict information about the asset, which will be included in the report generated after the entire testing process.
[0092] Network communication topology builder Device Manager 100 first performs network reconnaissance and creates a network communication topology graph for the discovered devices using Topology Builder 106. The network communication topology defines how nodes or devices connect and communicate with each other. This can be a useful step, and the topology graph also provides a comprehensive view of the network infrastructure to ensure devices are functioning properly. A directed graph can be used to keep Topology Builder 106 simple. To build the topology graph, network traffic from the environment is captured and transformed into a graph format with nodes and edges, where directed edges represent the source and destination of packets. The resulting graph provides a visual representation of network traffic and its connections, allowing auditors or engineers to understand a high-level view of the environment. Using network domain addresses as input, this framework tool begins by capturing network traffic in the form of a PCAP file over a specific period, then parses it and transforms it into a graph. For each new IP address in the PCAP file, a node is created in the graph for the source and destination IP addresses, with arrows connecting the nodes to indicate the direction of packet flow; if a node already exists, connections are marked between existing nodes.
[0093] Figure 4 This is an example illustration of a network packet flow topology graph 400 according to one possible embodiment. Once the topology builder 106 plots all entries into a graph, the entire communication topology can be visualized, as shown in topology graph 400. Traditionally, this approach only captures the network within a limited time period and is static during the testing period. At least some disclosed embodiments can provide dynamic creation of topology graphs, changing the topology graph according to the auditor's needs.
[0094] Analysis The next part of Device Manager 100 identifies the devices it finds in the ship's bridge network and profiles them based on device characteristics using a classifier (such as a random forest classifier or other classifiers).
[0095] Dataset For maritime bridge device identification, dataset 112 contains data on specific characteristics of these devices, enabling the differentiation between maritime bridge-specific devices and general IT / OT devices. Since datasets for profiling maritime equipment are not publicly available, a new dataset was created to test this framework. Because real-time scanning of bridges on operational vessels is challenging, destructive, and risky, data was collected from the Cyber-SHIP laboratory at Plymouth University, a cyber-physical testbed. Cyber-SHIP is a maritime-network research facility that configures real maritime hardware as an electrically accurate representation of a testable vessel bridge. In the experiments, the equipment and software are configured to act as bridges on vessels, so the collected network data is not simulated and therefore has high fidelity.
[0096] Since the framework is validated using collected data, the scope of the dataset can be determined before data collection: (1) how much data is required, (2) what types of data will be collected, and (3) what the expected output is. One aspect of identifying marine equipment is understanding how it differs from IT equipment in terms of operation and setup. To illustrate this, the collected data attributes, such as characteristics, are the equipment, different port numbers, operating system, IP address, equipment type, and manufacturer. Other characteristics may also be collected, as disclosed above. For the port number attribute, a list of common vulnerable network ports is considered. Several top open TCP ports include port 80 (Hypertext Transfer Protocol or HTTP), 23 (Telnet), 21 (File Transfer Protocol or FTP), 22 (Secure Shell or SSH), 25 (Simple Mail Transfer Protocol or SMTP), 445 (Microsoft SMB), and 53 (Domain), and UDP ports include port 139 (Netbios-SS), port 445 (Microsoft-DS), port 161 (Simple Network Management Protocol or SNMP), port 123 (Network Time Protocol or NTP), and so on. This initially resulted in 19 port numbers as attributes in the dataset, and this number will change as more information is acquired. These attributes are represented in the database with values of "1" or "0" based on whether the port is open or closed.
[0097] The accuracy of predictions depends on historical data, and missing values can affect the results, just as when filling in a dataset. Therefore, it's possible for some values to be missing. For example, historical data is scarce during system upgrades. Because this can lead to difficulties in predicting accuracy, these values are marked as "unknown" when Device Manager 100 lacks confidence. Auditors can review these values later if necessary, and the chances of the profiler tool misleading the auditor are reduced. Furthermore, the different types of data contained in the dataset can be considered. Port numbers are numeric, while operating system data and manufacturer data are alphanumeric categorical data. To ensure data consistency, encoder 114 can be used to encode categorical data into dummy variables, which can be optimized for classifier 116 in machine learning.
[0098] Classifier settings Classifiers, such as random forest classifiers, can be useful for classifying devices on ship networks for several reasons. First, due to the limited amount of data available for profiling ship systems, classifier 116 can work on finite datasets while still achieving high accuracy. While this works on IT / IoT devices, the disclosed embodiments can work on less conventional systems in ship bridging environments. For example, a random forest classifier is a supervised learning algorithm that is an aggregation of multiple decision trees. Decision tree classifiers are used in the disclosed embodiments because they can be simple and good classification algorithms. However, they can overfit. Overfitting occurs when a model tries to fit the training data to improve accuracy—that is, it tries to memorize the entire training data, making it unstable when new data is introduced. Random forests can correct this aspect of decision trees. During the process of building and splitting nodes in the trees, a random forest generates multiple decision trees based on random sample sizes and random numbers of features. The sum of the outputs of all the created decision trees is then computed to classify the data, thus eliminating any chance of bias and overfitting.
[0099] In addition to the physical hardware in the Cyber-SHIP lab, data can be collected from the configured network using virtual machines running Kali Linux OS or any other OS. This machine can virtually connect to the lab ship network automatically drawn by Topology Builder 106. The resulting topology can be used to initiate a network port scan using scanning tools (such as those in Feature Extractor 110) to identify hosts in the network and their communication protocols. The scanning tools can determine which devices are active and running. Once a ping scan is complete, an open port and service scan is performed on each device from the host list. The results obtained can be filtered and encoded into a dataset populated by the results from the scanning tools and the Topology Builder.
[0100] Model building To build and test accurate models, datasets can be divided into training and testing subsets. Training data is the initial set of data fed into the model to learn and find patterns among data. In other words, it's historical data used to teach the model to make accurate predictions. Test data is the set of data used to measure or validate the model's accuracy. It's unseen data that can be fed into the model to validate it. In an illustrative example, the `train test split()` function from the Sklearn library is used to split the dataset into training and testing sets in a 70:30 ratio. That is, 70% of the data is split into the training set, and 30% is reserved for testing. The model's input is attributes selected from the dataset, and the output is the "device" attribute. This is also useful when historical data is limited.
[0101] The random forest classifier is initially built with 100 n-estimates, where the number of n-estimates represents the number of decision trees to build before averaging all outputs and making predictions. Model fit measures how well a model performs on data similar to the training data, and models can be well-fitted, overfitted, or underfitted. A well-fitted model provides accurate predictions or outputs, while an overfitted model is too close to the training data, and an underfitted model is not a good fit at all. An example random state with a value of 42 can be provided as a seed for randomness to ensure that the split dataset is the same for each execution. Once the model has been trained and fitted, the model's accuracy score is obtained using the reserved test value and predicted value of the "device" attribute. Note that any specific numbers mentioned in different embodiments are for initial testing and any numbers can be used depending on the desired implementation.
[0102] Classification For each host found, a model is created in order of IP address and trained and fitted. Once the model is fitted and tuned using hyperparameters (described in more detail below), it is ready for profiling. When a new host is identified in network 102 using topology builder 106, the host's details and characteristics are identified and extracted. This information is written to dataset 112, with the "device" attribute value set to "virtual" or a similar value. This information is then encoded as a numerical value and presented as a list as input to the model. This input list is then fed into classifier 116 to make predictions about the device type, and the output is the value of the "device" attribute for dataset 112. Once this value is predicted, the "virtual" value in the data frame is replaced with the predicted output and then written to dataset 112. In this way, dataset 112 continues to grow, thus enabling learning. This process is repeated for all identified hosts, and at the end of the profiling of all devices, the average accuracy score of the model is calculated.
[0103] Results and findings from the analysis Assets can be profiled, and the results can be displayed in a way that technical auditors / testers and engineers / seafarers can understand. To facilitate this, the results can be automatically generated in PDF files. The experimental setup in the CyberSHIP lab has a bridged network with an average of 30 devices, varying by plus or minus 2 depending on the configuration. The domain addresses of the bridged network are entered into Device Manager 100, and the entire automated profiling process takes approximately 40-45 minutes to complete and generate a PDF report. The example machine executing the Device Manager 100 tool is a Kali Linux virtual machine (VM) with 2048 MB of basic memory, while the Windows machine hosting the Kali VM is a Microsoft Windows 11 OS with 32 GB of RAM. The results and analysis from the profiler are automatically generated and visually presented in the report using the graphs and charts described below. In creating this model, we learned that marine equipment differs from IT equipment in characteristics, such as different open ports for functions. We also found that certain equipment from specific companies has dedicated open ports for configuration and settings. The following results were obtained by analyzing the histograms generated by the profiler.
[0104] • The web configuration server for most devices is hosted on port 80, and 20 out of 29 devices have port 80 open.
[0105] • All USR IoT serial-to-IP converters have port 1501 open, which is assigned to the satellite data acquisition system 3, while Moxa Technologies converters have port 4000 open, as well as other ports for web configuration, such as port 80.
[0106] • Another interesting finding is that VDR has all the same open ports as a Windows PC, including port 3389 for the Remote Desktop Protocol and port 445 for Windows SMB. This means that VDR may behave similarly to a PC, and vulnerabilities and exploits applicable to Windows systems may also affect this system.
[0107] • All Moxa serial-to-IP converters have port 4900 for device firmware upgrades. Several firmware-related vulnerabilities for Moxa NPort devices are published in the Common Vulnerabilities and Exposures (CVE) database, including vulnerabilities that can be created and sent via the firmware upgrade port.
[0108] • Port 10010 of navigation devices such as AIS transponders and weather fax receivers manufactured by Furuno Electric is open for broadcasting AIS and NMEA messages.
[0109] Parameter tuning and verification Decision tree algorithms can easily fit to extreme cases, and random forest classifiers can mitigate these extremes to some extent by adding randomness, but may not completely eliminate them. To strike a balance between overfitting and underfitting, parameters affecting the model's accuracy and performance can be tuned and optimized—this is known as hyperparameter tuning. Some hyperparameters used for tuning include the n estimator (the number of decision trees in the forest), maximum depth (the maximum number of levels allowed in a tree), minimum split samples (the minimum number of samples required to split a node), minimum leaf samples (the minimum number of samples at a leaf node), and maximum features (the maximum number of features used when splitting a node). Choosing values for these hyperparameters can be done through experimentation, trying random values and default values to see how the model performs under these settings. This process can be tedious and time-consuming. Therefore, another approach is to use validation methods such as K-fold cross-validation and validation curves, which help identify the optimal hyperparameters for the model and diagnose fitting problems. Validation curves visually plot the performance metric or accuracy score of a given model against a range of selected parameters on training and test data. Analyzing this graph can help identify parameters that may be causing underfitting or overfitting of the model.
[0110] First, a model was built using 100 n-estimates (default values), with all other hyperparameters set to their default values. Using 100 n-estimates, the model achieved an average accuracy score of 0.988905, and the entire process of classifying all devices discovered in the network took 46 minutes. To further refine and tune the model to ensure higher accuracy scores, validation curves were plotted for different parameters, taking into account both the fitting problem and the unique characteristics of the devices. If both the training and validation score lines are low, the model may be underfitting, while if the training score is high and the validation score is low, the model may be overfitting. Therefore, the optimal values of the hyperparameters are likely the points where the distance between these lines is shortest and the accuracy is highest. The hyperparameters will be described in further detail below.
[0111] • n-estimate: The n-estimate defines the number of decision trees built for the forest. For cross-validation and to identify possible values for the n-estimate, validation curves were plotted using two cross-folds, considering n-estimate values of 10, 25, 50, 100, and 150. In both plots, the accuracy score of the cross-validation curves reaches its maximum at a value of 50 and then slowly decreases to a stable value. It can be noted that the accuracy score does not change after a certain n-estimate value, meaning that changing the n-estimate value does not affect the accuracy score and may indicate overfitting. The accuracy value changes at 25 n-estimates, thus this value is considered possible for the model and does not cause overfitting. Choosing a lower value for the n-estimate may reduce computation time but also affect the accuracy score.
[0112] • Maximum Depth: The maximum depth indicates the maximum number of levels a decision tree can have. If set to the default value, the model will split until a node reaches 100% purity or all its data belongs to the same category. To identify possible values for the maximum depth parameter, validation curves were plotted using two cross-folds, considering maximum depth values of 5, 10, 15, 20, and 25. After a maximum depth of ten, both training accuracy and validation scores increased sharply and then stabilized. Therefore, ten can be considered a possible value for the maximum depth parameter, as any larger value may not affect the model's accuracy.
[0113] • Minimum Leaf Samples: The minimum number of samples a leaf node must have is the minimum number of samples required to be at a leaf node. If, after splitting a node, an internal leaf node has fewer than this value, it will not be considered a leaf node, and its parent node will be considered a leaf node. This value helps limit the size of the tree and the number of levels it can grow. Validation curves using values 2, 4, 6, 8, and 10 are plotted. The default value for minimum leaf samples in Sklearn is 1, meaning that leaf nodes must have at least one sample. The plotted graph shows that the accuracy score decreases continuously as the minimum leaf sample value increases. An acceptable accuracy score can be achieved when the minimum leaf sample value is set to two, so this value can be chosen as a possible option.
[0114] • Minimum Splitting Samples: Similar to minimum leaf samples, minimum splitting samples represents the minimum number of samples required for a split to occur at a node. After splitting a node, if the number of samples in an internal leaf node is less than this value, the internal node will not be split. Otherwise, splitting will occur repeatedly until the node is clean. This parameter is also used to limit the growth of the tree and avoid overfitting. Validation curves using values 2, 4, 6, 8, and 10 are plotted. Similar to the previous plots, the validation curve for this parameter decreases as the hyperparameter value increases. The default value for minimum splitting samples in Sklearn is two, and this value is reserved as a possible value because the accuracy score is acceptable compared to using other values.
[0115] Using hyperparameter tuning to validate and optimize models is useful. A new random forest classifier model was created using hyperparameters derived from the tuning scheme. Its accuracy score was 0.9867, while the model with 100 estimators achieved an accuracy score of 0.98899. The new model completed the profiling process in an average of 37 minutes, while the first model with 100 estimators took 46 minutes. The results show that when the estimator value is reduced to 25, the accuracy drops to a minimum, but the process is completed much faster. Therefore, classifier models can be selected based on the model's purpose, considering both accuracy score and time, while avoiding fitting problems.
[0116] Visualization and Reporting Penetration testers typically generate security reports manually, highlighting the test environment, its inherent risks, vulnerabilities, and possible mitigation measures. Device Manager 100 automates this process, and the quality of information can be the same as, or even better than, the manual approach. Automation makes the job easier for penetration testers / auditors, allowing them to focus on other aspects of their work. Often, typical reports are highly technical and may be difficult for individuals with limited knowledge of the system to understand. Therefore, when deciding what information to include in maritime cybersecurity reports, Device Manager 100 considers its scope and audience, as the field of maritime cybersecurity is still relatively new. Reports that are visually comprehensive and include important facts about the ship's environment are more effective at conveying information. Furthermore, images and graphs can quickly communicate messages to non-cyber-aware audiences such as seafarers or ship engineers. To make reports more user-friendly, tables, graphs, bar charts, and pie charts can be used for presentation.
[0117] Figure 5This is an example illustration of an asset counting histogram 500 according to one possible embodiment. This graph shows the number of assets of each type. After profiling and forecasting all hosts with the "device" value in the category, the count of all assets within each device category is displayed as histogram 500. By using this histogram 500, auditors and engineers can verify the number of devices connected to the network based on the type of device. This is also useful for simultaneously showing changes in a vessel over the years. For example, the number of IoT devices added to older vessels may increase to improve surveillance and other capabilities.
[0118] Figure 6 This is an example illustration of a port count distribution diagram 600 drawn according to one possible embodiment. Figure 600 shows an overview of the number of devices with specified ports open. Figure 600 assists in visualizing and reviewing the most frequently open ports in devices, as well as unintentionally open ports that may be opened for testing or auditing purposes. For example, Figure 600 shows the count of open ports in each device, where the X-axis represents the port number and the Y-axis depicts the number of devices with open ports.
[0119] Figure 7 This is an example illustration of an open port heatmap 700 according to one possible embodiment. A heatmap is a visual representation of data using different colors. This color-coding technique helps users quickly and easily understand complex information. When combined with appropriate color scales and used based on similarity, users will be able to see new patterns and structures that are not visible otherwise. The open port heatmap illustrates which ports are open on each device. The X-axis of map 700 represents different ports, while the Y-axis depicts previously analyzed assets and predicted device types. "Open" ports in a device are encoded with a specific shade or color, while "closed" ports are encoded with a different shade or color. Therefore, the characteristics of various devices can be understood by relating similarities between device types and between devices manufactured by the same company.
[0120] Additional Examples In one possible embodiment, the topology builder 106 can be static, meaning that users may not be able to interactively edit the topology, but can review and evaluate it based on a graph. The topology builder 106 creates a topology for a given time period and then generates a network graph with the configuration at the time of test execution. If the user repeats the test at some later time, the topology builder is executed again, generating a new graph based on that configuration. Furthermore, if devices in the topology graph are not connected to any other devices and are not communicating, then these devices may be missing from the topology graph. One solution to this problem is to obtain an Address Resolution Protocol (ARP) table containing a list of all devices, then plot them as single idle nodes in the graph, and probe open ports to gather more information. This embodiment can be used in ships and other fields where significant changes often occur before or after scheduled modifications or maintenance, meaning that topology updates can be planned in advance.
[0121] In one possible embodiment regarding the profiling phase, a public dataset on maritime equipment may be lacking. This can be addressed by accessing a hardware testbed equipped with vessel systems in a ship bridging configuration. Therefore, data can still be collected from real systems to construct this dataset, which can then be validated. During the initial phase, even if the classification process is automated, human supervisors can validate the collected data to ensure accuracy. Once the dataset is created, the model profils the hosts discovered during each test execution, returning the results to the dataset, allowing it to continuously grow. Including a large amount of data from a variety of devices can be useful, depending on the number of devices available to the researcher. This can be addressed by collecting data from real-time networks on ships and in other fields.
[0122] In one possible implementation, the limited amount of available data can introduce overfitting or underfitting into the classifier model. Random forest classifiers can be used instead of decision trees by incorporating randomness and reducing fitting, but they may not completely eliminate it. In random forest models, there may be a trade-off between accuracy and computation time. When constructing a forest using a large number of trees, completing the model may take longer, but the accuracy score may increase as a result. Considering that accuracy values stop improving after a certain threshold, fewer decision trees than that threshold can be used. Since trees are sensitive to parameter values, other parameters can also be tuned and optimized. When there are many devices in the configuration, the profiler can be quite time-consuming; therefore, reducing computational power and resources may be a viable solution in such cases.
[0123] in conclusion The embodiments provide Device Manager 100, an asset profiler that users can use to manage their asset inventory and comply with regulations and requirements, such as those in IACS UR 26 and UR 27. In embodiments with a given network configuration, the tool automatically constructs a topology map of communication flows and uses a random forest classifier algorithm to intelligently identify devices or assets on the bridge. To ensure that potential users (i.e., seafarers and engineers) understand the results, Device Manager 100 also provides detailed information about (one or more) devices and their (one or more) networks. Shipboard crew and engineers can use the graphs and diagrams generated by the tool to better understand their networks and systems, manage assets more efficiently, and better inform maintenance and security efforts. Generally, this domain-focused tool is more accurate in classifying bridge equipment compared to similar works designed for IT / IoT environments. This is likely due to the large number of customized and novel system solutions in maritime space and other unique domains, requiring the classifier to handle more unique attributes. The disclosed embodiments are useful in the broader maritime, aviation, manufacturing, cyber-physical topics of cybersecurity, and other fields.
[0124] Security testers and auditors can use this tool on board vessels to gather situational awareness information about the systems and environments in which they operate. Device Manager 100 reduces time and effort required for manual system inspections that could otherwise take days instead of minutes, especially when panels need to be removed to access hidden components. Device Manager 100 provides an automated, non-intrusive tool that accelerates the testing process and requires less specialized maritime knowledge. Automated asset detection and classification can be even more beneficial in future work, as penetration testers can leverage these capabilities to build specific exploits for the systems and networks they target. The security testing framework can provide exploit modules for IT devices and OT systems such as SCADA components. A detailed vessel-based asset inventory can also help select the correct and most appropriate exploit or test type for the devices. A vessel-based ethical penetration testing tool can be built to extend the embodiments of this disclosure. This can also aid in network risk assessments, allowing those responsible for devices / assets to identify operationally critical devices / assets and ensure they are updated and patched, and that appropriate security controls are in place.
[0125] Hardware Overview According to one possible embodiment, the techniques described herein are implemented by one or more dedicated computing devices. These dedicated computing devices may be hardwired to execute these techniques; or they may include digital electronic devices, such as one or more application-specific integrated circuits (ASICs) or field-programmable gate arrays (FPGAs), which are persistently programmed to execute these techniques; or they may include one or more general-purpose hardware processors programmed to execute these techniques according to program instructions in firmware, memory, other storage devices, or combinations thereof. Such dedicated computing devices may also combine custom hardwired logic, ASICs, or FPGAs with custom programming to implement these techniques. These dedicated computing devices may be desktop computer systems, portable computer systems, handheld devices, networking devices, or any other devices that include hardwired and / or program logic to implement these techniques.
[0126] For example, Figure 8 This is a block diagram illustrating a computer system 800 on which various aspects of the illustrative embodiments may be implemented. The computer system 800 includes a bus 802 or other communication mechanism for conveying information, and a hardware processor 804 coupled to the bus 802 for processing information. The hardware processor 804 may be, for example, a general-purpose microprocessor.
[0127] Computer system 800 also includes main memory 806, such as random-access memory (RAM) or other dynamic storage devices, coupled to bus 802, for storing information and instructions to be executed by processor 804. Main memory 806 can also be used to store temporary variables or other intermediate information during the execution of instructions to be executed by processor 804. When such instructions are stored in non-transitory storage media accessible to processor 804, computer system 800 becomes a special-purpose machine customized to execute the operations specified in the instructions.
[0128] The computer system 800 also includes a read-only memory (ROM) 808 or other static storage device coupled to the bus 802 for storing static information and instructions for the processor 804. A storage device 810, such as a disk, optical disk, or solid-state drive, coupled to the bus 802, is provided for storing information and instructions.
[0129] Computer system 800 may be coupled via bus 802 to display 812 for displaying information to the computer user. Input device 814, including alphanumeric and other keys, is coupled to bus 802 for conveying information and command selection to processor 804. Another type of user input device is cursor control device 816, such as a mouse, trackball, touchscreen, touchpad, and / or cursor arrow keys, for conveying directional information and command selection to processor 804 and for controlling cursor movement on display 812. Such input devices typically have two degrees of freedom on two axes—a first axis (e.g., x) and a second axis (e.g., y)—allowing the device to specify a position in a plane.
[0130] Computer system 800 may implement the techniques described herein using custom hardwired logic, one or more ASICs or FPGAs, firmware, and / or program logic, which, in combination with the computer system, enable or program the computer system 800 to be a special-purpose machine. According to one embodiment, computer system 800 executes the techniques described herein in response to processor 804 executing one or more sequences of one or more instructions contained in main memory 806. Such instructions may be read into main memory 806 from another storage medium (e.g., storage device 810). Execution of the sequence of instructions contained in main memory 806 causes processor 804 to perform the process steps described herein. In alternative embodiments, hardwired circuitry may be used in place of or in combination with software instructions.
[0131] The term "storage medium" is used herein to refer to any non-transitory medium that stores data and / or instructions that cause a machine to operate in a particular manner. Such storage media may include non-volatile media and / or volatile media. Non-volatile media include, for example, optical discs, magnetic disks, or solid-state drives, such as storage device 810. Volatile media include dynamic memory, such as main memory 806. Common forms of storage media include, for example, floppy disks, flexible disks, hard disks, solid-state drives, magnetic tape or any other magnetic data storage medium, CD-ROMs, any other optical data storage media, any physical medium with a perforated pattern, RAM, PROMs and EPROMs, FLASH-EPROMs, NVRAMs, any other memory chips or magnetic tape cassettes and / or any other storage medium.
[0132] Storage media and transmission media are distinct, but can be used in conjunction with each other. Transmission media participate in transferring information between storage media. For example, transmission media include coaxial cables, copper wires, and optical fibers, including the conductors that form bus 802. Transmission media can also take the form of sound waves or light waves, such as those generated during radio wave and infrared data communication.
[0133] Various forms of media can participate in delivering one or more sequences of one or more instructions to processor 804 for execution. For example, instructions may initially be carried on a disk or solid-state drive of a remote computer. The remote computer may load the instructions into its dynamic memory and transmit the instructions via a modem over a telephone line or over a network. A receiver local to computer system 800, such as a modem, can receive data. In one example, the receiver may use an infrared transmitter to convert the data into an infrared signal. An infrared detector can receive the data carried in the infrared signal, and appropriate circuitry can place the data on bus 802. Bus 802 transports the data to main memory 806, where processor 804 retrieves the instructions and executes them. Instructions received by main memory 806 may optionally be stored on storage device 810 before or after execution by processor 804.
[0134] Computer system 800 also includes a communication interface 818 coupled to bus 802. Communication interface 818 provides bidirectional data communication coupling with network link 820 connected to local network 822. For example, communication interface 818 may be an Integrated Services Digital Network (ISDN) card, cable modem, satellite modem, or modem to provide data communication connectivity with a corresponding type of telephone line. As another example, communication interface 818 may be a local area network (LAN) card to provide data communication connectivity to a compatible LAN. Wireless links, such as links to wireless local area networks (WLANs) or cellular networks, may also be implemented. In any such implementation, communication interface 818 transmits and receives electrical signals, electromagnetic signals, radio signals, optical signals, and / or other signals carrying digital data streams representing various types of information.
[0135] Network link 820 typically provides data communication with other data devices over one or more networks. For example, network link 820 can provide connectivity to a host 824 or a data device operating through a local network 822 and an Internet Service Provider (ISP) 826. The ISP 826 then provides data communication services through a global packet data communication network now commonly referred to as the “Internet” 828. Both local network 822 and Internet 828 use electrical, electromagnetic, or optical signals carrying digital data streams. Signals through various networks, as well as signals on network link 820 and signals through communication interface 818 (which carry digital data to and from computer system 800), are example forms of transmission media.
[0136] Computer system 800 can send messages and receive data, including program code, via one or more networks, network links 820, and communication interfaces 818. In the Internet example, server 830 can send the code of a requested application via the Internet 828, ISP 826, local network 822, and communication interface 818.
[0137] The received code may be executed by processor 804 upon receipt and / or stored in storage device 810 or other non-volatile storage device for later execution.
[0138] Software Overview Figure 9 This is a block diagram of a basic software system 900 that can be used to control the operation of computer system 800. The software system 900 and its components, including their connections, relationships, and functions, are intended to be exemplary only and are not intended to limit the implementation of one or more example embodiments. Other software systems suitable for implementing one or more example embodiments may have different components, including components with different connections, relationships, and functions.
[0139] A software system 900 is provided to guide the operation of the computer system 800. The software system 900 may be stored in system memory (RAM) 806 and on a fixed storage device (e.g., hard disk or flash memory) 810, and includes a kernel or operating system (OS) 910.
[0140] OS 910 manages low-level aspects of computer operations, including managing process execution, memory allocation, file input and output (I / O), and device I / O. One or more applications (represented as 902A, 902B, 902C...902N) can be "loaded" (e.g., transferred from fixed storage device 810 to memory 806) for execution by system 900. Applications or other software intended for use on computer system 800 can also be stored as a set of downloadable computer-executable instructions, for example, for downloading and installing from an Internet location (e.g., a web server, app store, or other online service).
[0141] Software system 900 includes a graphical user interface (GUI) 915 for receiving user commands and data graphically (e.g., "point and click" or "touch gestures"). This input can then be manipulated by system 900 according to instructions from operating system 910 and / or (one or more) applications 902. GUI 915 also displays the results of actions from OS 910 and (one or more) applications 902, allowing the user to provide additional input or terminate the session (e.g., log off).
[0142] OS 910 can execute directly on the bare hardware 920 of computer system 800 (e.g., one or more processors 804). Alternatively, a super-supervisor or virtual machine monitor (VMM) 930 can be plugged between the bare hardware 920 and OS 910. In this configuration, VMM 930 acts as a software "buffer" or virtualization layer between OS 910 and the bare hardware 920 of computer system 800.
[0143] VMM 930 instantiates and runs one or more virtual machine instances (“guest machines”). Each guest machine includes a “guest” operating system, such as OS 910, and one or more applications, such as Application(s)902, which are designed to run on the guest operating system. VMM 930 presents a virtual operating platform for the guest operating system and manages the execution of the guest operating system.
[0144] In some cases, VMM 930 can allow a guest operating system to run as if it were running directly on the bare hardware 920 of the computer system 800. In these cases, the same version of the guest operating system configured to run directly on the bare hardware 920 can also run on VMM 930 without modification or reconfiguration. In other words, in some situations, VMM 930 can provide complete hardware and CPU virtualization for a guest operating system.
[0145] In other cases, the guest operating system can be specifically designed or configured to run on VMM 930 for improved efficiency. In these cases, the guest operating system is "aware" that it is running on the virtual machine monitor. In other words, in some situations, VMM 930 can provide paravirtualization to the guest operating system.
[0146] A computer system process includes allocated hardware processor time and allocated memory (physical and / or virtual). The allocated memory is used to store instructions executed by the hardware processor, data generated by the execution of those instructions, and / or to store hardware processor state (e.g., register contents) during allocated hardware processor time when the computer system process is not running. Computer system processes run under the control of the operating system and can also run under the control of other programs executing on the computer system.
[0147] cloud computing The term "cloud computing" is generally used in this article to describe a computing model that enables on-demand access to a shared pool of computing resources, such as computer networks, servers, software applications, and services, and allows for the rapid provisioning and release of resources with minimal management effort or service provider interaction.
[0148] Cloud computing environments (sometimes called cloud environments, or the cloud itself) can be implemented in various ways to best meet different requirements. For example, in a public cloud environment, the underlying computing infrastructure is owned by an organization that provides its cloud services to other organizations or the general public. In contrast, a private cloud environment is generally intended for use by a single organization or for internal use only. A community cloud is designed to be shared by several organizations within a community; while a hybrid cloud includes two or more types of clouds (e.g., private, community, or public clouds) that are combined through data and application portability.
[0149] Generally, cloud computing models enable the shift of some responsibilities, previously provided by an organization's own IT departments, to service layers within the cloud environment for use by consumers (which, depending on the public / private nature of the cloud, can be internal or external to the organization). Depending on the specific implementation, the precise definitions of the components or features provided by each cloud service layer, or within each cloud service layer, may vary, but common examples include: Software as a Service (SaaS), where consumers use software applications running on cloud infrastructure, while the SaaS provider manages or controls the underlying cloud infrastructure and applications; and Platform as a Service (PaaS), where consumers use software programming languages and development tools supported by the PaaS provider to develop, deploy, and otherwise control their own applications, while the PaaS provider manages or controls other aspects of the cloud environment (i.e., everything below the runtime execution environment). Infrastructure as a Service (IaaS) allows consumers to deploy and run any software application and / or configure processing, storage, networking, and other basic computing resources, while the IaaS provider manages or controls the underlying physical cloud infrastructure (i.e., everything below the operating system layer). Database as a Service (DBaaS) allows consumers to use database servers or database management systems running on cloud infrastructure, while the DBaaS provider manages or controls the underlying cloud infrastructure, applications, and servers, including one or more database servers.
[0150] In the foregoing specification, embodiments of the invention have been described with reference to numerous specific details, which may vary depending on the implementation. Therefore, the specification and drawings should be considered exemplary rather than restrictive. The sole and exclusive reference to the scope of the invention, and what the applicant intends the scope of the invention to be, is the literal and equivalent scope of the set of claims in the specific form claimed in this application, including any subsequent corrections.
Claims
1. A method comprising: Perform a first query on the devices in the operational technology equipment network, wherein the operational technology equipment is operated using physical processes; Based on the first query, the first characteristic information of the device is collected; Generate a first instance of a device record that includes the first feature information; A second query is selected based on a first instance of the device record, the second query being selected from a query set of characteristics of operating technology devices operating in a specific physical domain; Execute the second query on the device; The second feature information of the device is collected based on the second query; The device record is updated based on the second feature information to generate an updated device record; and The type of the device is predicted based on the updated device records.
2. The method of claim 1, further comprising: Predict the first type of the device based on the first instance recorded by the device; and The first instance of the device record is updated based on the predicted first type of the device. Selecting the second query includes selecting the second query based on a first instance of the updated device record.
3. The method of claim 2, wherein Predicting the first type of the device includes predicting that the first type of the device is an intermediate device that converts a first communication protocol of the connected devices into a second communication protocol. Generating the second query includes generating the second query based on the first communication protocol and the second communication protocol. Collecting the second feature information includes collecting the second feature information of the connected device based on the second query. Updating the device record includes storing connection device records based on the second feature information, and Predicting the type of the device includes predicting the type of the connected device based on the connected device record.
4. The method of claim 1, further comprising selecting the first query from a query set of features of operating technology equipment operating in the particular physical domain. in, Each instance recorded by the device includes a predetermined set of features of the operating technology device operating in the specific physical domain, and Each instance recorded by the device indicates whether the features of the predetermined feature set exist.
5. The method of claim 4, wherein, The features in the feature set include protocol features corresponding to the communication protocol and multiple port features, each of which corresponds to a different port number.
6. The method of claim 1, wherein, The second query is a query of the communication protocol of the operating technology equipment.
7. The method of claim 6, wherein, The specific physical domain in question is ships. The device is a ship operation technology device coupled to the bridge of the ship, and The second query is a query of the communication protocol of the ship operation technology equipment.
8. The method of claim 7, wherein, The communication protocols for the ship operation technology equipment include the National Marine Electronics Association (NMEA) protocol or the Automatic Identification System (AIS) protocol.
9. The method of claim 1, wherein, The query set is constructed to inquire about the characteristics of operating technology devices that operate in a specific physical domain.
10. The method of claim 1, further comprising: Create a network communication flow diagram for the network of operating technology devices, wherein the network communication flow diagram includes: A first network connection using a first communication protocol between a first device and an intermediate device, and A second network connection using a second communication protocol between the intermediate device and the second device, wherein the second communication protocol is different from the first communication protocol.
11. The method of claim 1, wherein, The prediction is performed by a classifier trained on data collected from a set of cyber-physical testbeds, which are based on marine hardware equipment configured as a vessel bridge.
12. A network device manager for a marine vessel bridge, the device manager comprising: Feature extractor, the feature extractor: A first query is performed on the devices in the network of operational technology devices coupled to the marine vessel bridge, wherein the operational technology devices are operated in conjunction with the physical processes of the marine vessel. Information about the characteristics of the device is collected based at least on the first query, and Generate a device record that includes the features in a feature set; A database storing the records of the device; An encoder that generates the first encoded data based on the feature set; as well as A classifier that predicts the type of device based on the first encoded data by comparing it with stored device feature data. The device manager updates the feature set in the device record with information about the predicted device type. Wherein, the feature extractor: A second query is selected based on the predicted equipment type information. This second query is selected from a query set of characteristics of maritime vessel operation technology equipment. The second query is executed on the device, and Information on additional characteristics of the device is collected based on the second query. The device manager updates the feature set in the device record with the information of the additional features. The encoder generates second encoded data based on a feature set including the additional features, and The classifier predicts the updated type of the device based on the second encoded data by comparing the second encoded data with stored device feature data.
13. One or more storage media, the media storing instructions that, when executed by one or more processors, cause to perform the method as described in any one of claims 1-11.