Method for automatically identifying internet of things protocol

By performing in-depth analysis and training on the communication code streams of IoT devices using a large language model, a protocol recognition model is generated, solving the problem of identifying new protocols in IoT systems. This enables efficient and accurate automatic protocol identification and parsing, adapting to the diversity and rapid changes of protocols, and reducing operation and maintenance costs.

CN120956819APending Publication Date: 2025-11-14JIANGXI FASHION TECH
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202511493026.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-20
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Existing technologies are unable to effectively identify new protocols when faced with the diversity and rapid changes of IoT devices, resulting in limited scalability and adaptability of IoT systems. Traditional methods are inefficient and prone to errors.

Method used

A large language model is used to perform in-depth analysis of the communication code streams of IoT devices. Through data collection, cleaning, labeling and training, a protocol recognition model is generated, and parsing scripts and documents are automatically generated to support the learning and recognition of new protocols.

Benefits of technology

It improves the efficiency and accuracy of protocol identification, reduces manual intervention, adapts to the diversity and rapid changes of protocols, supports the identification of new protocols, reduces operation and maintenance costs, and improves system deployment and operation and maintenance efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120956819A_ABST
    Figure CN120956819A_ABST
Patent Text Reader

Abstract

The invention discloses a method for automatically identifying an internet of things protocol, which comprises the following steps of: establishing a data collection module, collecting a multi-dimensional data source of internet of things equipment, cleaning the collected multi-dimensional data source to remove noise data, converting the cleaned data into a uniform standard format, and carrying out multi-dimensional data fusion training, so as to automatically identify the internet of things protocol. The model can learn different protocol structures and semantics, not only can identify standard protocols such as Modbus and CAN, but also can process customized protocols through adaptive learning, and also supports learning identification of new protocols and rapid change requirements of adaptive protocols, and the method can automatically generate analysis scripts and documents, directly provide support for equipment interconnection and data analysis, shorten the process period, and improve the efficiency. The deployment operation and maintenance efficiency of the Internet of Things system is improved, new protocol data can be supplemented, data processing and model training process optimization models can be repeated, the architecture does not need to be transformed, the operation and maintenance cost is reduced, it is ensured that the system adapts to the latest protocol, and the long-term development requirement is met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of Internet of Things (IoT) technology, and specifically relates to a method for automatic identification of IoT protocols. Background Technology

[0002] With the rapid development of IoT technology, various IoT devices are widely used in numerous fields such as industry, agriculture, smart homes, and healthcare, greatly promoting the intelligent transformation of various industries. Communication between IoT devices relies on various protocols, including standardized protocols such as Modbus, CAN, MQTT, and HTTP, as well as a large number of protocols customized by different manufacturers according to their own needs. According to relevant data, by the end of 2024, the number of connected IoT devices worldwide had exceeded 50 billion. With such a huge number of devices, the complexity of their protocols has also increased dramatically.

[0003] In the Internet of Things (IoT) field, existing technologies primarily rely on traditional manual or semi-automated methods. For example, relying on protocol documents and manual parsing often depends on the experience and knowledge base of technical personnel, resulting in low efficiency and a high risk of errors, especially when dealing with a large number of devices and multiple protocols, where manual parsing becomes a bottleneck. Some systems use predefined rules or protocol templates to identify and parse device communication protocols. These rules are typically written based on common protocol formats (such as Modbus and CAN), and can perform preliminary protocol identification based on the code stream format, data fields, and message patterns. However, they often struggle when encountering complex, non-standard device protocols. Furthermore, some existing technologies attempt to parse protocols by pre-installing protocol stack software libraries. However, this static parsing method proves inadequate in the face of constantly emerging new or unknown protocols, failing to properly identify them and severely limiting the scalability and adaptability of IoT systems. For example, in smart home scenarios, newly launched smart appliances may use innovative communication protocols that traditional protocol stack software libraries cannot recognize and parse, preventing these new devices from successfully integrating into existing smart home systems. Therefore, we need to provide a method for automatic IoT protocol identification. Summary of the Invention

[0004] The purpose of this invention is to provide a method for automatic identification of IoT protocols. This method utilizes a large language model to perform deep analysis and understanding of device communication streams, automatically identifying protocol types and structures. This significantly reduces manual intervention and improves the efficiency and accuracy of protocol identification. By training on a large amount of device protocol data, the model can identify and understand various common and customized protocols, addressing the diversity and rapid changes in device protocols. It also supports the learning and identification of new and unknown protocols. This solves the problem mentioned in the background art, where existing technologies attempt to parse protocols by pre-installing protocol stack software libraries. However, this static parsing method is inadequate in the face of constantly emerging new or unknown protocols, failing to properly identify them and severely limiting the scalability and adaptability of IoT systems.

[0005] To achieve the above objectives, the present invention adopts the following technical solution: a method for automatic identification of Internet of Things (IoT) protocols, comprising the following steps: A data collection module is established to collect multi-dimensional data sources from IoT devices. The collected multi-dimensional data sources are cleaned to remove noisy data, and the cleaned data is converted into a unified standard format. Key features related to protocol parsing are extracted from the data, and a standardized input dataset is generated. The protocol documents, log files, and protocol parsing scripts in the generated standardized input dataset are labeled and used for data annotation and training to form structured training data. The structured training data is divided into a training set and a validation set, and a pre-trained large language model is selected as the base model. The base model is then fine-tuned using the training set to obtain the trained protocol recognition model. The standardized input dataset sent in real time by IoT devices is input into the trained protocol recognition model. Based on the recognition and parsing results of the protocol recognition model, the corresponding protocol parsing script and parsing document are automatically generated. The new device protocol data is added to the multidimensional data source. The data collection, data labeling and training operations described above are repeated to iteratively optimize the protocol recognition model in order to adapt to changes in protocol format and the recognition requirements of new protocols.

[0006] Preferably, the pre-trained large language model is a BERT model or a GPT series model; the training set is used to fine-tune the basic model.

[0007] Preferably, the protocol identification model analyzes and infers the communication data stream, identifies the protocol type to which the data stream belongs, and parses the field content, data type, checksum, and data length information.

[0008] Preferably, the base model is fine-tuned during training, with the goal of minimizing the loss function. The training effect of the model is verified and the model parameters are optimized using the validation set. The loss function is either the mean squared error function or the cross-entropy function, resulting in a trained protocol recognition model.

[0009] Preferably, the tagging process includes tagging the protocol's message header, field content, data type, checksum, and data length information to form structured training data.

[0010] Preferably, the multidimensional data source includes at least device communication code stream data, protocol documents, log files and protocol parsing scripts, and covers standard protocols and customized protocols. The standard protocols include Modbus protocol, CAN protocol, MQTT protocol and HTTP protocol.

[0011] Preferably, the key features related to protocol parsing include at least the protocol message header, field content, and data length information, and the data format after data cleaning and format conversion is any one or more combinations of JSON, XML, or CSV formats.

[0012] Preferably, the protocol parsing script is a JavaScript script or a Lua script, and the generated parsing document is used for subsequent data analysis and device interconnection operations of IoT devices.

[0013] Preferably, when the protocol identification model is iteratively optimized, the supplementary new device protocol data needs to undergo data cleaning, format conversion and feature extraction operations to form standardized data before being integrated into a multi-dimensional data source.

[0014] Preferably, the data collection module supports both the collection of real-time communication data streams from IoT devices and the offline collection of stored historical protocol documents, log files, and protocol parsing scripts.

[0015] Technical effects and advantages of the present invention: The method for automatic identification of Internet of Things protocols proposed in this invention has the following advantages compared with the prior art: 1. This invention performs in-depth analysis and understanding of device communication code streams through a large language model, automatically identifies protocol types and their structures, greatly reduces manual intervention, and improves the efficiency and accuracy of protocol identification. By training on a large amount of device protocol data, the model can identify and understand a variety of common and customized protocols, cope with the diversity and rapid changes of device protocols, and support the learning and identification of new and unknown protocols. 2. This invention utilizes multi-dimensional data fusion training, enabling the model to learn different protocol structures and semantics. It can recognize standard protocols such as Modbus and CAN, and also handle customized protocols through adaptive learning. Furthermore, it supports the learning and recognition of new protocols, adapting to rapidly changing protocol requirements. This method can automatically generate parsing scripts and documents, directly supporting device interconnection and data analysis, shortening process cycles, improving the deployment and maintenance efficiency of IoT systems, supplementing new protocol data, optimizing repetitive data processing and model training processes, eliminating the need for architecture modifications, reducing maintenance costs, ensuring system compatibility with the latest protocols, and meeting long-term development needs.

[0016] Other features and advantages of the invention will be set forth in the following description, and will be apparent in part from the description, or may be learned by practicing the invention. The objects and other advantages of the invention can be realized and obtained by means of the structures pointed out in the description and the drawings. Attached Figure Description

[0017] Figure 1 This is a flowchart of the steps of the present invention. Detailed Implementation

[0018] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The specific embodiments described herein are merely used to explain the present invention and are not intended to limit the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0019] This invention provides, for example Figure 1 The method for automatic identification of IoT protocols, as shown, includes the following steps: A data collection module is established to collect multi-dimensional data sources from IoT devices. The collected multi-dimensional data sources are cleaned to remove noisy data, and the cleaned data is converted into a unified standard format. Key features related to protocol parsing are extracted from the data, and a standardized input dataset is generated. The protocol documents, log files, and protocol parsing scripts in the generated standardized input dataset are labeled and used for data annotation and training to form structured training data. The structured training data is divided into a training set and a validation set, and a pre-trained large language model is selected as the base model. The base model is then fine-tuned using the training set to obtain the trained protocol recognition model. The standardized input dataset sent in real time by IoT devices is input into the trained protocol recognition model. Based on the recognition and parsing results of the protocol recognition model, the corresponding protocol parsing script and parsing document are automatically generated. The new device protocol data is added to the multidimensional data source. The data collection, data labeling and training operations described above are repeated to iteratively optimize the protocol recognition model in order to adapt to changes in protocol format and the recognition requirements of new protocols.

[0020] The pre-trained large language model is either a BERT model or a GPT series model; the base model is fine-tuned using the training set. Specifically, when fine-tuning the basic model using the training set, the goal is to learn the structural features, field meanings, encoding methods, and semantic information of IoT protocols. The training set comes from multi-dimensional data sources that have been labeled—including communication code stream data of IoT devices, protocol documents, log files, and structured data corresponding to protocol parsing scripts. During the fine-tuning process, by inputting message samples of different protocol formats, the basic model gradually masters the differentiated features of standard protocols (such as Modbus, CAN, MQTT, and HTTP protocols) and customized protocols, and ultimately has the ability to identify and parse IoT protocols.

[0021] The protocol identification model analyzes and infers the communication data stream, identifies the protocol type of the data stream, and parses the field content, data type, checksum, and data length information. Specifically, when the protocol identification model analyzes and infers communication data streams, it relies on the structural features, field meanings, encoding methods, and semantic information of different protocols (including standard protocols and customized protocols) learned during the model training phase. Standard protocols include Modbus, CAN, MQTT, and HTTP. The communication data streams encompass various types of raw communication data sent in real-time by IoT devices. The model first identifies the protocol type of the data stream through real-time decomposition and feature matching, and then further parses the field content, data type, checksum, and data length information. The parsed information serves as the basis for automatically generating corresponding protocol parsing scripts (such as JavaScript or Lua scripts) and parsing documents, providing accurate data support for subsequent data analysis and device interconnection operations of IoT devices.

[0022] The base model is fine-tuned during training, with the goal of minimizing the loss function. The training effect of the model is verified and the model parameters are optimized using the validation set. The loss function is either the mean squared error function or the cross-entropy function, resulting in a trained protocol recognition model. Specifically, the training data used is structured training data that has undergone preprocessing and labeling. This data originates from multi-dimensional data sources of IoT devices, including device communication code stream data, protocol documents, log files, and protocol parsing scripts. It covers standard protocols such as Modbus, CAN, MQTT, and HTTP, as well as custom protocols from various manufacturers. During training, the goal is to minimize the loss function, which is either the mean squared error function or the cross-entropy function. The model's internal parameters are continuously adjusted by calculating the deviation between the protocol recognition results output by the model and the labeled protocol structure (including message headers, field content, data types, etc.) in the training data. At the same time, the model's training effect is verified in stages through the validation set. After each verification, the model parameters are further optimized based on the verification results (such as protocol recognition accuracy and field parsing accuracy). If the verification finds that the model's recognition accuracy for a specific protocol (especially a custom protocol) is insufficient, training samples for the corresponding protocol are added and fine-tuned until the model's protocol recognition and parsing performance in the validation set meets the preset threshold. Finally, a protocol recognition model with the ability to stably recognize multiple IoT protocols is obtained.

[0023] The tagging process includes tagging the protocol's message header, field content, data type, checksum, and data length information to form structured training data; Specifically, the labeling process uses protocol documents, log files, and protocol parsing scripts from the standardized input dataset as processing objects. The protocol documents include protocol description documents for IoT devices provided by various manufacturers; the log files cover the original communication stream logs and parsed protocol logs recorded by the IoT platform's acquisition program; and the protocol parsing scripts include JavaScript and Lua scripts on the IoT platform. During the labeling operation, the protocol message header must be accurately marked according to preset labeling rules, including key identifier fields such as protocol identifier and version information, field content, the specific meaning and value range of data fields, data types (e.g., integer, floating-point, character), checksums (including the corresponding numerical values ​​and checksum rules), and data length information, including the overall message length and the independent length of each field. Through this labeling process, unstructured protocol-related data is transformed into structured training data containing clear protocol feature labels. This structured training data must simultaneously cover data samples of standard protocols such as Modbus, CAN, MQTT, and HTTP, as well as custom protocols from various manufacturers, providing high-quality training data with clear feature labels for subsequent fine-tuning of the basic model.

[0024] The multidimensional data source includes at least device communication code stream data, protocol documents, log files and protocol parsing scripts, and covers standard protocols and customized protocols. The standard protocols include Modbus protocol, CAN protocol, MQTT protocol and HTTP protocol. Specifically, the device communication stream data consists of raw communication data packets generated during real-time interaction between IoT devices, covering various communication data such as inter-device command transmission and data feedback; the protocol documents are official protocol specification documents provided by various IoT device manufacturers, containing core information such as protocol format specifications, field definitions, and communication processes; the log files come from the log system of the IoT platform's data acquisition program, including both unparsed raw stream logs and protocol data logs that have undergone preliminary parsing; and the protocol parsing scripts are executable scripts used in the IoT platform to parse specific protocols, specifically including JavaScript and Lua scripts. These multi-dimensional data sources cover both standard and custom protocols. Standard protocols include Modbus (commonly used for communication in industrial automation equipment), CAN (often used for interconnection of devices in automotive electronics and industrial control), MQTT (suitable for remote data transmission between IoT devices), and HTTP (commonly used for data interaction between IoT devices and cloud platforms). Custom protocols are non-standardized communication protocols developed by various manufacturers based on the functional characteristics of their own devices. By collecting multi-dimensional data covering both types of protocols, comprehensive and diverse data support is provided for subsequent data preprocessing and model training, ensuring that the trained protocol recognition model has broad protocol adaptability.

[0025] The key features related to protocol parsing include at least the protocol message header, field content, and data length information. The data format after data cleaning and format conversion is any one or more combinations of JSON, XML, or CSV formats. Specifically, the protocol message header is used to quickly locate the core identifiers of the protocol, such as the protocol type identifier and version number field, providing a basis for the preliminary judgment of the protocol type. The field content covers the specific values ​​or instruction information corresponding to each functional module in the protocol data, which is the core of parsing the device's communication intent. The data length information includes the overall message length and the independent length of each field, which is used to verify data integrity and decompose the boundaries of different fields.

[0026] During the data cleaning and format conversion process, the collected multidimensional data sources, including device communication code stream data, protocol documents, log files, and protocol parsing scripts, first undergo noise removal operations. This includes filtering invalid and redundant data and correcting errors or missing values ​​generated during data transmission to ensure data accuracy. Then, the cleaned data is converted into any combination of JSON, XML, or CSV formats. JSON format is suitable for storing highly structured protocol field data, facilitating rapid feature extraction by the model. XML format is suitable for preserving hierarchical relationships in protocol documents, such as the hierarchical relationship between protocol modules and subfields, adapting to the structural parsing needs of complex protocols. CSV format facilitates batch processing of massive code stream data in log files, improving data preprocessing efficiency. Ultimately, a standardized input dataset with a unified format and clearly defined features is formed, providing a high-quality data foundation for subsequent data annotation and model training.

[0027] The protocol parsing script is a JavaScript script or a Lua script, and the generated parsing document is used for subsequent data analysis and device interconnection operations of IoT devices. Specifically, the protocol parsing script is either a JavaScript script or a Lua script. The script's generation is directly based on the recognition and parsing results output by the protocol identification model. The model first identifies the protocol type of the real-time communication data stream of the IoT device, then parses out key information such as the protocol's message header, field content, data type, checksum, and data length. Based on this information, the system automatically generates an executable parsing script adapted to the corresponding protocol. The JavaScript script is suitable for protocol parsing scenarios on IoT web platforms, while the Lua script is suitable for local protocol data processing scenarios on lightweight IoT devices. The generated parsing document includes protocol type descriptions, field meaning interpretations, data format specifications, and a guide to using the parsing script. This document can be used for subsequent data analysis operations on IoT devices, providing data analysts with a clear basis for interpreting protocol data. Furthermore, it supports device interconnection operations, lowering the technical threshold for device interconnection in IoT systems and improving system deployment and maintenance efficiency.

[0028] When the protocol recognition model is iteratively optimized, the supplemented new device protocol data needs to undergo data cleaning, format conversion and feature extraction operations to form standardized data and then be integrated into a multi-dimensional data source. Specifically, the supplementary new device protocol data originates from the real-time communication scenarios of newly added IoT devices. This includes communication stream data generated by devices with customized protocols from new manufacturers, new protocol documentation provided by the corresponding device manufacturers, communication log files generated after the new devices connect to the IoT platform, and preliminary parsing scripts written for temporary parsing of the new protocols. The new device protocol data must strictly adhere to the same preprocessing standards as the initial multidimensional data source, undergoing data cleaning, format conversion, and feature extraction operations in sequence. The data cleaning stage removes invalid and redundant information and noisy data from the new protocol data to ensure data accuracy. The format conversion stage converts the cleaned new protocol data into JSON format. The data can be in any of the following formats: XML or CSV, ensuring data format compatibility. During the feature extraction stage, key features relevant to parsing are extracted from the new protocol data, including at least the message header, field content, and data length information of the new protocol. For some complex new protocols, additional data type and checksum information need to be extracted. After the above preprocessing forms standardized data, it is integrated into the original multidimensional data source, further expanding the protocol coverage of the multidimensional data source. This provides more comprehensive training samples to support subsequent repeated data annotation and model fine-tuning training, ensuring that the iterative protocol recognition model can adapt to the recognition needs of new device protocols and improving the model's dynamic adaptability to changes in IoT protocols.

[0029] The data collection module supports both the collection of real-time communication data streams from IoT devices and the offline collection of stored historical protocol documents, log files, and protocol parsing scripts. Specifically, the data collection module establishes data interaction channels with IoT devices, IoT platforms, and storage servers through preset acquisition interfaces to acquire data sources in multiple scenarios. Among these, the acquisition of real-time communication data streams from IoT devices involves real-time monitoring of communication links between devices to synchronously capture the raw communication code stream data generated by the devices during normal operation. This type of real-time data can be directly used for real-time protocol identification and verification in the subsequent model inference stage, and can also be added to the training dataset to improve the model's adaptability to dynamic protocol scenarios. The offline collection of stored historical protocol documents, log files, and protocol parsing scripts involves accessing the historical database of the IoT platform, the archived document storage path provided by the device manufacturer, and the script backup directory of the local server to batch obtain previously accumulated protocol documentation, device communication logs, and verified and effective protocol parsing scripts (such as JavaScript scripts and Lua scripts). This type of offline data can enrich the sample coverage of multi-dimensional data sources, especially providing key data support for the model to learn the characteristics of early protocol versions and protocol variants in special scenarios. Furthermore, during real-time and offline data collection, the data collection module performs preliminary classification and labeling of the acquired data, which facilitates classification processing and feature extraction in the subsequent data preprocessing stage, ensuring that the multidimensional data source can efficiently serve the generation of standardized input datasets.

[0030] The data collection module transmits real-time bitstream data and historical document data to the data preprocessing module using a "classified transmission channel." Real-time data is transmitted through a high-priority channel to ensure a latency of ≤100ms. Historical data is transmitted through a batch channel. After cleaning and format conversion, the data preprocessing module generates a "data quality report" (including indicators such as data integrity and noise removal rate). Only when the report meets the standards, and the integrity is ≥95%, is the standardized dataset transmitted to the model training module. After the protocol recognition and inference module outputs the recognition results, it synchronously feeds them back to the data collection module to mark "high-value data" (such as new protocol data that is successfully recognized for the first time), providing a data priority basis for subsequent model iterations.

[0031] During data cleaning, a "rule-based filtering and anomaly detection algorithm" is used. First, obviously invalid data, such as bitstreams with a length less than the minimum protocol message length, is filtered out using preset rules. Then, an isolated forest algorithm is used to detect abnormal data. During feature extraction, "keyword matching and regular expressions" are used to extract the identifier field from the protocol message header, and "semantic word segmentation and entity recognition" are used to extract key information from the field content.

[0032] In addition, the present invention also provides a terminal device. The method for automatic identification of Internet of Things protocols involved in this embodiment is mainly applied to the terminal device, which can be a PC, a portable computer, a mobile terminal or other device with display and processing functions.

[0033] Specifically, the terminal device may include a processor (e.g., CPU), a communication bus, a user interface, a network interface, and memory. The communication bus is used to enable communication between these components; the user interface may include a display screen or an input unit such as a keyboard; the network interface may optionally include a standard wired interface or a wireless interface (such as a Wi-Fi interface); the memory may be high-speed RAM or stable non-volatile memory, such as disk storage, and may also optionally be a storage device independent of the aforementioned processor.

[0034] The memory stores a readable storage medium, which in turn stores an automatic IoT protocol identification program. The processor can call the automatic IoT protocol identification program stored in the memory and execute the automatic IoT protocol identification method provided in this embodiment of the invention.

[0035] Understandably, a readable storage medium can be a tangible device capable of holding and storing instructions for use by an instruction execution device. A computer-readable storage medium can be, for example—but not limited to—an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital multifunction disc (DVD), memory sticks, floppy disks, mechanical encoding devices, such as punch cards or recessed protrusions storing instructions thereon, and any suitable combination thereof. The computer-readable storage medium as used herein is not to be construed as a transient signal itself, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses through fiber optic cables), or electrical signals transmitted through wires.

[0036] The computer-readable program instructions described herein can be downloaded from computer-readable storage media to various computing / processing devices, or downloaded via a network, such as the Internet, local area network, wide area network, and / or wireless network, to an external computer or external storage device. The network may include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to the computer-readable storage media in the respective computing / processing device.

[0037] Computer program instructions used to perform operations of this disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, status setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk, C++, etc., and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The computer-readable program instructions may execute entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuitry, such as programmable logic circuitry, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), is personalized by utilizing the status information of the computer-readable program instructions to implement various aspects of this disclosure.

[0038] Finally, it should be noted that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for automatic identification of Internet of Things (IoT) protocols, characterized in that, Includes the following steps: Establish a data collection module to collect multi-dimensional data sources from IoT devices, clean the collected multi-dimensional data sources to remove noisy data, and convert the cleaned data into a unified standard format. The key features related to protocol parsing in the data are extracted, and a standardized input dataset is generated from the collected data. The protocol documents, log files, and protocol parsing scripts in the generated standardized input dataset are labeled and used for data annotation and training to form structured training data. The structured training data is divided into a training set and a validation set, and a pre-trained large language model is selected as the base model. The base model is then fine-tuned using the training set to obtain the trained protocol recognition model. The standardized input dataset sent in real time by IoT devices is input into the trained protocol recognition model. Based on the recognition and parsing results of the protocol recognition model, the corresponding protocol parsing script and parsing document are automatically generated. The new device protocol data is added to the multidimensional data source. The data collection, data labeling and training operations described above are repeated to iteratively optimize the protocol recognition model in order to adapt to changes in protocol format and the recognition requirements of new protocols.

2. The method for automatic identification of IoT protocols according to claim 1, characterized in that: The pre-trained large language model is a BERT model or a GPT series model; the training set is used to fine-tune the basic model.

3. The method for automatic identification of IoT protocols according to claim 1, characterized in that: The protocol identification model analyzes and infers the communication data stream, identifies the protocol type to which the data stream belongs, and parses the field content, data type, checksum, and data length information.

4. The method for automatic identification of Internet of Things protocols according to claim 1, characterized in that: The base model is fine-tuned during training, with the goal of minimizing the loss function. The training effect of the model is verified and the model parameters are optimized using the validation set. The loss function is either the mean squared error function or the cross-entropy function, resulting in a trained protocol recognition model.

5. The method for automatic identification of Internet of Things protocols according to claim 1, characterized in that: The tagging process includes tagging the protocol's message header, field content, data type, checksum, and data length information to form structured training data.

6. The method for automatic identification of Internet of Things protocols according to claim 1, characterized in that: The multidimensional data source includes at least device communication code stream data, protocol documents, log files and protocol parsing scripts, and covers standard protocols and customized protocols. The standard protocols include Modbus protocol, CAN protocol, MQTT protocol and HTTP protocol.

7. The method for automatic identification of Internet of Things protocols according to claim 1, characterized in that: The key features related to protocol parsing include at least the protocol message header, field content, and data length information. The data format after data cleaning and format conversion is any one or more combinations of JSON, XML, or CSV formats.

8. The method for automatic identification of Internet of Things protocols according to claim 1, characterized in that: The protocol parsing script is a JavaScript script or a Lua script, and the generated parsing document is used for subsequent data analysis and device interconnection operations of IoT devices.

9. The method for automatic identification of Internet of Things protocols according to claim 1, characterized in that: When the protocol recognition model is iteratively optimized, the supplementary new device protocol data needs to undergo data cleaning, format conversion and feature extraction operations to form standardized data and then be integrated into a multi-dimensional data source.

10. The method for automatic identification of Internet of Things protocols according to claim 1, characterized in that: The data collection module supports both the collection of real-time communication data streams from IoT devices and the offline collection of stored historical protocol documents, log files, and protocol parsing scripts.

Citation Information

Patent Citations

  • Industrial control protocol identification method based on fusion BERT and CNN network

    CN116389614A

  • Automatic internet of things protocol adaptation method and system based on large model

    CN119211393A

  • Internet of Things equipment long protocol text sequence modeling method and device, equipment and medium

    CN120455555A

  • Industrial control protocol intelligent analysis method based on large model

    CN120471045A