A method for deriving network protocol message structure based on RFC correlation analysis

Through the RFC association analysis method, the correlation relationship between network protocols among multiple RFCs is analyzed and derived, which solves the problem of insufficient universality and accuracy of network protocol packet structure analysis in the prior art, and realizes the functions of automated derivation and synchronous update.

CN116233263BActive Publication Date: 2025-05-13INSTITUTE OF INFORMATION ENGINEERING CHINESE ACADEMY OF SCIENCES
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310153726.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-13
Publication Date
2025-05-13
Estimated Expiration
2043-02-13

AI Technical Summary

Technical Problem

The prior art is difficult to effectively analyze and deduce the network protocol packet structure, especially when the update, substitution and extension relationships between multiple RFCs are complex, resulting in insufficient universality and accuracy.

Method used

Using an RFC association analysis method, the effective network packet structure that can be received by the specified network protocol is automatically derived by analyzing and associating all valid RFCs within a certain time period. The method includes steps such as preprocessing stage, RFC association extraction, RFC association derivation, parameter structure extraction and correlation analysis between parameters.

Benefits of technology

It realizes the automated derivation of network protocol packet structure, with stronger versatility and higher degree of automation, and can synchronously update the RFC update situation or the RFC range that users pay attention to, to meet the needs of programming, testing and security analysts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116233263B_ABST
    Figure CN116233263B_ABST
Patent Text Reader

Abstract

The invention discloses a network protocol message structure derivation method based on RFC association analysis, and the steps include: 1) crawling RFCs to be analyzed and IANA registration information from the network; obtaining the RFC corresponding to each RFC number, checking whether each parameter attribute information contains a specified keyword, and if not, adding the corresponding parameter attribute information and RFC to a list to be analyzed; 2) extracting the association relationship between RFCs based on the acquired data; 3) performing RFC association derivation according to the association relationship, and determining the chapter in the RFC where each parameter attribute is located; 4) extracting a structured description of protocol parameters from each RFC in the list to be analyzed; 5) performing association analysis between protocol parameters, and obtaining the combination relationship and position between each pair of protocol parameters; 6) combining the protocol parameters according to the combination relationship and position between the protocol parameters, and generating the message structure of the network protocol to be analyzed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of program analysis, and in particular to a method for deriving a network protocol message structure, specifically a method for deriving a network protocol message structure based on RFC association analysis. The method can assist programmers, testers, and security analysts in understanding the design of network protocols, constructing test cases, and discovering security issues caused by inconsistencies between design and code implementation. Background Art

[0002] A network protocol is a set of rules, standards or conventions established for data exchange in a computer network. The makers and maintainers of network protocols will provide protocol specifications (Request for Comments, RFC for short) described in natural language to guide the development of protocol software. RFC specifies the acceptable message structure of the protocol (including the composition of parameters in the message and the length, type, value range, relationship between parameters, etc. of each parameter), the response that should be taken after receiving the message, etc. The network protocol message structure described in RFC is the main basis for programmers and testers to design key data structures and construct test cases. However, the description of a network protocol message structure is usually scattered in different RFCs in different forms. With the continuous updating and improvement of network protocol functions, new RFCs will be continuously released. New RFCs can be an extension of old RFCs or replace old RFCs. Therefore, in addition to considering the different messages and parameter structures described in a single RFC and the relationship between them, the analysis and derivation of the network protocol message structure must also consider the update, replacement and extension relationship between RFCs, as well as the resulting impact on the message structure.

[0003] Due to the large number of RFCs, manual understanding of RFCs and extraction of message structures inevitably involve omissions, while existing technologies for automatically deriving network protocol message structures face challenges in terms of generality and accuracy. Representative technologies include: (1) Methods based on reverse engineering and dynamic testing: Analyze messages collected during normal work through reverse engineering, analyze their attributes and generate structures to be verified, and then use dynamic testing methods to send the structures to be verified to communication devices to verify the correctness of the structures through interaction with them. This method mainly faces the problem of insufficient generality. When it is necessary to analyze the message structure of another protocol, it is necessary to perform reverse engineering analysis again and build a communication environment. (2) Network protocol entity extraction method based on small sample learning: Extract parameters and description information from the RFC document of the network protocol and process them into a text set. Use the text set to train a potential network protocol entity classifier and a network protocol entity accurate recognition model. Then, these two parts are merged into a protocol entity extraction model to extract network protocol entities in the network protocol (entities refer to fields and description information in the message structure). However, this method does not consider the association between the extracted entities, so it is impossible to obtain a complete protocol message structure. Summary of the invention

[0004] In view of the problems existing in the prior art, the purpose of the present invention is to provide a method for deriving the network protocol message structure based on RFC association analysis, which automatically derives the valid network message structure that can be received by the protocol by analyzing and associating all valid RFCs or user-specified RFCs of a specified network protocol within a certain time period.

[0005] The basic process of this method is as follows:

[0006] 1) Preprocessing stage: crawl the RFC to be analyzed and IANA (Internet Assigned Numbers Authority) registration information from the Internet to prepare for subsequent association analysis and structure derivation. Among them, IANA registration information describes the protocol parameters (parameter name, RFC, description, etc.), parameter attributes (value range, value description and RFC corresponding to each value, etc.), but the above information is maintained manually and may contain errors or inconsistencies (for example, IANA and RFC have different names for the same protocol parameter, IANA information is not registered in time, etc.);

[0007] 2) RFC association extraction: extracting the update, replacement and extension relations between RFCs from the preprocessed data;

[0008] 3) RFC association derivation: Through heuristic derivation, the association relationships (especially extended relationships) extracted in the previous step are refined, possible errors in IANA information maintenance are corrected, and more fine-grained association relationships are obtained;

[0009] 4) Parameter structure extraction from a single RFC: Extract the structured description of the parameter from a single RFC, as well as the information of each structure in the structured description of the parameter (name, length, value range, etc.);

[0010] 5) Parameter association analysis: Analyze different parameters within the same RFC and the parameter association between multiple RFCs, with the aim of determining their combination relationship (parallel, inclusion, instantiation, etc.) and the position of the combination;

[0011] 6) Message structure combination: Complete the message structure combination based on the analysis and deduction results.

[0012] The technical solution of the present invention is:

[0013] A method for deriving a network protocol message structure based on RFC association analysis, the steps of which include:

[0014] 1) For the network protocol to be analyzed, query the protocol parameters involved in the network protocol in the IANA database, and then query the RFC number and parameter attribute information of the protocol parameter in the IANA database according to the name of each protocol parameter; obtain the RFC corresponding to each RFC number, check whether each parameter attribute information contains the specified keyword, and if so, directly proceed to step 6); if not, add the corresponding parameter attribute information and RFC to the to-be-analyzed list TODO, and then proceed to steps 2 to 6);

[0015] 2) extracting association relationships between RFCs based on the data obtained in step 1); the association relationships include update relationships, replacement relationships, and extension relationships;

[0016] 3) Perform RFC association derivation based on the extracted association relationship to determine the section in the RFC where each parameter attribute is located;

[0017] 4) Extract the structured description of protocol parameters from each RFC in the to-be-analyzed list TODO;

[0018] 5) According to the structured description of the protocol parameters extracted in step 4) and the section of the RFC where each parameter attribute is located determined in step 3), the association analysis between the protocol parameters is performed to obtain the combination relationship and position between each pair of protocol parameters;

[0019] 6) Combining the protocol parameters according to the combination relationship and position between the protocol parameters to generate the message structure of the network protocol to be analyzed.

[0020] Furthermore, the combination relationship includes a parallel relationship, an inclusion relationship, and an instantiation relationship.

[0021] Furthermore, the identification condition of the parallel relationship is: when parameter A and parameter B are listed in a table at the same time, and the chapters corresponding to parameter A and parameter B are at the same level, parameters A and B are called parallel parameters, and the parameters corresponding to the table are used as the parent parameters of parameters A and B; the identification condition of the inclusion relationship is: if the name of parameter B appears in the structure of parameter A, and the chapter corresponding to parameter B is in the chapter corresponding to parameter A, then parameter A contains parameter B, parameter A is the including parameter, and parameter B is the included parameter; the identification condition of the instantiation relationship is: if a parameter m in the structure of parameter A has multiple values ​​and descriptions corresponding to each value, and the value description of parameter B is the same as a value description of parameter m, then parameter B is an instantiation of parameter A, parameter A is called an abstract parameter, and parameter B is called an instance parameter.

[0022] Furthermore, in step 6), the method for generating the message structure of the network protocol to be analyzed is: first, obtaining the root document of the network protocol to be analyzed according to the update relationship, then obtaining the first parameter and the corresponding structure that are not in the combination relationship from the root document, and taking the chapter where the corresponding structure is located as the starting chapter; then obtaining the root structure of the data packet from the starting chapter; and then iterating steps a to b:

[0023] a. Traverse each parameter i in the root structure and parameter structure, find the parameter associated with the parameter i in the combination relationship, and store the associated parameters according to the association relationship;

[0024] b. Traverse each parameter j in the root structure and parameter structure. If the parameter j is a parameter registered in IANA, add an extended attribute to the parameter j to save the parameter attribute registered by IANA for the parameter j and the specific attribute information; then for each parameter attribute of the parameter j, obtain the description corresponding to the parameter attribute from the chapter corresponding to the parameter attribute; if there is a parameter structure in the description, store it in the structure attribute of the parameter j, otherwise set the structure attribute of the parameter j to empty, and store the chapter number and RFC number corresponding to the parameter attribute in the information attribute of the parameter j;

[0025] After the iterative analysis is completed, if there are parameters that are not combined into the data packet, they will be returned to the user, who will specify whether these parameters need to be included in the data packet and the specific location to put them in.

[0026] Furthermore, in step a, the method for storing associated parameters according to the association relationship is: for parameters of parallel relationship, they are extended to the extended structure attributes of the corresponding upper-level parameters; for parameters of inclusion relationship, the included parameters are placed in the structure attributes of the corresponding containing parameters; for parameters of instantiation relationship, the instance parameters are placed in the extended structure attributes of the abstract parameters respectively, and the instance type corresponding to each instance parameter is identified.

[0027] Furthermore, the method for determining the chapter in the RFC where each parameter attribute is located is as follows: first, compare the similarity of the name of the parameter attribute with all the titles of each chapter in the RFC to obtain the chapter that each parameter attribute preliminarily matches; then determine the chapter corresponding to the parameter attribute based on the preliminary matching results of two adjacent parameter attributes: for attribute j, its adjacent attributes are i and k, and the serial numbers of the preliminary matching chapters of attribute ijk are Ni, Nj, and Nk respectively. The chapter after Ni and the chapter before Nk will be used as chapters to be confirmed. If the attribute value and name of attribute j exist in the contents of the two chapters to be confirmed, the chapter to be confirmed will be determined as the chapter corresponding to attribute j. If it exists in both chapters, the chapter after Ni will be used as the chapter corresponding to attribute j; if it does not exist in both chapters, the preliminary matching chapter of attribute j will be determined as the chapter corresponding to attribute j.

[0028] Furthermore, the structured description includes multiple structures and corresponding structure information, and the structure information includes name, length, and value range.

[0029] A server, characterized in that it comprises a memory and a processor, wherein the memory stores a computer program, the computer program is configured to be executed by the processor, and the computer program comprises instructions for executing each step in the above method.

[0030] A computer-readable storage medium stores a computer program, wherein the computer program implements the steps of the above method when executed by a processor.

[0031] The advantages of the present invention are as follows:

[0032] The present invention automatically derives the valid network message structure that can be received by the specified network protocol by analyzing and associating all valid RFCs or user-specified RFCs within a certain period of time. Compared with similar technologies, the present invention effectively uses RFC's description of knowledge related to network protocols, and has stronger versatility; it does not require manual participation and has a higher degree of automation; at the same time, it supports synchronous updating of the derived message structure according to the update status of RFC or the scope of RFC that users are concerned about, meeting the needs of programmers, testers and security analysts. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1 The figure is a flow chart of the method of the present invention.

[0034] Figure 2 This is a partial RFC relationship diagram for DHCP. DETAILED DESCRIPTION

[0035] The present invention is further described in detail below in conjunction with the accompanying drawings. The examples given are only used to explain the present invention but not to limit the scope of the present invention.

[0036] Now combined with the process Figure 1 , the detailed steps of the present invention are described as follows:

[0037] 1. Preprocessing stage.

[0038] According to the selected protocol, query the IANA database for all protocol parameters that have been registered for the protocol. These protocol parameters consist of the values ​​used in the protocol. Using the crawler tool, record these protocol parameter names, the RFC number that registers the protocol parameter, and the attributes under the protocol parameter and their related information (including the value of the attribute, the name or description of the attribute, and the RFC that registers the attribute to IANA) as Dataset, and record all the RFCs involved as Files separately. At the same time, when obtaining the protocol parameter information, check whether the name or description of the attribute contains keywords such as unassigned, removed, deprecated, reserved or experimental use. Attributes containing these keywords do not need to be further analyzed (steps 2-5), and the processing in step 6 is the same as that of other attributes that do not contain these keywords; add the parameter attributes that do not contain such keywords and the corresponding RFCs to the list to be analyzed, recorded as TODO.

[0039] 2. RFC association extraction.

[0040] This step extracts three types of relationships between RFCs: update, replacement, and extension. A Updated RFC B It means that document A has revised the specified paragraph or sentence in document B; RFC A Replaces RFC B It means that document A completely replaces document B, making document B no longer in use; RFC A Extended RFC B It means that document A supplements the protocol parameters and structure described in document B, for example, it defines new attributes.

[0041] In the three types of associations, update and replacement associations can be obtained directly on the document header: For example, if RFCA Updated RFC B , then RFC A The header will be marked with "Updates: RFC B If it is a replacement, it will be marked with "Obsoletes:RFC B Therefore, analyze the RFCs in all files and obtain the updates between these RFCs according to the RFC document header (including direct updates and indirect updates, for example, if the RFC A Update RFC B , RFC B Update RFC C , then RFC A Indirectly update RFC C ), alternative association.

[0042] The extended association can be extracted based on the Dataset in step 1. From the Data, the RFC that registers the protocol parameter and the RFC that registers the attribute for the parameter can be obtained. In this case, the latter is an extension of the former.

[0043] 3.RFC association derivation.

[0044] The extended relationship in step 2 describes the relationship at the granularity of the RFC document. In this step, a more detailed association will be derived to determine the chapter where each parameter attribute is located: the association between the parameter attribute and the specific chapter of the document where the attribute is located. Through such a more fine-grained association, the document content directly related to the attribute can be obtained.

[0045] First, the name of the attribute is compared with all the titles in the corresponding RFC section for similarity. A preliminary matching result is obtained by using fuzzy matching methods such as unifying the capitalization of letters and considering common word inflections. However, since both IANA and RFC are maintained manually, their descriptions of the same attribute may be inconsistent, so the accuracy of the preliminary matching result is not high, so it needs to be corrected.

[0046] There are two main bases for correcting the matching results: the matching results of adjacent attributes and the correlation between the chapter content and the attributes. The former is based on the rule that the attributes that are adjacent in IANA have corresponding chapters that are also adjacent in RFC; the principle of the latter is that the attribute information (including the attribute value and the attribute name) must appear in the chapter corresponding to the attribute. Therefore, the verification method is to obtain the chapter to be confirmed based on the chapter corresponding to the adjacent attribute, and then find the attribute information in the chapter to be confirmed. The chapter that meets these two bases is confirmed as the final result. The attribute and its corresponding chapter are recorded as LinkInfo. The specific verification method is: for the attribute j to be verified, if ijk is an adjacent attribute, the serial numbers of the matching chapters of these attributes are Ni, Nj, and Nk respectively. Then the next chapter of Ni and the previous chapter of Nk will be used as chapters to be confirmed (for example, the next chapter of chapter 4.3 is 4.4, and the previous chapter of chapter 5.1.2 is 5.1.1), and the chapters to be confirmed meet the first credential; that is, the two chapters to be confirmed are obtained through the first credential, and the basis of the first credential is that for the adjacent attributes i, j, and k when registered with IANA, the corresponding chapter serial numbers are I, J, and K, and IJK are adjacent; therefore, for attribute j, its two chapters to be confirmed are the next chapter of I and the previous chapter of K. For the chapters to be confirmed, the value and name of the attribute are found in the text of the chapter. If found, it is considered to meet the second credential. For the chapters that meet both credentials, it is confirmed as the final result. If there is no chapter that meets both credentials at the same time, the preliminary matching result is used as the final result.

[0047] 4. Extraction of parameter structure in a single RFC.

[0048] This step will extract the protocol parameters, parameter attributes and their corresponding structures in each RFC in the TODO, which are usually described in structured diagrams in RFC.

[0049]

[0050] In the graphic structure, "|" is used to separate different fields, and the length is described by the position occupied or the subsequent brackets. By identifying the boundaries of the field, the field and the length of the field can be extracted. In the graphic structure, the fields are separated horizontally using "+-" and vertically using "|", and the graphic structure has a digital ruler. The length of the Version field in the structure can be extracted to be 8. The length is marked with brackets after the Optional Parameters field, which indicates that the length of this field is not fixed.

[0051] The table format is usually used to enumerate the values ​​of the Type attribute, or other protocol parameters that only have values ​​but no structure. The table format is mainly identified by some table features. The text fragments are very neat in the vertical direction. The first row is used as the header and is separated from other rows by spaces (or other symbols). By identifying the header and columns, we can get information such as the value 0x00 represents no tunnel information present, and the RFC document where the value is located is RFC6514. The extracted result is recorded as StructInfo.

[0052]

[0053] 5. Correlation analysis between protocol parameters.

[0054] This step extracts three types of associations between protocol parameters: parallel, inclusive, and instantiated. In the parallel relationship, parameter A and parameter B are parallel parameters at the same level in the message structure, and then the text description is used to determine whether these parameters can appear in the same message at the same time; in the inclusive relationship, parameter B is a supplement to parameter A, that is, there is a structure in parameter A that is omitted in the description of parameter A, and parameter B subdivides the structure rather than supplementing parameter A; in the instantiation relationship, a field in the structure of parameter A (such as the type field) takes a specific value, and other related structures of parameter A become the form of parameter B.

[0055] The identification condition of parallel association is that when parameter A and parameter B are listed at the same time in a table-like text, and the chapters corresponding to parameter A and parameter B are at the same level, in this case, parameter A and parameter B are judged to be parameters in a parallel relationship. In addition, according to the text content of the chapter where the table is located, it is judged whether parameter A and parameter B can appear in the same message at the same time. The parameters corresponding to the table are called upper-level parameters, and parameters A and B are called parallel parameters.

[0056] The recognition condition of the inclusion association is that the structure of parameter A is known, the name of parameter B appears in the structure of parameter A, and the chapter corresponding to parameter B is in the chapter corresponding to parameter A. At this time, the recognition result of parameter B is directly updated to the extraction result of parameter A. Parameter A is called the inclusion parameter, and parameter B is called the included parameter.

[0057] The identification condition of instantiation association is that in the structure of parameter A, there is a field with multiple values ​​and the descriptions corresponding to the values. If the description of parameter B is the same as one of the descriptions corresponding to these values, then parameter B is considered to be an instantiation of parameter A. The identification result of parameter B is extended to the corresponding field of parameter A. Parameter A is called an abstract parameter and parameter B is called an instance parameter.

[0058] First, perform association analysis on all documents in TODO, then update the association according to RFC, and perform association analysis on cross-document parameters.

[0059] 6. Message structure combination.

[0060] According to the RFC update relationship, the root document is obtained (usually the earliest document among all RFCs that are still in use for the network protocol, which describes the basic functions of the network protocol and the basic structure of the transmission message), and the starting chapter of the root structure is obtained according to the parameter relationship in the root document (according to the field association in the root document, the parameters and corresponding structures that are not in parallel, included, or instantiated relationships are obtained, and the chapter where such a structure first appears in the root document is used as the starting chapter by default), and the root structure of the data packet is obtained from the starting chapter of the root structure, including information such as parameter name and parameter length, and then iterative combination operations are performed to expand all parameters, parameter attributes (including possible values ​​and corresponding structures), etc. to the root structure according to the association:

[0061] a. Traverse each parameter in the root structure and parameter structure, find the parameters associated with the parameter in the association result of step 5, and store the associated parameters in different extensions according to the type of association. For parallel parameters, extend them to the extended structure attribute extend_struct of the parent parameter; for parameters of inclusion relationship, put the included parameters directly into the structure attribute struct of the corresponding field of the containing parameter; for parameters of instantiation relationship, put the instance parameters into the extend_struct attribute of the abstract parameter respectively, and identify the instance type corresponding to each instance parameter.

[0062] b. If the parameter is a parameter registered in IANA, then add an extend attribute to the parameter, which stores the attribute registered by IANA for the parameter and the specific information of the attribute, including the attribute name, attribute value, attribute value description, and attribute value-related RFC. At the same time, for each parameter attribute, obtain the description corresponding to the parameter attribute according to the previously matched chapter: if there is a detailed structure, store it in the struct extension, otherwise the struct structure is set to empty. At the same time, store the chapter number and RFC number corresponding to the parameter attribute in the information structure info of the parameter attribute.

[0063] c. Iteratively analyze the chapters and associations corresponding to the extended parameters and parameter attributes.

[0064] After the iterative analysis is completed, if there are parameters that are not combined into the data packet, they are returned to the user, who specifies whether these parameters need to be included in the data packet and the specific location to put them in.

[0065] After all parameters are combined, the message structure file of the selected protocol is obtained; the final message structure is saved in the file. If the user specifies the analysis scope, the structures that do not have protocol parameters and fields within the scope from the root structure to the bottom structure are removed.

[0066] A specific embodiment is described below, taking the analysis of DHCP attributes as an example.

[0067] In step 1, according to the specified scope, the crawler obtains the information of the parameter and its subordinate attributes. Name stores the parameter name, refer stores the RFC file that registers the parameter, and link stores the attribute information of the parameter. In each attribute, Tag stores the attribute itself (according to IANA, it may be Code, Value, etc.), Name is the name of the attribute (according to IANA's writing method, it may be Description, Meaning, etc.), and Reference is the RFC that registers the attribute. At the same time, determine whether it is a valid attribute based on the attribute name. Save the information obtained by the crawler and the list of RFCs involved in the valid attributes.

[0068] In step 2, the relationship between the RFC to which the parameter belongs and the RFC to which the attribute under the parameter belongs is constructed through the crawler information saved in step 1, and a directed graph of the relationship between RFCs is generated. At the same time, the explicit update and replacement relationships between RFC documents are obtained according to the document header, and all RFCs updated by the RFC in the list are obtained according to the valid RFC list saved in step 1. The RFC relationship graph and the results of each RFC update are saved, and some of the relationships are as follows: Figure 2 shown.

[0069] In step 3, according to the RFC association relationship in step 2, for each parameter, its subordinate attributes are matched with the chapters in the corresponding RFC to obtain the chapter number corresponding to the attribute. First, a preliminary matching result is obtained through direct matching, such as the preliminary matching results of some attributes of the DHCP Option parameter shown in Table 1. After all the attributes under the Option parameter have matching results, verification is performed. For example, the attribute name of the attribute value 18 in Table 1 is Extension File, and the direct matching result is RFC2132, which does not correspond to a specific chapter. At this time, the attribute value 17 matches the chapter 3.19 of RFC2132, and the attribute value 19 matches the chapter 4.1 of RFC2132. Therefore, the alternative chapters for the attribute value 18 are 3.20 and 4.0. Finally, the attribute value 18 is found in chapter 3.20, so the chapter corresponding to Extension File is corrected to 3.20.

[0070] Table 1. Results of obtaining fine-grained association relationships (partially)

[0071]

[0072] In step 4, for the RFCs in the RFC list involved in the valid attributes in step 1, the corresponding RFCs are first reorganized according to the title hierarchy to facilitate the subsequent analysis of the text range corresponding to each chapter. Then the structural information of the parameters in the RFC is extracted and identified with chapter titles, chart titles, etc. Finally, the single document structure analysis results are saved.

[0073] In step 5, the association between the structural information in the single document obtained in step 4 is analyzed, and the corresponding structure is adjusted. According to the RFC update result in step 2, the parameters in the RFC are matched and analyzed with the parameter structure in the RFC it updates to obtain the parameter relationship between the cross-documents. Finally, the structural analysis result of the single document is updated, and the parameter relationship between the cross-documents is saved.

[0074] In step 6, specify the root document of DHCP as RFC2131, and the chapter where the basic structure is located is ProtocolSummary. According to the single document result obtained in step 5, obtain the corresponding structure, iteratively analyze each parameter in the structure, and according to the cross-document association result obtained in step 5, know that the options parameter in the structure is the parameter BOOTP Vendor Extensions and DHCP Options registered in IANA, so save the attributes under the parameter to the extend structure of the options parameter, and then obtain the structure of these attributes in the corresponding RFC chapter according to the matching result of step 3. Continue iterating the analysis until there is no extensible part in the structure.

[0075] After the iteration is completed, check whether the IANA parameters in the range specified in step 1 are all included in the data packet. If not, generate a configuration file to be filled in. The user confirms whether the parameters need to be associated with the data packet. If so, specify the location of the association.

[0076] Finally, a data packet structure of the specified network protocol is generated and saved in a file.

[0077] Although the specific embodiments of the present invention are disclosed for the purpose of illustration, the purpose is to help understand the content of the present invention and implement it accordingly, those skilled in the art will understand that various substitutions, changes and modifications are possible without departing from the spirit and scope of the present invention and the appended claims. Therefore, the present invention should not be limited to the content disclosed in the best embodiment, and the scope of the present invention is subject to the scope defined in the claims.

Claims

1. A method for deriving a network protocol message structure based on RFC association analysis, the steps of which include: 1) For the network protocol to be analyzed, query the protocol parameters involved in the network protocol in the IANA database, and then query the RFC number and parameter attribute information of the protocol parameter in the IANA database according to the name of each protocol parameter; obtain the RFC corresponding to each RFC number, and check whether each parameter attribute information contains the specified keyword. If it does, proceed directly to step 6); if not, add the corresponding parameter attribute information and RFC to the to-be-analyzed list TODO, and then proceed to steps 2 to 6); 2) extracting association relationships between RFCs based on the data obtained in step 1); the association relationships include update relationships, replacement relationships, and extension relationships; 3) Perform RFC association deduction based on the extracted association relationship to determine the section in the RFC where each parameter attribute is located; 4) Extract the structured description of protocol parameters from each RFC in the TODO list to be analyzed; 5) According to the structured description of the protocol parameters extracted in step 4) and the section of the RFC where each parameter attribute is located determined in step 3), perform correlation analysis between the protocol parameters to obtain the combination relationship and position between each pair of protocol parameters; 6) Combining the protocol parameters according to the combination relationship and position between the protocol parameters to generate the message structure of the network protocol to be analyzed; wherein the method for generating the message structure of the network protocol to be analyzed is: firstly obtaining the root document of the network protocol to be analyzed according to the update relationship, and then obtaining the first parameter, the included parameter or the abstract parameter and the corresponding structure that are not in the combination relationship from the root document, and taking the chapter where the corresponding structure is located as the starting chapter; Then, the root structure of the data packet is obtained from the starting section; then steps a to b are iterated; after the iterative analysis is completed, if there are parameters that are not combined into the data packet, they are returned to the user, and the user specifies whether these parameters need to be put into the data packet and the specific location to put them; a traversing the root structure, each parameter i in the parameter structure, looking for parameters associated with the parameter i in the combination relationship, and storing the associated parameters according to the association relationship; b. Traverse each parameter j in the root structure and parameter structure. If the parameter j is a parameter registered in IANA, add an extended attribute to the parameter j to save the parameter attribute registered by IANA for the parameter j and the specific attribute information; then for each parameter attribute of the parameter j, obtain the description corresponding to the parameter attribute from the chapter corresponding to the parameter attribute; if there is a parameter structure in the description, store it in the structure attribute of the parameter j, otherwise set the structure attribute of the parameter j to empty, and store the chapter number and RFC number corresponding to the parameter attribute in the information attribute of the parameter j.

2. The method according to claim 1, characterized in that The combination relationship includes a parallel relationship, an inclusion relationship, and an instantiation relationship.

3. The method according to claim 2, characterized in that The identification condition of the parallel relationship is: when parameter A and parameter B are listed in a table at the same time, and the chapters corresponding to parameter A and parameter B are at the same level, parameters A and B are called parallel parameters, and the parameters corresponding to the table are used as the parent parameters of parameters A and B; the identification condition of the inclusion relationship is: if the name of parameter B appears in the structure of parameter A, and the chapter corresponding to parameter B is in the chapter corresponding to parameter A, then parameter A contains parameter B, parameter A is the including parameter, and parameter B is the included parameter; the identification condition of the instantiation relationship is: if a parameter m in the structure of parameter A has multiple values ​​and descriptions corresponding to each value, and the value description of parameter B is the same as a value description of parameter m, then parameter B is an instantiation of parameter A, parameter A is called an abstract parameter, and parameter B is called an instance parameter.

4. The method according to claim 1, characterized in that: In step a, the method for storing associated parameters according to the association relationship is: for parameters of parallel relationship, they are expanded into the extended structure attributes of the corresponding superior parameters; for parameters of inclusion relationship, the included parameters are placed into the structure attributes of the corresponding containing parameters; for parameters of instantiation relationship, the instance parameters are placed into the extended structure attributes of the abstract parameters respectively, and the instance type corresponding to each instance parameter is identified.

5. The method according to claim 1, characterized in that The method for determining the chapter in the RFC where each parameter attribute is located is as follows: first, compare the name of the parameter attribute with all the titles of each chapter in the RFC for similarity, and obtain the chapter that each parameter attribute preliminarily matches; then determine the chapter corresponding to the parameter attribute based on the preliminary matching results of two adjacent parameter attributes: for attribute j, its adjacent attributes are i and k, and the serial numbers of the preliminary matching chapters of attribute ijk are Ni, Nj, and Nk respectively. The chapter after Ni and the chapter before Nk will be used as chapters to be confirmed. If the attribute value and name of attribute j exist in the contents of the two chapters to be confirmed, the chapter to be confirmed will be determined as the chapter corresponding to attribute j. If it exists in both chapters, the chapter after Ni will be used as the chapter corresponding to attribute j; if it does not exist in both chapters, the preliminary matching chapter of attribute j will be determined as the chapter corresponding to attribute j.

6. The method according to claim 1, 2 or 3, characterized in that: The structured description includes multiple structures and corresponding structure information, and the structure information includes name, length, and value range.

7. A server, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program, the computer program is configured to be executed by the processor, and the computer program comprises instructions for executing each step in the method according to any one of claims 1 to 6.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Network physical topology discovering method and network management server based on SNMP

    CN101764709A

  • Systems and methods of building sequenceable test methodologies

    US20160077944A1