Operation system version similarity evaluation method and system based on knowledge graph
By building an operating system knowledge graph and using the RotatE embedding algorithm, the difficulties of similarity and difference evaluation in operating system version evaluation are solved, and a more accurate and intuitive version evaluation is achieved.
Patent Information
- Application Number
- CN202510204629.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-24
- Publication Date
- 2025-06-13
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
It is difficult for prior art to accurately evaluate the similarity and differences of operating system versions, especially when the complex associations between software packages and mirrors are underutilized.
By building an operating system knowledge graph, the RotatE embedding algorithm is used to vector represent entities, and the similarity degree of the two operating system versions is evaluated through a similarity function.
A more accurate and intuitive similarity evaluation of the operating system version is achieved, making full use of the complex relationship between the software package and the mirror, and improving the accuracy and rationality of the evaluation.
Smart Images

Figure CN120144172A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of entity vector representation in computer knowledge graphs, and particularly discloses a method and system for evaluating the similarity of operating system versions based on a knowledge graph. Background Art
[0002] In the Linux operating system, the distribution version is a complex concept. Narrowly speaking, it generally refers to the image provided by the distributor, while in a broad sense, it also includes the software repository provided to users. The system image is fixed after release, while the software repository is constantly changing due to reasons such as continuously fixing software defects, adding new functions, and fixing security vulnerabilities. When these changes accumulate to a certain extent or time, a new system image is released, which promotes the version evolution and update of the operating system. The analysis of operating system version evolution is an important research work in the field of operating systems, which can help downstream distributors clarify the development trend of the operating system, help users select appropriate system versions, evaluate the costs of migration and upgrade, etc. The evaluation of the similarity of operating system versions is exactly an important research content among them.
[0003] Operating system distributors usually provide various software to users in the form of software source repositories. The software repository organizes these software packages in the form of a directory tree and records detailed attribute information through the Packages file. Taking the Uubntu system as an example, Figure 1 shows the basic organizational structure of the software source repository. To evaluate the differences and similarities between two versions, we can start from the Packages file to obtain the basic information of all software packages. The traditional method is based on simple data statistics, comparing multiple indicators such as the total number of software packages, the number of software with the same name, and the number of software with the same name but different versions. Using this information to reflect the similarity and difference between two versions is not intuitive and it is difficult to calculate the differences. At the same time, this comparison method only uses relatively superficial information and ignores various association relationships between entities such as software packages and images, such as the dependency relationship and conflict relationship between binary software packages, the support relationship between the image and the repository, the inclusion relationship between the image and the binary software package, etc. Therefore, the results obtained also lack accuracy and rationality.
[0004] A knowledge graph is a technical method that uses a graph model to describe knowledge and construct the associated relationships between all things. It aims to identify, discover, and infer the complex relationships between things and concepts from data, and it is a computable model of the relationships between things. The common storage methods of knowledge graphs are to use databases, including relational databases, graph model databases, triple databases, etc. In these mainstream storage forms, information is still expressed in a symbolic or discrete form, and semantic calculations are often not possible. Knowledge graph embedding technology maps entities and relationships into a low-dimensional continuous vector space. This not only facilitates corresponding calculations, but also enables the vectors to contain rich semantic information during the mapping process. In addition, it can be combined with newer technologies such as machine learning to use algorithms with higher complexity to improve the calculation accuracy.
[0005] The training of knowledge graph embedding methods needs to be based on supervised learning, that is, a training data set needs to be provided. Common embedding algorithms include the following categories: transfer distance models, semantic matching models, and models considering additional information, etc. Classic algorithms include TransE, TransH, RotatE, RESCAL, DistMul, etc. Each algorithm has its own defects and applicable scenarios. Among them, the RotatE algorithm model is relatively more complex than algorithms such as TransE and TransH, but it has better performance in the training of complex relationships.
[0006] As mentioned above, in the Linux operating system, the generalized definition of the operating system version includes the system image and the software source repository that supports it behind. And the software repository is always changing due to reasons such as security updates, introduction of new functions, and bug fixes. Therefore, due to the complexity of its composition and the diversity of changes in the operating system version, it is difficult to describe its similarity or difference with a single metric or piece of data using a simple statistical method. Using multiple data for comparison is not intuitive and clear enough, which adds difficulty to the version evolution analysis of the operating system. Knowledge graph embedding technology has been applied in many fields currently. It can map the content of the knowledge graph into a vector space, use vectors to represent entities and relationships, and compared with traditional discrete representation methods, it is easier to integrate complex information and improve the calculation efficiency and the accuracy of the results. Therefore, the knowledge graph embedding technology can be applied to the operating system field to provide a new technical solution for version evolution analysis and similarity evaluation. Summary of the Invention
[0007] The present invention provides a method and system for evaluating the similarity of operating system versions based on a knowledge graph, aiming to solve at least one of the above defects.
[0008] According to one aspect of the present invention, it relates to a method for evaluating the similarity of operating system versions based on a knowledge graph, including the following steps:
[0009] Construct an operating system knowledge graph based on the package information in the target operating system repository and the package list information in the target operating system image;
[0010] Use the RotatE embedding algorithm to train the operating system knowledge graph to obtain the vector representation of the target operating system entity;
[0011] According to the vector representation of the target operating system entity, use the entity similarity function to evaluate the similarity between two operating system versions.
[0012] The package information includes the Packages file and the repository source organizational structure information, and the package list information includes the filesystem.manifest package list file. The steps to construct the operating system knowledge graph based on the package information in the target operating system repository and the package list information in the target operating system image include:
[0013] Obtain all Packages files and the repository source organizational structure information in the target operating system repository;
[0014] Obtain the filesystem.manifest package list file in the target operating system image;
[0015] Parse the Package file and the filesystem.manifest package list file, create entities and relationships in the graph database, and construct the operating system knowledge graph.
[0016] Furthermore, the steps to parse the Package file and the filesystem.manifest package list file, create entities and relationships in the graph database, and construct the operating system knowledge graph include:
[0017] Obtain the attribute information of all packages from the Packages file. The attribute information includes the first attribute information, the second attribute information, and the third attribute information; use the first attribute as the attribute field to describe the package entity and write it into the first file; use the second attribute information to establish the relationship between the binary installation package and the source code package, and use the third attribute information to establish the dependency and conflict relationships between binary packages or between binary packages and virtual packages, and write them into the second file; the first attribute information includes the software name, software version, software maintainer, and software function description, the second attribute information includes the source code package name, and the third attribute information includes the dependency relationship and the conflict relationship; the first file is the pkgNode.csv file, and the second file is the relationship file;
[0018] Parse the filesystem.manifest package list file to obtain the package names and software version information included in the operating system version, create an operating system image node, and write the support relationship between the image node and the repository source suitNode, as well as the inclusion relationship between the image node and the packages, into a third file, which is a relationship csv file;
[0019] Use the import tool to import the file information of the first file, the second file, and the third file into the graph database to construct an operating system knowledge graph.
[0020] Further, the steps of parsing the Package file and the filesystem.manifest package list file and creating entities and relationships in the graph database to construct the operating system knowledge graph also include:
[0021] According to the repository source organizational structure, create the suitNode.csv, componentNone.csv node files and the inclusion relationship file with the binary packages.
[0022] Further, the steps of training the operating system knowledge graph using the RotatE embedding algorithm to obtain the vector representation of the target operating system entity include:
[0023] Query all node and relationship names in the knowledge graph and write them into the entities.dict and relations.dict files respectively as entity and relationship dictionaries;
[0024] Obtain all relationships in the knowledge graph and write the above data into the train.txt, test.txt, and valid.txt files respectively according to the ratio of 8:1:1 as the training set, test set, and validation set;
[0025] Take the entities.dict, relations.dict files, train.txt, test.txt, and valid.txt files as parameters and pass them into the RotatE embedding model. Model various relationship patterns through the training set in train.txt, and correct and evaluate the modeling results through the test set and validation set to obtain the vector representations of relationships and nodes.
[0026] According to one aspect of the present invention, there is also provided an operating system version similarity evaluation system based on a knowledge graph, including:
[0027] A construction module for constructing an operating system knowledge graph based on the package information in the target operating system repository and the package list information in the target operating system image;
[0028] An acquisition module, which is used to train an operating system knowledge graph by using the RotatE embedding algorithm to obtain a vector representation of a target operating system entity;
[0029] An evaluation module, which is used to evaluate the similarity between two operating system versions according to the vector representation of the target operating system entity by using an entity similarity function.
[0030] Furthermore, the construction module includes:
[0031] A first acquisition unit, which is used to acquire all Packages files and repository source organization structure information in a target operating system repository;
[0032] A second acquisition unit, which is used to acquire a filesystem.manifest package list file in a target operating system image;
[0033] A parsing unit, which is used to parse the Package file and the filesystem.manifest package list file, create entities and relationships in a graph database, and construct an operating system knowledge graph.
[0034] Furthermore, the parsing unit includes:
[0035] An acquisition subunit, which is used to acquire attribute information of all software packages from the Packages file. The attribute information includes first attribute information, second attribute information, and third attribute information; the first attribute is used as an attribute field for describing a software package entity and written into a first file; the second attribute information is used to establish a relationship between a binary installation package and a source package, and the third attribute information is used to establish a dependency and conflict relationship between binary packages or between a binary package and a virtual package and written into a second file; the first attribute information includes software name, software version, software maintainer, and software function description, the second attribute information includes source package name, and the third attribute information includes dependency relationship and conflict relationship; the first file is a pkgNode.csv file, and the second file is a relationship file;
[0036] A parsing subunit, which is used to parse the filesystem.manifest package list file to obtain software package names and software version information included in an operating system version, create an operating system image node, and write the support relationship between the image node and a repository source suitNode, and the inclusion relationship between the image node and a software package into a third file. The third file is a relationship csv file;
[0037] A construction subunit, which is used to import file information of the first file, the second file, and the third file into a graph database by using an import tool to construct an operating system knowledge graph.
[0038] Further, the parsing unit further includes:
[0039] A creating subunit, configured to create suitNode.csv, componentNone.csv node files and the inclusion relationship file between the binary packages according to the organizational structure of the repository source.
[0040] Further, the obtaining module includes:
[0041] A first writing unit, configured to query all nodes and relationship names in the knowledge graph, and write them into the entities.dict and relations.dict files respectively as entity and relationship dictionaries;
[0042] A second writing unit, configured to obtain all relationships in the knowledge graph, and write the above data into the train.txt, test.txt and valid.txt files respectively according to the ratio of 8:1:1 as the training set, test set and validation set;
[0043] A third obtaining unit, configured to take the entities.dict, relations.dict files, train.txt, test.txt and valid.txt files as parameters and pass them into the RotatE embedding model, model various relationship patterns through the training set in train.txt, and correct and evaluate the modeling results through the test set and validation set to obtain the vector representations of relationships and nodes.
[0044] The beneficial effects achieved by the present invention are:
[0045] The present invention provides a method and system for evaluating the similarity of operating system versions based on a knowledge graph. An operating system knowledge graph is constructed based on the software package information in the target operating system repository and the software package list information in the target operating system image. The RotatE embedding algorithm is used to train the operating system knowledge graph to obtain the vector representation of the target operating system entity. According to the vector representation of the target operating system entity, the similarity function of the entity is used to evaluate the similarity degree between two operating system versions. The method and system for evaluating the similarity of operating system versions based on the knowledge graph provided by the present invention construct an operating system knowledge graph by using the information of the software source repository and the image, fully considering the various attribute values of the software and the relationships between them, which is more accurate and reasonable than the traditional method of comparing software names and version differences. Representing the operating system image entity with low-dimensional vectors and using an evaluation function to calculate the similarity between two versions is more intuitive and easier to compare than the traditional multi-index description of differences, and also provides data support for the evolution analysis of operating system versions. It can realize the similarity evaluation of different operating system versions, quantify the differences between operating system versions, make full use of the relationships and attribute information between operating system versions and the software packages that make them up, provide support for the evolution analysis of operating system versions, and has the advantages of reliability, intuitiveness, and easy comparative analysis. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1 It is a schematic flowchart of the method for evaluating the similarity of operating system versions based on the knowledge graph of the present invention;
[0047] Figure 2 It is an organizational structure diagram of the software repository of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0048] In order to better understand the above technical solutions, the following will describe the above technical solutions in detail in conjunction with the accompanying drawings of the specification and specific embodiments.
[0049] As Figure 1 and Figure 2 shown, a first embodiment of the present invention proposes a method for evaluating the similarity of operating system versions based on a knowledge graph, including:
[0050] Step S100: Construct an operating system knowledge graph based on the software package information in the target operating system repository and the software package list information in the target operating system image.
[0051] Extract all software package information from the target operating system repository, extract the software package list information of the build version from the target operating system image, and construct an operating system knowledge graph based on the above extracted information.
[0052] Step S200: Train the operating system knowledge graph using the RotatE embedding algorithm to obtain the vector representation of the target operating system entity.
[0053] Train the operating system knowledge graph using the RotatE embedding algorithm to obtain the vector representation of the target operating system entity.
[0054] Step S300: According to the vector representation of the target operating system entity, use the entity similarity function to evaluate the similarity between two operating system versions.
[0055] Based on the vector representation of the target operating system entity obtained in Step S200, use the entity similarity function to evaluate the similarity between two operating system versions.
[0056] In this embodiment, the construction of the operating system knowledge graph is the foundation, and the modeling of entities and relationships is related to the accuracy of the calculation results. Taking the Ubuntu22.04 image and the Ubuntu20.04 image as examples, there are hundreds of software packages in the most original systems. Therefore, to calculate the similarity of the images, it is necessary to compare the software package situations included in the images, and the information of these packages can be obtained from the filesystem.manifest file. Just from the comparison of the package lists, there may be differences in the presence or absence of several software packages, or they may all exist, but only the versions are inconsistent. However, these software packages are also complex individuals, and they are not isolated nodes. Different versions may mean various complex situations such as different provided symbols, changed dependency relationships, and changed internal binary files. Therefore, it is necessary to add the software package information to the evaluation system to improve the accuracy and rationality of the comparison. As known from the background technology section, the software repository contains software list files, and parsing these files can obtain information such as the name, version number, dependency relationship, conflict relationship, and file size of the software packages. Therefore, the software repository is also incorporated into the evaluation system, and nodes such as iso, suit, component, pkg, and source will be created, and various types of relationships such as compilation relationship, conflict relationship, dependency relationship, inclusion relationship, and support relationship will be introduced to form a knowledge graph in the field of operating systems.
[0057] Neo4j is a mature graph database system. The free community version can be downloaded, which provides sound functions such as creating a graph database, data import and export, and query, and is used as the system tool for storing the knowledge graph in this embodiment.
[0058] Further, please refer to Figures 1 to 2, for the operating system version similarity evaluation method based on knowledge graph proposed in this embodiment, the software package information includes the Packages file and the repository source organizational structure information, and the software package list information includes the filesystem.manifest package list file. Step S100 includes:
[0059] Step S110, obtain all Packages files and the repository source organizational structure information in the target operating system repository.
[0060] Step S120, obtain the filesystem.manifest package list file in the target operating system image.
[0061] Step S130, parse the Package file and the filesystem.manifest package list file, create entities and relationships in the graph database, and construct the operating system knowledge graph.
[0062] Preferably, see Figures 1 to 2 , for the operating system version similarity evaluation method based on knowledge graph proposed in this embodiment, step S130 includes:
[0063] Step S131, obtain the attribute information of all software packages from the Packages file. The attribute information includes the first attribute information, the second attribute information, and the third attribute information; use the first attribute as the attribute field for describing the software package entity and write it into the first file; use the second attribute information to establish the relationship between the binary installation package and the source code package, and use the third attribute information to establish the dependency and conflict relationships between binary packages or between binary packages and virtual packages, and write them into the second file; the first attribute information includes the software name, software version, software maintainer, and software function description, the second attribute information includes the source code package name, and the third attribute information includes the dependency relationship and the conflict relationship; the first file is the pkgNode.csv file, and the second file is the relationship file.
[0064] Obtain the attribute information of all packages from the Packages file, including software name, software version, source package name, software maintainer, software function description, dependency relationship, conflict relationship, etc. Among them, attributes such as software name, software version, software maintainer, and software function description are written as attribute fields describing the package entity into the pkgNode.csv file, while the source package name is used to establish the relationship between the binary installation package and the source package, and the dependency relationship and conflict relationship fields are used to establish the dependency and conflict relationships between binary packages or between binary packages and virtual packages, and are written into the relationship file. In addition, create node files such as suitNode.csv and componentNone.csv and inclusion relationship files with binary packages according to the repository source organization structure. Taking the pkgNode.csv file as an example, the node file format is as follows:
[0065] packageID:LABEL, pkgname, version, priority, maintainer, architecture, depends, predepends
[0066] pkg2024-01-10 17:01:35.690084, Package, libnvidia-common-390, 390.132-0ubuntu2, optional, Ubuntu Core Developers <ubuntu-devel-discuss@lists.ubuntu.com>, all, null,
[0067] pkg2024-01-10 17:01:35.690195, Package, libnvidia-common-418, 430.50-0ubuntu3, optional, Ubuntu Core Developers <ubuntu-devel-discuss@lists.ubuntu.com>, all, libnvidia-common-430,
[0068] pkg2024-01-10 17:01:35.690259, Package, libnvidia-common-430, 430.50-0ubuntu3, optional, Ubuntu Core Developers <ubuntu-devel-discuss@lists.ubuntu.com>, all, null,
[0069] pkg2024-01-10 17:01:35.690321, Package, libnvidia-common-430, 440.82+really.440.64-0ubuntu6, optional, Ubuntu Core Developers <ubuntu-devel-discuss@lists.ubuntu.com>, all, libnvidia-common-440,
[0070] ………………
[0071] Taking the dependsRel.csv file as an example, the format of the relationship file is as follows:
[0072] Type, fpackageID:START_ID, tpackageID:END_ID, rversion, rlimitdepends, 21, 20, null, null
[0073] depends, 22, 21, 0.0.5+16.04.20160307-0kord1, >=
[0074] depends, 27, 25, null, null
[0075] depends, 29, 28, null, null
[0076] depends, 34, 35, null, null.
[0077] Step S132: Parse the filesystem.manifest package list file to obtain the package names and software version information included in the operating system version, create an operating system image node, and write the support relationship between the image node and the repository source suitNode, as well as the inclusion relationship between the image node and the package, into a third file, which is a relationship csv file.
[0078] Parse the package list file in the operating system image to obtain the package names and software version information included in the operating system version, create an operating system image node, and write the support relationship between the image node and the repository source suitNode, the inclusion relationship between the image node and the package, etc. into the relationship csv file.
[0079] Step S133: Use the import tool to import the file information of the first file, the second file, and the third file into the graph database to construct an operating system knowledge graph.
[0080] Use the import tool to import the information of all the above entity csv files and relationship csv files into the Neo4j graph database to construct an operating system knowledge graph.
[0081] After the knowledge graph is created, the required basic information has been gathered in the graph database, and it is still in the state of discrete symbolic information representation. Compared with the traditional method, this embodiment transforms the similarity evaluation of operating system versions into a computational problem, uses knowledge graph embedding technology to realize the vector representation of discrete information of entities and relationships, and further performs computational operations on the vectors. The RotatE embedding model defines each relationship as a rotation from the source entity to the target entity in the complex vector space, can model and infer various relationship patterns, uses the information in the graph to train the vectors of two operating system versions, and further substitutes them into the similarity function to obtain the result.
[0082] Furthermore, see Figures 1 to 2 , the method for evaluating the similarity of operating system versions based on the knowledge graph proposed in this embodiment, step S200 includes:
[0083] Step S210: Query all node and relationship names in the knowledge graph and write them into the entities.dict and relations.dict files respectively as entity and relationship dictionaries.
[0084] Query all node and relationship names in the knowledge graph and write them into the entities.dict and relations.dict files respectively as entity and relationship dictionaries. For example, the format of the relations.dict file is as follows:
[0085] osInclude
[0086] depends
[0087] conflicts
[0088] supportedBy
[0089] build
[0090] suitInclude
[0091] ……
[0092] Step S220: Obtain all the relationships in the knowledge graph and write the above data into the train.txt, test.txt, and valid.txt files respectively according to the ratio of 8:1:1 as the training set, test set, and validation set.
[0093] Obtain all the relationships in the knowledge graph and write the above data into the train.txt, test.txt, and valid.txt files according to the ratio of 8:1:1 as the training set, test set, and validation set. The formats of the above three files are <head, relationship, tail>, as follows:
[0094] Ubuntu22.04 osInclude libgdm6
[0095] Ubuntu22.04 osInclude apt
[0096] Ubuntu22.04 osInclude openssl
[0097] quickcal depends python3
[0098] caja depends libgail-3-0
[0099] Ubuntu22.04 supportedBy mainComp
[0100] ………
[0101] Step S230: Pass the entities.dict, relations.dict files, train.txt, test.txt, and valid.txt files as parameters into the RotatE embedding model, model various relationship patterns through the training set in train.txt, and correct and evaluate the modeling results through the test set and validation set to obtain the vector representations of relationships and nodes.
[0102] Pass the above files as parameters into the RotatE embedding model, model various relationship patterns through the training set in train.txt, and correct and evaluate the modeling results through the test set and validation set to obtain the vector representations of relationships and nodes.
[0103] Please see Figure 1 and Figure 2, the present invention also relates to an operating system version similarity evaluation system based on a knowledge graph, including a construction module, an acquisition module, and an evaluation module. Among them, the construction module is used to construct an operating system knowledge graph based on the package information in the target operating system repository and the package list information in the target operating system image. The acquisition module is used to train the operating system knowledge graph using the RotatE embedding algorithm to obtain the vector representation of the target operating system entity. The evaluation module is used to evaluate the similarity degree of two operating system versions according to the vector representation of the target operating system entity by using the entity similarity function.
[0104] The construction module extracts all package information from the target operating system repository, extracts the package list information of the build version from the target operating system image, and constructs an operating system knowledge graph according to the above extracted information.
[0105] The acquisition module trains the operating system knowledge graph using the RotatE embedding algorithm to obtain the vector representation of the target operating system entity.
[0106] The evaluation module evaluates the similarity degree of two operating system versions according to the vector representation of the target operating system entity obtained by the acquisition module by using the entity similarity function.
[0107] Further, refer to Figure 1 and Figure 2 , for an operating system version similarity evaluation system based on a knowledge graph provided in this embodiment, the construction module includes a first acquisition unit, a second acquisition unit, and an analysis unit. Among them, the first acquisition unit is used to acquire all Packages files and repository source organization structure information in the target operating system repository. The second acquisition unit is used to acquire the filesystem.manifest package list file in the target operating system image. The analysis unit is used to analyze the Package file and the filesystem.manifest package list file, create entities and relationships in the graph database, and construct an operating system knowledge graph.
[0108] Preferably, refer to Figure 1 and Figure 2, the operating system version similarity evaluation system based on a knowledge graph provided in this embodiment, the parsing unit includes an acquisition subunit, a parsing subunit, and a construction subunit. Among them, the acquisition subunit is used to obtain the attribute information of all software packages from the Packages file. The attribute information includes first attribute information, second attribute information, and third attribute information; the first attribute is used as an attribute field for describing the software package entity and written into the first file; the second attribute information is used to establish the relationship between the binary installation package and the source code package, and the third attribute information is used to establish the dependency and conflict relationships between binary packages or between binary packages and virtual packages and written into the second file; the first attribute information includes software name, software version, software maintainer, and software function description, the second attribute information includes the source code package name, and the third attribute information includes dependency relationships and conflict relationships; the first file is the pkgNode.csv file, and the second file is the relationship file. The parsing subunit is used to parse the filesystem.manifest package list file to obtain the software package names and software version information included in the operating system version, create an operating system image node, and write the support relationship between the image node and the repository source suitNode, as well as the inclusion relationship between the image node and the software package into the third file, and the third file is the relationship csv file. The construction subunit is used to import the file information of the first file, the second file, and the third file into the graph database using the import tool to construct an operating system knowledge graph.
[0109] The acquisition subunit obtains the attribute information of all software packages from the Packages file, including software name, software version, source code package name, software maintainer, software function description, dependency relationship, conflict relationship, etc. Among them, attributes such as software name, software version, software maintainer, and software function description are written into the pkgNode.csv file as attribute fields for describing the software package entity, while the source code package name is used to establish the relationship between the binary installation package and the source code package, and the dependency relationship and conflict relationship fields are used to establish the dependency and conflict relationships between binary packages or between binary packages and virtual packages and written into the relationship file. In addition, node files such as suitNode.csv and componentNone.csv and inclusion relationship files with binary packages are created according to the repository source organizational structure.
[0110] The parsing subunit parses the package list file in the operating system image to obtain the software package names and software version information included in the operating system version, creates an operating system image node, and writes the support relationship between the image node and the repository source suitNode, the inclusion relationship between the image node and the software package, etc. into the relationship csv file.
[0111] The construction subunit uses the import tool to import the information of all the above entity csv files and relationship csv files into the Neo4j graph database to construct an operating system knowledge graph.
[0112] Further, see Figure 1 and Figure 2 , in the operating system version similarity evaluation system based on a knowledge graph provided in this embodiment, the parsing unit further includes a creation subunit, and the creation subunit is used to create the suitNode.csv and componentNone.csv node files and the inclusion relationship file between the binary packages according to the repository source organizational structure.
[0113] Further, see Figure 1 and Figure 2 , in the operating system version similarity evaluation system based on a knowledge graph provided in this embodiment, the acquisition module includes a first writing unit, a second writing unit, and a third acquisition unit. Among them, the first writing unit is used to query all the node and relationship names in the knowledge graph and write them into the entities.dict and relations.dict files respectively as entity and relationship dictionaries. The second writing unit is used to obtain all the relationships in the knowledge graph and write the above data into the train.txt, test.txt, and valid.txt files respectively according to the ratio of 8:1:1 as the training set, test set, and validation set. The third acquisition unit is used to pass the entities.dict, relations.dict files, train.txt, test.txt, and valid.txt files as parameters into the RotatE embedding model, model various relationship patterns through the training set in train.txt, and correct and evaluate the modeling results through the test set and validation set to obtain the vector representations of relationships and nodes.
[0114] The first writing unit queries all the node and relationship names in the knowledge graph and writes them into the entities.dict and relations.dict files respectively as entity and relationship dictionaries.
[0115] The second writing unit obtains all the relationships in the knowledge graph and writes the above data into the train.txt, test.txt, and valid.txt files respectively according to the ratio of 8:1:1 as the training set, test set, and validation set.
[0116] The third acquisition unit passes the above files as parameters into the RotatE embedding model, models various relationship patterns through the training set in train.txt, and corrects and evaluates the modeling results through the test set and validation set to obtain the vector representations of relationships and nodes.
[0117] The method and system for evaluating the similarity of operating system versions based on a knowledge graph provided in this embodiment, compared with the prior art, construct an operating system knowledge graph based on the software package information in the target operating system repository and the software package list information in the target operating system image; use the RotatE embedding algorithm to train the operating system knowledge graph to obtain the vector representation of the target operating system entity; according to the vector representation of the target operating system entity, use the similarity function of the entity to evaluate the similarity degree of two operating system versions. The method and system for evaluating the similarity of operating system versions based on a knowledge graph provided in this embodiment construct an operating system knowledge graph by using the information of the software source repository and the image, fully considering the various attribute values of the software and the relationships between them, which is more accurate and reasonable than the traditional method of comparing software names and version differences; represent the operating system image entity with a low-dimensional vector, and use the evaluation function to calculate the similarity of two versions, which is more intuitive and easier to compare than the traditional description of differences by multiple indicators, and also provides data support for the evolution analysis of operating system versions; can realize the evaluation of the similarity of different operating system versions, quantify the differences between operating system versions, make full use of the relationships and attribute information between operating system versions and the software packages that make them up, provide support for the evolution analysis of operating system versions, and have the advantages of reliability, intuitiveness, and easy comparative analysis.
[0118] Although the preferred embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications to these embodiments once they learn the basic creative concept. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications falling within the scope of the present invention. Obviously, those skilled in the art can make various changes and variations to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.
Claims
1. A method for evaluating operating system version similarity based on knowledge graph, characterized in that: The following steps are involved: Based on the software package information in the target operating system repository and the software package list information in the target operating system image, an operating system knowledge graph is constructed; Using the RotatE embedding algorithm to train the operating system knowledge graph to obtain a vector representation of the target operating system entity; According to the vector representation of the target operating system entity, the similarity function of the entity is used to evaluate the similarity between the two operating system versions.
2. The method for evaluating operating system version similarity based on knowledge graph according to claim 1, characterized in that: The software package information includes Packages file and warehouse source organizational structure information, the software package list information includes filesystem.manifest package list file, and the step of constructing the operating system knowledge graph based on the software package information in the target operating system warehouse and the software package list information in the target operating system image includes: Get all Packages files and warehouse source organizational structure information in the target operating system warehouse; Get the filesystem.manifest package list file in the target operating system image; Parse the Package file and the filesystem.manifest package list file, create entities and relationships in the graph database, and build an operating system knowledge graph.
3. The operating system version similarity evaluation method based on knowledge graph according to claim 2, characterized in that: The steps of parsing the Package file and the filesystem.manifest package list file, creating entities and relationships in a graph database, and constructing an operating system knowledge graph include: Obtain attribute information of all software packages from the Packages file, wherein the attribute information includes first attribute information, second attribute information, and third attribute information; write the first attribute into the first file as an attribute field describing the software package entity; use the second attribute information to establish a relationship between a binary installation package and a source package, and use the third attribute information to establish a dependency and conflict relationship between binary packages or between binary packages and virtual packages, and write them into the second file; the first attribute information includes software name, software version, software maintainer, and software function description, the second attribute information includes source package name, and the third attribute information includes dependency and conflict relationships; the first file is a pkgNode.csv file, and the second file is a relationship file; Parse the filesystem.manifest package list file, obtain the software package name and software version information contained in the operating system version, create an operating system mirror node, and write the support relationship between the mirror node and the warehouse source suitNode, and the inclusion relationship between the mirror node and the software package into a third file, wherein the third file is a relationship csv file; An import tool is used to import the file information of the first file, the second file and the third file into the graph database to construct an operating system knowledge graph.
4. The method for evaluating operating system version similarity based on knowledge graph according to claim 3, characterized in that: The step of parsing the Package file and the filesystem.manifest package list file, creating entities and relationships in a graph database, and building an operating system knowledge graph also includes: According to the warehouse source organizational structure, create the suitNode.csv, componentNone.csv node files and the inclusion relationship files between the binary packages.
5. The method for evaluating operating system version similarity based on knowledge graph according to claim 4, characterized in that: The step of training the operating system knowledge graph using the RotatE embedding algorithm to obtain a vector representation of the target operating system entity includes: Query all node and relationship names in the knowledge graph and write them into entities.dict and relations.dict files as entity and relationship dictionaries respectively; Get all the relationships in the knowledge graph and write the above data into the train.txt, test.txt and valid.txt files in a ratio of 8:1:1 as the training set, test set and validation set respectively; The entities.dict, relations.dict files, train.txt, test.txt and valid.txt files are passed as parameters to the RotatE embedding model. Various relationship patterns are modeled through the training set in train.txt, and the modeling results are corrected and evaluated through the test set and validation set to obtain the vector representation of relationships and nodes.
6. An operating system version similarity evaluation system based on knowledge graph, characterized in that: include: A construction module is used to construct an operating system knowledge graph based on the software package information in the target operating system repository and the software package list information in the target operating system image; An acquisition module, used to train the operating system knowledge graph using a RotatE embedding algorithm to obtain a vector representation of a target operating system entity; The evaluation module is used to evaluate the similarity between two operating system versions using an entity similarity function based on the vector representation of the target operating system entity.
7. The operating system version similarity evaluation system based on knowledge graph according to claim 6, characterized in that: The building blocks include: The first acquisition unit is used to acquire all Packages files and warehouse source organizational structure information in the target operating system warehouse; The second acquisition unit is used to acquire the filesystem.manifest package list file in the target operating system image; The parsing unit is used to parse the Package file and the filesystem.manifest package list file, create entities and relationships in the graph database, and build an operating system knowledge graph.
8. The operating system version similarity evaluation system based on knowledge graph according to claim 7, characterized in that: The parsing unit comprises: The acquisition subunit is used to obtain the attribute information of all software packages from the Packages file, wherein the attribute information includes first attribute information, second attribute information and third attribute information; the first attribute is used as an attribute field describing the software package entity and is written into the first file; the second attribute information is used to establish the relationship between the binary installation package and the source code package, and the third attribute information is used to establish the dependency and conflict relationship between binary packages or between binary packages and virtual packages, and is written into the second file; the first attribute information includes the software name, software version, software maintainer and software function description, the second attribute information includes the source code package name, and the third attribute information includes the dependency relationship and the conflict relationship; the first file is a pkgNode.csv file, and the second file is a relationship file; A parsing subunit is used to parse the filesystem.manifest package list file, obtain the software package name and software version information contained in the operating system version, create an operating system mirror node, and write the support relationship between the mirror node and the warehouse source suitNode, and the inclusion relationship between the mirror node and the software package into a third file, wherein the third file is a relationship csv file; A construction subunit is used to use an import tool to import the file information of the first file, the second file and the third file into the graph database to construct an operating system knowledge graph.
9. The operating system version similarity evaluation system based on knowledge graph according to claim 8, characterized in that: The parsing unit also includes: A subunit is created to create the suitNode.csv and componentNone.csv node files and the inclusion relationship files between the binary package and the suitNode.csv and componentNone.csv node files according to the warehouse source organization structure.
10. The operating system version similarity evaluation system based on knowledge graph according to claim 9, characterized in that: The acquisition module comprises: The first writing unit is used to query all node and relationship names in the knowledge graph and write them into entities.dict and relations.dict files as entity and relationship dictionaries respectively; The second writing unit is used to obtain all the relations in the knowledge graph and write the above data into the train.txt, test.txt and valid.txt files in a ratio of 8:1:1 as the training set, test set and validation set respectively; The third acquisition unit is used to pass entities.dict, relations.dict files, train.txt, test.txt and valid.txt files as parameters into the RotatE embedding model, model various relationship patterns through the training set in train.txt, and calibrate and evaluate the modeling results through the test set and validation set to obtain vector representations of relationships and nodes.
Citation Information
Patent Citations
Linux ecological dependency graph construction method based on graph database and application
CN117407047A
Container mirror image similarity evaluation method based on knowledge graph
CN118689589A
Automated Knowledge Graph Based Regression Scope Identification
US20230305815A1