Information processing device and method
Patent Information
- Application Number
- PCT/JP2025/012539
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-03-27
- Publication Date
- 2026-10-01
Smart Images

Figure JP2025012539_01102026_PF_FP_ABST
Abstract
Description
Information processing device and method
[0001] One aspect of this invention relates to an information processing device and a method.
[0002] For each of several different objects, there may be situations where identification information for objects such as equipment and goods, as well as connection relationships between objects, exist in multiple different data sources. If the identification information is assigned independently across multiple data sources, the correspondence between the identification information of objects in different data sources—that is, the combinations of identification information for the same object—is not self-evident. Therefore, when integrating and processing information held in multiple data sources using a computer, it becomes necessary to associate the identification information of the same object.
[0003] In contrast, a method is known that uses the connection relationships between objects in each data source when associating identification information of the same object (see, for example, Patent Document 1).
[0004] International Publication No. 2021 / 124525
[0005] The technologies described above may result in different approaches to the relationships between objects depending on the data source, which can lead to different sets of adjacent objects. In this case, directly using the relationship relationships in one data source as a clue to match object identification information may lead to incorrect matching.
[0006] This invention has been made in consideration of the above circumstances and provides a technique for reducing errors in associating objects with each other between data sources that represent different connection relationships.
[0007] To solve the above problems, one embodiment of the information processing apparatus according to the present invention comprises a first acquisition unit, a second acquisition unit, a derivation unit, a mapping unit, and an output unit. The first acquisition unit acquires first connection information from a first data source, which includes first identification information indicating each object within a target to which a plurality of objects are connected, and represents a first type of connection relationship for each of the objects. The second acquisition unit acquires second connection information from a second data source, which includes second identification information indicating each object within the target, and represents a second type of connection relationship for each of the objects. The derivation unit derives third connection information, which includes the first identification information and represents a second type of connection relationship assumed from the first type of connection relationship, based on the first connection information. The mapping unit maps the first identification information and the second identification information based on the third connection information and the second connection information. The output unit outputs the result of the mapping.
[0008] According to one aspect of this invention, identification information of the same type of second-type connection relationship is associated with a third connection information representing a second-type connection relationship assumed from a first-type connection relationship represented by first connection information obtained from a first data source, and a second connection information obtained from a second data source representing a second-type connection relationship. This reduces errors in associating objects between data sources that represent different connection relationships.
[0009] In other words, according to one aspect of this invention, it is possible to reduce errors in associating objects with each other between data sources that represent different connection relationships.
[0010] Figure 1 is a block diagram showing an example of the configuration of a system equipped with an information processing device according to the first embodiment of the present invention. Figure 2 is a schematic diagram illustrating an example of a first data source shown in Figure 1. Figure 3 is a schematic diagram illustrating an example of a second data source shown in Figure 1. Figure 4 is a schematic diagram illustrating an example of a third data source shown in Figure 1. Figure 5 is a schematic diagram illustrating an example of a processing unit shown in Figure 1. Figure 6 is a flowchart illustrating an example of operation in the first embodiment. Figure 7 is a flowchart illustrating an example of derivation operation in Figure 6. Figure 8 is a schematic diagram illustrating the flowchart in Figure 7. Figure 9 is a schematic diagram illustrating the flowchart in Figure 7. Figure 10 is a schematic diagram illustrating the flowchart in Figure 7. Figure 11 is a flowchart illustrating an example of mapping operation in Figure 6. Figure 12 is a schematic diagram illustrating the flowchart in Figure 11. Figure 13 is a schematic diagram illustrating a first modified example of the processing unit in the first embodiment. Figure 14 is a flowchart illustrating an example of mapping operation in the first modified example of the first embodiment. Figure 15 is a flowchart illustrating an example of the correspondence operation in a second modification of the first embodiment. Figure 16 is a schematic diagram illustrating a comparative example to the first embodiment. Figure 17 is a schematic diagram illustrating the comparative example of Figure 16. Figure 18 is a schematic diagram illustrating the comparative example of Figure 16. Figure 19 is a schematic diagram illustrating the comparative example of Figure 16. Figure 20 is a flowchart illustrating an example of the derivation operation in the second embodiment. Figure 21 is a schematic diagram illustrating the flowchart of Figure 20. Figure 22 is a schematic diagram illustrating the flowchart of Figure 20. Figure 23 is a schematic diagram illustrating the flowchart of Figure 20. Figure 24 is a schematic diagram illustrating the flowchart of Figure 20. Figure 25 is a schematic diagram illustrating the flowchart of Figure 20. Figure 26 is a schematic diagram illustrating an example of the processing unit in the third embodiment. Figure 27 is a schematic diagram illustrating an example of operation in the third embodiment. Figure 28 is a schematic diagram illustrating an example of operation in the third embodiment.Figure 29 is a schematic diagram illustrating an example of operation in the third embodiment. Figure 30 is a schematic diagram illustrating an example of operation in the fourth embodiment. Figure 31 is a schematic diagram illustrating an example of operation in the fourth embodiment. Figure 32 is a schematic diagram illustrating an example of operation in the fourth embodiment.
[0011] An embodiment of this invention will be described below with reference to the drawings. In the following description, a typical example will be given using two data sources, but the invention is not limited to this, and three or more arbitrary data sources may be used. Furthermore, in the following description, the information processing device is implemented as a personal computer, but the invention is not limited to this, and at least a part of the device may be implemented as a server computer installed on the Web or in the cloud.
[0012] <First Embodiment> Figure 1 is a block diagram showing an example of the configuration of a system equipped with an information processing device according to the first embodiment of the present invention, and Figures 2 to 4 are schematic diagrams illustrating an example of a first to third data source. The information processing device 1 shown in Figure 1 includes a program storage unit 10, a data storage unit 20, a processing unit 30, a communication interface 50, and an input / output interface 60.
[0013] Multiple data source DSs, such as the first data source DS1, the second data source DS2, the third data source DS3, etc., are wirelessly or wiredly connected to the communication interface 50 via the network Nw. When not specifically identified, each of the first data source DS1, the second data source DS2, the third data source DS3, etc., is referred to as "each data source DS." Typically, each data source DS is the first data source DS1 and the second data source DS2, respectively.
[0014] Each data source DS1, DS2, DS3, ... is a storage device that stores identification information for each object, information representing the connection relationships between objects, and additional information about each object, as shown in Figures 2 to 4. The information representing the connection relationships may also be called connection information or connection relationship data. The identification information, connection information representing the connection relationships, and additional information stored in each data source DS1, DS2, DS3, ... are all information about the same object to which each of the multiple objects is connected.
[0015] Here, an object is something that can be connected to other objects via a route such as a cable, and for example, equipment or goods are available for use as needed. Objects such as equipment and goods may include structures such as racks or devices such as switches. In the following explanation, we will use "equipment" as an object, and the structures and devices housed within the equipment object will also be called "internal equipment."
[0016] Object identification information is information that uniquely identifies an object, and may include, for example, the object's name or number, as appropriate.
[0017] As connection information between objects, graph data representing the connection relationships between components (e.g., ports) contained within an object (e.g., a device) can be used. As graph data, for example, an undirected graph is used, where object identification information is used as vertices and vertices in connection (adjacency relationships) are connected by edges. However, directed graphs may also be used. Note that even if multiple objects are connected to the same object, different connection relationships may be represented because each data source DS1, DS2, DS3, ... has a different concept of connection relationships between objects. For example, for the same object, physical connection relationships such as racks (non-switchable connection relationships) and logical connection relationships such as switches (switchable connection relationships) have different connection relationships. Hereafter, physical connection relationships will be called first-type connection relationships, and logical connection relationships will be called second-type connection relationships. However, the types of connection relationships are not limited to these. Also, there are dependencies between connection relationships held in one data source DS and connection relationships held in other data source DSs. For example, the logical connectivity of a switch can be determined by its physical connectivity.
[0018] Additional information may include attributes such as the location of the object (equipment or goods), identification information of internal equipment (contained structures or machinery), and inspection results of the object. Note that this additional information is optional and may be omitted.
[0019] Returning to Figure 1, the input / output interface 60 is connected to an input device 3 and a display device 4 via wireless or wired communication. The input device 3 consists of, for example, a keyboard and a mouse, and is used to input various information and commands to the processing unit 30 in response to user operations. The display device 4 is controlled by the processing unit 30 and is used to display data in progress or data resulting from processing.
[0020] The program storage unit 10 combines, for example, a non-volatile memory such as an SSD that can be written to and read at any time, and a volatile memory such as RAM (Random Access Memory), as storage media. The program storage unit 10 stores, for example, middleware such as an OS (Operating System), as well as application programs necessary to execute various controls according to one embodiment. Hereafter, the OS and each application program will be collectively referred to as a program. The program may be installed on the computer in advance from a network or a non-transitory computer-readable storage medium, or it may be recorded on the computer in advance. In any case, the program is executed by the processor and causes the computer to function as an information processing device 1 according to each embodiment and each modified example.
[0021] The data storage unit 20 combines, for example, a non-volatile memory such as an SSD that allows writing and reading at any time, and a volatile memory such as RAM (Random Access Memory) as a storage medium. The data storage unit 20 includes a rule storage unit 21, an acquired data storage unit 22, a temporary storage unit 23, and a result storage unit 24.
[0022] The rule storage unit 21 stores the rules set by the processing unit 30 in response to the user's operation of the input device 3. A rule is a set of rules for deriving a second connection information stored in another data source DS from a first connection information stored in a certain data source DS.
[0023] Here, we will show an example of a rule for deriving the expected connection relationships of a certain logical layer from the connection relationships of the physical layer.
[0024] <Example of a rule> This rule includes rules A-1 and A-2. This rule is used in the process of deriving equipment and connection relationships at the logical layer according to the type of components (e.g., racks and communication devices) contained inside the equipment (e.g., building) which is an object indicated by the identification information. ・Rule A-1: Rule A-1 concerns the case of equipment (racks, etc.) where the connection relationship between ports is determined by the connection information of the physical layer data source DS. Rule A-1 includes rules (A-1-1) to (A-1-3).
[0025] (A-1-1) The equipment in question shall be excluded.
[0026] (A-1-2) The connection relationship between the equipment in question and other equipment is also excluded.
[0027] (A-1-3) Add a connection relationship to each combination of other equipment to which the equipment is connected, in place of the excluded connection relationship. Rule A-2: Rule A-2 concerns the process of replacing connection relationships related to internal equipment with connection relationships related to equipment that are objects indicated by the identification information, and omitting internal equipment. Rule A-2 is as follows:
[0028] (A-2) If the internal equipment is subject to only Rule A-1, the equipment itself, which is the object indicated by the identification information, is excluded.
[0029] The above is an example of a rule. Note that this rule may also be called a derived rule. The rule storage unit 21 is an example of a storage unit that stores rules in advance for deriving logical connection relationships from physical connection relationships in response to user operations.
[0030] The acquired data storage unit 22 stores data acquired by the processing unit 30 from each data source DS1, DS2, DS3, ... in response to the user's operation of the input device 3.
[0031] The temporary storage unit 23 temporarily stores data that is being processed by the processing unit 30.
[0032] The result storage unit 24 stores the data of the processing results from the processing unit 30.
[0033] As shown in Figure 5, the processing unit 30 includes a data acquisition unit 31, a connection relationship derivation unit 32, and a mapping processing unit 33. Since the data acquisition unit 31, connection relationship derivation unit 32, and mapping processing unit 33 are all included in the processing unit 30, each of these can be replaced with "processing unit 30." Furthermore, "processing unit 30" can be replaced with "processing circuitry." The processing unit 30 sequentially executes the processing of the data acquisition unit 31, the connection relationship derivation unit 32, and the mapping processing unit 33. During this process, the processing unit 30 makes the data used in the processing of each preceding unit available for use in the processing of each subsequent unit as appropriate. For example, the data acquired by the data acquisition unit 31 is available for use in the subsequent connection relationship derivation unit 32 and mapping processing unit 33 as appropriate. Furthermore, the processing unit 30 is not limited to the processing of the data acquisition unit 31, the connection relationship derivation unit 32, and the mapping processing unit 33, but can also perform arbitrary processing in response to user operations, such as setting or changing rules, or displaying data in the process of processing.
[0034] Each of the data acquisition unit 31, connection relationship derivation unit 32, and mapping processing unit 33 is implemented by having the hardware processor of the processing unit 30 execute an application program stored in the program storage unit 10. Note that some or all of the data acquisition unit 31, connection relationship derivation unit 32, and mapping processing unit 33 may be implemented using processing circuits that include integrated circuits such as LSIs (Large Scale Integration) or ASICs (Application Specific Integrated Circuits), which are hardware circuits.
[0035] Here, the data acquisition unit 31 acquires identification information, connection information, and additional information from each data source DS individually or collectively in response to the user's operation of the input device 3, and stores the acquired information in the acquired data storage unit 22. For example, for a target to which multiple objects are connected, the data acquisition unit 31 acquires connection information from each data source DS that includes identification information indicating each object within the target and represents the connection relationship of each object. For example, the data acquisition unit 31 acquires first connection information from the first data source DS1 that includes first identification information indicating each object within the target and represents the first type of connection relationship of each object. Also, for example, the data acquisition unit 31 acquires second connection information from the second data source DS2 that includes second identification information indicating each object within the target and represents the second type of connection relationship of each object. The acquired first connection information and second connection information are stored in the acquired data storage unit 22. Note that the data acquisition unit 31 is an example of the first and second acquisition units.
[0036] The connection relationship derivation unit 32, from among the connection information obtained from each data source DS, assumes a connection relationship represented by one connection information and derives new connection information representing the assumed connection relationship. For example, based on the first connection information, the connection relationship derivation unit 32 derives third connection information that includes first identification information and represents a second type of connection relationship assumed from a first type of connection relationship. For example, the connection relationship derivation unit 32 may derive third connection information from first connection information based on rules in the rule storage unit 21. Specifically, the first type of connection relationship may be a physical connection relationship. The second type of connection relationship may be a logical connection relationship. That is, the connection relationship derivation unit 32 may derive third connection information that represents a logical connection relationship assumed from the physical connection relationship represented by first connection information based on rules in the rule storage unit 21. Note that the connection relationship derivation unit 32 is an example of a derivation unit.
[0037] The mapping processing unit 33 associates the first identification information with the second identification information based on the third connection information and the second connection relationship. The mapping processing unit 33 also outputs the result of the association and displays it on the display device 4. Note that the mapping processing unit 33 is an example of a mapping unit and an output unit.
[0038] Next, the operation of the information processing device configured as described above will be explained using Figures 6 to 12. The following explanation will use the case where the first connection information in the first data source DS1 represents the connection relationship at the physical layer, and the second connection information in the second data source DS2 represents the connection relationship at the logical layer as an example. Note that the connection relationship at the physical layer refers to a physical connection relationship, such as a fixed connection relationship by a rack. The connection relationship at the logical layer refers to a logical connection relationship, such as a switchable connection relationship by a switch.
[0039] First, the data acquisition unit 31 within the processing unit 30 acquires first connection information from the first data source DS1 for a target to which multiple objects are connected, which includes first identification information indicating each object within the target and represents the physical connection relationship of each object. The data acquisition unit 31 also stores the acquired first connection information in the acquired data storage unit 22.
[0040] Similarly, the data acquisition unit 31 acquires second connection information from the second data source DS2, which includes second identification information indicating each object within the target and represents the logical connection relationship of each object. The data acquisition unit 31 also stores the acquired second connection information in the acquired data storage unit 22. In this state, the processing unit 30 executes steps ST10 to ST50 as shown in Figure 6.
[0041] In step ST10, the data acquisition unit 31 within the processing unit 30 acquires the connection relationship at the physical layer by reading the first connection information in the acquired data storage unit 22.
[0042] After step ST10, in step ST20, the connection relationship derivation unit 32 derives the assumed connection relationships at the logical layer based on the connection relationships at the physical layer. Specifically, the connection relationship derivation unit 32 derives the connection relationships based on rules A-1 and A-2 in the rule storage unit 21. This step ST20 is performed by steps ST21 to ST26, as shown in Figures 7 to 10.
[0043] In step ST21, the connection relationship derivation unit 32 identifies internal equipment and connection relationships that are unnecessary in the logical layer, based on the connection relationships at the physical layer, as shown in Figure 8.
[0044] In Figure 8, Equipment (a) to Equipment (ko) are equipment (objects) that house devices and are identified by the identification information of the first connection information. Equipment (a) to Equipment (ka) and Equipment (ke) to Equipment (ko) house internal equipment (such as switches) whose connection relationships are not determined at the physical layer, as shown by the shaded rectangles. Equipment (u) to Equipment (ku) house internal equipment (such as racks) whose connection relationships are determined at the physical layer, as shown by the white rectangles. The straight lines connecting the internal equipment of adjacent equipment represent connection relationships at the core wire level, and each line represents a single line that is representative of a collection of multiple core wires. The cylinder superimposed on the core wire (straight line) represents the collection of multiple core wires (such as tape or cable).
[0045] In the example shown in Figure 8, the connection relationship derivation unit 32 identifies the internal equipment and its connection relationships, indicated by white rectangles, as unnecessary internal equipment and connection relationships in the logical layer.
[0046] In step ST22, the connection relationship derivation unit 32 excludes the identified internal equipment and connection relationships based on rule A-1, as shown in Figure 9.
[0047] In step ST23, the connection relationship derivation unit 32 adds a connection relationship to replace the excluded internal equipment and connection relationship, based on rule A-1. For example, if the internal equipment (such as a rack) of equipment (u) that relays the connection between equipment (a) and equipment (ke) and its connection relationship are excluded, a connection relationship that directly connects equipment (a) and equipment (ke) is added.
[0048] In step ST24, the connection relationship derivation unit 32 replaces the connection relationships of the internal equipment with the connection relationships of equipment that have identification information, based on rule A-2, and omits the internal equipment.
[0049] In step ST25, the connection relationship derivation unit 32 excludes equipment if there is equipment for which all internal equipment has been excluded, based on rule A-2. For example, equipment (ki) and equipment (ku) for which all internal equipment has been excluded are excluded.
[0050] In step ST26, the connection relationship derivation unit 32 assumes the processing results applied up to rule A-2 as connection relationships at the logical layer, as shown in Figure 10. Thus, the connection relationship derivation unit 32 derives the assumed connection relationships at the logical layer. Step ST20 ends upon completion of step ST26.
[0051] Returning to Figure 6, in step ST30, the data acquisition unit 31 acquires the actual connection relationship at the logical layer by acquiring the second connection information in the acquired data storage unit 22.
[0052] After step ST30, in step ST40, the correspondence processing unit 33 associates equipment with assumed connection relationships at the logical layer with equipment with actual connection relationships at the logical layer. This step ST40 is performed by steps ST41 to ST43, as shown in Figures 11 and 12. In Figure 12, the upper row shows the assumed connection relationships, and the lower row shows the actual connection relationships.
[0053] In step ST41, the correspondence processing unit 33 calculates the degree of similarity between the derived assumed connection relationship and the actual connection relationship of the logical layer.
[0054] In step ST42, the correspondence processing unit 33 calculates the similarity of the names of the equipment between the first data source DS1 and the second data source DS2.
[0055] In step ST43, the correspondence processing unit 33 associates the names of the equipment between the first data source DS1 and the second data source DS2 based on the similarity scores calculated in steps ST41 and ST42.
[0056] In the example in Figure 12, "Equipment (a)" and "Equipment (A)", "Equipment (b)" and "Equipment (B)", "Equipment (c)" and "Equipment (C)", etc., each represent the name of the same equipment in both data sources DS1 and DS2.
[0057] Here, the degree of similarity between the names of the equipment in both data sources DS1 and DS2 for the same equipment is the same.
[0058] Furthermore, the similarity of the names of different pieces of equipment in both data sources DS1 and DS2 (for example, "Equipment (a)" and "Equipment (b)") is the same in both cases, and is lower than the similarity of names for the same piece of equipment.
[0059] In this situation, in the example in Figure 12, the assumed connection relationships of the logic layer and the actual connection relationships are the same. For example, in the assumed connection relationships, the set of adjacent equipment to equipment (ke) is {equipment (a), equipment (o), equipment (ka)}, and the set of adjacent equipment to equipment (ko) is {equipment (i), equipment (u), equipment (e)}. In contrast, in the actual connection relationships, the set of adjacent equipment to equipment (ke) is {equipment (a), equipment (o), equipment (ka)}, and the set of adjacent equipment to equipment (ko) is {equipment (i), equipment (u), equipment (e)}.
[0060] Therefore, the correspondence between equipment (ke) and equipment (ke), and between equipment (ko) and equipment (ko), as per the correct answer, will have a higher degree of similarity than any other correspondence. In this way, the correspondence processing unit 33 associates equipment with the assumed connection relationship with equipment with the actual connection relationship of the logical layer. Step ST40 ends when step ST43 is completed.
[0061] Returning to Figure 6, in step ST50, the data acquisition unit 31 outputs the associated results and displays them on the display device 4.
[0062] As described above, according to the first embodiment, the data acquisition unit 31 acquires first connection information from the first data source DS1, which includes first identification information indicating each object within a target to which multiple objects are connected, and represents a first type of connection relationship for each of the objects. The data acquisition unit 31 acquires second connection information from the second data source DS2, which includes second identification information indicating each object within the target, and represents a second type of connection relationship for each of the objects. The connection relationship derivation unit 32 derives third connection information, which includes first identification information and represents a second type of connection relationship assumed from the first type of connection relationship, based on the first connection information. The mapping processing unit 33 maps the first identification information and the second identification information based on the third connection information and the second connection relationship. The mapping processing unit 33 outputs the result of the mapping. In this way, based on the third connection information representing the second type of connection relationship that can be assumed from the first type of connection relationship represented by the first connection information obtained from the first data source DS1, and the second connection information obtained from the second data source DS2 representing the second type of connection relationship, identification information of the same type of second type connection relationship is associated with each other. This reduces errors in associating objects between data sources that represent different types of connection relationships.
[0063] Furthermore, according to the first embodiment, the rule storage unit 21 stores rules in advance for deriving a second type of connection relationship from a first type of connection relationship in response to user operations. The connection relationship derivation unit 32 derives third connection information from first connection information based on these rules. In this way, the rules used for derivation are prepared in advance by the user based on what kind of connection information is held for the combination of data sources DS1 and DS2 that are the subject of derivation. Therefore, when deriving a second type of connection relationship from a first type of connection relationship, the correct connection relationship can be easily assumed. Also, even if the correct connection relationship cannot be assumed, the user can modify the rules to make the correct connection relationship achievable.
[0064] Furthermore, according to the first embodiment, the first type of connection relationship is a physical connection relationship. The second type of connection relationship is a logical connection relationship. The connection relationship derivation unit 32 derives third connection information that represents a logical connection relationship assumed from the physical connection relationship represented by the first connection information, based on a rule. Therefore, by deriving a logical connection relationship assumed from a physical connection relationship and associating identification information between the same logical connection relationships, it is possible to reduce errors in associating objects between data sources representing physical connection relationships and data sources representing logical connection relationships.
[0065] <First Modification> In the first embodiment, the mapping was performed based on the similarity between the assumed connection relationships in the logical layer and the actual connection relationships in the logical layer, but the embodiment is not limited to this. Specifically, the mapping may be performed based on the cost of editing the assumed connection relationships in the logical layer to the actual connection relationships in the logical layer. For example, the processing unit 30 further includes a cost calculation unit 34, as shown in Figure 13.
[0066] Here, the cost calculation unit 34 edits the expected second type of connection relationship for each combination of the first identification information and the second identification information to the second type of connection relationship represented by the second connection information, calculates the cost corresponding to the editing, and identifies the combination that minimizes the cost. The mapping processing unit 33 then associates the first identification information and the second identification information with the identified combination. A logical connection relationship is used as an example of the second type of connection relationship.
[0067] Specifically, the cost calculation unit 34 edits the assumed graph Ge, which represents the assumed connection relationships of the logic layer derived in step ST20, and edits it so that it is identical in type to the actual graph Ga, which represents the actual connection relationships held in the data source DS2 of the logic layer. That is, the cost calculation unit 34 calculates a first cost g related to editing such as adding or deleting vertices and edges of the assumed graph Ge. 0 We find the correspondence x between vertices that minimizes (x) as a combinatorial optimization problem. That is, x is the first cost g 0(x) is a correspondence between the hypothetical graph Ge and the actual graph Ga such that it is the edit distance. x is a vector variable whose elements are scalar variables that take the value of 0 or 1 depending on whether, for example, a vertex in the hypothetical graph Ge is associated with a vertex in the actual graph Ga. In this case, the first cost g 0 For each edit that makes up (x), the following costs are used: - The cost of deleting a vertex from the assumed graph Ge is either a positive value or not allowed (cost is ∞). - The cost of adding a vertex to the assumed graph Ge is either a positive value or not allowed (cost is ∞). - The cost of deleting an edge from the assumed graph Ge is either a positive value or not allowed (cost is ∞). - The cost of adding an edge to the assumed graph Ge is either a positive value or not allowed (cost is ∞).
[0068] As shown in Figure 14, the cost calculation unit 34 executes step ST40, which includes steps ST41A to ST42A. That is, the cost calculation unit 34 calculates the first cost g related to editing between the assumed graph Ge and the actual graph Ga. 0 (x) is calculated (step ST41A). The cost calculation unit 34 also calculates the first cost g 0 The correspondence between vertices that minimizes (x) is identified (step ST42A). As a result, the correspondence processing unit 33 obtains the correspondence result of identification information between objects corresponding to the correspondence result between vertices.
[0069] According to the first modification described above, the cost calculation unit 34 edits the assumed logical connection relationship for each combination of the first identification information and the second identification information to the logical connection relationship represented by the second connection information, calculates the cost corresponding to the editing, and identifies the combination that minimizes the cost. The correspondence processing unit 33 associates the first identification information and the second identification information with the identified combination. This makes it possible to obtain the same effect as in the first embodiment. Furthermore, according to the first modification, correspondence between identification information can be performed using the dissimilarity (cost related to editing) between the graph structures corresponding to the assumed connection relationship and the actual connection relationship, respectively.
[0070] <Second Modification> Similar to the first modification, the second modification executes cost-based association through a configuration including a cost calculation unit 34.
[0071] Here, the cost calculation unit 34 calculates each cost g constituting the second cost for association between vertices as follows separately 1 (x), g 2 (x), ... are calculated separately. Further, the cost calculation unit 34 calculates the total overall cost f(x) obtained by summing the aforementioned first cost g 0 (x) and the second cost g 1 (x)+g 2 (x)+g 3 (x)+..., and obtains the association x that minimizes the overall cost f(x) as a combinatorial optimization problem.
[0072] f(x)=g 0 (x)+g 1 (x)+g 2 (x)+g 3 (x)+... It should be noted that the minimum value of the first cost g 0 (x) corresponds to the edit distance of an unlabeled graph. A label means identification information.
[0073] The second cost g 1 (x)+g 2 (x)+g 3 (x)+... is the dissimilarity between individual (labels of) vertices, and corresponds to the label replacement cost.
[0074] The minimum value of the overall cost f(x) corresponds to the edit distance of a labeled graph.
[0075] For each cost constituting the second cost, for example, the following may be used.
[0076] - The similarity of facility names held in each data source DS is obtained, and a monotonically decreasing value based on this similarity may be used as the cost g 1 (x).
[0077] - For additional information held as additional information in each data source DS, such as facilities and types, an evaluation value of the similarity between additional information representing attributes other than names and the degree of matching with expected relationships is obtained, and a monotonically decreasing value based on these evaluation values is set as the cost g 2(x) = g 2,1 (x) + g 2,2 It may also be used as (x) + .... An example of calculating evaluation values for additional information will be described in the fifth embodiment.
[0078] - In addition to the data source DS, calculate the co-occurrence-based correspondence evaluation value for the set of equipment identification information recorded in multiple installments, and assign a value that decreases monotonically to this value g. 3 It may also be used as (x). An example of calculating the evaluation value for co-occurrence will be described in the fourth embodiment.
[0079] In the case of the second variation, the method of mapping is not determined solely by the first cost related to editing between graphs, so the addition and deletion of vertices and edges are allowed (the cost is not set to infinity). Also, the total cost is given as the sum of the individual costs, but it may also be the product of the individual costs (equivalent to taking the logarithm). Furthermore, when finding the mapping between vertices that minimizes the total cost as a combinatorial optimization problem, approximation methods may be used. For example, the search space may be limited, or instead of the first cost related to editing, g may be used for the set of vertices adjacent to each vertex in the assumed graph Ge and the actual graph Ga. 1 (x) and g 2 Calculate the same value as (x) and then g 0 It may also be used as an approximation of (x).
[0080] As shown in Figure 15, the cost calculation unit 34 executes step ST40, which includes steps ST41B to ST44B. Specifically, the cost calculation unit 34 calculates a first cost related to editing between the assumed graph Ge and the actual graph Ga (step ST41B). The cost calculation unit 34 also calculates a second cost related to the correspondence between vertices in each graph Ge and Ga (step ST42B). Based on the first and second costs, the cost calculation unit 34 calculates the total cost (step ST43B). The cost calculation unit 34 also identifies the correspondence between vertices that minimizes the total cost (step ST44B). As a result, the correspondence processing unit 33 obtains the correspondence result of identification information between objects corresponding to the correspondence result between vertices.
[0081] The same effects as those of the first embodiment can be obtained with the second modified example described above. In addition, with the second modified example, since the mapping can be performed based on evaluations other than similarity regarding the string and connection relationships of the identification information, the validity of the mapping results can be improved. For example, various evaluations of the mapping other than similarity regarding the string and connection relationships, such as evaluation values for additional information or evaluation values for mapping based on co-occurrence, can be reflected in the selection of highly valid mappings.
[0082] This effect will be further explained using the comparative examples shown in Figures 16 to 19.
[0083] First, the first connection information obtained from the first data source DS1 represents the actual connection relationship at the physical layer, using identification information from equipment (a) to equipment (ko), as shown in the upper part of Figure 16. Similarly, the second connection information obtained from the second data source DS2 represents the actual connection relationship at a certain logical layer, using identification information from equipment (a) to equipment (ko), as shown in the lower part of Figure 16.
[0084] In this case, the correct correspondence between the identification information is as shown in part in Figure 17, between the first data source DS1 and the second data source DS2, with equipment(a) and equipment(a), equipment(b) and equipment(b), and so on. The similarity of the equipment names all exceed 0.8, and the total value is 4.39. Similarly, the similarity of the connection relationships all exceed 0.8, and the total value is 4.5.
[0085] In this comparative example, the first step involves matching based on either the similarity of the string of identification information or the similarity of the connection relationship. The remaining items that could not be matched are then matched using the other method in the second step to obtain the final matching result.
[0086] Figure 18 shows the case where the correspondence was performed in the previous step based on the similarity of the names of the equipment, and Figure 19 shows the case where the correspondence was performed in the previous step based on the similarity of the connection relationships.
[0087] In the example shown on the left side of Figure 18, if there are multiple matching candidates among those with a similarity score of 0.9 or higher for the equipment names, the one with the highest similarity is selected. As a result, equipment (u) is matched to equipment (e), as shown by the dashed line.
[0088] In the example shown on the right side of Figure 18, if there are multiple possible match candidates for an item that has not been matched in the previous step, and the similarity of the connection relationship is 0.9 or higher, the one with the highest similarity is selected. As a result, equipment (e) is matched to equipment (u), as shown by the dashed line. Note that the similarity of the connection relationship indicated by symbol *1 for equipment (a) through equipment (u) that were matched in the previous step does not need to be calculated, but the values are shown for reference.
[0089] In the example shown in Figure 18, the sum of the similarity scores for the equipment names is 4.18, and the sum of the similarity scores for the connection relationships is 4.26. Both sums are smaller than the sums for the correct answer shown in Figure 17. In other words, the comparative example shown in Figure 18 yields a correspondence result with a lower similarity score than the correct answer.
[0090] On the other hand, in the example shown on the left side of Figure 19, if there are multiple matching candidates among those with a similarity score of 0.9 or higher, the one with the highest similarity is selected. As a result, equipment (e) is matched to equipment (u), as shown by the dashed line.
[0091] In the example shown on the right side of Figure 19, if there are multiple matching candidates among those with a similarity score of 0.9 or higher for equipment names that were not matched in the previous step, the one with the highest similarity is selected. As a result, equipment (a) is matched to equipment (d), as shown by the dashed line. Similarly, equipment (u) is matched to equipment (a), as shown by the dashed line. Note that for equipment (b), equipment (d), and equipment (e) that were matched in the previous step, the similarity scores of the equipment names indicated by the symbol *1 do not need to be calculated, but the values are shown for reference.
[0092] In the example shown in Figure 19, the sum of the similarity scores for the equipment names is 4.21, and the sum of the similarity scores for the connection relationships is 3.69. Both sums are smaller than the sums for the correct answer shown in Figure 17. In other words, the comparative example shown in Figure 19 yields a correspondence result with a lower similarity score than the correct answer.
[0093] As shown in Figures 18 and 19, in the comparative example, the combinations whose mapping was determined in the preceding step include those with low similarity used in the subsequent step, resulting in a low validity of the mapping result. In such a comparative example, even if there are other mappings with higher overall validity in terms of both identification information and connection relationships, the mapping based on similarity in the subsequent step is not changed for the combination whose mapping was determined in the preceding step, or the method for doing so is unclear, so a mapping with higher overall validity cannot be obtained. This concludes the explanation of the comparative example.
[0094] In contrast, according to the second modified example, unlike the comparative example which prioritizes the evaluation of either identification information or connection relationships when performing the mapping, the mapping can be performed based on an overall evaluation (overall cost) that integrates the evaluation of connection relationships and the evaluation of identification information, etc. Therefore, according to the second modified example, a highly valid mapping result can be obtained from the overall perspective of identification information and connection relationships.
[0095] <Third Modification> The third modification, like the first or second modification, performs cost-based mapping with a configuration that includes a cost calculation unit 34. However, the third modification concerns cases where there are three or more data source DSs that are the target of object identification information mapping.
[0096] In this case, the cost calculation unit 34 associates the i-th data source DSi and the j-th data source DSj among the three or more data sources DS with respect to x i,j The cost of f i,j (x i,j ) is calculated as follows. In addition, the cost calculation unit 34 calculates the cost f as shown in equation (1). i,j (x i,j The sum of the above is calculated. Alternatively, the cost calculation unit 34 calculates the cost f in a simplified manner as shown in equation (2). i,j(x i,j Find the sum of a part of ).
[0097]
[0098] Subsequently, the cost calculation unit 34 calculates the cost f i,j (x i,j The correspondence x that minimizes the sum of the sums 1,2 , x 1,3 ...is identified. As a result, the correspondence processing unit 33 obtains the correspondence result of identification information between objects corresponding to the correspondence result between vertices.
[0099] Even with the third modification described above, the same effects as the first embodiment can be obtained.
[0100] <Second Embodiment> The second embodiment is a modification of the first embodiment and relates to the processing of the connection relationship derivation unit 32 when the connection destination of the internal equipment can be switched.
[0101] Here, in addition to the functions described above, the connection relationship derivation unit 32 derives third connection information to represent the connection relationships in which an object can be logically connected to two or more objects, if one of the objects can be logically connected to two or more objects.
[0102] Furthermore, the rule storage unit 21 stores rules B-1, B-2, and B-3. Here are some examples of rules B-1, B-2, and B-3: • Rule B-1: Same as rule A-1. • Rule B-2: In the case of equipment (such as switches) where the connection relationship between ports is not determined by the data source information at the physical layer, this rule is stored as an additional connection relationship for each combination of other equipment to which the equipment is connected. • Rule B-3: Same as rule A-2.
[0103] The other configurations are the same as in the first embodiment.
[0104] Next, an example of the operation of the information processing device configured as described above will be explained using Figures 20 to 25. The following explanation will mainly describe step ST20, which differs from that of the first embodiment. Furthermore, this example of operation represents a case where the first connection information in the first data source DS1 and the second connection information in the second data source can switch the connection destination of the internal equipment.
[0105] Now, after step ST10 is executed as described above, step ST20 will begin.
[0106] In step ST20, the connection relationship derivation unit 32 derives the assumed connection relationships at the logical layer based on the connection relationships at the physical layer. Specifically, the connection relationship derivation unit 32 derives the connection relationships based on rules B-1, B-2, and B-3 in the rule storage unit 21. This step ST20 is performed by steps ST21 to ST23, ST23C, and ST24 to ST26, as shown in Figures 20 to 25.
[0107] In step ST21, the connection relationship derivation unit 32 identifies internal equipment and connection relationships that are unnecessary in the logical layer, based on the connection relationships at the physical layer, as shown in Figure 21. Note that in Figure 21, only equipment (e) and equipment (o) are shown in detail, illustrating their internal equipment and connection relationships. In the example in Figure 21, the connection relationship derivation unit 32 identifies the internal equipment and its connection relationships, shown as white rectangles, for equipment (e) and equipment (o) among the connection relationships at the physical layer as internal equipment and connection relationships that are unnecessary in the logical layer.
[0108] In step ST22, the connection relationship derivation unit 32 excludes the identified internal equipment and connection relationships based on rule B-1, as shown in Figure 22. Note that connection relationships that may be excluded in the rules for deriving the assumed connection relationships of the logic layer but cannot be definitively determined may be identified and retained.
[0109] In step ST23, the connection relationship derivation unit 32 adds connection relationships to replace the excluded internal equipment and connection relationships based on rule B-1.
[0110] In step ST23C, the connection relationship derivation unit 32 adds the necessary connection relationships to the logical layer for internal equipment whose connection destination can be switched, based on rule B-2, as shown by the dashed line in Figure 23. Specifically, for example, if one of the equipment has internal equipment that can be logically connected to two or more equipment, the connection relationship derivation unit 32 derives third connection information to represent the connection relationships that connect each of the two or more objects that are the connection destinations of that object to each other.
[0111] In step ST24, the connection relationship derivation unit 32, based on rule B-3, replaces the connection relationships of the internal equipment with the connection relationships of equipment that have identification information, as shown in Figure 24, and omits the internal equipment.
[0112] In step ST25, the connection relationship derivation unit 32, based on rule B-3, excludes any equipment that has been excluded from all internal equipment.
[0113] In step ST26, the connection relationship derivation unit 32 assumes the processing result as a connection relationship at the logical layer, as shown in the upper part of Figure 25. Thus, the connection relationship derivation unit 32 derives the assumed connection relationship at the logical layer. Step ST20 ends upon completion of step ST26.
[0114] After step ST20, in step ST30, the data acquisition unit 31 acquires the second connection information in the acquired data storage unit 22, thereby obtaining the actual connection relationships at the logical layer, as shown in the lower part of Figure 25.
[0115] After step ST30, step ST40, which includes steps ST41 to ST43, is executed.
[0116] In step ST41, the mapping processing unit 33 calculates the similarity between the derived assumed connection relationships and the actual connection relationships of the logical layer. The similarity of connection relationships is calculated by assigning a weight to each element so that if an element in the set of adjacent equipment is an element resulting from an additional connection relationship under Rule B-2, it will have a different weight than other elements. For example, the set of adjacent equipment to equipment (ka) is {equipment (a) [0.2], equipment (i) [0.2], equipment (e) [1], equipment (ki) [0.2], equipment (ke) [1]} (the values in brackets represent the weight of the element). For example, elements of connection relationships added under Rule B-2 are assigned a lower weight compared to other elements.
[0117] In step ST42, the correspondence processing unit 33 calculates the similarity of the names of the equipment between the first data source DS1 and the second data source DS2.
[0118] In step ST43, the correspondence processing unit 33 associates the names of the equipment between the first data source DS1 and the second data source DS2 based on the similarity scores calculated in steps ST41 and ST42.
[0119] As described above, according to the second embodiment, if an object among the objects can logically connect to two or more objects, the connection relationship derivation unit 32 derives third connection information that includes connection relationships connecting each of the two or more objects to which the object is connected. This reduces errors in associating objects with each other between data sources, even when an object can switch its connection destination, by deriving the expected logical connection relationships.
[0120] <First Modification of the Second Embodiment> The first modification of the second embodiment is the same as the first modification of the first embodiment. However, the first cost g 0For each edit that makes up (x), the following costs are applied: - The cost of deleting a vertex from the assumed graph Ge is either a positive value or not allowed (cost is ∞). - The cost of adding a vertex to the assumed graph Ge is either a positive value or not allowed (cost is ∞). - The cost of deleting an edge from the assumed graph Ge is either a positive value or not allowed (cost is ∞). - The cost of adding an edge to the assumed graph Ge is zero if the edge was added according to rule B-2, and a positive value or not allowed (cost is ∞) for other edges.
[0121] Furthermore, in the rules for deriving the assumed connection relationships of the logical layer, equipment and connection relationships that may be excluded but cannot be definitively determined are identified and retained.
[0122] Accordingly, the cost of deleting potentially excluded vertices and edges is set to zero, and the first cost related to editing is calculated.
[0123] Even with the first modified example described above, the same effects as those of the second embodiment can be obtained.
[0124] <Third Embodiment> The third embodiment is a modification of the second embodiment and relates to a process for notifying missing connection relationships by comparing multiple connection relationships.
[0125] Accordingly, as shown in Figure 26, the processing unit 30 includes a missing data notification unit 35 in addition to the parts described above.
[0126] The omission notification unit 35 notifies the user of the portion of the connection relationship represented by the third connection information and the connection relationship represented by the second connection information that is present in one connection relationship but not in the other, as a omission. For example, the omission notification unit 35 notifies the user of the portion of the connection relationship represented by the third connection information that is present in the connection relationship represented by the third connection information but not in the connection relationship represented by the second connection information, as a omission.
[0127] The other configurations are the same as in the second embodiment.
[0128] With the above configuration, as described above, the connection relationship derivation unit 32 derives third connection information representing the expected connection relationships at the logical layer based on first connection information representing the actual connection relationships at the physical layer, as shown in Figure 27.
[0129] Furthermore, the correspondence processing unit 33, based on the third connection information and the second connection information, associates the identification information of each piece of equipment having an assumed connection relationship at the logical layer with the identification information of each piece of equipment having an actual connection relationship at the logical layer, as shown in Figure 28.
[0130] Subsequently, the missing information notification unit 35 notifies the user of the missing location, which is the part of the connection relationship represented by the third connection information that is present in one of the connection relationships represented by the second connection information but not in the other. For example, the missing information notification unit 35 notifies the user of the missing location, which is the part of the connection relationship represented by the second connection information that is present in the connection relationship represented by the second connection information but not in the connection relationship represented by the third connection information. In this case, as shown in the upper part of Figure 29, the missing information notification unit 35 notifies the user of the missing location by displaying a graph structure on the display device 4 that includes the missing location of the actual connection relationship at the physical layer.
[0131] For example, the missing information notification unit 35 notifies the user of a missing location if it is in a connection relationship represented by the third connection information but not in a connection relationship represented by the second connection information. In this case, as shown in the lower part of Figure 29, the missing information notification unit 35 notifies the user of the missing location by displaying a graph structure on the display device 4 that includes the missing location of the actual connection relationship at the logical layer.
[0132] As described above, according to the third embodiment, the missing data notification unit 35 notifies the user of the portion of the connection relationship represented by the third connection information and the connection relationship represented by the second connection information that is present in one connection relationship but not in the other, as a missing portion. This makes it possible to notify the user of missing portions for the connection relationships represented by the connection information in the data source DS, in addition to the effects of the second embodiment.
[0133] <Fourth Embodiment> The fourth embodiment is a second modification of the first embodiment, and is an example of calculating an evaluation value regarding the co-occurrence of identification information by the cost calculation unit 34. In the fourth embodiment, for ease of understanding, the identification information of each object is expressed simply as shown below.
[0134] In other words, the identification information for each object in the first data source DS1 is replaced with equipment (a), equipment (b), equipment (c), equipment (d), equipment (e), ... a 1 a 2 a 3 a 4 a 5 ...and so on.
[0135] Similarly, the identification information for each object in the second data source DS2 is replaced with equipment (a), equipment (b), equipment (c), equipment (d), equipment (e), ... by b. 1 , b 2 , b 3 , b 4 , b 5 ...and so on.
[0136] The identification information for each object in the third data source DS3 is replaced with c instead of Equipment (A), Equipment (B), Equipment (C), Equipment (D), Equipment (E), ... 1 , c 2 , c 3 , c 4 , c 5 ...and so on.
[0137] Furthermore, as shown in Figure 30, the identification information sets #1, #2, #3, ... between each data source DS1 to DS3 are assumed to be pre-generated and stored in the temporary storage unit 23. In Figure 30, identification information set #1 is {a 1 a 2 a 3 , b 2 , b 3 , b 4 , c 1 , c 4 , c 5} contains the identification information indicated by "1". Similarly, sets #2 and #3 also contain the identification information indicated by "1".
[0138] Here, the cost calculation unit 34 first reads multiple sets of identification information from the temporary storage unit 23. Then, the mapping processing unit 33 determines which identification information is included in each of the read sets. Based on the determination result, it generates a distribution list information in which, for example, "1" is set if the identification information is included and "0" is set otherwise, and stores the generated distribution list information in the temporary storage unit 23.
[0139] Figure 30 shows an example of distribution list information. In this example, the presence or absence of identification information a1-a5, b1-b5, and c1-c5 is determined for each of the data sources DS1, DS2, and DS3, and the results of that determination are displayed in a list with a distribution of "1" and "0".
[0140] Several methods can be considered for determining co-occurrence relationships. In this embodiment, two determination methods will be described as Example 1 and Example 2, respectively.
[0141] (Example 1) The cost calculation unit 34 reads distribution list information from the temporary storage unit 23. Based on the read distribution list information, it identifies combinations of identification information that satisfy the judgment conditions as data in a co-occurrence relationship. The judgment conditions are stored in advance in the rule storage unit 21.
[0142] First, the conditions that a set of identification information must satisfy are as follows: (1) The set of identification information includes the identification information of all data source DSs. (2) The identification information of the same data source DS for the same object is always the same (it is possible to determine that it is the same). (3) For any set of identification information, the identification information is recorded simultaneously only for some of the multiple objects being targeted. (4) The objects for which the identification information is recorded simultaneously differ depending on the set of identification information.
[0143] Of these conditions, condition (2) will be used depending on how the set of identification information was prepared, and it is difficult to determine whether or not it is met from the actual set of identification information itself, so the correspondence processing unit 33 does not make a determination regarding (2) above.
[0144] The cost calculation unit 34 checks for the presence or absence of each identification information in each data source DS1 to DS3, identifies identification information from different data sources DS that match as data satisfying a co-occurrence relationship, and sets the cost evaluation value to zero. In other words, if the distribution patterns of "1" and "0" match across sets of identification information, the evaluation value is zero.
[0145] For example, in the example shown in Figure 30, identification information a is obtained across all data sources DS1, DS2, and DS3. 1 ~a 5 , b 1 ~b 5 , c 1 ~c 5 This shows the distribution patterns of "1" and "0". Here, for example, we compare the distribution patterns of "1" and "0" in any two pieces of identification information between data sources DS1 and DS2. Then, as a result of this comparison, the evaluation value is set to zero for correspondences of identification information where the distribution patterns of "1" and "0" are the same. For example, identification information a 2 , b 2 For the matching of "1" and "0", the evaluation value is set to zero because the distribution patterns of "1" and "0" match. Furthermore, for identification information where even one distribution pattern of "1" and "0" does not match, the evaluation value is set to -∞. For example, identification information a 1 , b 1 Regarding the correspondence, since the distribution patterns of "1" and "0" do not match, the evaluation value is set to -∞.
[0146] In this way, an evaluation value (or cost) can be obtained for any pair of objects. In the process of searching for the pair x that minimizes the overall cost f(x), a value g that monotonically decreases from the evaluation value based on co-occurrence is obtained. 3 When calculating (x), a monotonically decreasing value is added to or multiplied by the evaluation value for the correspondence between the two objects.
[0147] The cost calculation unit 34 calculates the value g as described above. 3 (x) and the value g 3 The combination of identification information corresponding to the evaluation value used in the calculation of (x) is associated with and stored in the temporary storage unit 23.
[0148] (Example 2) In Example 2, conditions (1) and (2) of the set of identification information are relaxed so that co-occurrence relationships can be determined even if there are omissions or inconsistencies due to noise in some of the identification information.
[0149] First, the conditions that the set of identification information must satisfy in Example 2 are as follows: (1)' The set of identification information includes identification information from at least two or more data source DSs. (2)' The identification information from the same data source DS for the same object generally matches, but there are discrepancies due to omissions or noise in some of the identification information. (3) In any set of identification information, the identification information is recorded simultaneously only for some of the multiple objects being targeted. (4) The objects on which the identification information is recorded simultaneously differ depending on the set of identification information.
[0150] Furthermore, regarding the identification information shown in (2)', it is difficult to determine from the actual set of identification information itself whether there are any omissions or noises; therefore, the correspondence processing unit 33 does not make a determination regarding (2)'.
[0151] The correspondence processing unit 33 first reads the determination conditions to be used in Example 2 from the rule storage unit 21.
[0152] The cost calculation unit 34 then reads the distribution list information from the temporary storage unit 23. The matching processing unit 33 then determines whether the identification information in the distribution list information satisfies the judgment conditions, and if it is determined that it does, it calculates an evaluation value based on co-occurrence based on the distribution list information.
[0153] Figures 31 and 32 are schematic diagrams illustrating an example of operation in the cost calculation unit.
[0154] For example, in the example shown in Figure 31, identification information b in identification information set #4 3 Identification information c in identification information set #12 1 , and identification information a in set of identification information #21 2shows a case where identification information that should originally exist is regarded as "none (0)" due to, for example, the influence of deletion or noise.
[0155] FIG. 32 illustrates the identification information a of each data source DS1, DS2, DS3 1 to a 5 , b 1 to b 5 , c 1 to c 5 the relationship between each item, where the denominator represents the number of times the identification information of both data sources DS are included in the set of identification information, and the numerator represents the number of times the presence or absence of the identification information of both data sources DS matches. Specifically, for example, the identification information a 1 to a 5 , b 1 to b 5 the relationship between each item is indicated by an evaluation value based on the distribution pattern of "1" and "0" in sets #1 to #8 shown in FIG. 31. As this evaluation value, among the fractions shown in FIG. 32, only the numerator indicating the number of matches in the presence or absence of identification information may be used, or numerator / denominator indicating the proportion of matches in the presence or absence of identification information may be used. In the following example, numerator / denominator indicating the proportion of matches in the presence or absence of identification information is used as the evaluation value.
[0156] For example, regarding the association between identification information a 1 and b 1 , since the proportion of matching distribution patterns of "1" and "0" in sets #1 to #8 is 3 / 8, the evaluation value is set to 3 / 8. Further, for example, regarding the association between identification information a 2 and b 2 , all distribution patterns of "1" and "0" match, so the evaluation value is set to 8 / 8. Furthermore, identification information a 1 and b 3 has matching distribution patterns of "1" and "0" for 7 times out of the 8 sets #1 to #8, and the evaluation value is 7 / 8. The reason why the evaluation value of identification information a 1 and b 3 is not 8 / 8 is that, as described above, in identification information set #4, the identification information b that should originally exist 3 is regarded as "none (0)" due to the influence of noise or the like.
[0157] In this way, an evaluation value (or cost) can be obtained for any pair of objects. In the process of searching for the pair x that minimizes the overall cost f(x), when calculating a value g3(x) that is monotonically increasing based on the co-occurrence evaluation value, a value that is monotonically decreasing is added to or multiplied by the evaluation value for the pair of objects.
[0158] The cost calculation unit 34 stores the value g3(x) calculated as described above and the combination of identification information corresponding to the evaluation value used in calculating the value g3(x) in the temporary storage unit 23, associating them with each other.
[0159] <Fifth Embodiment> The fifth embodiment is a second modification of the first embodiment, and is an example of calculating evaluation values such as the similarity and degree of agreement of additional information (attribute information) by the cost calculation unit 34. In the fifth embodiment, from the viewpoint of facilitating understanding, as shown below, the additional information between each data source DS is a combination that is related to each other and whose relationship is modeled by equations, sequential formulas, logical formulas, etc. As a model, for example, formulation by equation as shown in the following equation can be used. However, it is not limited to this, and the relationship between dependent and explanatory variables may be expressed and formulated using relational expressions such as sequential formulas and logical formulas.
[0160]
[0161] However, Y: Additional information from one of the data sources DS, X 1 , X 2 : Additional information from the other data source DS (or a value calculated from that additional information), ε: Error.
[0162] The following are examples 1 and 2 of the relational expressions used in the formulation.
[0163] (Example 1) The first data source DS1 relates to the physical equipment of the communication facility, and the object identification information is the "building name," and the additional information is the "building's accommodation equipment" as well as the "physical distance" between adjacent buildings.
[0164] The second data source, DS2, pertains to fixed assets of communication equipment, with the object identification information being "building name," and the additional information being "building address" as well as "building latitude and longitude."
[0165] In this case, the additional information "physical distance" from the first data source DS1 can be modeled based on the additional information "latitude and longitude of the building" from the second data source DS2, as shown in the following relationship.
[0166]
[0167] However, Y: physical distance between buildings, X 1 : Latitude difference between the two buildings, X 2 : Difference in longitude between two buildings, ε ~ N(10,0): Error, a 1 = 40409785.305, a 2 = 27076647.025.
[0168] (Example 2) The first data source DS1 relates to the physical equipment of the communication facility, and the object identification information is "building name," with additional information including "building accommodation equipment" and "importance of accommodation equipment."
[0169] The second data source, DS2, relates to communication services for communication equipment, and the object identification information is "building name," with additional information being "building capacity."
[0170] In this case, the additional information "building capacity" from the second data source DS2 can be modeled based on the additional information "importance of the accommodation equipment" from the first data source DS1 and the "number of adjacent buildings" obtained from the connection relationships, as shown in the following relational expression.
[0171]
[0172] However, Y: the building's capacity, X 1 : Number of adjacent buildings, X 2 : Importance of containment device, ε ~ N(150,0): Error, a 1 = 90, a 2 = 100.
[0173] The above explains the relational expressions used in the formulation. These relational expressions are stored, for example, in the rule storage unit 21.
[0174] Accordingly, the cost calculation unit 34 calculates how well the relationships between the values of the additional information, with respect to combinations of additional information from different data sources whose relationships have been modeled, match the predictions made by the model.
[0175] Specifically, the cost calculation unit 34 uses the relational expression in the rule storage unit 21 to calculate additional information (X) from one of the data sources DS. 1 , X 2 The cost calculation unit 34 calculates a predicted value of additional information (Y) from the other data source DS. The cost calculation unit 34 also calculates evaluation values such as similarity and degree of agreement based on the difference (= actual value - predicted value) between the actual value of the additional information (Y) from the other data source DS and the calculated predicted value. For example, the evaluation value can be obtained by a calculation such as |actual value - difference| / actual value.
[0176] The cost calculation unit 34 stores the obtained evaluation value and the combination of identification information corresponding to the combination of additional information used to calculate the evaluation value in the temporary storage unit 23, associating them. This makes it possible to use the evaluation value according to the combination of identification information when using the evaluation value of the additional information when calculating the overall cost.
[0177] <Other Embodiments> The functional configuration of the information processing device 1 and its peripheral devices, their processing procedures and contents, the types and uses of data and information, etc., can be modified in various ways without departing from the spirit of this invention.
[0178] For example, in the first embodiment, the system is provided with a data acquisition unit 31 and a processing unit 30 capable of acquiring identification information, connection information, and additional information from each data source DS, but it is not limited to this. For example, instead of the data acquisition unit 31, the processing unit 30 may include an identification information acquisition unit that acquires identification information from each data source DS, a connection information acquisition unit that acquires connection information from each data source DS, and an additional information acquisition unit that acquires additional information from each data source DS.
[0179] For example, in the first embodiment, logical connection relationships were derived from physical connection relationships based on rules in the rule storage unit 21, but the embodiment is not limited to this. For example, connection relationships can also be derived by executing a program that describes a procedure similar to that of the rules.
[0180] Furthermore, in some of the modifications of the first embodiment, a cost calculation unit 34 is provided separately from the correspondence processing unit 33, but the system is not limited to this, and the correspondence processing unit 33 may also include the cost calculation unit 34.
[0181] Furthermore, for example, in the fifth embodiment, a pre-modeled relational expression was used, but the system is not limited to this. For example, the system may include a modeling unit that generates relational expressions between additional information by extracting dependent and explanatory additional information from the additional information that represents numerical values using regression analysis with the least squares method. In the case of Example 1 in the fifth embodiment, by using regression analysis with the least squares method on "physical distance," the difference in "latitude and longitude of the building" can be extracted as an explanatory element. Also, in the case of Example 2 in the fifth embodiment, by using regression analysis with the least squares method on "building capacity," "importance of the accommodation equipment" and "number of adjacent buildings" can be extracted as explanatory elements. Thus, the system may be configured so that the modeling unit generates relational expressions between the additional information, and the cost calculation unit 34 uses the generated relational expressions. It should be noted that the modeling unit may not be limited to generating relational expressions; it may also assist in the creation of relational expressions. For example, the modeling unit may assist in the creation of relational expressions by extracting explanatory elements.
[0182] Although embodiments of this invention have been described in detail above, the above description is merely illustrative in all respects. It goes without saying that various improvements and modifications can be made without departing from the scope of this invention. In other words, when implementing this invention, specific configurations may be adopted as appropriate depending on the embodiment.
[0183] In short, this invention is not limited to the embodiments described above, and in the implementation stage, the components can be modified and materialized without departing from the gist of the invention. Furthermore, various inventions can be formed by appropriately combining the multiple components disclosed in the embodiments. For example, some components may be deleted from all the components shown in the embodiments. Moreover, components from different embodiments may be appropriately combined.
[0184] 1... Information processing device 3... Input device 4... Display device 10... Program storage unit 20... Data storage unit 21... Rule storage unit 22... Acquired data storage unit 23... Temporary storage unit 24... Result storage unit 30... Processing unit 31... Data acquisition unit 32... Connection relationship derivation unit 33... Mapping processing unit 34... Cost calculation unit 35... Missing data notification unit 50... Communication I / F 60... Input / Output I / F DS1... First data source DS2... Second data source DS3... Third data source
Claims
1. An information processing device comprising: a first acquisition unit that acquires first connection information from a first data source, which includes first identification information indicating each object within the target and represents a first type of connection relationship for each of the objects connected to the target; a second acquisition unit that acquires second connection information from a second data source, which includes second identification information indicating each object within the target and represents a second type of connection relationship for each of the objects; a derivation unit that derives third connection information from the first connection information, which includes the first identification information and represents a second type of connection relationship assumed from the first type of connection relationship; an association unit that associates the first identification information with the second identification information based on the third connection information and the second connection information; and an output unit that outputs the result of the association.
2. The information processing apparatus according to claim 1, wherein the first type of connection relationship is a physical connection relationship, the second type of connection relationship is a logical connection relationship, and the derivation unit derives the third connection information which represents a logical connection relationship assumed from the physical connection relationship represented by the first connection information.
3. The information processing apparatus according to claim 1, comprising: a cost calculation unit that, for each combination of the first identification information and the second identification information, edits the assumed second type of connection relationship to the second type of connection relationship represented by the second connection information, calculates a cost corresponding to the editing, and identifies the combination that minimizes the cost; and the matching unit associates the first identification information and the second identification information with the identified combination.
4. A method executed by an information processing device, comprising: obtaining first connection information from a first data source, which includes first identification information indicating each object within a target to which a plurality of objects are connected, and which represents a first type of connection relationship between each object; obtaining second connection information from a second data source, which includes second identification information indicating each object within the target, and which represents a second type of connection relationship between each object; deriving third connection information from the first connection information, which includes the first identification information and represents a second type of connection relationship assumed from the first type of connection relationship; associating the first identification information with the second identification information based on the third connection information and the second connection information; and outputting the result of the association.