Data fusion method and electronic equipment

By dividing and combining data from multiple departments into nodes and performing data fusion based on similarity and correlation information, the problem of cross-departmental data silos is solved and the efficiency and accuracy of data fusion are improved.

CN120744822APending Publication Date: 2025-10-03LCFC HEFEI ELECTRONICS TECH
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510855698.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-24
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

Since each department is responsible for different tasks during the joint collaboration process, data becomes isolated, which hinders the sharing of cross-departmental data information and reduces the efficiency of inter-departmental collaboration.

Method used

By obtaining multiple data to be fused, dividing them into node blocks based on the attribute information of the nodes, and combining the node blocks according to the similarity and association information, data fusion is performed to ensure that the nodes in the combination have a certain similarity, and only the nodes with a certain similarity are fused.

Benefits of technology

It improves the efficiency and accuracy of data fusion, reduces mis-fusion caused by differences in data descriptions across departments, ensures that each node in the fused data uniquely corresponds to an entity, and eliminates redundancy and ambiguity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120744822A_ABST
    Figure CN120744822A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of data processing, and provides a data fusion method and electronic device.The data fusion method comprises the steps that multiple pieces of to-be-fused data are obtained, the to-be-fused data comprise at least two nodes, first to-be-fused data are divided into at least one node block on the basis of attribute information of all the nodes in the first to-be-fused data, and the at least one node block is used for fusing the first to-be-fused data; the first to-be-fused data is any to-be-fused data in the multiple to-be-fused data; determining N first combinations from M node blocks corresponding to the multiple pieces of to-be-fused data; wherein the first combination comprises at least two node blocks, the similarity between the node blocks in the first combination meets a first preset condition, and M and N are positive integers; and based on the attribute information of the nodes in the first combinations, performing data fusion on the nodes in the corresponding first combinations to obtain target fusion data. According to the data fusion method, cross-department data information sharing can be realized, and the cooperation efficiency among departments is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing technology, and in particular to a data fusion method and electronic equipment. Background Art

[0002] Work demands necessitate collaboration between multiple departments. However, because each department has different responsibilities and focuses during this collaborative process, and often uses different standards for describing the same thing, data becomes siloed across departments, hindering cross-departmental data and information sharing and reducing interdepartmental collaboration efficiency. Summary of the Invention

[0003] The present application provides a data fusion method and electronic device to at least solve the above technical problems existing in the related art.

[0004] In a first aspect of the present application, a data fusion method is provided, comprising:

[0005] Acquire a plurality of data to be fused, wherein the data to be fused includes at least two nodes; the nodes are used to indicate entities, and the nodes include attribute information of the entities;

[0006] Based on the attribute information of each node in the first data to be fused, the first data to be fused is divided into at least one node block; the first data to be fused is any data to be fused among the multiple data to be fused;

[0007] Determining N first combinations from the M node blocks corresponding to the plurality of data to be fused; wherein the first combination includes at least two node blocks, the similarity between the node blocks in the first combination satisfies a first preset condition, and M and N are positive integers;

[0008] Based on the attribute information of the nodes in each first combination, data of each node in the corresponding first combination is fused to obtain target fused data; the entities indicated by each node in the target fused data are different.

[0009] In one possible implementation manner, the attribute information of the entity includes functional attributes; and dividing the first data to be fused into at least one node block includes:

[0010] Determining similarities between the nodes in the first data to be fused based on functional attributes of the nodes in the first data to be fused;

[0011] Nodes whose similarities satisfy a second preset condition are divided into the same node block to obtain at least one node block.

[0012] In one possible implementation, the node further includes association information between the entity and other entities;

[0013] The determining N first combinations from the M node blocks corresponding to the plurality of data to be fused includes:

[0014] Determining first description information describing the semantics of the node block based on the attribute information and association information included in each node in each node block;

[0015] Determining the similarity between the corresponding node block and other node blocks according to the first description information of each node block;

[0016] The first node block and the second node block are combined to obtain the first combination, the similarity between the second node block and the first node block satisfies the first preset condition, the first node block is any node block among the M node blocks, and the second node block is a node block among the M node blocks except the first node block.

[0017] In one possible implementation, the step of fusing data of the nodes in the corresponding first groups based on the attribute information of the nodes in the first groups includes:

[0018] For any first combination, determining second description information describing the semantics of the corresponding node based on the association information and attribute information included in each node in the first combination;

[0019] Determining the similarity between the corresponding node and other nodes based on the second description information of each node;

[0020] Data of the first node and the second node are fused; the similarity between the first node and the second node satisfies a third preset condition, the first node is any node in the first combination, and the second node is a node in the first combination other than the first node.

[0021] In one possible implementation manner, the attribute information includes an attribute name and an attribute value; and the step of fusing data of the first node and the second node includes:

[0022] Determining, based on a first attribute name in first attribute information and a second attribute name in second attribute information, whether an attribute indicated by the first attribute information is the same as an attribute indicated by the second attribute information, wherein the first attribute information is any attribute information in the first node and the second attribute information is any attribute information in the second node;

[0023] In response to the attribute indicated by the first attribute information being the same as the attribute indicated by the second attribute information, determining a third attribute name corresponding to a target attribute based on the first attribute name and the second attribute name, and determining a third attribute value corresponding to the target attribute based on the first attribute value of the first attribute information and the second attribute value of the second attribute information, the target attribute being the attribute indicated by the first attribute information and the second attribute information;

[0024] The first attribute information and the second attribute information are merged into target attribute information, where the target attribute information includes the third attribute name and the third attribute value.

[0025] In one possible implementation manner, determining a third attribute value corresponding to the target attribute based on the first attribute value of the first attribute information and the second attribute value of the second attribute information includes:

[0026] determining the first attribute value as the third attribute value;

[0027] Alternatively, the second attribute value is determined as the third attribute value;

[0028] Alternatively, information generated based on the first attribute value and the second attribute value is determined as the third attribute value.

[0029] In one possible implementation, each node block includes level information;

[0030] The determining N first combinations from the M node blocks corresponding to the plurality of data to be fused includes:

[0031] From the M node blocks, P node blocks indicated by the hierarchy information as the first hierarchy are determined, and the N first combinations are determined from the P node blocks, where P is a positive integer.

[0032] In one possible implementation, the step of fusing data of the nodes in the corresponding first groups based on the attribute information of the nodes in the first groups includes:

[0033] For any first combination, determining an upper-level node block of the corresponding node block based on the level information of each node block in the first combination, wherein the corresponding node block belongs to the upper-level node block;

[0034] Data fusion is performed on the nodes in each upper-level node block corresponding to the first combination.

[0035] In one possible implementation manner, the data to be fused is a knowledge graph.

[0036] A second aspect of the present application provides a data fusion device, comprising:

[0037] A data acquisition module is used to acquire a plurality of data to be fused, wherein the data to be fused includes at least two nodes; a node is used to indicate an entity, and a node includes attribute information of the entity;

[0038] A node division module, configured to divide the first data to be fused into at least one node block based on attribute information of each node in the first data to be fused; the first data to be fused is any data to be fused among the plurality of data to be fused;

[0039] A first combination module is configured to determine N first combinations from the M node blocks corresponding to the plurality of data to be fused; wherein the first combination includes at least two node blocks, the similarity between the node blocks in the first combination satisfies a first preset condition, and M and N are positive integers;

[0040] The data fusion module is used to fuse the data of each node in the corresponding first combination based on the attribute information of the nodes in each first combination to obtain target fused data; the entities indicated by each node in the target fused data are different.

[0041] According to a third aspect of the present application, an electronic device is provided, including:

[0042] at least one processor; and

[0043] a memory communicatively connected to at least one processor; wherein,

[0044] The memory stores instructions that can be executed by at least one processor. The instructions are executed by the at least one processor so that the at least one processor can execute the data fusion method of the present application.

[0045] In a fourth aspect of the present application, a non-transitory computer-readable storage medium storing computer instructions is provided, where the computer instructions are used to enable a computer to execute the data fusion method of the present application.

[0046] The data fusion method and electronic device provided by the present application, the data fusion method adopted by the present application first divides the nodes into node blocks. Then, the node blocks whose similarity meets the first preset condition are combined, and then the nodes in the combination are fused. Because there is a certain similarity between the node blocks in the combination, it means that there is a certain similarity between the nodes in the node blocks in the combination. Therefore, the data fusion method provided by the present application does not need to fuse all nodes, but only needs to perform a fusion operation on nodes with a certain similarity, thereby improving the efficiency of data fusion. In addition, because the nodes with a certain similarity are fused, the accuracy of data fusion can also be improved.

[0047] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present application, nor is it intended to limit the scope of the present application. Other features of the present application will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] The above and other objects, features and advantages of the exemplary embodiments of the present application will become readily understood by reading the detailed description below with reference to the accompanying drawings. In the accompanying drawings, several embodiments of the present application are shown in an illustrative and non-limiting manner, in which:

[0049] In the drawings, the same or corresponding reference numerals denote the same or corresponding parts.

[0050] Figure 1 A schematic diagram of the steps of the data fusion method according to an embodiment of the present application is shown;

[0051] Figure 2 A schematic diagram of node division of the data fusion method according to an embodiment of the present application is shown;

[0052] Figure 3 A schematic diagram of a process flow of the data fusion method according to an embodiment of the present application is shown;

[0053] Figure 4 A schematic diagram showing a flow chart of node block combination of a data fusion method according to an embodiment of the present application is shown;

[0054] Figure 5 Another node division diagram of the community discovery algorithm according to an embodiment of the present application is shown;

[0055] Figure 6 Another flow chart of the data fusion method according to an embodiment of the present application is shown;

[0056] Figure 7 Another flow chart of the data fusion method according to the embodiment of the present application is shown;

[0057] Figure 8 Another flowchart of the node block combination of the data fusion method according to an embodiment of the present application is shown;

[0058] Figure 9 Another flow chart of the data fusion method according to the embodiment of the present application is shown;

[0059] Figure 10 A schematic structural diagram of a data fusion device according to an embodiment of the present application is shown;

[0060] Figure 11 A schematic diagram of the structure of an electronic device according to an embodiment of the present application is shown. DETAILED DESCRIPTION

[0061] In order to make the purpose, features, and advantages of this application more obvious and easy to understand, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without making creative efforts shall fall within the scope of protection of this application.

[0062] A data fusion method and electronic device provided by an embodiment of the present application are described below with reference to the accompanying drawings.

[0063] The data fusion method provided in the embodiment of the present application can be implemented by an electronic device, that is, the execution subject of the data fusion method can be an electronic device. The electronic device can be a smart terminal or a server; wherein the smart terminal can be a mobile phone, a personal digital assistant (PAD), a tablet computer, a laptop computer, a desktop computer and other devices. The server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery network (CDN) services, as well as basic cloud computing services such as big data and artificial intelligence platforms.

[0064] like Figure 1 As shown, the embodiment of the present application provides a data fusion method, including:

[0065] S101, the electronic device obtains a plurality of data to be fused, where the data to be fused includes at least two nodes; the node is used to indicate an entity, and the node includes attribute information of the entity.

[0066] As an example, the multiple data to be fused may be data corresponding to various departments of a company, with each department having corresponding data to be fused. For a further example, one data to be fused may be data from the company's software department, and another data to be fused may be data from the company's hardware department.

[0067] For example, taking the hardware department as an example, the data to be merged in the hardware department may include multiple nodes indicating entities such as processors, memory, storage devices, and input devices. For example, the data to be merged in the hardware department may include node A indicating a processor, node B indicating a memory, node C indicating a switch device, and node D indicating a solenoid valve.

[0068] A node may be data indicating an entity, and may include one or more attribute information of the entity.

[0069] In this application, the attribute information of an entity may include a name attribute, a function attribute, a size attribute, a data capacity, a data processing speed attribute, etc. For example, for a node A indicating a processor in the data to be fused, the name attribute of the node A may be the processor, the function attribute may be the function of the processor, such as the image processing function; the size attribute may indicate the size of the processor, and the data processing speed attribute may indicate the speed at which the processor processes signals or images.

[0070] The electronic device obtains multiple data to be fused, which may be that the electronic device extracts all data from the reports of various departments, thereby obtaining multiple data to be fused. Alternatively, the electronic device obtains the data to be fused based on the data in the databases of various departments according to the received data fusion instruction. The data fusion instruction may be an instruction generated based on a user input requirement. The user input requirement may be a requirement to improve the computer's operating performance. The electronic device can respond to the data fusion instruction and read data from various databases that can improve the computer's operating performance. For example, data that can change the computer's operating performance, such as node A indicating the processor and node B indicating the memory, is obtained, thereby forming data to be fused consisting of at least two nodes.

[0071] Exemplarily, the data to be fused may be a knowledge graph. A knowledge graph can display the attribute information of entities in a visual graphic, and a knowledge graph is a semantic network between entities. Each department corresponds to a piece of data to be fused, and when the data to be fused is a knowledge graph, the electronic device can construct a knowledge graph for each department, combining the data of the department (for example, task execution-related data in a joint collaboration process), to obtain the data to be fused corresponding to the department.

[0072] S102, the electronic device divides the first data to be fused into at least one node block based on attribute information of each node in the first data to be fused; the first data to be fused is any data to be fused among the multiple data to be fused.

[0073] In the present application, the electronic device can divide multiple nodes into a node block based on the attribute information of each node in the first data to be fused, thereby realizing the division of the first data to be fused into at least one node block. As an example, multiple nodes with similar attribute information can be divided into a node block. Among them, multiple nodes have similar attribute information, which means that the similarity of the attribute information of the multiple nodes is greater than or equal to a preset threshold. The preset threshold is usually a value greater than 0 and less than 1. As an example, the preset threshold can be 0.8. It should be pointed out that node block division based on the attribute information of each node can realize the division of multiple nodes with similar attributes into a node block.

[0074] For example, Figure 2 As shown, the nodes can be divided according to the name attribute in the attribute information of the nodes. For example, because the name attribute of the E node whose name attribute is the central processing unit and the name attribute of the A node whose name attribute is the processor have a similarity of 0.9, which is greater than the preset threshold value of 0.8. Therefore, the E node and the A node can be divided into the same node block. Among them, the preset threshold value can be set as needed. In this application, the first data to be fused can also be divided based on other attribute information of each node, which is not limited in this application. For example, the C node indicates that the functional attribute of the switch device is a device for controlling the switch, and the D node indicates that the functional attribute of the solenoid valve is a device for controlling the on and off of the fluid. The similarity of the functional attributes of the C node and the D node is 0.8. That is, the similarity of the functional attributes of the C node and the D node is equal to the preset threshold value of 0.8. Therefore, the C node and the D node can be divided into the same node block.

[0075] For example, Figure 3 As shown, the electronic device can divide the nodes in the data to be fused 1 into node blocks A1, B1, ..., N1. The electronic device can divide the nodes in the data to be fused 2 into node blocks A2, B2, ..., N2.

[0076] S103, the electronic device determines N first combinations from M node blocks corresponding to the multiple data to be fused; wherein the first combination includes at least two node blocks, the similarity between the node blocks in the first combination meets a first preset condition, and M and N are positive integers.

[0077] In the present application, after node division of multiple data to be fused, M node blocks can be obtained. Similarity calculation is performed between each node block in the M node blocks. Multiple node blocks whose similarity between each node block meets the first preset condition are combined to obtain N first combinations. The first combination may include at least two node blocks. Among them, the first preset condition is the judgment criterion based on which N first combinations are determined from the M node blocks corresponding to the multiple data to be fused. The role of the first preset condition is to measure whether the similarity between the node blocks reaches the first threshold for combination, so as to ensure that the combined node blocks have the basis for data fusion. The first preset condition is that the similarity between the node blocks is greater than or equal to the first threshold. The first threshold is a pre-set value (usually 0.8 to 0.95, such as 0.85). For example, the first threshold can be 0.85. If the similarity between node block 1 and node block 2 is greater than or equal to 0.85, node block 1 and node block 2 are combined to obtain the first combination. The first threshold can be set according to actual needs and is not limited here.

[0078] For example, Figure 4As shown, the data to be fused 1 includes a node block A1, and the node block A1 includes an A node indicating a processor and an E node indicating a central processing unit; the data to be fused 2 includes a node block A2, and the node block A2 includes an F node indicating a processor and a G node indicating an image processor. In this application, the similarity between the node block A1 and the node block A2 can be calculated. If the similarity between the node block A1 and the node block A2 meets the first preset condition, the node block A1 and the node block A2 are combined into a first combination. It can be understood that if there is a node block A3 in the data to be fused 3, and the similarity between A3 and the node block A1 and the node block A2 meets the first preset condition, the node block A1, the node block A2 and the node block A3 are combined into a first combination (considering the limitation of space, the data to be fused 3 and the node block A3 are in Figure 4 does not appear in ).

[0079] Among them, in the scheme of calculating the similarity between each node block, the description information of each node block can be obtained first. The description information is information that can describe the node block. For example, the description information can be a natural language text that describes the category characteristics of the node block. The description information can be generated by extracting information and semantic understanding based on the attribute information of the entities included in all nodes in the node block. The electronic device can determine the similarity between the two node blocks by calculating the similarity between the description information of each two node blocks. For example, if the similarity between the description information of node block A1 and the description information of node block A2 is greater than or equal to 0.85, it is considered that the similarity between node block A1 and node block A2 is greater than or equal to 0.85, and node block A1 and node block A2 constitute a first combination. The present application can also calculate the similarity between each node block in other ways, for example, by determining the similarity between each node block through the similarity between a certain attribute information of all nodes in the node block, and the present application is not limited here.

[0080] S104, the electronic device performs data fusion on each node in the corresponding first combination based on the attribute information of the nodes in each first combination to obtain target fusion data; the entities indicated by each node in the target fusion data are different.

[0081] After the electronic device determines the first combination, it can calculate the similarity of the attribute information of each two nodes in the first combination, compare the similarity of the attribute information of each two nodes with the second threshold, and perform data fusion on multiple nodes whose attribute information similarity is greater than or equal to the second threshold, thereby obtaining target fusion data. The second threshold is a pre-set value (usually a decimal between 0 and 1, such as 0.9), which is used to measure the similarity of the attribute information (such as name, function, association, etc.) of two nodes. The data fusion operation is triggered only when the similarity of the attribute information between the nodes is greater than or equal to the second threshold. The second threshold can be the same as or different from the first threshold. For example, the second threshold can be set to 0.9. When the similarity between the attribute information is greater than 0.9, data fusion is performed on each node corresponding to the attribute information.

[0082] For example, Figure 4 As shown, the first combination includes node block A1 and node block A2. The A node indicating the processor in node block A1 and the F node indicating the processor in node block A2 have the same name attribute, and it can be considered that the similarity between the name attributes of node A and node F is 1, which is greater than 0.9. Therefore, node A and node F are fused and merged into one node. The nodes that cannot be fused in node block A1 and node block A2 are retained. Therefore, after data fusion of the nodes of each node block in the first combination, the entities indicated by the obtained nodes are different.

[0083] The data fusion method provided in the embodiment of the present application is that after the electronic device obtains multiple data to be fused, the data to be fused is divided based on the attribute information of each node in the data to be fused to obtain at least one node block. From the multiple node blocks obtained from the multiple data to be fused, the node blocks whose similarity meets the first preset condition are determined, and the node blocks that meet the first preset condition constitute a first combination. According to the attribute information of the nodes in the first combination, data fusion is performed on each node in the corresponding first combination to obtain target fused data. In the present application, the nodes with similar attributes are first divided in the way of dividing the nodes into node blocks, and then the nodes in the node blocks that meet the first preset condition are fused to achieve hierarchical filtering of the nodes, reduce the computational complexity of directly processing a large number of nodes, and improve the fusion efficiency. In addition, the present application can avoid the misfusion of cross-departmental data to be fused due to description differences, thereby improving the fusion accuracy. In the target fused data finally obtained, each node uniquely corresponds to an entity, eliminating redundancy and ambiguity.

[0084] In some embodiments, the attribute information of the entity includes functional attributes; dividing the first data to be fused into at least one node block includes:

[0085] Determining similarities between the nodes in the first data to be fused based on functional attributes of the nodes in the first data to be fused;

[0086] Nodes whose similarities satisfy a second preset condition are divided into the same node block to obtain at least one node block.

[0087] In this application, the electronic device can divide the node blocks according to the functional attributes of the entities included in the nodes. Specifically, the similarity between the nodes is determined by the similarity of the functional attributes of each node. If the similarity between the nodes meets the second preset condition, the corresponding nodes are divided into the same node block. The second preset condition is that the similarity between the nodes is greater than or equal to a preset threshold. Figure 4 For example, if the preset threshold is 0.8, the functional attribute of the central processing unit indicated by the E node in the data to be fused 1 is the function of processing signals, the functional attribute of the processor indicated by the A node is the function of controlling and processing signals, and the similarity between the functional attributes of the central processing unit included in the E node and the functional attributes of the processor included in the A node is 0.95, which is greater than the preset threshold of 0.8. Therefore, the E node and the A node can be divided into the same node block A1. The B node in the data to be fused 1 indicates a memory, and its functional attribute is a storage function. The similarity between the functional attributes of the A node and the functional attributes of the B node is 0.75, which is less than the preset threshold of 0.8. Therefore, the A node and the B node will not be divided into the same node block. In this way, each node in the data to be fused is divided multiple times to obtain a node block consisting of the E node and the A node.

[0088] The present application may also adopt a community discovery algorithm to divide the nodes in the fused data. The community discovery algorithm is an algorithm that has been disclosed in the relevant technology.

[0089] As an embodiment, the data to be integrated can be the knowledge graphs of various departments. The electronic device can use a community discovery algorithm to divide the knowledge graph. The community discovery algorithm can identify node clusters (communities) with strong correlations, thereby dividing the knowledge graph into node blocks with clear structure and semantic cohesion. In the community discovery algorithm, modularity is the core indicator for evaluating the quality of network community division. Modularity quantifies the compactness and significance of the community structure by comparing the actual connection density within the community with the expected connection density in the random network, thereby judging the rationality of the node block division. Dividing the knowledge graph through the community discovery algorithm can improve the accuracy and efficiency of data division.

[0090] When an electronic device uses a community discovery algorithm to partition nodes in a knowledge graph, it first treats each node in the knowledge graph as a community, adds a certain community to the neighboring community, and calculates the modularity of all newly formed communities. The difference between the modularity of all newly formed communities and the modularity of the nodes before the partition is calculated, and the communities with positive modularity differences are determined. A positive modularity difference means that the modularity of the community after the partition has a modularity gain compared to the modularity of the community before the partition. The two communities with the largest modularity gain are merged. If the number of communities in the knowledge graph is greater than 1, the iteration continues (i.e., the communities in the knowledge graph are added to the neighboring communities, and the two communities with the largest modularity gain in the newly formed communities are merged). The modularity value of each community obtained after the knowledge graph partition is traversed. If the modularity value is already the maximum value (i.e., the modularity value of the new community formed by further partitioning does not increase), then the current community partition is the optimal community partition of the knowledge graph.

[0091] The following combination Figure 5 Give an example to illustrate the community discovery algorithm. Figure 5 In the example, the knowledge graph includes multiple nodes, and the node a, node b, node c, node d, and node e are taken as examples to illustrate the node block division process based on the community discovery algorithm. The electronic device adds node a to neighboring node c, neighboring node e, and neighboring node b to form multiple initial node blocks, calculates the difference between the modularity of each initial node block and the value of the modularity of node a before division, determines the initial node block in which the modularity difference is a positive value (modularity gain), and then determines the initial node block with the largest modularity gain, and merges the two nodes with the largest modularity gain. For example, if the modularity gain of the initial node block formed by node a and node c is the largest, then node a and node c are divided into one node block, and thus a division is completed.

[0092] Continue iterating and try adding node b to the node block consisting of neighboring nodes a and c to obtain an initial node block. Add node b to node d to obtain an initial node block. Calculate the modularity of these two initial node blocks and the difference between the modularity of these two initial node blocks and the modularity of node b before the division. Determine the initial node block in which the modularity difference is positive (modularity gain), and then determine the initial node block with the largest modularity gain. If the modularity gain of the initial node block formed after node b is added to node d is the largest, then add node b to node d to form a node block.

[0093] This partitioning is then repeated for all nodes in the knowledge graph until the modularity gain no longer increases, i.e., the modularity of the node block in the knowledge graph is maximized. The node partitioning result with the maximum modularity is the optimal node block partitioning.

[0094] In the embodiments of the present application, electronic devices can use a community discovery algorithm to partition nodes in the data to be fused based on their functional attributes, thereby achieving semantic clustering of the data to be fused. This effectively solves the problem of assigning nodes to appropriate blocks during the process of partitioning nodes into blocks, significantly improving the accuracy and efficiency of subsequent data fusion.

[0095] In an embodiment of the present application, the electronic device can first determine whether different nodes in the same data to be fused indicate the same entity, that is, determine whether there are identical nodes indicating the same entity in each obtained node block. If so, the repeated identical nodes representing the same entity are deleted, and only one is retained. In an embodiment of the present application, the electronic device performs a node division process multiple times, and when the modularity can no longer be increased, the community discovery is considered to be completed, and the node blocks obtained by dividing the nodes in each data to be fused are obtained.

[0096] For example, Figure 6 As shown, the electronic device uses a community discovery algorithm to divide the nodes in the data to be fused multiple times, dividing the node E indicating the central processing unit, the node A indicating the processor, and the node B indicating the memory into a node block. Node E and node A have similar name attributes, so they are merged, and either node is deleted, leaving the other node. Assuming that the merge is to node A, the resulting node block includes node A and node B indicating the memory.

[0097] The community discovery algorithm can also further optimize the division by merging, thereby improving the accuracy and rationality of node division.

[0098] In some embodiments, the node also includes association information between the entity and other entities;

[0099] Determining N first combinations from M node blocks corresponding to the plurality of data to be fused includes:

[0100] Determining first description information describing the semantics of the node block based on the attribute information and association information included in each node in each node block;

[0101] Determining the similarity between the corresponding node block and other node blocks according to the first description information of each node block;

[0102] The first node block and the second node block are combined to obtain a first combination, the similarity between the second node block and the first node block satisfies a first preset condition, the first node block is any node block among the M node blocks, and the second node block is a node block among the M node blocks except the first node block.

[0103] In the present application, the node also includes the association information between the entity and other entities. The association information between the entity and other entities can be the connection relationship between the entity and other entities. For example, node A also includes the connection relationship between the processor and other devices, such as the connection relationship between the processor and the memory, the connection relationship between the processor and the solenoid valve, etc. After the nodes in the multiple data to be fused are divided by the community discovery algorithm, the present application obtains M node blocks. According to the attribute information and association information of each node in each node block, the first description information describing the semantics of the node block can be obtained, and the similarity between the corresponding node block and other node blocks can be determined according to the similarity between the first description information of each node block. If the similarity between the first node block and the second node block is greater than or equal to the first threshold, it means that the first node block is similar to the second node block, and the first node block and the second node block are combined to obtain a first combination.

[0104] In an embodiment of the present application, after the electronic device inputs the attribute information and associated information included in all nodes in a node block into the large model, it can obtain the first description information of the node block output by the large model. The large model in the embodiment of the present application is a large language model (LLM). The first description information of the node block is a natural language text describing the semantics of the node block obtained by natural language processing (NLP) of the LLM.

[0105] In an embodiment of the present application, after obtaining the first description information of each node block, the electronic device can use a cosine similarity algorithm to calculate the similarity of the first description information of each node block. The electronic device compares the similarity of the description information of each node block with a first threshold. If the similarity value of the description information of two or more node blocks is greater than or equal to the first threshold, the two or more node blocks are combined to obtain a first combination. Otherwise, it is considered that the node blocks cannot be combined.

[0106] During implementation, the description information of each node block can be represented by a vector. For example, the term frequency-inverse document frequency (TF-IDF) algorithm, the word vector (Word to Vector, Word2Vec) model, etc. are used to represent the description information of the node block as a vector, that is, a description vector. In the field of data fusion, a description vector is a technical means of converting natural language text (such as the semantic description of a node or node block) into a numerical vector that can be processed by a computer. The role of the description vector is to convert semantic information into structured data so that semantic similarity can be calculated by a mathematical algorithm. In this application, the cosine similarity algorithm is used to calculate the similarity between vectors, that is, the description information of the node blocks. The calculation result of the cosine similarity algorithm can effectively reflect the degree of similarity between the description vectors of the node blocks. The node blocks whose calculated similarity is greater than or equal to the first threshold are combined to obtain a first combination. Among them, cosine similarity is an algorithm in the prior art, and this application will not go into details here.

[0107] In an embodiment of the present application, after the electronic device vectorizes the description information of the node block, it uses the cosine similarity algorithm to calculate the similarity between the description information of the node blocks, which can effectively solve the problem of calculating the similarity between the description information of the node blocks. It has the advantages of accurate calculation and high efficiency, quickly determines the first combination, and improves the efficiency of data fusion.

[0108] In an embodiment of the present application, the electronic device uses LLM to obtain the first description information of each node block, and uses the cosine similarity algorithm to calculate the similarity of the description information of each node block, thereby determining the N first combinations of the M node blocks based on the similarity of the description information of the node blocks. Because the first combination is determined based on the similarity between the description information generated after summarizing and analyzing the attribute information and associated information included in each node block. Therefore, there must be similarity between certain nodes in the first combination, so when further data fusion is performed on the nodes, the corresponding nodes to be fused can be determined in the first combination. The present application first determines the node blocks with similar semantics as the first combination, and then uses the nodes in the first combination for fusion, which can avoid the mistaken fusion of nodes with different descriptions but the same name, that is, avoid the fusion of data with the same form but different meanings. It improves the accuracy of semantic alignment, ensures the accuracy of data fusion, and provides a guarantee for the accurate and efficient fusion between subsequent data to be fused.

[0109] In some embodiments, data fusion is performed on each node in the corresponding first group based on the attribute information of the node in each first group, including:

[0110] For any first combination, determining second description information describing the semantics of the corresponding node based on the association information and attribute information included in each node in the first combination;

[0111] Determining the similarity between the corresponding node and other nodes based on the second description information of each node;

[0112] The first node and the second node are data-fused; the similarity between the first node and the second node satisfies a third preset condition, the first node is any node in the first combination, and the second node is a node in the first combination other than the first node.

[0113] It is understandable that in an embodiment of the present application, the electronic device inputs the association information and attribute information included in each node in the first combination into the LLM, and obtains the second description information generated by the LLM to describe the semantics of each node. The second description information is a natural language text that can comprehensively describe the semantics of each node. The present application uses a cosine similarity algorithm to calculate the similarity between the second description information of the nodes, compares the similarity between the second description information of the nodes with a second threshold, and performs data fusion on the nodes corresponding to the second description information whose similarity is greater than or equal to the second threshold.

[0114] The present embodiment takes into account the existence of heterogeneous synonymous entities. These entities have different names but the same functional attributes. They can essentially be considered the same entity, differing only in name attributes due to different departments. For example, a device called a central processing unit (CPU) in the hardware department might be called a processor in other departments, or even a controller in some departments. However, a CPU, processor, and controller are all devices used for processing and control and can be considered the same entity. Therefore, when performing data fusion on each node, not only should the similarity of the node names be compared, but the node associations and other attribute information should also be considered. In the present embodiment, the association information and attribute information included in each node in the first group are input into the macro model, and the macro model outputs second description information for each node in the first group. This allows the macro model to comprehensively determine whether each node in the first group represents the same entity, even if the node names differ. When the similarity between the description information of each node is greater than or equal to a second threshold, meaning that the corresponding two or more nodes meet a third preset condition, data fusion can be performed on the corresponding nodes. The third preset condition is that the similarity between the two nodes is greater than a similarity threshold. The third preset condition ensures that the fused entities are unique and accurate by quantifying node-level semantic consistency. The third preset condition is a pre-set similarity threshold (usually 0.9 to 0.99, such as 0.95), which serves as the basis for node-level fusion. When the cosine similarity between the second description information of two nodes is greater than or equal to the third preset condition, the data fusion operation is triggered.

[0115] For example, it is taken as an example to determine whether the node X and the node Y in the first combination meet the third preset condition. Figure 7 As shown, first, the electronic device inputs the association information and attribute information included in node X and the association information and attribute information included in node Y into the LLM, respectively. The LLM outputs the second description information of node X and the second description information of node Y in the first combination. The similarity between the second description information of node X and the second description information of node Y is calculated using cosine similarity. If the similarity is greater than or equal to a second threshold, nodes X and Y in the first combination are considered to represent the same entity, and data fusion of nodes X and Y is performed. If the similarity is less than the second threshold, nodes X and Y are considered to represent different entities, and data fusion of nodes X and Y cannot be performed.

[0116] For example, Figure 4 As shown, the first combination includes node block A1 and node block A2. Node block A1 includes an E node indicating a central processing unit, and node block A2 includes an F node indicating a processor and a G node indicating an image processor. The electronic device inputs the association information and attribute information of the E node into the LLM to obtain the second description information of the E node. The electronic device inputs the association information and attribute information of the F node into the LLM to obtain the second description information of the F node. The cosine similarity algorithm is used to calculate the similarity between the first description information of the E node and the second description information of the F node. If the similarity is greater than or equal to the second threshold, it can be determined that the entities indicated by the E node and the F node are heteromorphic synonymous entities, that is, the same entity. Data fusion can be performed on the E node and the F node.

[0117] After determining that each node indicates the same entity, the embodiment of the present application performs data fusion on the nodes indicating the same entity to improve the accuracy and consistency of the data. Based on the accuracy and consistency of the data, the embodiment of the present application can more efficiently perform the fusion operation on the nodes indicating the same entity, avoiding the interference of redundant data and improving the accuracy of data fusion.

[0118] In some embodiments, the attribute information includes an attribute name and an attribute value; and fusing data of the first node with the second node includes:

[0119] Determining, based on a first attribute name in the first attribute information and a second attribute name in the second attribute information, whether an attribute indicated by the first attribute information is the same as an attribute indicated by the second attribute information, wherein the first attribute information is any attribute information in the first node and the second attribute information is any attribute information in the second node;

[0120] In response to the attribute indicated by the first attribute information being the same as the attribute indicated by the second attribute information, determining a third attribute name corresponding to the target attribute based on the first attribute name and the second attribute name, and determining a third attribute value corresponding to the target attribute based on the first attribute value of the first attribute information and the second attribute value of the second attribute information, the target attribute being the attribute indicated by the first attribute information and the second attribute information;

[0121] The first attribute information and the second attribute information are merged into target attribute information, where the target attribute information includes a third attribute name and a third attribute value.

[0122] In the present application, the attribute information of a node includes an attribute name and an attribute value. For example, the attribute name includes a size attribute, and the attribute value is length 16 and width 25. The attribute name can also be data capacity, and the attribute value is 1GB. In the scheme of data fusion of the first node and the second node, the first attribute name of the first attribute information of the first node can be compared with the second attribute name of the second attribute information of the second node to determine whether the attribute indicated by the first attribute name is the same as the attribute indicated by the second attribute name. If they are the same, the third attribute name corresponding to the target attribute after the final fusion can be determined based on the first attribute name and the second attribute name. And based on the first attribute value of the first attribute information and the second attribute value of the second attribute information, the third attribute value corresponding to the target attribute after the final fusion is determined.

[0123] Determining whether the attribute indicated by the first attribute name is identical to the attribute indicated by the second attribute name can be done by inputting the first and second attribute names into an LLM, obtaining third description information output by the LLM, and determining whether the first and second attribute names are identical based on the similarity between their respective third description information. If the similarity between the third description information of the first and second attribute names is greater than or equal to a third threshold, the first and second attribute names are confirmed to be identical. The third threshold can be considered the minimum threshold at which the third description information of each attribute name can be compared for similarity, ensuring that the corresponding attribute names are identical. The third threshold is used to measure whether the semantic similarity between nodes within the same first group meets the fusion condition. Essentially, it quantifies node-level semantic consistency to ensure that the fused entity is unique and accurate. The third threshold is a pre-set similarity value (typically 0.9 to 0.99, such as 0.95) and is a specific implementation of the third pre-set condition. When the cosine similarity between the second description information of two nodes is greater than or equal to the third threshold, the data fusion operation is triggered. If the similarity between the third description information of the first attribute name and the second attribute name is less than a third threshold, the first attribute name and the second attribute name are determined to be different attribute names. The first attribute name and the second attribute name are retained when data fusion is performed on the first node and the second node.

[0124] For example, when data is merged between a B node indicating a storage device and an H node indicating a memory bank, the first attribute name of the B node is a storage attribute, and the first attribute value is 512G; the second attribute name of the H node is a memory attribute, and the second attribute value is 1T; the attributes of the B node and the H node are essentially storage attributes, so the first attribute name of the B node is considered to be the same as the second attribute name of the H node. The storage attribute can be determined as the third attribute name after the data of the B node and the H node are merged. Based on the first attribute value of 512G and the second attribute value of 1T, the third attribute value is determined to be 1T or 512G. The fourth attribute name of the B node is a name attribute, and the fifth attribute name of the H node is an age attribute. It is determined that the attributes of the fourth attribute name of the B node are essentially different from the fifth attribute name and other attribute names of the H node. Therefore, when the data of the B node and the H node are merged, the third attribute name, the fourth attribute name, and the fifth attribute name are retained.

[0125] This application uses a hierarchical approach to attribute names and values ​​to semantically identify entities with different names for the same attribute across different departments. After semantic matching, entities with different names are merged to the same attribute name. This ensures that the merged attribute name uniquely corresponds to the entity. This application also further merges the attribute values ​​of entities to achieve unified standards across departments, improving data standardization and usability.

[0126] In some embodiments, determining a third attribute value corresponding to a target attribute based on a first attribute value of the first attribute information and a second attribute value of the second attribute information includes:

[0127] determining the first attribute value as the third attribute value;

[0128] Alternatively, the second attribute value is determined as the third attribute value;

[0129] Alternatively, information generated based on the first attribute value and the second attribute value is determined as the third attribute value.

[0130] In this application, the electronic device may determine the first attribute value as the third attribute value, may determine the second attribute value as the third attribute value, or may determine the third attribute value based on information generated by the first attribute value and the second attribute value.

[0131] For example, if it is an age attribute, the first attribute value is 25 years old and the second attribute value is 26 years old, then 25 years old can be determined as the third attribute value, or 26 years old can be determined as the third attribute value, or the average of the first attribute value and the second attribute value, 25.5 years old, can be determined as the third attribute value.

[0132] In this application, when electronic devices determine the attribute values ​​of nodes after fusion, they can determine them by selecting attribute values ​​or generating attribute values ​​by combining them. This flexible processing of attribute values ​​in this application achieves the standardization and consistency of node attributes, eliminates data conflicts due to different attribute values ​​during data fusion, and improves data standardization.

[0133] In some embodiments, each node block includes level information;

[0134] Determining N first combinations from M node blocks corresponding to the plurality of data to be fused includes:

[0135] From the M node blocks, P node blocks indicated by the level information as the first level are determined, and N first combinations are determined from the P node blocks, where P is a positive integer.

[0136] In the present application, each node block includes hierarchical information, and the hierarchical information is used to identify the level to which the node block belongs (such as the first level, the second level, etc.). Different hierarchical information indicates that the data range included in the node block is different. In the present application, the first level can be used as the level with the smallest data range of the node block. When determining the first combination, the P node blocks of the first level can be combined to obtain the first combination. Among them, the data range of the node block of the second level is larger than the data range of the node block of the first level. Among them, the first level can also be an intermediate level. That is, the first level can have an upper level, and the first level can also have a lower level, which is not limited in the present application.

[0137] After obtaining M node blocks, the present application determines which of the M node blocks are the first-level node blocks, and combines the node blocks among the P node blocks that meet the first preset condition to obtain a first combination. It should be noted that the node blocks in the first combination can all be the first level in their respective levels, or they can be different levels. For example, one node block in the first combination belongs to the first level, and the other node block belongs to the second level.

[0138] For example, Figure 8As shown, node block C1 has an upper-level node block D1, and node block C2 has an upper-level node block D2. The similarity between node block C1 and node block C2 is calculated. If the similarity between node block C1 and node block C2 meets the first preset condition, it means that the similarity between the first description information of node block C1 and node block C2 is greater than or equal to the first threshold value, then node block C1 and node block C2 are combined to obtain a first combination. If node block C1 and node block C2 can be combined, then node block D1 is the upper level of node block C1, which means that the first description information of node block D1 has a certain similarity with the first description information of node block C1. Similarly, the first description information of node block D2 has a certain similarity with the first description information of node block C2. In order to be able to fuse the nodes on a larger scale and ensure the accuracy of node fusion, the present application can also fuse the nodes in node block D1 and node block D2.

[0139] This application prioritizes combining the lowest-level node blocks through hierarchical information to narrow the scope of nodes that need to be fused and improve data fusion efficiency.

[0140] In some embodiments, data fusion is performed on each node in the corresponding first group based on the attribute information of the node in each first group, including:

[0141] For any first combination, determining an upper-level node block of the corresponding node block based on the level information of each node block in the first combination, wherein the corresponding node block belongs to the upper-level node block;

[0142] Data fusion is performed on the nodes in each upper-level node block corresponding to the first combination.

[0143] In this application, if Figure 9 As shown, after determining the first combination, based on the hierarchical information of each node block in the first combination, it is determined whether the corresponding node block has an upper-level node block. If the corresponding node block has an upper-level node block, data fusion is performed on the nodes in the upper-level node block of the corresponding node block. If the corresponding node block does not have an upper-level node block, data fusion is directly performed on the nodes in the node block with the nodes in the upper-level node block of the corresponding node block.

[0144] For example, continue as Figure 8 As shown, the first combination includes node block C1 and node block C2. If node block C1 has an upper-level node block D1 and node block C2 has an upper-level node block D2, then when performing node data fusion, data fusion can be performed on the nodes in the upper-level node block D1 and the upper-level node block D2 to fuse nodes within a larger data range and improve the accuracy of node data fusion. If node block C2 does not have an upper-level node block, data fusion is directly performed on the nodes in node block C2 and the nodes in node block D1.

[0145] When the node block includes multiple levels of information, data fusion can be performed on the nodes of the higher-level node blocks corresponding to the first-level node blocks in the first combination. Data fusion can also be performed on the upper-level node blocks corresponding to the first-level node blocks in the first combination. For example: node block D1 also has an upper-level node block E1, and node block D2 also has an upper-level node block E2, then the nodes in node block E1 and node block E2 corresponding to node block C1 and node block C2 in the first combination can be fused. However, fusing the highest-level node blocks corresponding to each node block in the first combination will cause the fused data to be too large. Therefore, the present application can be set to fuse the nodes in the upper-level node blocks of the node blocks in the first combination.

[0146] The embodiments of the present application can more precisely divide and match nodes by distinguishing the hierarchical information of node blocks. Determining the first combination from the lower-level node blocks in different hierarchies ensures the accuracy of determining nodes indicating the same entity. By fusing the nodes in the upper-level node blocks, data fusion can be achieved on a larger scale for the nodes, achieving more accurate and efficient fusion of data to be fused from multiple departments.

[0147] like Figure 10 As shown, an embodiment of the present application provides a data fusion device, comprising:

[0148] The data acquisition module 1001 is used to acquire a plurality of data to be fused, wherein the data to be fused includes at least two nodes; a node is used to indicate an entity, and a node includes attribute information of the entity;

[0149] The node division module 1002 is configured to divide the first data to be fused into at least one node block based on attribute information of each node in the first data to be fused; the first data to be fused is any data to be fused among the multiple data to be fused;

[0150] A first combination module 1003 is configured to determine N first combinations from the M node blocks corresponding to the plurality of data to be fused; wherein the first combination includes at least two node blocks, the similarity between the node blocks in the first combination satisfies a first preset condition, and M and N are positive integers;

[0151] The data fusion module 1004 is configured to fuse the data of each node in the corresponding first combination based on the attribute information of the nodes in the first combination to obtain target fused data; each node in the target fused data indicates a different entity.

[0152] In the embodiment of the present application, a plurality of data to be fused is acquired by a data acquisition module 1001. A node partitioning module 1002 is used to partition the first data to be fused into at least one node block based on the attribute information of each node in the first data to be fused. The first data to be fused is any data to be fused from the plurality of data to be fused. A first combination module 1003 is used to determine N first combinations from the M node blocks corresponding to the plurality of data to be fused. Finally, a data fusion module 1004 is used to fuse the nodes in the corresponding first combinations based on the attribute information of the nodes in each first combination to obtain target fused data. Each node in the target fused data indicates a different entity.

[0153] Figure 11 A schematic block diagram of an example electronic device 800 that can be used to implement embodiments of the present application is shown. The electronic device 800 includes a computing unit 801 that can perform various appropriate actions and processes based on a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. The RAM 803 may also store various programs and data required for the operation of the device 800. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.

[0154] Various components in device 800 are connected to I / O interface 805, including an input unit 806, such as a keyboard, mouse, etc.; an output unit 807, such as various types of displays, speakers, etc.; a storage unit 808, such as a magnetic disk, optical disk, etc.; and a communication unit 809, such as a network card, modem, wireless communication transceiver, etc. The communication unit 809 allows device 800 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0155] The computing unit 801 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 801 performs the various methods and processes described above, such as the data fusion method. For example, in some embodiments, the data fusion method can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded into the RAM 803 and executed by the computing unit 801, one or more steps of the data fusion method described above can be performed. Alternatively, in other embodiments, the computing unit 801 can be configured to perform the data fusion method by any other suitable means (e.g., via firmware).

[0156] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features being referred to. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one such feature. Throughout the description of this application, "plurality" means two or more, unless otherwise specifically defined.

[0157] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.

Claims

1. A data fusion method, characterized in that: include: Acquire a plurality of data to be fused, wherein the data to be fused includes at least two nodes; the nodes are used to indicate entities, and the nodes include attribute information of the entities; Based on the attribute information of each node in the first data to be fused, dividing the first data to be fused into at least one node block; The first data to be fused is any data to be fused among the multiple data to be fused; Determining N first combinations from the M node blocks corresponding to the plurality of data to be fused; wherein the first combination includes at least two node blocks, the similarity between the node blocks in the first combination satisfies a first preset condition, and M and N are positive integers; Based on the attribute information of the nodes in each first combination, data of each node in the corresponding first combination is fused to obtain target fused data; the entities indicated by each node in the target fused data are different.

2. The data fusion method according to claim 1, characterized in that: The attribute information of the entity includes functional attributes; and dividing the first data to be fused into at least one node block includes: Determining similarities between the nodes in the first data to be fused based on functional attributes of the nodes in the first data to be fused; Nodes whose similarities satisfy a second preset condition are divided into the same node block to obtain at least one node block.

3. The data fusion method according to claim 1, characterized in that: The node also includes association information between the entity and other entities; The determining N first combinations from the M node blocks corresponding to the plurality of data to be fused includes: Determining first description information describing the semantics of the node block based on the attribute information and association information included in each node in each node block; Determining the similarity between the corresponding node block and other node blocks according to the first description information of each node block; The first node block and the second node block are combined to obtain the first combination, the similarity between the second node block and the first node block satisfies the first preset condition, the first node block is any node block among the M node blocks, and the second node block is a node block among the M node blocks except the first node block.

4. The data fusion method according to claim 3, characterized in that: The step of fusing data of the nodes in the corresponding first combination based on the attribute information of the nodes in the first combination includes: For any first combination, determining second description information describing the semantics of the corresponding node based on the association information and attribute information included in each node in the first combination; Determining the similarity between the corresponding node and other nodes based on the second description information of each node; Data of the first node and the second node are fused; the similarity between the first node and the second node satisfies a third preset condition, the first node is any node in the first combination, and the second node is a node in the first combination other than the first node.

5. The data fusion method according to claim 4, characterized in that: The attribute information includes an attribute name and an attribute value; The fusing data of the first node and the second node includes: Determining, based on a first attribute name in first attribute information and a second attribute name in second attribute information, whether an attribute indicated by the first attribute information is the same as an attribute indicated by the second attribute information, wherein the first attribute information is any attribute information in the first node and the second attribute information is any attribute information in the second node; In response to the attribute indicated by the first attribute information being the same as the attribute indicated by the second attribute information, determining a third attribute name corresponding to a target attribute based on the first attribute name and the second attribute name, and determining a third attribute value corresponding to the target attribute based on the first attribute value of the first attribute information and the second attribute value of the second attribute information, the target attribute being the attribute indicated by the first attribute information and the second attribute information; The first attribute information and the second attribute information are merged into target attribute information, where the target attribute information includes the third attribute name and the third attribute value.

6. The data fusion method according to claim 5, characterized in that: The determining, based on the first attribute value of the first attribute information and the second attribute value of the second attribute information, a third attribute value corresponding to the target attribute includes: determining the first attribute value as the third attribute value; Alternatively, the second attribute value is determined as the third attribute value; Alternatively, information generated based on the first attribute value and the second attribute value is determined as the third attribute value.

7. The data fusion method according to claim 1, characterized in that: Each node block includes hierarchical information; The determining N first combinations from the M node blocks corresponding to the plurality of data to be fused includes: From the M node blocks, P node blocks indicated by the hierarchy information as the first hierarchy are determined, and the N first combinations are determined from the P node blocks, where P is a positive integer.

8. The data fusion method according to claim 7, characterized in that: The step of fusing data of the nodes in the corresponding first combination based on the attribute information of the nodes in the first combination includes: For any first combination, determining an upper-level node block of the corresponding node block based on the level information of each node block in the first combination, wherein the corresponding node block belongs to the upper-level node block; Data fusion is performed on the nodes in each upper-level node block corresponding to the first combination.

9. The data fusion method according to any one of claims 1 to 8, characterized in that: The data to be fused is a knowledge graph.

10. An electronic device, characterized in that: include: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor. The instructions are executed by the at least one processor to enable the at least one processor to perform the data fusion method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Multi-source data-based knowledge fusion method

    CN108647318A

  • A data fusion method and device for a knowledge graph

    CN109739939A

  • Data fusion method and device, computer equipment and computer storage medium

    CN110580304A

  • Multi-source data fusion method and device, computer equipment and storage medium

    CN118709142A

  • Data processing method and device, equipment and storage medium

    CN119884498A