An Ontology-Based Composite Semantic Relevance Quantification Method and System

By preprocessing the composite semantics and using virtual nodes, the errors and difficulties in the calculation of the correlation of compound semantics and node sets in the prior art are solved, and a more accurate and applicable semantic correlation quantization and privacy protection effect is achieved.

CN119761383BActive Publication Date: 2025-06-20BEIJING ANDY TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510273422.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-10
Publication Date
2025-06-20
Estimated Expiration
2045-03-10

AI Technical Summary

Technical Problem

The existing ontology-based semantic correlation measurement scheme has errors in dealing with compound semantics and uncommon semantics, and it is difficult to calculate the correlation of node sets, and it is impossible to effectively protect the collective privacy of multiple semantics.

Method used

By preprocessing the composite semantics, it is broken down into simple semantics, and using virtual nodes to represent the node set, a many-to-many mapping relationship is used to quantify the mutual inference probability and information volume between nodes, and calculate the correlation between virtual nodes.

Benefits of technology

It reduces the calculation error in semantic correlation quantization, can effectively calculate the correlation of node sets, supports the allocation of sensitivity of composite semantics to multiple nodes in node sets, and improves the accuracy and applicability of privacy protection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119761383B_ABST
    Figure CN119761383B_ABST
Patent Text Reader

Abstract

The present invention discloses an ontology-based method and system for quantifying composite semantic relevance, including: S1, preprocessing the first composite semantics and the second composite semantics and respectively mapping them to the first node set and the second node set of the ontology in the semantic network; S2, setting a first virtual node and a second virtual node for the first composite semantics and the second composite semantics respectively; S3, quantifying the first relevance of the first virtual node to all nodes in the first node set; S4, quantifying the second relevance between the nodes in the first node set and the nodes in the second node set; S5, quantifying the third relevance of the nodes in the second node set to the second virtual node; S6, quantifying the fourth relevance of the first virtual node to the second virtual node, that is, the relevance between the first composite semantics and the second composite semantics. The present invention solves the problem that it is difficult to map real semantics to semantic nodes and the problem of mutual conversion between the relevance of real composite semantics and the relevance of node sets.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of semantic relevance quantification, and specifically relates to a method and system for quantifying composite semantic relevance based on ontology. Background Art

[0002] While big data brings convenience to people's lives, it also brings the problem of information overload. Now people can use algorithms to collect the maximum amount of users' personal information and interaction information to predict users' future interest preferences, so as to achieve the purpose of providing precise services for different users. However, in many cases, there are sensitive information in a large amount of user information that poses a threat to personal privacy, which means that attackers have more ways to attack privacy than ever before.

[0003] Suppose there is a need to protect specific privacy attributes, and different attributes have different protection intensity requirements. Then the privacy protection task needs to find other information with a relatively high correlation with specific privacy information. The current mainstream solutions are to quantify relevance based on datasets and to quantify relevance based on semantics. Quantifying relevance based on semantics includes quantification based on corpus and quantification based on ontology.

[0004] The existing ontology-based semantic relevance measurement scheme uses a method of reasoning along semantic relationships from known nodes in a semantic network until the target node is reached. At this time, the relevance between this node and the target node is equal to the reasoning contribution of this node to the target node.

[0005] The main problems existing in the existing technical solutions are that most semantic relevance schemes focus on the connections between nodes, and quantify the degree of association of real-world semantics by calculating the degree of association between a single node and the target node in the graph structure. This method has the following specific problems:

[0006] (1) There are composite semantics and rare semantics in real-world semantics. It is difficult to find completely corresponding nodes for these semantics in the semantic network, or only marginal nodes with very few connection relationships with other nodes can be found. The existing technology can only map these semantics to the nodes with the closest semantics, which will cause a large error in semantic relevance quantification.

[0007] (2) The existing technology only considers the degree of association between nodes and does not consider the degree of association between node sets. There are composite semantics in real-world semantics that need to be mapped to multiple nodes at the same time, and the relevance of these semantics should be represented by the relevance of the node set. The existing technology only considers the one-to-one node relevance, and there is no definite scheme for studying the relevance of node sets.

[0008] (3) The target privacy semantics of the privacy protection task often have more than one, and may be a set of multiple semantics. Different privacy semantics have different sensitivities and require different levels of protection. Even if the sensitivity of the realistic composite semantics has been set, there is currently no technology that clearly explains how to distribute the sensitivity of the composite semantics to multiple nodes in the node set. Summary of the Invention

[0009] In view of the deficiencies in the prior art, the present invention provides an ontology-based method and system for quantifying the relevance of composite semantics, which solves the problem that realistic semantics are difficult to map to semantic nodes, solves the relevance calculation of the node set, and solves the mutual conversion problem between the relevance of realistic composite semantics and the relevance of the node set.

[0010] The present invention provides an ontology-based method for quantifying the relevance of composite semantics, and the method includes:

[0011] S1. After preprocessing the first composite semantics and the second composite semantics, map them to the first node set and the second node set of the ontology in the semantic network respectively, where the relevance between the first composite semantics and the second composite semantics is the relevance from the first node set to the second node set;

[0012] S2. Set a first virtual node and a second virtual node for the first composite semantics and the second composite semantics respectively, where the relevance between the first composite semantics and the second composite semantics is the relevance from the first virtual node to the second virtual node, and the virtual node is a hyponym of all nodes in the corresponding node set;

[0013] S3. Quantify the first relevance from the first virtual node to all nodes in the first node set, where the relevance between nodes is the mutual inference probability between nodes;

[0014] S4. Based on the graph features of the ontology and the information content of the nodes in the first node set and the second node set, quantify the second relevance between the nodes in the first node set and the nodes in the second node set;

[0015] S5. Based on the information content shared between the nodes in the second node set, quantify the third relevance from the nodes in the second node set to the second virtual node;

[0016] S6. Based on the transitivity of the inference probability, the first relevance, the second relevance, and the third relevance, quantify the fourth relevance from the first virtual node to the second virtual node, that is, the relevance between the first composite semantics and the second composite semantics.

[0017] Preferably, step S1 includes:

[0018] S11. Decompose the first composite semantics and the second composite semantics into simple semantics that can be directly mapped to semantic nodes. Before decomposition, it also includes semantic fusion of rare words, where the simple semantics are nouns or adjectives.

[0019] S12. Map the simple semantics corresponding to the first composite semantics and the second composite semantics to the first node set and the second node set of the semantic network respectively by calling the lexical node mapping method of the ontology.

[0020] Preferably, the relationship between the node set and the node is represented as the relationship between the composite semantics and the sub-semantics. The existence of the composite semantics means that all the sub-semantics exist simultaneously; the virtual node refers to a node that does not exist in the ontology but can correspond to the real composite semantics; the virtual node is the hyponym of all the nodes in the node set, and the occurrence of this virtual node means that all the nodes in the corresponding node set occur; the correlation between the two node sets can be converted into the correlation between the two virtual nodes.

[0021] Preferably, the correlation between nodes is the mutual inference probability between nodes , if the semantics represented by node can infer the semantics represented by node with a probability of , then the correlation between node and node is:

[0022] = (1);

[0023] The inference probability should satisfy the following two conditions:

[0024] ① Transitivity: Assume there exists a node , then there is:

[0025] = (2);

[0026] ② Probability additivity: Assume there exists a node , then there is:

[0027] = (3);

[0028] The virtual node is a hyponym of all nodes in the node set. The occurrence of this virtual node indicates the occurrence of all nodes in the corresponding node set, that is, the relevance of the virtual node pointing to all nodes in the node set is 1. In step S3, the first relevance of the first virtual node to all nodes in the first node set is 1.

[0029] Preferably, in step S4, the following method is used to quantify the second relevance of the nodes in the first node set to the nodes in the second node set, that is, to quantify the relevance between the nodes in the ontology. The relevance between nodes is the probability that the starting node infers the ending node through the relationship between the nodes:

[0030] S41. Perform a graph traversal based on the starting node, set the longest path between the starting node and the ending node, use a stack to record the current path, and set two sets to record the nodes and node relationships of the current path respectively;

[0031] S42. When the ending node is traversed at the specified path length, record the nodes and the relationships between the nodes in the stack in the two sets. After the traversal is completed, construct the sensitive value transfer subgraph of the current path;

[0032] S43. Calculate the relevance between nodes in the sensitive value transfer subgraph of each path. When a node has relationships with multiple nodes, the calculation formula for the relevance between nodes of each path is:

[0033] (4);

[0034] In formula (4), represents the path weight from node to node , that is, the inference probability and also the relevance, represents the set of all hyponyms of node , represents the information content of the parent node , is the information content of the target node , and the information content of the node is determined based on the graph features of the ontology or based on the corpus;

[0035] S44. If there is a single path between the starting node and the ending node, the relevance between the starting node and the ending node is the product of all path weights on this path, and the calculation formula is:

[0036] (5);

[0037] In formula (5), represents the path from the starting node to the ending node The relevance of, where k represents a node to node The length of the longest path between, represents from node to the adjacent node The path weight of, where i and k are positive integers;

[0038] S45. If there are multiple paths between the starting node and the ending node in the same sensitive value transfer subgraph, the relevance between the starting node and the ending node is the sum of the relevance deduced from the multiple paths, and the relevance does not exceed 1. The calculation formula is:

[0039] (6);

[0040] In formula (6), represents from the starting node to the ending node The relevance of, where k represents a node to node The length of the longest path between, q represents the starting node to the ending node The number of paths existing in the same sensitive value transfer subgraph, represents from node to the adjacent node The path weight of, where i, j, k, and q are positive integers.

[0041] Preferably, in step S5, the following method is used to quantify the third relevance of the nodes in the second node set to the second virtual node:

[0042] When there are two nodes in the node set the relevance of node to the virtual node is the relevance between the two nodes The calculation formula is:

[0043] (7);

[0044] When there are at least three nodes in the node set, the calculation formula for the relevance of node to the virtual node is:

[0045] (8);

[0046] In formulas (7) and (8), assuming that the compound semantics corresponding to event is and the simple semantics corresponding to events A and B are , composite semantics The corresponding node set is , event The corresponding semantics is , semantics The corresponding node set is , Indicates the probability that event S occurs when event A occurs, Indicates the probability that event Occurs when event A occurs. The degree of mutual influence among multiple nodes in the node set is related to the amount of information shared between the nodes. The amount of information shared between the nodes is represented by the least common ancestor node of the nodes, The semantics represented by , k represents to The path length of

[0047] Preferably, in step S5, the correlation between the first composite semantics and the second composite semantics is the correlation from the first node set to the second node set, and is also the correlation from the first virtual node to the second virtual node. The calculation formula is: ;

[0048] In formula (9), Represents the first composite semantics, Represents the second composite semantics, Represents the first node set, Represents the second node set, Represents the first virtual node, Represents the second virtual node.

[0049] Based on the same inventive concept, the present invention also provides an ontology-based composite semantics correlation quantification system, and the system includes:

[0050] A semantic node mapping module, configured to preprocess the first composite semantics and the second composite semantics and map them to the first node set and the second node set of the ontology in the semantic network respectively, where the correlation between the first composite semantics and the second composite semantics is the correlation from the first node set to the second node set;

[0051] A virtual node setting module, configured to set a first virtual node and a second virtual node for the first composite semantics and the second composite semantics respectively, where the correlation between the first composite semantics and the second composite semantics is the correlation from the first virtual node to the second virtual node, and the virtual node is a hyponym of all nodes in the corresponding node set;

[0052] The first quantization module is used to quantify the first correlation of the first virtual node to all nodes in the first node set, where the correlation between nodes is the mutual inference probability between nodes;

[0053] The second quantization module is used to quantify the second correlation between the nodes in the first node set and the nodes in the second node set based on the graph features of the ontology, the information content of the nodes in the first node set and the second node set;

[0054] The third quantization module is used to quantify the third correlation of the nodes in the second node set to the second virtual node based on the information content shared between the nodes in the second node set;

[0055] The fourth quantization module is used to quantify the fourth correlation of the first virtual node to the second virtual node, that is, the correlation between the first composite semantics and the second composite semantics, based on the transitivity of the inference probability, the first correlation, the second correlation, and the third correlation.

[0056] Preferably, the semantic node mapping module is specifically configured to:

[0057] Decompose the first composite semantics and the second composite semantics into simple semantics that can be directly mapped to semantic nodes. Before decomposition, semantic fusion of rare words is also included, where the simple semantics are nouns or adjectives;

[0058] Map the simple semantics corresponding to the first composite semantics and the second composite semantics to the first node set and the second node set of the semantic network respectively by calling the lexical node mapping method of the ontology.

[0059] Preferably, the relationship between the node set and the node is represented as the relationship between the composite semantics and the sub-semantics. The existence of the composite semantics means that all sub-semantics exist simultaneously; the virtual node refers to a node that does not exist in the ontology but can correspond to the real composite semantics; the virtual node is the hyponym of all nodes in the node set, and the occurrence of this virtual node means that all nodes in the corresponding node set occur; the correlation between two node sets can be converted into the correlation between two virtual nodes.

[0060] Compared with the prior art, the beneficial effects of the present invention are:

[0061] 1. The present invention solves the problem that the real semantics and semantic nodes cannot be fully mapped. By preprocessing the semantics, the semantic granularity is reconstructed, so that all semantics are mapped to the semantic nodes in a many-to-many relationship, reducing the calculation error.

[0062] 2. The present invention solves the technical shortage of quantifying the correlation of node sets, extends the one-to-one correlation quantification of nodes in the existing solution to the quantification of many-to-many relationships, and solves the problem of cyclic traversal required by the existing technology through the idea of virtual nodes.

[0063] 3. The present invention solves the problem of mutual mapping between composite semantic correlation and node set correlation, solves the problem that it is difficult to transition between node correlation and semantic correlation, and makes the technology more in line with the actual privacy protection requirements. BRIEF DESCRIPTION OF THE DRAWINGS

[0064] Figure 1 It is a schematic flowchart of a method for quantifying composite semantic correlation based on ontology provided by the present invention;

[0065] Figure 2 It is a schematic diagram of a virtual node provided by the present invention;

[0066] Figure 3 It is a schematic diagram of the transitivity of inference probability provided by the present invention;

[0067] Figure 4 It is a schematic diagram of the additivity of probabilities of inference probability provided by the present invention;

[0068] Figure 5 It is a schematic structural diagram of a system for quantifying composite semantic correlation based on ontology provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0069] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0070] The present invention will be further described in detail below with reference to the accompanying drawings.

[0071] As Figure 1 shown, an embodiment of the present invention provides a method for quantifying composite semantic correlation based on ontology, and the method includes:

[0072] S1. After preprocessing the first composite semantics and the second composite semantics, map them to the first node set and the second node set of the ontology in the semantic network respectively, where the correlation between the first composite semantics and the second composite semantics is the correlation from the first node set to the second node set;

[0073] S2. Set a first virtual node and a second virtual node for the first composite semantics and the second composite semantics respectively. Among them, the relevance between the first composite semantics and the second composite semantics is the relevance from the first virtual node to the second virtual node, and the virtual node is the hyponym of all nodes in the corresponding node set.

[0074] S3. Quantify the first relevance of the first virtual node to all nodes in the first node set, where the relevance between nodes is the mutual inference probability between nodes.

[0075] S4. Based on the graph features of the ontology, the information content of the nodes in the first node set and the second node set, quantify the second relevance of the nodes in the first node set to the nodes in the second node set.

[0076] S5. Based on the information content shared among the nodes in the second node set, quantify the third relevance of the nodes in the second node set to the second virtual node.

[0077] S6. Based on the transitivity of the inference probability, the first relevance, the second relevance, and the third relevance, quantify the fourth relevance of the first virtual node to the second virtual node, that is, the relevance between the first composite semantics and the second composite semantics.

[0078] In the embodiment of the present invention, step S1 includes:

[0079] S11. Decompose the first composite semantics and the second composite semantics into simple semantics that can be directly mapped to nodes by semantics. The simple semantics are nouns or adjectives. Before decomposition, semantic fusion of rare words is also included.

[0080] It is difficult to directly find completely corresponding nodes in the ontology for real semantics. Preprocessing the real semantics can solve this problem. The main work of preprocessing is semantic decomposition. There are a large number of words with a composite relationship of multiple semantics. The semantics of such words are regarded as a set of multiple semantics. Facing relatively complex semantics, the attribute semantics are extended from being mapped to one node to being mapped on a node set through semantic decomposition. In the field of privacy, most privacy is represented by nouns and adjectives. Therefore, the default for semantic decomposition is to only retain adjectives and nouns. For example, "unmarried woman" is decomposed into a semantic set {"unmarried", "woman"}.

[0081] In addition, preprocessing also includes semantic fusion of rare words. Some words are too rare to be recorded in the semantic network. By fusing the semantics, the semantics can be made more abstract, so as to be mapped to nodes with more connection relationships.

[0082] S12. Map the simple semantics corresponding to the first composite semantics and the second composite semantics to the first node set and the second node set of the semantic network respectively by invoking the lexical node mapping method of the ontology.

[0083] ;

[0084] ;

[0085] Among them, represents the first composite semantics, represents the second composite semantics, represents the composite semantics corresponding to the first node set, represents the composite semantics corresponding to the second node set, represents the disassembled simple semantics, represents the disassembled simple semantics, represents the simple semantics mapped to the nodes of the semantic network, represents the simple semantics mapped to the nodes of the semantic network.

[0086] The early technical solution was to perform graph traversal from a single node, reach another specified node within the specified path range, and calculate the correlation between the two nodes through these paths. This method can only calculate the correlation between single nodes in the ontology. To apply this method to node sets, the node set can be replaced by a single node by abstracting a virtual node.

[0087] A virtual node refers to a node that does not exist in the ontology but can correspond to the actual composite semantics. For example, the composite semantics "unmarried woman" does not have a corresponding node in the ontology, and it can be assumed that it can be mapped to a virtual node , whose represented semantics is unmarried_female.n, and is the hyponym node of the node unmarried.a.01 mapped by the semantics "unmarried" and the node female.n.02 mapped by the semantics "female".

[0088] As Figure 2 shown, the hypernym and hyponym are the most common relationship in the ontology. For example, car.n.01 and motor_vehicle.n.01 are a pair of hyponym and hypernym, and this pair of relationships can be understood as an "is-a" relationship, such as "car is a kind of motor_vehicle".

[0089] In the embodiments of the present invention, the relationship between the node set and the nodes is represented as the relationship between the composite semantics and the sub-semantics. The existence of the composite semantics means the simultaneous existence of all the sub-semantics; the virtual node refers to a node that does not exist in the ontology but can correspond to the real composite semantics; the virtual node is the hyponym of all the nodes in the node set, and the occurrence of the virtual node means the occurrence of all the nodes in the corresponding node set; the correlation between two node sets can be converted into the correlation between two virtual nodes, as follows.

[0090] ;

[0091] Among them, represents the correlation, represents the first node set, represents the correlation of the second node set, represents the first virtual node, represents the correlation of the second virtual node.

[0092] In the embodiments of the present invention, the correlation between nodes is the mutual inference probability between nodes , if the node represents the semantics that can infer the semantics represented by the node with a probability of , then the correlation between the node and the node is:

[0093] = (1);

[0094] As Figures 3-4 shown, the inference probability should satisfy the following two conditions:

[0095] ① Transitivity ( Figure 3 ): Assume that there is a node , then there is:

[0096] = (2);

[0097] ② Probability additivity ( Figure 4 ): Assume that there is a node , then there is:

[0098] = (3);

[0099] In the ontology, the path weight of different path relationships between nodes is the inference probability as shown in Table 1.

[0100] Table 1

[0101] ;

[0102] The virtual node is a hyponym of all nodes in the node set. The occurrence of this virtual node indicates the occurrence of all nodes in the corresponding node set, that is, the relevance between the virtual node and all nodes in the node set is 1. In step S3, the first relevance between the first virtual node and all nodes in the first node set is 1. ;

[0103] In the embodiment of the present invention, in step S4, the following method is used to quantify the second relevance between the nodes in the first node set and the nodes in the second node set, that is, to quantify the relevance between the nodes in the ontology. The relevance between nodes is the probability that the starting node infers the ending node through the relationship between the nodes:

[0104] S41. Perform a graph traversal according to the starting node, set the longest path between the starting node and the ending node, use a stack to record the current path, and set two sets to record the nodes and node relationships of the current path respectively;

[0105] S42. When the ending node is traversed under the specified path length, record the nodes and the relationships between the nodes in the stack in the two sets. After the traversal is completed, construct the sensitive value transfer subgraph of the current path;

[0106] S43. Calculate the relevance between nodes in the sensitive value transfer subgraph of each path. When a node has relationships with multiple nodes, the calculation formula for the relevance between nodes in each path is:

[0107] (4);

[0108] In formula (4), represents the path weight from node to node , that is, the inference probability and also the relevance, represents the set of all hyponyms of node , represents the information amount of the parent node , is the information amount of the target node . The information amount of a node is determined based on the graph features of the ontology or based on the corpus;

[0109] S44. If there is a single path between the starting node and the ending node, the relevance between the starting node and the ending node is the product of all path weights on this path, and the calculation formula is:

[0110] (5);

[0111] In formula (5), represents the relevance from the starting node to the ending node , k represents the node to node The longest path length between them, represents from node to the adjacent node The path weight of, i and k are positive integers;

[0112] S45. If there are multiple paths between the starting node and the ending node in the same sensitive value transfer subgraph, the relevance between the starting node and the ending node is the sum of the relevance inferred from multiple paths, and the relevance does not exceed 1. The calculation formula is:

[0113] (6);

[0114] In formula (6), represents the relevance from the starting node to the ending node , k represents the node to node The longest path length between them, q represents the starting node to the ending node The number of paths existing in the same sensitive value transfer subgraph, represents from node to the adjacent node The path weight of, i, j, k, and q are positive integers.

[0115] In the ontology, the inference weight for the superordinate term to infer the subordinate term is the weight value of the information amount. This method is not applicable to virtual nodes because the information amount of virtual nodes cannot be obtained through the structure of the ontology. For calculating the semantic relevance of complex events, this method needs to be improved.

[0116] In the embodiment of the present invention, in step S5, the third relevance of the nodes in the second node set to the second virtual node is quantified by the following method:

[0117] When there are two nodes in the node set , the node to the virtual node The relevance of is the relevance between the two nodes . The calculation formula is:

[0118] (7);

[0119] When there are at least three nodes in the node set, the node To the virtual node The calculation formula for the relevance is as follows:

[0120] (8);

[0121] In formulas (7) and (8), assume that the composite semantics corresponding to event is , the simple semantics corresponding to events A and B are , the node set corresponding to the composite semantics is , the semantics corresponding to event is , the node set corresponding to the semantics is , represents the probability that event S occurs when event A occurs, represents the probability that event occurs when event A occurs. The degree of mutual influence among multiple nodes in the node set is related to the amount of information shared between the nodes. The amount of information shared between the nodes is represented by the least common ancestor node of the nodes, the semantics represented by is to is the path length.

[0122] Specifically, assume there is a complex event , the semantics corresponding to event is , where a and b are the semantics corresponding to events A and B respectively. Set the virtual node corresponding to the semantics s as , the node set is { }, and are the nodes corresponding to the semantics a and b respectively, then the relevance between and is:

[0123] ;

[0124] represents the probability that event S occurs when event A occurs. According to the above formula, it can be deduced that:

[0125] ;

[0126] The problem is transformed into calculating , that is, when there are two nodes in the node set, the relevance between the node and the virtual node is the relevance between the two nodes.

[0127] For example, if "single female" = "single" ∩ "female", then the virtual node can be set as unmarried_female.n, and calculating NR(female.n.02 unmarried_female.n) can be converted to calculating NR(female.n.02 unmarried.a.01).

[0128] After expanding to multiple nodes, assuming that the semantics represents event S, and the corresponding node set is , The corresponding true semantics is , and the represented event is , so we can get:

[0129] ;

[0130] Here is a recursive algorithm with relatively high complexity. Adopting a simplified idea, only the relationship between the hypernym and the hyponym is considered. The degree of mutual influence among multiple nodes in the node set is related to the amount of information shared between the nodes, and the shared information can be represented by the least common ancestor node of the nodes, and the semantics represented by . In this way, the calculation of can be simplified. Assuming that the path length from to is k, the formula (8) can be obtained.

[0131] In the embodiment of the present invention, in step S5, the correlation between the first composite semantics and the second composite semantics is the correlation from the first node set to the second node set, and also the correlation from the first virtual node to the second virtual node. The calculation formula is:

[0132] ;

[0133] In formula (9), represents the first composite semantics, represents the second composite semantics, represents the first node set, represents the second node set, represents the first virtual node, represents the second virtual node.

[0134] Specifically, the ultimate goal of the present invention is to calculate the semantic correlation . If there is no composite semantics, that is, it can directly map to the nodes in the ontology, then The semantic relevance is the node relevance of the nodes mapped within the ontology:

[0135] ;

[0136] Among them, are respectively the nodes that can be mutually mapped with the semantics within the ontology.

[0137] If is a composite semantics, and the semantics of needs to be mapped through the node set within the ontology, then a virtual node can be abstracted according to the above method to map with the semantics . At this time, the semantic relevance of is:

[0138] ;

[0139] If are all composite semantics, calculating the relevance of the composite semantics is to calculate the relevance of the node set. The relevance of the node set can be abstracted as the relevance of the virtual node, that is:

[0140] ;

[0141] As shown in Figure 5 , an embodiment of the present invention further provides an ontology-based composite semantic relevance quantification system, and the system includes:

[0142] A semantic node mapping module 100, configured to preprocess the first composite semantics and the second composite semantics and map them to the first node set and the second node set of the ontology in the semantic network respectively, where the relevance between the first composite semantics and the second composite semantics is the relevance from the first node set to the second node set;

[0143] A virtual node setting module 200, configured to set a first virtual node and a second virtual node for the first composite semantics and the second composite semantics respectively, where the virtual node is the hyponym of all nodes in the corresponding node set;

[0144] A first quantification module 300, configured to quantify the first relevance of the first virtual node to all nodes in the first node set, where the relevance between nodes is the mutual inference probability between nodes;

[0145] A second quantification module 400, configured to quantify the second relevance of the nodes in the first node set to the nodes in the second node set based on the graph features of the ontology and the information content of the nodes in the first node set and the second node set;

[0146] The third quantization module 500 is configured to quantify the third correlation between the nodes in the second node set and the second virtual node based on the information amount shared by the nodes in the second node set;

[0147] The fourth quantization module 600 is configured to quantify the fourth correlation between the first virtual node and the second virtual node, that is, the correlation between the first composite semantics and the second composite semantics, based on the transitivity of the inference probability, the first correlation, the second correlation, and the third correlation.

[0148] In an embodiment of the present invention, the semantic node mapping module 100 is specifically configured to:

[0149] Decompose the first composite semantics and the second composite semantics into simple semantics that can be directly mapped between semantics and nodes. The simple semantics are nouns or adjectives. Before decomposition, semantic fusion of rare words is also included;

[0150] Respectively map the simple semantics corresponding to the first composite semantics and the second composite semantics to the first node set and the second node set of the semantic network by invoking the lexical node mapping method of the ontology.

[0151] In an embodiment of the present invention, the relationship between the node set and the node is represented as the relationship between the composite semantics and the sub-semantics. The existence of the composite semantics means that all the sub-semantics exist simultaneously; the virtual node refers to a node that does not exist in the ontology but can correspond to the real composite semantics; the virtual node is the hyponym of all the nodes in the node set, and the occurrence of the virtual node means that all the nodes in the corresponding node set occur; the correlation between two node sets can be converted into the correlation between two virtual nodes.

[0152] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A composite semantic relevance quantification method based on ontology, characterized in that: The method comprises: S1. Mapping the first composite semantics and the second composite semantics to the first node set and the second node set of the ontology in the semantic network after preprocessing, respectively, wherein the correlation between the first composite semantics and the second composite semantics is the correlation between the first node set and the second node set; S2, setting a first virtual node and a second virtual node for the first composite semantics and the second composite semantics respectively, wherein the correlation between the first composite semantics and the second composite semantics is the correlation between the first virtual node and the second virtual node, and the virtual node is a hyponym of all nodes in the corresponding node set; S3. quantify a first correlation between the first virtual node and all nodes in the first node set, wherein the correlation between nodes is a mutual inference probability between nodes; S4. quantifying a second correlation between the nodes in the first node set and the nodes in the second node set based on the graph features of the ontology and the information amounts of the nodes in the first node set and the second node set; S5. quantifying a third correlation between the nodes in the second node set and the second virtual node based on the amount of information shared between the nodes in the second node set; S6. Based on the transitivity of the inference probability, the first correlation, the second correlation and the third correlation, quantify the fourth correlation between the first virtual node and the second virtual node, that is, the correlation between the first compound semantics and the second compound semantics.

2. The ontology-based composite semantic relevance quantification method according to claim 1, characterized in that: Step S1 includes: S11, decomposing the first compound semantics and the second compound semantics into simple semantics that can be directly mapped to nodes, and before decomposing, the process also includes semantic fusion of uncommon words, wherein the simple semantics are nouns or adjectives; S12. Map the simple semantics corresponding to the first compound semantics and the second compound semantics to the first node set and the second node set of the semantic network respectively by calling the vocabulary node mapping method of the ontology.

3. The ontology-based composite semantic relevance quantification method according to claim 1, characterized in that: The relationship between a node set and a node is expressed as the relationship between a composite semantics and a sub-semantic. The existence of a composite semantics indicates that all sub-semantics exist simultaneously. A virtual node refers to a node that does not exist in the ontology but can correspond to the actual composite semantics. A virtual node is a hyponym of all nodes in a node set. The occurrence of a virtual node indicates the occurrence of all nodes in the corresponding node set. The correlation between two node sets can be converted into the correlation between two virtual nodes.

4. The ontology-based composite semantic relevance quantification method according to claim 1, characterized in that: The correlation between nodes is the mutual inference probability between nodes. The semantics represented can be used to infer the node The probability of the semantics represented is , then the node With Node The correlation is: = (1); Inference Probability The following two conditions should be met: ① Transitivity: Assume that there is a node , then: = (2); ②Probability additivity: Assume that there are nodes , then: = (3); The virtual node is a hyponym of all nodes in the node set. The occurrence of the virtual node indicates the occurrence of all nodes in the corresponding node set, that is, the correlation between the virtual node and all nodes in the node set is 1; in step S3, the first correlation between the first virtual node and all nodes in the first node set is 1.

5. The ontology-based composite semantic relevance quantification method according to claim 1, characterized in that: In step S4, the second correlation between the nodes in the first node set and the nodes in the second node set is quantified by the following method, that is, the correlation between the nodes in the ontology is quantified, and the correlation between the nodes is the probability that the starting node infers the end node through the relationship between the nodes: S41, perform graph traversal according to the starting node, set the longest path from the starting node to the end node, use a stack to record the current path, and set two sets to respectively record the nodes and node relationships of the current path; S42, after traversing to the end node under the specified path length, recording the nodes in the stack and the relationship between the nodes in the two sets, and constructing the sensitive value transfer subgraph of the current path after the traversal is completed; S43. Calculate the correlation between nodes in the sensitive value transfer subgraph of each path. When a node has a relationship with multiple nodes, the calculation formula for the correlation between nodes of each path is: (4); In formula (4), Represents a slave node To Node The path weight is the inference probability, that is, the correlation. Representation Node The set of all hyponyms of Represents the parent node The amount of information, The target node The amount of information of a node is determined based on the graph features of the ontology or based on a corpus; S44. If there is a single path between the starting node and the end node, the correlation between the starting node and the end node is the product of all path weights on the path, and the calculation formula is: (5); In formula (5), From the starting node To the end node The correlation of k represents the node To Node The longest path length between Represents a slave node To the adjacent node Path weight, i and k are positive integers; S45. If there are multiple paths between the start node and the end node in the same sensitive value transfer subgraph, the correlation between the start node and the end node is the sum of the correlations inferred from the multiple paths, and the correlation does not exceed 1. The calculation formula is: (6); In formula (6), From the starting node To the end node The correlation of k represents the node To Node The longest path length between them, q represents the starting node To the end node The number of paths that exist in the same sensitive value transfer subgraph, Represents a slave node To the adjacent node The path weight of , i, j, k, q are positive integers.

6. The ontology-based composite semantic relevance quantification method according to claim 1, characterized in that: In step S5, the third correlation between the nodes in the second node set and the second virtual node is quantified by the following method: When there are two nodes in the node set When the node With virtual nodes The correlation is two nodes The correlation between them is calculated as: (7); When the node set has at least three nodes, the node With virtual nodes The calculation formula of the correlation is: ; (8) In formulas (7) and (8), it is assumed that the event The corresponding composite semantics is , the simple semantics corresponding to events A and B are , composite semantics The corresponding node set is ,event The corresponding semantics is , semantics The corresponding node set is , represents the probability of event S occurring when event A occurs, Indicates that event A occurs The probability of occurrence, the degree of mutual influence of multiple nodes in a node set is related to the amount of information shared between the nodes, and the amount of information shared between the nodes is determined by the minimum common ancestor node of the nodes. express, The semantics represented is , k represents arrive The path length.

7. The ontology-based composite semantic relevance quantification method according to claim 6, characterized in that: In step S6, the correlation between the first composite semantics and the second composite semantics is the correlation between the first node set and the second node set, and is also the correlation between the first virtual node and the second virtual node, and is calculated as follows: ; In formula (9), represents the first composite semantics, represents the second composite semantics, represents the first node set, represents the second node set, represents the first virtual node, Represents the second virtual node.

8. An ontology-based composite semantic relevance quantification system, used to implement the method according to any one of claims 1 to 7, characterized in that: The system comprises: A semantic node mapping module, used for mapping the first composite semantics and the second composite semantics to the first node set and the second node set of the ontology in the semantic network after preprocessing, wherein the correlation between the first composite semantics and the second composite semantics is the correlation between the first node set and the second node set; A virtual node setting module, used to set a first virtual node and a second virtual node for the first composite semantics and the second composite semantics, respectively, wherein the correlation between the first composite semantics and the second composite semantics is the correlation between the first virtual node and the second virtual node, and the virtual node is a hyponym of all nodes in the corresponding node set; A first quantization module, configured to quantify a first correlation between the first virtual node and all nodes in the first node set, wherein the correlation between nodes is a mutual inference probability between the nodes; A second quantification module, configured to quantify a second correlation between nodes in the first node set and nodes in the second node set based on the graph features of the ontology and the information amounts of the nodes in the first node set and the second node set; A third quantization module, configured to quantify a third correlation between the nodes in the second node set and the second virtual node based on the amount of information shared between the nodes in the second node set; The fourth quantization module is used to quantify the fourth correlation between the first virtual node and the second virtual node, that is, the correlation between the first composite semantics and the second composite semantics, based on the transitivity of the inference probability, the first correlation, the second correlation and the third correlation.

9. The ontology-based composite semantic relevance quantification system according to claim 8, characterized in that: The semantic node mapping module is specifically used for: Decomposing the first compound semantics and the second compound semantics into simple semantics that can be directly mapped to nodes, and before decomposing, the method further includes semantic fusion of uncommon words, wherein the simple semantics are nouns or adjectives; The simple semantics corresponding to the first compound semantics and the second compound semantics are respectively mapped to the first node set and the second node set of the semantic network by calling the vocabulary node mapping method of the ontology.

10. The ontology-based composite semantic relevance quantification system according to claim 8, characterized in that: The relationship between a node set and a node is expressed as the relationship between a composite semantics and a sub-semantic. The existence of a composite semantics indicates that all sub-semantics exist simultaneously. A virtual node refers to a node that does not exist in the ontology but can correspond to the actual composite semantics. A virtual node is a hyponym of all nodes in a node set. The occurrence of a virtual node indicates the occurrence of all nodes in the corresponding node set. The correlation between two node sets can be converted into the correlation between two virtual nodes.

Citation Information

Patent Citations

  • User privacy data protection method and device based on semantic reasoning, electronic equipment and storage medium

    CN112580097A

  • Scoring dictionary construction method and system based on English vocabulary linguistic attribute prediction

    CN117874242A