Large-scale ontology matching method based on double-population adaptive evolution mechanism
Through the method of the adaptive evolution mechanism of the two populations, combined with multi-coding strategies and adaptive step size adjustment, the problem of blocking and similarity characteristics separation in large-scale ontology matching is solved, and efficient and high-quality ontology matching effect is achieved.
Patent Information
- Application Number
- CN202510337149.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-07-22
AI Technical Summary
The existing large-scale ontology matching method splits the correlation between ontology blocking and similarity feature construction, resulting in poor matching effect. The traditional blocking strategy weakens the integrity and consistency of ontology structure and increases the difficulty of similarity feature construction.
Using a method based on the adaptive evolution mechanism of the two-population population, the collaboration between population P1 and task-optimized population P2 is achieved through block construction, combining multi-coding strategies and adaptive step size adjustment to achieve dynamic integration and optimization of ontology blocking and similarity characteristics.
The efficiency and quality of large-scale ontology matching are significantly improved. Through knowledge interaction and adaptive step size adjustment, the diversity and convergence of the algorithm are dynamically balanced, and blocked results with reasonable structure, optimal scale and complete semantic information are generated.
Smart Images

Figure CN120354141A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of knowledge engineering, and particularly relates to a large-scale ontology matching method based on a dual-population adaptive evolution mechanism. Background Art
[0002] An ontology is a shared conceptual model used to describe the knowledge system of a specific domain. It defines the core concepts within that domain, the relationships between these concepts, and the attributes of each concept. Ontology matching refers to automatically or semi-automatically identifying the similarities or correspondences between different ontologies. This process aims to achieve effective data sharing, information exchange, and reasoning between ontologies from different sources. A large-scale ontology refers to an ontology with a wide coverage, containing a large number of concepts, classes, instances, and relationships in practical applications. With the continuous increase of domain knowledge and information volume, how to establish accurate corresponding relationships between large-scale ontologies has become a technical problem to be solved.
[0003] An effective method for dealing with large-scale ontology matching is to split the large-scale source ontology and target ontology into several sub-ontology modules through ontology partitioning, and then combine similarity feature matching for the corresponding sub-modules, transforming the overall task into a matching problem between multiple small-scale sub-modules. Existing large-scale ontology matching methods regard ontology chunking and similarity feature construction as two independent processes. First, they use a static chunking strategy based on simple rules to complete ontology partitioning, and then optimize similarity features on the basis of a fixed sub-ontology structure. However, multiple similarity feature construction schemes under the same chunking strategy may produce completely different matching results, and the same similarity feature construction scheme will also show significant quality differences due to changes in the chunking strategy. When ontology chunking and similarity feature construction are carried out independently, the combination of the optimal chunking strategy and the optimal similarity feature construction method does not necessarily achieve the optimization of the overall matching effect, breaking the association between the preprocessing method and the automation model in the large-scale ontology matching problem, resulting in difficulty in achieving the best matching effect. In addition, the chunking strategy based on simple rules weakens the integrity and consistency of the ontology structure, further increasing the difficulty of similarity feature construction and bringing additional obstacles to subsequent optimization. Summary of the Invention
[0004] The purpose of the present invention is to provide a large-scale ontology matching method based on a dual-population adaptive evolution mechanism to solve the deficiencies in the above background art.
[0005] To achieve the above object, the present invention adopts the following technical solutions:
[0006] A large-scale ontology matching method based on a dual-population adaptive evolution mechanism, the specific steps are as follows:
[0007] S1: A large-scale ontology matching model based on dual-population division of labor, including constructing population P1 by partitioning and task-optimizing population P2. Population P1 is responsible for obtaining the optimal sub-ontology partitioning construction information, and population P2 is responsible for optimizing the sub-ontology partitioning matching tasks;
[0008] S2: Dual-population individual representation based on multi-coding strategies. Population P1 integrates the core parameters of the partitions through a tree structure, involving newly defined function sets and terminal sets. The function sets include diffusion anchors, diffusion structures, diffusion depths, and diffusion scales, which are used to describe the diffusion characteristics of individuals. The terminal set is used to input candidate anchors, where the anchor candidate set is formally integrated and structurally independent; Population P2 uses the GER / ADF chromosome form for encoding and decoding;
[0009] S3: Matching knowledge interaction based on dual-population collaboration, including explicit interaction between populations and implicit interaction within populations. Explicit interaction is to feedback the matching results of population P2 to population P1 to assist population P1 in dynamically adjusting the partitioning construction parameters. Implicit interaction improves the matching efficiency and quality through knowledge transfer between multiple tasks within population P2;
[0010] S4: Partitioning parameter adjustment based on adaptive step size, which adjusts the selection probability of partitioning construction parameters through an adaptive step size, and dynamically balances the diversity of the population and the convergence of the algorithm during the optimization process.
[0011] Preferably, the specific process of the large-scale ontology matching model based on dual-population division of labor in step S1 is as follows:
[0012] (1) Configure the control parameters of populations P1 and P2, such as population size, number of iterations, crossover probability, and mutation probability;
[0013] (2) Randomly initialize the tree-structured individuals in population P1. The individuals divide the large-scale source ontology and target ontology according to grammar rules to generate corresponding partitioning results, calculate the partitioning similarity matrix under different similarity techniques, and use it as the input of population P2; Randomly initialize the chromosome individuals in population P2 and evaluate the matching quality of the sub-partitions;
[0014] (3) Execute evolution. Generate the next-generation individuals of populations P1 and P2 through selection, crossover, and mutation operators under different grammar rules, update the partitioning construction results of population P1 and the partitioning matching results of population P2, and evaluate the matching quality of the sub-partitions again;
[0015] (4) Adjust the selection probability of the partitioning construction parameters according to the elite solution and the adaptive step size, execute knowledge interaction, and generate the next-generation population according to the updated selection probability and the roulette wheel method;
[0016] (5) Return to step (3) until the termination condition is met, output the elite solution, and obtain a block strategy and optimal matching result for the source ontology and the target ontology.
[0017] Preferably, the specific contents of the dual-population individual characterization based on the multi-coding strategy in step S2 include:
[0018] Population P1 integrates the core parameters of building blocks into the syntax tree. It adopts a multi-level tree encoding structure and a bottom-up tree decoding method, including a function set and a terminal set:
[0019] (1) The function set is used to store the diffusion functions required in the block process. The specific contents are:
[0020] Diffusion anchors refer to key pairs randomly selected from a set of potential candidate points with high confidence values in context relevance and matching, which are used as the basis for promoting block expansion;
[0021] Diffusion structure refers to the structural paradigm adopted by a certain anchor point when constructing an extended block with a core concept, including three association modes: inheritance relationship, composition relationship and equivalence relationship;
[0022] Diffusion depth refers to the minimum number of steps required to propagate along the predetermined diffusion path from the diffusion anchor point to the entity to be diffused. It comprehensively considers the structural correlation and the influence of the common ancestor between the two, and includes two core dimensions: the structural depth and path length difference between the diffusion anchor point and the entity to be diffused relative to their common ancestor;
[0023] The diffusion scale refers to the maximum number of entities that can be accommodated in a given block;
[0024] (2) The terminal set is used to input candidate anchor points, and its specific contents are:
[0025] According to the number of contextual concepts, the entities in the source ontology and the target ontology are sorted respectively, and the entities with the highest ranking are selected from the largest to the smallest to form the anchor training set. Then the SMOA similarity between the source entity and the target entity in the anchor training set is calculated in turn, and the entity pairs with a similarity value of 1.0 are added to the candidate anchors;
[0026] The chromosome gene bits of population P2 are partitioned with "basic operation operator", "operation constant", "ADFs" and "ADFs virtual input value" as basic units.
[0027] Preferably, the anchor point candidate set in step S2 is fused in form and independent in structure, and the specific contents are:
[0028] Formally, fusion refers to all anchor points obtained under all structures as the only anchor point candidate set under the current ontology level;
[0029] Structurally independent means that the anchor candidate sets obtained under different structures do not affect each other, are independent of each other, and allow the same anchor to appear in the anchor candidate sets of different structures.
[0030] Preferably, the specific content of the step S3 based on the matching knowledge interaction of double-population collaboration includes:
[0031] Feedback the matching results of the tasks of population P2 to population P1 to induce population P1 to adjust the parameter information for constructing the sub-ontology blocks; according to the correlation between multiple tasks in population P2, divide population P2 into several task groups, and within each task group, establish a knowledge sharing mechanism among individuals. The specific steps are as follows:
[0032] Create G task groups; obtain the performance capabilities of each individual in the population on all subtasks and sort them according to the correlation; for a certain task group, obtain the ability vectors of each individual in the population within this task group and determine the corresponding task identification bit results; use the tournament selection mechanism to add M individuals to each task group, where:
[0033] T represents the total number of tasks.
[0034] Preferably, the performance ability of the individual in a certain subtask in the step S3 is represented by the approximate evaluation result of the individual on this task, denoted as F β ; the F β values of the individual on all tasks are sorted in descending order to form the ability vector of the individual, denoted as v i .
[0035] Preferably, the specific content of the step S4 based on the block parameter adjustment with an adaptive step size includes:
[0036] Addition and deletion operations of diffusion anchors: If the current individual performs better than the previous generation of elite solutions, then randomly select some anchors from the remaining anchors in the anchor candidate set according to the step size σ and supplement them to the current diffusion anchors; conversely, randomly delete some anchors from the current diffusion anchors according to the step size σ to reduce the number of diffusion anchors;
[0037] Selection probability of diffusion structures: If the current individual performs better than the previous generation of elite solutions, then increase the selection probability of the current diffusion structure by σ, and at the same time reduce the selection probabilities of the other two diffusion structures by σ / 2 respectively; conversely, reduce the selection probability of the current diffusion structure by σ, and at the same time increase the selection probabilities of the other two diffusion structures by σ / 2 respectively;
[0038] Selection probability of diffusion depth: If the current individual performs better than the elite solution of the previous generation, the selection probability of the current diffusion depth is increased by σ, while the selection probabilities of the remaining possible diffusion depths are each decreased by σ / d; conversely, the selection probability of the current diffusion depth is decreased by σ, while the selection probabilities of the remaining possible diffusion depths are each increased by σ / d; d represents the maximum depth allowed for diffusion of the current anchor point.
[0039] Ratio adjustment of diffusion scale: If the current individual performs better than the elite solution of the previous generation, the current block size m is adjusted to (1 + σ)m; conversely, the current block size m is adjusted to (1 - σ)m.
[0040] Preferably, in step S4, the adaptive step size σ changes with the replacement of the elite solution of the population, and the specific content includes:
[0041] In the initial or middle stage of population evolution, each time an elite solution is replaced, the step size parameter σ is decreased by 0.01; if there is no replacement of the elite solution for 10 consecutive generations, the step size parameter σ is increased by 0.01.
[0042] In the late stage of population evolution, no incremental adjustment is made to σ until the termination condition is met.
[0043] Preferably, the change basis of the adaptive step size σ in step S4 also includes attenuation with the increase of the population evolution generation. Define the first 70% of the maximum population evolution generation as the initial and middle stages of evolution; when the evolution generation exceeds 70% of the maximum population evolution generation, it is considered to enter the late stage of evolution.
[0044] Compared with the prior art, the beneficial effects achieved by the present invention are:
[0045] (1) Adopting a dual-population mechanism, integrating ontology partitioning and similarity feature construction into an optimization model. In each iteration, the ontology partitioning method will be deduced and updated, and similarity features will be dynamically constructed based on the current partitioning results, thus ensuring a tight association between the two. In addition, the ontology partitioning and similarity feature construction are not a simple one-way information transfer process, but rather promote each other through knowledge interaction. The construction result of the similarity feature provides a basis for parameter selection for the next-generation ontology partitioning strategy and adjusts the search direction accordingly; while the high-quality ontology partitioning result helps to enhance the potential of similarity feature construction. This mechanism also considers multi-task collaboration within a single population, significantly improving the efficiency and quality of large-scale ontology matching.
[0046] (2) In the process of knowledge interaction, the step size information is not fixed, but is adaptively adjusted according to the improvement of the elite solution and the increase of the evolution generation. The adaptive adjustment avoids the limitations of subjectively setting parameters, effectively balancing the diversity of genetic information and the leading role of elite information, and optimizing the global search ability and convergence effect of the algorithm.
[0047] (3) Integrate and partition parameters using a syntax tree structure. Driven by the ontology matching quality, transform the traditional manually enumerated partition rules combination into an optimization problem. Embed this optimization process into the automated model to achieve fully dynamic ontology partition construction; this method can generate partition results with reasonable structures, optimal scales, and complete semantic information. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1 is a flowchart of an embodiment of the present invention;
[0049] Figure 2 is a flowchart of population iteration optimization of an embodiment of the present invention;
[0050] Figure 3 is a schematic diagram of the coding of partition individuals in an embodiment of the present invention;
[0051] Figure 4 is a schematic diagram of anchor point selection in an embodiment of the present invention;
[0052] Figure 5 is a schematic diagram of the chromosome structure of matching individuals in an embodiment of the present invention;
[0053] Figure 6 is a schematic diagram of double-population knowledge transfer in an embodiment of the present invention;
[0054] Figure 7 is a comparison schematic diagram between an embodiment of the present invention and other ontology partitioning methods in terms of average partition scale and the number of unassociated entities. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0055] The technical solutions of the present invention will be clearly and completely described below with reference to the drawings and embodiments.
[0056] The ontology matching data set used in this embodiment is six large-scale publicly available test data sets provided by OAEI official. The relevant descriptions are shown in Table 1.
[0057] Table 1 Descriptions of OAEI Anatomy and Large Biomedical Ontologies
[0058]
[0059]
[0060] As Figure 1 shown, a large-scale ontology matching method based on a double-population adaptive evolution mechanism is as follows:
[0061] S1: A large-scale ontology matching model based on double-population division of labor, including the construction of population P1 for sub-ontology partitioning and population P2 for task optimization. Population P1 is responsible for obtaining the optimal sub-ontology partitioning information, and population P2 is responsible for optimizing the sub-ontology partitioning matching tasks. The population iterative optimization process is as follows Figure 2 shown, and the specific content includes:
[0062] (1) Configure the control parameters of populations P1 and P2. In this embodiment, the population size = 100, the number of iterations = 2000, the crossover probability = 0.8, and the mutation probability = 0.1 are set;
[0063] (2) Randomly initialize the tree-structured individuals in population P1. The individuals divide the large-scale source ontology and target ontology according to the grammar rules, generate the corresponding partitioning results, calculate the partitioning similarity matrix under different similarity technologies, and use it as the input of population P2; randomly initialize the chromosome individuals in population P2 and evaluate the matching quality of the sub-partitions. The evaluation formula for the matching quality is:
[0064]
[0065] where, |O 1,map |, |O 2,map | respectively represent the number of correctly matched entities in the source ontology O1 and the target ontology O2; |O1|, |O2| respectively represent the total number of entities in the source ontology O1 and the target ontology O2; represents the number of matching pairs found by the algorithm; Recall p is the recall rate index, used to evaluate the integrity of the matching results; Precision p is the precision rate index, which integrates the search precision and the one-to-one matching principle, and is used to evaluate the accuracy of the matching results; Fβ is a comprehensive evaluation index that comprehensively considers the recall rate and the precision rate;
[0066] (3) Execute evolution. Generate the next-generation individuals of populations P1 and P2 through selection, crossover, and mutation operators under different grammar rules, update the partitioning construction results of population P1 and the partitioning matching results of population P2, and evaluate the matching quality of the sub-partitions again;
[0067] (4) Adjust the selection probability of the partitioning construction parameters according to the elite solution and the adaptive step size, and execute knowledge interaction. Population P1 uses the roulette wheel method to select half of the individuals from the population after crossover and mutation, which is dominated by the evolutionary strategy; the other half of the individuals are generated according to the updated selection probability and are dominated by the selection probability. The two together constitute the next-generation population; population P2 uses the roulette wheel method to select the next-generation population;
[0068] (5) Return to step (3) until the termination condition is met, output the elite solution, and obtain a partitioning strategy and the optimal matching result of the source ontology and the target ontology.
[0069] S2: Dual-population individual representation based on multi-coding strategy. Population P1 integrates the core parameters of the divided blocks through a tree structure, involving newly defined function sets and terminal sets. The function set includes diffusion anchor points, diffusion structures, diffusion depths, and diffusion scales, which are used to describe the diffusion characteristics of individuals. The terminal set is used to input candidate anchor points. Population P2 uses the GER / ADF chromosome form for encoding and decoding. The specific content includes:
[0070] Population P1 integrates the core parameters of the divided blocks into the syntax tree. As Figure 3 shown, it adopts a multi-level tree-like coding structure and a bottom-up tree-like decoding method, including function sets and terminal sets:
[0071] (1) The function set is used to store the diffusion functions required during the division process. The specific content is:
[0072] The diffusion anchor point refers to a key pair randomly selected from a set of potential candidate points, which has a high confidence value in terms of context relevance and matching degree, and is used as the basis for promoting the expansion of the divided block.
[0073] The diffusion structure refers to the structural paradigm adopted by an anchor point when constructing an extended divided block with a certain anchor point as the core concept, including three association modes: inheritance relationship, composition relationship, and equivalence relationship.
[0074] The diffusion depth refers to the minimum number of steps required to propagate along a predetermined diffusion path from the diffusion anchor point to the entity to be diffused, taking into account the structural association between the two and the influence of the common ancestor, including two core dimensions: the structural depth and the path length difference between the diffusion anchor point and the entity to be diffused relative to their common ancestor.
[0075] The diffusion scale refers to the maximum number of entities that can be accommodated within a given divided block.
[0076] In this embodiment, the information parameters of different diffusion functions are initialized as follows:
[0077] Diffusion anchor point: The anchor point obtained from each branch under the ontology structure is sorted according to its in-degree and out-degree, and the first one is taken.
[0078] Diffusion structure probability: P 继承 = 1 / 3, P 组成 = 1 / 3, P 等价 = 1 / 3. The inheritance relationship, composition relationship, and equivalence relationship represent three structures of ontology diffusion.
[0079] Diffusion depth: 50% of the maximum diffusion depth that can be reached in the diffusion structure constructed by the selected diffusion anchor point.
[0080] Diffusion scale: |O1| represents the number of entities in the source ontology, |O2| represents the number of entities in the target ontology, and k represents the number of chunks.
[0081] (2) The terminal set is used to input candidate anchors, and the selection process of the anchors is as Figure 4 shown, and the specific content is:
[0082] According to the number of context concepts, the entities in the source ontology and the target ontology are sorted respectively, and the top 70% of the entities ranked from large to small are selected to form the anchor training set. Then, the SMOA similarity between the source entity and the target entity in the anchor training set is calculated in turn, and the entity pairs with a similarity value of 1.0 are added to the candidate anchors; the similarity measurement method based on SMOA is:
[0083] SMOA(e s ,e t ) = Com(e s ,e t ) - Dif(e s ,e t ) + Jaro(e s ,e t )
[0084] Among them, Com(e s ,e t ) is used to evaluate the similarity between entity e s and e t , Dif(e s ,e t ) is used to evaluate the dissimilarity between entity e s and e t , and Jaro(e s ,e t ) is used as a correction factor for the similarity evaluation between two entities.
[0085] The obtained candidate anchor set is formally integrated, that is, all the anchors obtained under all structures are used as the only candidate anchor set at the current ontology level; it is structurally independent, that is, the candidate anchor sets obtained under different structures do not affect each other, are independent of each other, and allow the same anchor to appear in the candidate anchor sets of different structures.
[0086] The chromosome gene positions of population P2 are partitioned with "basic operation operators", "operation constants", "ADFs", and "ADFs virtual input values" as basic units, as Figure 5As shown, the individual code is in the form of a chromosome, and its overall structure includes two parts: the main chromosome and the ADFs sub-chromosome. Each part consists of a head and a tail. The head provides the core coding information of the chromosome, such as the description of the problem, constraints, and objective function. The tail is used to represent variable values or decision variables of the optimization problem. The selection of each gene locus strictly follows the partitioning principle. The value range of the "basic operation operator" is [0, A], including the operators +, -, *, / , sin, cos, and ave, where ave represents the operation of taking the average value; the value range of "ADFs" is [A, B - 1], which is used to represent a group of genes or gene fragments to construct sub-functions or sub-structures, and its quantity can be set according to the specific problem requirements. In this embodiment, to ensure the effectiveness of the algorithm, only one "ADFs" is set, denoted as G1; the value range of the "operation constant" is [B, C - 1], which is used to represent the constants involved in the optimization operation process. In this embodiment, this operation constant refers to the similarity matrix of sub-ontology blocks obtained by different similarity measurement techniques; the value range of the "ADFs virtual input value" is [C, D - 1], which is used to represent the virtual input quantity required for the calculation of the sub-function ADFs. Based on the individual decoding method, a binary tree structure is adopted for decoding in this embodiment, so the ADFs virtual input quantity is set to two bits, denoted as a and b respectively. The values of A, B, C, and D are determined by the following formula:
[0087]
[0088] B = A + N a
[0089]
[0090] D = C + N g
[0091] Among them, F i represents the number of basic operation operators; N a represents the number of ADFs. In this embodiment, N a = 1; T i represents the number of operation constants; N g represents the number of ADFs virtual input quantities;
[0092] The length constraints of the chromosome head and tail follow the following formula:
[0093]
[0094] Among them, h represents the length of the head of the main chromosome; l represents the length of the tail of the main chromosome; ξ(a) represents the maximum number of received input parameters under all chromosome gene locus nodes. In this embodiment, ξ(a)=2;
[0095] The encoding length of an individual chromosome follows the following formula:
[0096] n = h + l + N a *(h' + l')
[0097] where n represents the encoding length of the individual chromosome; h and l respectively represent the head and tail lengths of the main chromosome. In this embodiment, h = 4 and l = 5; l' respectively represents the head and tail lengths of the ADFs sub-chromosome. In this embodiment, l' = 4;
[0098] After determining the individual chromosome length n and the chromosome gene partition range [0, D], n constants are randomly generated within this interval. According to the preset chromosome gene partition rule, each constant is mapped to the corresponding gene symbol to complete the chromosome encoding process of the individual. After completing the chromosome encoding, a decoding operation of the binary tree structure is performed on it. The decoding process is executed sequentially from top to bottom and from left to right according to the gene position order of the individual chromosome encoding.
[0099] S3: Matching knowledge interaction based on dual-population collaboration, the process is as Figure 6 shown, including explicit interaction between populations and implicit interaction within populations. Explicit interaction is to feedback the matching results of population P2 to population P1 to assist population P1 in dynamically adjusting the block construction parameters. Implicit interaction improves the matching efficiency and quality through knowledge transfer among multiple tasks within population P2. The specific content includes:
[0100] Feedback the matching results of the tasks of population P2 to population P1 to induce population P1 to adjust the parameter information for constructing sub-ontology blocks; according to the correlation between multiple tasks in population P2, population P2 is divided into several task groups. Within each task group, a knowledge sharing mechanism is established among individuals. The specific steps are as follows:
[0101] Create G task groups, obtain the ability vectors of each individual in population P2 on all subtasks, and sort them according to the correlation. The ability vector refers to the performance ability of each individual in the population on all tasks, and is sorted according to the correlation. Among them, the performance ability of an individual on a specific task is represented by its approximate evaluation result for this task, denoted as F β ; the F β values of each individual on all tasks are sorted in descending order to form the ability vector of this individual, denoted as v i , v i is determined by the following formula:
[0102]
[0103] where, represents the i-th individual r in the populationi After the performance on all tasks is sorted in descending order, the ranking of its performance on the j-th task, where k represents the number of sub-ontology matching tasks;
[0104] For a certain task group, obtain the task identification bit results corresponding to the ability vectors of each individual in the population on this task group, and add M individuals to each task group through tournament selection. In this embodiment:
[0105] T represents the number of sub-ontology matching tasks, and G = 10.
[0106] S4: Based on the adaptive step-size adjustment of the chunk parameters, adjust the selection probability of the chunk construction parameters through the adaptive step-size, and dynamically balance the diversity of the population and the convergence of the algorithm during the optimization process. The specific content includes:
[0107] Addition and deletion operations of diffusion anchor points: If the current individual performs better than the elite solution of the previous generation, then randomly select some anchor points from the remaining anchor points in the anchor point candidate set according to the step size σ and supplement them to the current diffusion anchor points; otherwise, randomly delete some anchor points from the current diffusion anchor points according to the step size σ to reduce the number of diffusion anchor points;
[0108] Selection probability of diffusion structure: If the current individual performs better than the elite solution of the previous generation, then increase the selection probability of the current diffusion structure by σ, and at the same time reduce the selection probabilities of the other two diffusion structures by σ / 2 respectively; otherwise, reduce the selection probability of the current diffusion structure by σ, and at the same time increase the selection probabilities of the other two diffusion structures by σ / 2 respectively;
[0109] Selection probability of diffusion depth: If the current individual performs better than the elite solution of the previous generation, then increase the selection probability of the current diffusion depth by σ, and at the same time reduce the selection probabilities of the other possible diffusion depths by σ / d respectively; otherwise, reduce the selection probability of the current diffusion depth by σ, and at the same time increase the selection probabilities of the other possible diffusion depths by σ / d respectively; d represents the maximum depth allowed for diffusion of the current anchor point;
[0110] Ratio adjustment of diffusion scale: If the current individual performs better than the elite solution of the previous generation, then adjust the current chunk scale m to (1 + σ)m; otherwise, adjust the current chunk scale m to (1 - σ)m.
[0111] In the initial or middle stage of population evolution, that is, in the first 70% of the maximum number of evolution generations, every time an elite solution replacement occurs, the step size parameter σ is reduced by 0.01; if there is no elite solution replacement for 10 consecutive generations, the step size parameter σ is increased by 0.01;
[0112] When the number of evolution generations exceeds 70% of the maximum number of evolution generations of the population, no incremental adjustment is made to σ until the termination condition is met. In this embodiment, σ = 0.05 is initialized.
[0113] As Figure 7 shown, compared with the other three typical ontology partitioning methods, in this embodiment, by adopting the block parameter integration strategy based on genetic programming and combining with the adaptive evolution mechanism under the interaction of double populations, the average block size is effectively reduced, and further the spatio-temporal complexity of the large-scale ontology matching problem is reduced. In addition, in the partitioning results obtained by the present invention, the number of unassociated entities is significantly reduced, effectively protecting the integrity of the ontology semantic information and having the potential to improve the subsequent ontology matching quality.
[0114] In this embodiment, the matching performance of the large-scale ontology matching method (BSAEM) based on the double-population adaptive evolution mechanism is compared with multiple mainstream matching systems on OAEI, including AML, LogMap, LogMapBio, ATMatch, ATBox, and Lily. The experimental data of BSAEM are the average values of the matching results obtained by the algorithm running independently 30 times.
[0115] Table 2 Comparison of f-measure between BSAEM and OAEI participants on all test cases
[0116] Dataset AML LogMap LogMapBio ATMatch ATBox Lily BSAEM MA-HA 0.94(+) 0.88(+) 0.89(+) 0.79(+) 0.8(+) 0.90(+) 0.96(0.01) FMA-NCI 0.93(+) 0.92(+) 0.91(+) 0.87(+) 0.87(+) 0.66(+) 0.96(0.03) FMA-SNOMED 0.84(+) 0.80(+) 0.80(+) 0.77(+) 0.78(+) NaN(+) 0.87(0.03) SNOMED-NCI 0.82(+) 0.79(+) 0.79(+) NaN(+) 0.67(+) NaN(+) 0.86(0.02) HP-MP 0.80(+) 0.82(+) 0.81(+) 0.45(+) 0.46(+) 0.66(+) 0.92(0.02) DOID-ORDO 0.76(+) 0.71(+) 0.73(+) NaN(+) 0.50(+) 0.67(+) 0.87(0.03) + / = / - 6 / 0 / 0 6 / 0 / 0 6 / 0 / 0 6 / 0 / 0 6 / 0 / 0 6 / 0 / 0
[0117] As shown in Table 2, the matching quality of BSAEM of the present invention is better than that of all OAEI participants in six large-scale test cases. For example, in the FMA-NCI test case, BSAEM performs excellently with a score of 0.96, significantly exceeding other technologies such as AML (0.93) and LogMap (0.92). The results of the HP-MP test case also highlight the advantages of BSAEM, with an F value of 0.92. This comparison result is also applicable to other test cases. In addition, the standard deviation of BSAEM is at most 0.03, indicating its high stability. These data further show that the present invention has broad application prospects and excellent practical performance when solving the large-scale ontology matching problem.
[0118] The above are only the preferred embodiments of the present invention. All equivalent replacements made according to the content of this specification and the drawings, or direct or indirect applications in the relevant technical fields, shall fall within the scope of patent protection of the present invention.
Claims
1. A large-scale ontology matching method based on a dual-population adaptive evolution mechanism, characterized in that The specific steps are as follows: S1: A large-scale ontology matching model based on the division of labor of two populations, including the block construction population P1 and the task optimization population P2. The population P1 is responsible for obtaining the optimal sub-ontology block construction information, and the population P2 is responsible for optimizing the sub-ontology block matching task; S2: The individual representation of the two populations based on multiple coding strategies. The population P1 integrates the core parameters of the blocks through a tree structure, involving newly defined function sets and terminal sets. The function set includes diffusion anchors, diffusion structures, diffusion depths, and diffusion scales, which are used to describe the diffusion characteristics of individuals. The terminal set is used to input candidate anchors, where the anchor candidate set is formally integrated and structurally independent; The population P2 uses the GER / ADF chromosome form for encoding and decoding; S3: The matching knowledge interaction based on the cooperation of the two populations, including explicit interaction between populations and implicit interaction within populations. The explicit interaction is to feedback the matching results of the population P2 to the population P1 to assist the population P1 in dynamically adjusting the block construction parameters. The implicit interaction improves the matching efficiency and quality through knowledge transfer between multiple tasks within the population P2; S4: The adjustment of block parameters based on an adaptive step size, which adjusts the selection probability of block construction parameters through an adaptive step size, and dynamically balances the diversity of the population and the convergence of the algorithm during the optimization process.
2. A large-scale ontology matching method based on a dual-population adaptive evolution mechanism according to claim 1, characterized in that The specific process of the large-scale ontology matching model based on the division of labor of two populations in step S1 is as follows: (1) Configure the control parameters of the populations P1 and P2, such as population size, number of iterations, crossover probability, and mutation probability; (2) Randomly initialize the tree-structured individuals in the population P1. The individuals divide the large-scale source ontology and target ontology according to grammar rules, generate corresponding block results, calculate the block similarity matrix under different similarity technologies, and use it as the input of the population P2; Randomly initialize the chromosome individuals in the population P2 and evaluate the matching quality of the sub-blocks; (3) Execute evolution, generate the next-generation individuals of the populations P1 and P2 through selection, crossover, and mutation operators under different grammar rules, update the block construction results of the population P1 and the block matching results of the population P2, and evaluate the matching quality of the sub-blocks again; (4) Adjust the selection probability of the block construction parameters according to the elite solution and the adaptive step size, execute knowledge interaction, and generate the next-generation population according to the updated selection probability and the roulette method; (5) Return to step (3) until the termination condition is met, output the elite solution, and obtain a block strategy and the optimal matching result of a source ontology and a target ontology.
3. A large-scale ontology matching method based on a dual-population adaptive evolution mechanism according to claim 1, characterized in that, The specific content of the individual representation of the two populations based on multiple coding strategies in step S2 includes: The population P1 integrates the core parameters of the constructed blocks in the syntax tree. It adopts a multi-level tree-like coding structure and a bottom-up tree-like decoding method, including a function set and a terminal set: (1) The function set is used to store the diffusion functions required during the block process. The specific content is: The diffusion anchor refers to a key pair randomly selected from a group of potential candidate points, which has a high confidence value in terms of context relevance and matching degree, and is used as the basis for promoting block expansion; Diffusion structure refers to the structural paradigm adopted by a certain anchor point when constructing an extended block with a core concept, including three association modes: inheritance relationship, composition relationship and equivalence relationship; Diffusion depth refers to the minimum number of steps required to propagate along the predetermined diffusion path from the diffusion anchor point to the entity to be diffused. It comprehensively considers the structural correlation and the influence of the common ancestor between the two, and includes two core dimensions: the structural depth and path length difference between the diffusion anchor point and the entity to be diffused relative to their common ancestor; The diffusion scale refers to the maximum number of entities that can be accommodated in a given block; (2) The terminal set is used to input candidate anchor points, and its specific contents are: According to the number of contextual concepts, the entities in the source ontology and the target ontology are sorted respectively, and the entities with the highest ranking are selected from the largest to the smallest to form the anchor training set. Then the SMOA similarity between the source entity and the target entity in the anchor training set is calculated in turn, and the entity pairs with a similarity value of 1.0 are added to the candidate anchors; The chromosome gene bits of population P2 are partitioned with "basic operation operator", "operation constant", "ADFs" and "ADFs virtual input value" as basic units.
4. A large-scale ontology matching method based on a dual-population adaptive evolution mechanism according to any one of claims 1 or 3, characterized in that The anchor point candidate set in step S2 is fused in form and independent in structure, and the specific contents are as follows: Formally, fusion refers to all anchor points obtained under all structures as the only anchor point candidate set under the current ontology level; Structurally independent means that the anchor candidate sets obtained under different structures do not affect each other and are independent of each other, and the same anchor point is allowed to appear in the anchor candidate sets of different structures.
5. A large-scale ontology matching method based on a dual-population adaptive evolution mechanism according to claim 1, characterized in that The specific contents of the matching knowledge interaction based on the dual population collaboration in step S3 include: Feedback the matching results of population P2 tasks to population P1 to induce population P1 to adjust the parameter information for constructing sub-ontology blocks; divide population P2 into several task groups according to the correlation between multiple tasks in population P2, and establish a knowledge sharing mechanism between individuals in each task group. The specific steps are as follows: Create G task groups; obtain the performance of each individual in the population on all subtasks and sort them by relevance; for a certain task group, obtain the capability vector of each individual in the population in the task group and determine the corresponding task identification result; use the competition selection mechanism to add M individuals to each task group, where: , where T represents the total number of tasks.
6. A large-scale ontology matching method based on a dual-population adaptive evolution mechanism according to claim 5, characterized in that, The performance ability of the individual in step S3 on a certain subtask is represented by the approximate evaluation result of the individual on this task, denoted as ; the values of the individual on all tasks are sorted in descending order to form the ability vector of the individual, denoted as .
7. A large-scale ontology matching method based on a dual-population adaptive evolution mechanism according to any one of claims 1 or 5, characterized in that The specific contents of the block parameter adjustment based on the adaptive step size in step S4 include: Addition and deletion operations of diffusion anchor points: If the current individual performance is better than the previous generation of elite solutions, some anchor points are randomly selected from the remaining anchor points in the anchor point candidate set according to the step size σ and added to the current diffusion anchor points; otherwise, some anchor points are randomly deleted from the current diffusion anchor points according to the step size σ to reduce the number of diffusion anchor points; Probability of selection of diffusion structure: If the current individual performance is better than the previous generation of elite solution, the selection probability of the current diffusion structure is increased by σ, and the selection probabilities of the other two diffusion structures are reduced by σ / 2 respectively; conversely, the selection probability of the current diffusion structure is reduced by σ, and the selection probabilities of the other two diffusion structures are increased by σ / 2 respectively; Selection probability of diffusion depth: If the current individual performs better than the elite solution of the previous generation, increase the selection probability of the current diffusion depth by σ, and at the same time decrease the selection probabilities of the remaining possible diffusion depths by σ / d respectively; otherwise, decrease the selection probability of the current diffusion depth by σ, and at the same time increase the selection probabilities of the remaining possible diffusion depths by σ / d respectively; d represents the maximum depth allowed for diffusion of the current anchor point. Proportion adjustment of diffusion scale: If the current individual performs better than the elite solution of the previous generation, adjust the current block size m to (1 + σ)m; otherwise, adjust the current block size m to (1 - σ)m.
8. A large-scale ontology matching method based on a dual-population adaptive evolution mechanism according to claim 7, characterized in that In step S4, the adaptive step size σ changes with the replacement of the elite solution of the population, and the specific content includes: In the initial or middle stage of population evolution, every time an elite solution is replaced, the step size parameter σ is decreased by 0.01; if there is no replacement of the elite solution for 10 consecutive generations, the step size parameter σ is increased by 0.
01. In the later stage of population evolution, no incremental adjustment is made to σ until the termination condition is met.
9. A large-scale ontology matching method based on a dual-population adaptive evolution mechanism according to claim 7, characterized in that The change basis of the adaptive step size σ in step S4 also includes attenuation with the increase of the population evolution generation. Define the first 70% of the maximum population evolution generation as the initial and middle stages of evolution; when the evolution generation exceeds 70% of the maximum population evolution generation, it is considered to enter the later stage of evolution.