A method for recommending missing information of requirements based on a domain model

CN115220695BActive Publication Date: 2026-09-25BEIHANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210687400.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-17
Publication Date
2026-09-25
Estimated Expiration
2042-06-17

AI Technical Summary

Technical Problem

[0003]本发明提供了一种基于领域模型的需求缺失信息的推荐方法,以解决现有技术中不能很好地进行需求缺失信息推荐的问题

Benefits of technology

[0030]本发明提出了一种自动构建需求和领域模型之间映射的方法,通过从需求中提取概念术语,并通过知识模型、需求模型以及对齐模型来将所提取的需求补充到原始的领域模型中,得到补全后的领域模型,实现将所提取的概念术语映射在补全后的领域模型,然就基于所述领域模型以及补全后的领域模型中出现的规律进行需求缺失信息的推荐,从而大大提高领域模型对需求的覆盖率,继而有效解决了现有不能很好地进行需求缺失信息推荐的问题。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115220695B_ABST
    Figure CN115220695B_ABST
Patent Text Reader

Abstract

The application discloses a requirement missing information recommendation method based on a domain model, extracts concept terms from requirements, and supplements the extracted requirements to an original domain model through a knowledge model, a requirement model and an alignment model to obtain a supplemented domain model, realizes mapping of the extracted concept terms on the supplemented domain model, and recommends requirement missing information based on rules in the domain model and the supplemented domain model, so that the coverage of the domain model on requirements is greatly improved, and the problem that existing requirement missing information recommendation cannot be well performed is effectively solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer technology, and in particular to a recommendation method based on missing requirement information from a domain model. Background Technology

[0002] Due to the limited domain knowledge of requirements analysts, short-term company plans, and the nature of changing requirements, it is very common for software requirements specifications to lack some basic or even critical functions. However, the functions and capabilities that a software product can provide constitute its core competitiveness, and the lack of key functions severely reduces the reputation and market share of a software product. Therefore, for organizations, identifying their missing functions is crucial, as it helps them to make sound and effective improvements. However, much existing research focuses on missing function detection based on requirements-based domain models, and for many domains, there are no domain models oriented towards software requirements. Therefore, how to implement a method for recommending missing requirements information has become an urgent problem to be solved. Summary of the Invention

[0003] This invention provides a recommendation method based on missing demand information from a domain model, in order to solve the problem that existing technologies cannot effectively perform recommendations based on missing demand information.

[0004] This invention provides a recommendation method based on missing requirement information from a domain model. The method includes: extracting terminology from preset requirements and constructing a mapping between the extracted terminology and entities in a domain model, wherein the domain model is a set of <entity, relation, entity> triples in an ontology or knowledge graph; embedding the domain model into a knowledge model and embedding the preset requirements into a requirement model, wherein both the knowledge model and the requirement model are vector space models; aligning the knowledge model and the requirement model using an alignment model, and supplementing the domain model with the requirement model through relation derivation. The missing relationships or entities in the model are identified to obtain a completed domain model. The patterns of the extracted term mappings appearing in the original and completed domain models are analyzed, and recommendations for missing requirement information are made based on these patterns. These patterns include the abstraction levels in the original and completed domain models, the distribution of mapped entities in the original and completed domain models, the types of entities in the original and completed domain models, the distribution of mapped entities in the original and completed domain models, the clusters to which the mapped entities belong, and the parent-child relationships between nodes in the original and completed domain models.

[0005] Optionally, the step of extracting terminology from preset requirements and constructing a mapping between terminology and entities in the domain model based on the extracted terminology includes: randomly selecting a preset number of requirements, denoted as R, from a preset set of requirements; extracting terminology from the selected requirements according to a preset text terminology extraction method; and representing the extracted terminology as RT; and constructing a mapping relationship between terminology and entities in the domain model based on the extracted terminology, denoted as S.

[0006] Optionally, the preset text term extraction method includes: a rule-based extraction method, a statistical extraction method, and a C-Value method, wherein the C-Value method is a term extraction method that combines rule extraction and statistical extraction.

[0007] Optionally, the step of extracting the terms of the requirements from the selected requirements according to a preset text term extraction method includes: selecting multi-word terms (MWTs) from the corpus using the C-Value method, that is, first obtaining all candidate words based on a series of language filters for nested noun selection and a stop word filter, and then assigning a term metric to the strings of all candidate words based on the total frequency of the candidate words, the frequency of the candidate words as part of other candidate terms, the number of candidate terms exceeding a preset length threshold, and the length of the candidate words, and selecting strings with a selection period higher than a predefined threshold as the terms of the requirements.

[0008] Optionally, the step of constructing the mapping between terms and entities in the domain model based on the extracted terms includes: constructing the mapping between terms and entities in the domain model based on the names of the extracted terms and the names of entities in the domain model, and when constructing the mapping between terms and entities in the domain model, de-synonyming processing is performed on terms and entities in the domain model that have synonyms.

[0009] Optionally, the de-synonymization process for terms and entities in the domain model that have synonymous relationships includes: calculating similarity based on word embeddings, detecting synonymous relationships between MWTs that have the same head or tail components, and removing terms and entities in the domain model that have synonymous relationships; wherein, the word embedding similarity calculation includes: based on the extracted multi-word term MWTs, representing terms, requirements, and domain models in the domain corpus with embedding vectors to calculate the similarity between terms, requirements, and domain models.

[0010] Optionally, if the term length between the head and the tail exceeds a preset length threshold, the method further includes: adding an M element between the head and the tail for expansion, and setting the multi-word term MWT = (E; M; T), where E is the head of the MWT, M is the middle part of the MWT, T is the tail of the MWT, and synonyms between the MWTs are extracted according to a basic preset rule.

[0011] Let two terms be MWT1 = (E1; M1; T1) and MWT2 = (E2; M2; T2), and let syn(MWT1, MWT2) be a synonym relationship between the two terms, expressed by the formula:

[0012]

[0013]

[0014]

[0015]

[0016] When two parts of two MWTs are equal, and one of them can be empty, the remaining part is a synonym, meaning the two terms are synonyms.

[0017] Optionally, the domain model is modeled as a triple (h, r, t), where h, t ∈ ε entity set and r ∈ r relation set. The entire fact triple in the domain model is labeled Δ, and the conditional probability of the triple (h, r, t) is defined as: in b is the bias constant. The knowledge model provides standardized definitions for Pr(r|h,t) and Pr(t|h,r):

[0018]

[0019] The step of embedding the domain model into the knowledge model includes: maximizing the conditional likelihood of triples existing in the domain model using the knowledge model.

[0020] The step of embedding the preset requirements into the requirement model includes: Requirement model: (w, r wv The term set in the requirements includes the multi-word term MWT extracted from the corpus, as well as nouns that appear as the only semantic unit in any requirement or window;

[0021] The probability that term w and term set label v appear simultaneously in the text window is defined as follows:

[0022]

[0023] Its loss function is:

[0024] Optionally, aligning the knowledge model and the requirement model using an alignment model includes: aligning based on identical expressions or synonyms between entities in the domain model and terms in the requirements, i.e., for (h, r, t) in Δ, if the entity name h is equal to or w is equal to in v. h Synonyms of, then generate a new triple (w h (r, t), while if entities t and w t The name and w in v t If they are the same or synonymous, then (h, r, w) is generated. t ) and (w h ,r,w t );

[0025] The loss function of the alignment model is:

[0026]

[0027] The likelihood function of the joint embedding learning model is defined as follows:

[0028] Optionally, the types of entities in the domain model are resource groups (Classes), attributes (ObjectProperties), data attributes (Data Properties), class instances (Named Individuals), and metadata (Annotation Properties).

[0029] The beneficial effects of this invention are as follows:

[0030] This invention proposes a method for automatically constructing a mapping between requirements and domain models. By extracting conceptual terms from requirements and supplementing the original domain model with the extracted requirements through a knowledge model, a requirement model, and an alignment model, a complete domain model is obtained. This achieves the mapping of the extracted conceptual terms onto the complete domain model. Then, based on the patterns appearing in the domain model and the complete domain model, missing requirement information is recommended, thereby greatly improving the coverage of requirements by the domain model and effectively solving the problem that existing methods cannot effectively recommend missing requirement information.

[0031] The above description is merely an overview of the technical solution of the present invention. In order to better understand the technical means of the present invention and to implement it in accordance with the contents of the specification, and in order to make the above and other objects, features and advantages of the present invention more apparent and understandable, specific embodiments of the present invention are described below. Attached Figure Description

[0032] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings:

[0033] Figure 1 This is a flowchart illustrating a recommendation method based on missing demand information from a domain model, provided in an embodiment of the present invention.

[0034] Figure 2 This is a schematic diagram of the overall process of the domain model-based missing information recommendation method provided in the embodiments of the present invention. Detailed Implementation

[0035] This invention addresses the problem of low coverage of requirements in existing domain models, which prevents effective recommendation of missing requirement information. It extracts conceptual terms from requirements and supplements the original domain model using a knowledge model, a requirement model, and an alignment model, resulting in a complete domain model. This maps the extracted conceptual terms to the complete domain model. Then, based on the patterns observed in both the complete and the existing domain models, it recommends missing requirement information, significantly improving the domain model's coverage of requirements and effectively solving the problem of ineffective recommendation of missing requirement information. The invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and do not limit the scope of the invention.

[0036] This invention provides a recommendation method based on missing requirement information from a domain model. (See also...) Figure 1 The method includes:

[0037] S101. Extract the terms of the requirements from the preset requirements, and construct the mapping between the terms and entities in the domain model based on the extracted terms.

[0038] Specifically, in this embodiment of the invention, a preset number of requirements, denoted as R, are randomly selected from a preset set of requirements. Conceptual terms of the requirements are extracted from the selected requirements according to a preset text terminology extraction method. The extracted conceptual terms are represented as RT. Based on the extracted conceptual terms, a mapping relationship between the conceptual terms and entities in the domain model is constructed, denoted as S.

[0039] Furthermore, the domain model in this embodiment of the invention can be obtained through networks or other means as needed. That is, the domain model in this embodiment of the invention is obtained based on needs, and then supplemented based on those needs to achieve better recommendations. Specifically, the domain model in this embodiment of the invention can be a set of <entity, relation, entity> triples from an ontology or knowledge graph;

[0040] In other words, not all domain models are useful, and a selection must be made. However, even for the selected domain model, due to differences in perspective, terminology, and functional coverage between the specific system's requirements and the domain model, there will inevitably be significant, even substantial, gaps in the coverage of functional information. Intuitively, embodiments of the present invention are based on the degree of matching (i.e., overlap) between requirements and the domain model. That is, to what extent can requirements be mapped into the domain model, the domain model is selected.

[0041] S102. Embed the domain model into the knowledge model, and embed the preset requirements into the requirement model;

[0042] In practical implementation, after obtaining the domain model, this embodiment of the invention needs to embed the domain model into the knowledge model, and then embed the requirements into the requirement model. It should be noted that both the knowledge model and the requirement model in this embodiment of the invention are vector space models. In other words, this embodiment of the invention embeds both the domain model and the requirements into their corresponding vector space models, thereby achieving the vectorization of the domain model and the requirements.

[0043] It should be noted that the knowledge model and demand model in the embodiments of the present invention are both pre-trained.

[0044] Specifically, in this embodiment of the invention, the domain model is modeled as a triple (h, r, t), where h, t ∈ ε (entity set) and r ∈ r (relation set), describing the relationship between h and t. The entire fact triple in the domain model is labeled Δ, and the conditional probability of the triple (h, r, t) is defined as:

[0045] in, b is a bias constant specified to adjust the scale for better numerical stability. The model standardizes the definitions of Pr(r|h,t) and Pr(t|h,r), and the probability of observing fact triples is defined as follows:

[0046] The goal of a knowledge model is to maximize the conditional likelihood of fact triplets existing in the knowledge graph:

[0047] Demand Model: If two terms w and v are consistent in a requirement, such as a text window or a requirement, there is a relationship r between them. wv Therefore, it is possible to declare (w, r) wv The triple of (v) is a fact. The term set in the requirement includes the initially extracted MWT, as well as the remaining nouns that appear as the sole semantic unit in any requirement or window. In this embodiment of the invention, the term set is labeled v.

[0048] The probability of terms w and v appearing simultaneously in a text window is defined as follows:

[0049] Its loss function is:

[0050] S103. Align the knowledge model and the demand model using the alignment model, and use the demand model to supplement the missing relationships or entities in the domain model through relation deduction to obtain the completed domain model.

[0051] Specifically, in this embodiment of the invention, the knowledge model and the requirement model are aligned by an alignment model, and the domain model is supplemented with requirements by relation derivation, so as to realize the correspondence between entities and requirements.

[0052] This can be understood as the alignment in this embodiment of the invention being performed based on the same expression or synonym relationship between entities in the domain model and terms in the requirements. Specifically, for (h, r, t) in Δ, if the entity name h is equal to or w in v h Based on the result of step 101, a new triple (w) is generated. h Similarly, if entities t and w t The name and w in v t If they are the same or synonymous, then (h, r, w) is generated. t ) and (w h ,r,w t ).

[0053] The loss function of the alignment model in this embodiment of the invention is:

[0054]

[0055] Considering the above three parts, the likelihood function of the joint embedding learning model is defined as:

[0056]

[0057] S104. Analyze the patterns of the extracted concepts and terms in the domain model and the completed domain model, and recommend missing demand information based on these patterns.

[0058] In this embodiment of the invention, the rules include the abstraction level in the domain model and the completed domain model, the distribution of mapped entities in the domain model and the completed domain model, the type of entities in the domain model and the completed domain model, the distribution of mapped entities in the domain model and the completed domain model, the cluster to which the mapped entities in the domain model and the completed domain model belong, and the parent-child relationship of the nodes in the domain model and the completed domain model.

[0059] Specifically, this step involves analyzing the patterns in which mapped entities appear in the domain model, including their level of abstraction in the model graph, their distribution, and other characteristics. This invention uses these patterns to narrow down the range of missing concepts recommended in the requirements.

[0060] From the perspective of domain model and requirements, based on four laws of entity overlap—namely, the distribution of mapped entities in the domain model, the type of entities in the domain model, the distribution of mapped entities in the domain model, and the cluster to which the mapped entities in the domain model belong—the search scope for missing clues in the domain model can be effectively reduced. Practice shows that as the scope of the domain model decreases, the F_2 can increase by 13%-24% in two domains beyond the original domain model. In particular, within the scope of family property, the regularity of the AHME metric proposed in this invention can achieve approximately 54%-70% improvement in F_2.

[0061] This invention relates to the distribution of mapped entities in the domain model. Typically, due to specific focuses, mapped entities only constitute a small portion of the domain model. For example, when implementing obstacle avoidance functionality for a drone, its battery control function may not be considered. Therefore, the scope of recommendations in the domain model should be narrowed to more accurately pinpoint the missing functionality targeted by the requirement. Thus, this invention achieves more accurate recommendations by analyzing the distribution patterns of mapped entities.

[0062] Specifically, the types of entities in the domain model of this invention are as follows: According to the OWL standard, elements in the domain model can generally be divided into several categories: Classes, Object Properties, Data Properties, Named Individuals, and Annotation Properties. Classes provide an abstraction mechanism to group resources with similar characteristics. Named Individuals represent objects in the domain, i.e., instances of classes. Properties (also known as Object Properties) represent a relationship between individuals. Data Properties link individuals to data values. Annotation Properties are metadata that can be used to interpret class, individual, object / data attributes.

[0063] Distribution of Mapping Entities in the Domain Model in This Invention: This invention attempts to observe the regularity of mapping by examining the distribution of mapping entities on the domain model graph. For this purpose, this invention selects and displays the subtrees of the original domain model, which contain the most and most concentrated mapping entities, highlighted with a yellow background;

[0064] The cluster to which the domain model mapping entity belongs in this embodiment of the invention: from Figure 2 As can be seen from the present invention, entities in the distributed range mapping domain model are often concentrated in child nodes under one or more intermediate nodes. Therefore, the present invention focuses specifically on the requirements of a specific version of a software product, rather than the entire domain. This observation inspires the prioritization of capturing existing software requirements when recommending missing features. With these items, the search scope for missing information can be greatly reduced.

[0065] In this invention, the domain model typically associates parent and child nodes in a tree structure. However, upon further analysis, it becomes apparent that many mapped entities do not have direct parent-child relationships. Nevertheless, they may share a common parent class or ancestor. To explore the key aspects of software requirements, it is necessary to find the common ancestor of the entities that are most frequently mapped. This ancestor should be far from the root; more specifically, its level should be as high as possible, and this invention aims to identify the root of this ancestor for entities missing from the tree. Therefore, this invention proposes a metric, AHME (Highest-Level Ancestor Containing the Most Mapped Entities), to locate entities that can be the parent or ancestor of the most mapped entities and are located at the highest level of the domain model. It can be defined by the following formula:

[0066]

[0067] MED represents the number of mapped entities belonging to the descendants of a node, ME represents the total number of mapped entities, and LEVEL(node) represents the level of that node. This invention calculates the AHME value for each node in the domain model, selects one or more of the highest AHME values ​​(excluding all mapped entities), and defines that node and all its child nodes as the required range in the domain model.

[0068] This invention designs an experiment by treating the remaining 30% of demands in two sets as missing demands and making recommendations based on the 70% of demands and a domain model. The invention evaluates the usefulness of these rules by calculating the degree to which concepts in the 30% of demands can be correctly recommended, using common metrics such as recall, precision, and F2, with recall being more important. Recall measures the degree to which correct missing information is automatically identified. Precision measures the proportion of correct information about missing demands among all automatically recommended entities. Taking into account random overlap between selected and remaining demands, this invention conducts 30 experiments to obtain the average of the metrics.

[0069] In general, the embodiments of this invention first construct a mapping between requirement concepts and domain models based on synonym identifiers, and then mine the occurrence patterns of the mapped concepts from the perspectives of requirements and domain models. This invention focuses on concept mapping and analysis by treating the distribution characteristics of requirement entities in the domain model and requirements themselves as "patterns," because this is the general principle currently used to detect missing functional information from requirements that depend on any external resources.

[0070] Specifically, the method described in this embodiment of the invention may include three stages: a mapping construction stage, a domain model completion stage, and a pattern analysis stage.

[0071] In the first stage, this invention extracts concepts from requirements and constructs mappings between these concepts and entities in the domain model. In the second stage, this invention completes a domain model with 70% random requirements through model alignment and relation derivation. In the third stage, this invention analyzes the regularity of concept mappings in the original domain model and the completed domain model. The procedure is as follows: Figure 2 As shown.

[0072] The first phase, the mapping construction phase, specifically includes:

[0073] Step 1: This invention randomly selects 70% of the requirements (denoted as R) from the entire set of requirements, and then extracts the requirement concept terms from these selected requirements;

[0074] Step Two: The mapping relationship (denoted as S) between these concepts and entities in the domain model is constructed. Note that this invention only selects Classes entities from the domain model because this invention aims to map conceptual classes to concepts in the requirements.

[0075] The second stage: the domain model completion stage, specifically includes:

[0076] Step 3: Embed the domain model and requirement model into the knowledge model and requirement model;

[0077] Step 4: Align them according to the alignment model of the present invention, and supplement the domain model (entities and relations) with requirements through relation derivation.

[0078] The third stage: the analysis of patterns in mapped entities, specifically includes:

[0079] Step 5: This invention analyzes the patterns of occurrence of mapped entities in the domain model from several different aspects, including their level of abstraction in the model graph, their distribution, and other characteristics. This invention aims to narrow down the scope of missing concepts recommended in the requirements, as not all missing concepts are valuable.

[0080] Step Six: Based on the patterns found in this invention, this invention presents preliminary ideas on missing information recommendation and uses 30% of the demand to verify the effectiveness of the method of this invention.

[0081] The advantages and positive effects of this invention lie in proposing a method for automatically constructing a mapping between NL requirements and RDF domain models, including concept extraction from NL requirements and synonym detection based on external supporting data (i.e., forums). For raw open data from both domains, this invention found that the domain model can cover an average of 31.7% of the requirements, which initially demonstrates the usefulness of domain models in recommending missing functional information.

[0082] From the perspective of domain model and requirements, this invention identifies four patterns of entity overlap. Experiments show that these rules can effectively reduce the search range for missing clues in the domain model. As the scope of the domain model decreases, F_2 can increase by 13%-24% in two domains beyond the original domain model. In particular, within the scope of family property, the regularity of the AHME metric proposed in this invention can achieve approximately 54%-70% of the F_2.

[0083] Considering the diversity of semantic relationships between requirements, this invention supplements the original domain model with known requirements through model alignment and relation derivation, aiming to better recommend missing requirements. By supplementing the domain model, the mapping rate for 70% of known requirements is improved to approximately 73.6% in both cases. Furthermore, the missing requirement recommendation based on the AHME model is improved, with F_2 gains of 23% and 34% respectively, exhibiting the same pattern.

[0084] The following will combine Figure 2 The methods described in the embodiments of the present invention will be explained and described in detail below:

[0085] See Figure 2 The method described in this embodiment of the invention includes the following steps:

[0086] Phase 1: Mapping Construction Phase.

[0087] Step 1: Extraction of entities and concepts;

[0088] The domain model in specific embodiments of this invention typically includes entities of different categories. In OWL ontology language rules, common categories include "Classes," "Object Properties," and "Named Individuals." To construct a conceptual entity mapping between the domain model and requirements, as shown in the figure, this invention extracts all entity concepts from the "Classes" of the domain model and represents them using classterms (CT). Furthermore, this invention needs to extract concepts from natural language requirements. Since "terms" and "concepts" are semantically similar, a term extraction method is used in this work. The concepts extracted from requirements in this invention are represented as RequirementTerms (RT).

[0089] By retrieving the topology of the domain model, entity concepts can be easily filtered and obtained from the domain model. This invention focuses more on concepts extracted from requirements.

[0090] This invention extracts text terms using rule-based extraction methods, statistical extraction methods, and a combination of both. Rule-based methods typically construct rules directly based on linguistic knowledge such as parts of speech and lexical patterns. Due to thorough analysis of corpora in similar domains, these methods often achieve good extraction accuracy. Generally, statistical methods, guided by statistical theory, study the statistical distribution characteristics of words in corpora, such as TF-IDF, Domain Relevance and Domain Consensus, Mutual Information, and Log-likehood. These methods are highly versatile, but this invention is not limited to a specific domain or corpus. However, the reliability of the analysis largely depends on the quality of the corpus. In this embodiment, a hybrid method is chosen because it offers the advantages of high accuracy and domain independence while combining the advantages of the two methods mentioned above. Specifically, this invention uses the C-Value method to combine grammatical rules and statistical information.

[0091] In simple terms, the C-Value implementation of this invention selects multi-word terms (MWTs) from a corpus through two steps. First, the invention obtains all candidate words based on a series of language filters for nested noun selection, such as (noun+noun)noun+noun, and a stop word filter. Then, by considering their total frequency of occurrence, their frequency as part of other longer candidate terms, the number of longer candidate terms, and their length (i.e., words), it assigns a term metric to all candidate strings. The concept of a requirement is used to define the final terms whose selection period exceeds a predefined threshold.

[0092] Step 2: Construct the mapping relationship between requirement terms and domain model entities.

[0093] First, this invention compares the names of entities and terms and establishes a direct mapping between them. However, due to the different scopes and contexts of domain models and software requirements, synonyms are also common. Therefore, this invention designs a method to detect synonym relationships.

[0094] Considering the sparse semantic information in software requirements and domain models, this invention requires finding a large amount of external domain data to connect the synonym relationships between them. Therefore, this invention combines domain entity names, high-frequency terms and entities in the requirements, and domain name models to perform extensive domain document crawling on Google. Finally, 10 documents are compiled from the search results of each case, with each corpus averaging 1000 sentences and 25,000 words.

[0095] When acquiring the domain corpus, this invention first extracts multi-word terms (MWTs) from it using the method in the first step. Embedding vectors are used to represent the terms, requirements, and domain models in the domain corpus. Finally, this invention calculates the similarity between them.

[0096] Most domain concepts consist of multiple terms. However, current research on term similarity largely focuses on single words, with few studies applicable to the detection of synonymous multi-word terms. This invention discovers a robust method based on word embedding similarity calculations to detect synonymous relationships between MWTs sharing the same head or tail components. However, the head(E) and tail(T) two-part MWT model is not suitable for longer terms. Therefore, this invention provides a simple extension by adding a middle(M) element to their definitions.

[0097] MWT = (E; M; T). In the definition of MWT above, this invention defines basic rules for synonym extraction between MWTs. The two terms are MWT1 = (E1; M1; T1) and MWT2 = (E2; M2; T2), and syn(MWT1, MWT2) indicates that there is a synonym relationship between the two terms, which can be expressed by the formula:

[0098]

[0099]

[0100]

[0101]

[0102] The four rules above follow a simple principle. Once two parts of two Multiword Terms are equal (one of which can be empty), the remaining parts are synonyms, and the two terms are synonyms. The synonymy between individual words in these three parts is determined according to a general dictionary, such as WordNet. Taking the first R1 as an example, given the head and tail of two MWTs, if they are not empty, and their head and middle parts are equal respectively, then the head is a synonym in a general dictionary. This invention considers these two key MWTs to be synonyms. The rules can be extended. For example, for very complex terms, these three parts can be represented as three more refined elements.

[0103] As with most word embedding-based similarity calculation methods, it is inevitable to select terms that have different meanings but appear frequently. Therefore, manual screening is required.

[0104] Phase Two: Domain Model Completion Phase

[0105] In this embodiment of the invention, the domain model and the requirement model are embedded into the knowledge model and the requirement model.

[0106] This invention proposes two hypotheses:

[0107] Assumption 1: Missing Domain Model Relationships. While the entities in the domain model are complete, some relationships between them may be missing. Specifically, all concepts in the requirements can be found in the domain model, either as identical expressions or as synonyms identified in stage i. However, additional relationships may exist between certain entities in the requirements.

[0108] Assumption 2: Domain model entities are missing. Specifically, some concepts in the requirements are missing from the domain model. In this case, the correspondence between these concepts will certainly also be lost.

[0109] Typically, the domain model and the requirement model are first modeled as fact triples (i.e., the Knowledge Model and the Requirement Model). This invention then adjusts their vectors, aligning them to a single vector by aligning the model.

[0110] Knowledge Model:

[0111] The domain model is modeled as a triple (h, r, t), where h, t ∈ ε (entity set) and r ∈ r (relation set), describing the relationship between h and t. The entire fact triple in the domain model is labeled Δ. The conditional probability of the triple (h, r, t) is defined as:

[0112]

[0113] in b is a bias constant specified to adjust the scale for better numerical stability. The model defines Pr(r|h,t) and Pr(t|h,r) using standardized definitions. The probability of observing a fact triple is defined as:

[0114]

[0115] The goal of a knowledge model is to maximize the conditional likelihood of fact triplets existing in the knowledge graph:

[0116]

[0117] Demand Model:

[0118] If two terms w and v are consistent in a requirement, such as a text window or a requirement, there is a relationship r between them. wv In other words, this invention can declare (w, r) wv The triple of (v) is a fact. The term set in the requirement includes the MWT extracted in step one, as well as the remaining nouns that appear as the only semantic unit in any requirement or window. This invention labels the term set as v.

[0119] The probability of terms w and v appearing simultaneously in a text window is defined as follows:

[0120]

[0121] Its loss function is:

[0122]

[0123] The alignment model according to the present invention aligns them, and the domain model (entities and relations) is supplemented with requirements through relational derivation.

[0124] Alignment Model:

[0125] Alignment is performed based on the identical expressions or synonyms between entities in the domain model and terms in the requirements. Specifically, for (h, r, t) in Δ, if the entity name h is equal to or w is equal to in v, then alignment is performed. h Based on the result of step one, a new triple (w) is generated. h Similarly, if entities t and w t The name and w in v t If they are the same or synonymous, then (h, r, w) is generated. t ) and (w h ,r,w t The loss function for the alignment model is:

[0126]

[0127] Considering the above three parts, the likelihood function of the joint embedding learning model is defined as:

[0128]

[0129] The goal of domain model completion is to add conceptual domain models of requirements that satisfy the following two conditions: 1) the concepts appear in the requirements but not in the domain model, and 2) each of their parent-child relationships (usually {hasSubClasses} or {subClassOf}) is with at least one entity domain model. To achieve this goal, this invention needs to explore the types of relationships between these requirement concepts and domain entities, as well as the types of relationships between newly added requirement concepts and existing added requirement concepts.

[0130] This invention can infer that in a triple (h, r, t), the relation vector r can be represented as r = (th). Considering that after alignment, each entity in the requirement and domain model is assigned a vector, this invention can infer the required additional relations based on this trigonometric law.

[0131] The third stage: the analysis of the patterns of mapped entities.

[0132] This invention analyzes the patterns of occurrence of mapped entities in domain models from several different perspectives, including their level of abstraction in the model graph, their distribution, and other characteristics. This invention aims to narrow down the scope of missing concepts recommended in requirements, since not all missing concepts are valuable.

[0133] The distribution of mapped entities in a domain model. Typically, due to specific focuses, mapped entities only occupy a small portion of the domain model. For example, when implementing obstacle avoidance functionality for a drone, its battery control function might not be considered. Therefore, the scope of recommendations in the domain model should be narrowed to more accurately pinpoint the missing functionality targeted by the requirement. Therefore, this invention analyzes the distribution patterns of mapped entities to achieve more accurate recommendations.

[0134] Types of Entities in a Domain Model: As mentioned in step one of this invention, according to the OWL standard, elements in a domain model can generally be divided into several categories: Classes, Object Properties, Data Properties, Named Individuals, and Annotation Properties. Classes provide an abstraction mechanism for grouping resources with similar characteristics. Named Individuals represent objects in the domain, i.e., instances of classes. Properties (also known as Object Properties) represent a relationship between individuals. Data Properties link individuals to data values. Annotation Properties are metadata that can be used to interpret classes, individuals, and object / data attributes.

[0135] Distribution of Mapping Entities in the Domain Model: This invention attempts to observe the regularity of mappings by examining their distribution on the domain model graph. For this purpose, this invention selects and displays the subtrees of the original domain model containing the most numerous and concentrated mapping entities, highlighted with a yellow background.

[0136] Cluster to which entities belong in the domain model mapping: As can be seen from the figure, entities in the distribution range mapping domain model are often concentrated in the child nodes under one or more intermediate nodes.

[0137] This phenomenon aligns with the assumption of this invention that a particular focus should always be placed on the requirements of a specific version of a software product, rather than the entire domain. This observation inspires this invention to first capture the focus of existing software requirements when recommending missing features. With these items, the scope of the search for missing information can be significantly reduced.

[0138] Domain models typically associate parent and child nodes in a tree structure. However, upon further analysis, this invention reveals that many mapped entities do not have direct parent-child relationships. Nevertheless, they may share a common parent class or ancestor. To explore the key aspects of software requirements, it is necessary to find the common ancestor of the entities that are most frequently mapped. This ancestor should be far from the root, and its level should be as high as possible (more specifically). This invention aims to identify missing entities in the tree and the root of this ancestor. Therefore, this invention proposes a metric, AHME (Highest-Level Ancestor Containing the Most Mapped Entities), to locate entities that can be the parent or ancestor of the most mapped entities and are located at the highest level of the domain model. This can be defined by the following formula:

[0139]

[0140] MED represents the number of mapped entities belonging to the descendants of a node, ME represents the total number of mapped entities, and LEVEL(node) represents the level of that node. This invention calculates the AHME value for each node in the domain model, selects one or more of the highest AHME values ​​(excluding all mapped entities), and defines that node and all its child nodes as the required range in the domain model.

[0141] Although preferred embodiments of the invention have been disclosed for illustrative purposes, those skilled in the art will recognize that various modifications, additions, and substitutions are possible, and therefore the scope of the invention should not be limited to the embodiments described above.

Claims

1. A recommendation method based on missing demand information from a domain model, characterized in that, include: The terminology of the requirements is extracted from the preset requirements, and a mapping between the terms and entities in the domain model is constructed based on the extracted terms. The domain model is a set of <entity, relation, entity> triples of ontology or knowledge graph. The domain model is embedded into the knowledge model, and the preset requirements are embedded into the requirement model, wherein both the knowledge model and the requirement model are vector space models; The knowledge model and the requirement model are aligned using an alignment model, and the missing relationships or entities in the domain model are supplemented by the requirement model through relation derivation to obtain a complete domain model. The extracted terminology mappings are analyzed to identify patterns in the domain model and the completed domain model, and recommendations for missing demand information are made based on these patterns. The rules include the abstraction level in the domain model and the completed domain model, the distribution of mapped entities in the domain model and the completed domain model, the type of entities in the domain model and the completed domain model, the distribution of mapped entities in the domain model and the completed domain model, the cluster to which the mapped entities in the domain model and the completed domain model belong, and the parent-child relationship of the nodes in the domain model and the completed domain model. The process of constructing a mapping between terms and entities in the domain model based on the extracted terms includes: constructing a mapping between terms and entities in the domain model based on the names of the extracted terms and the names of entities in the domain model, and de-synonyming of terms and entities in the domain model that have synonyms when constructing the mapping between terms and entities in the domain model. The step of extracting terminology from preset requirements and constructing a mapping between terminology and entities in the domain model based on the extracted terminology includes: randomly selecting a preset number of requirements from a preset set of requirements, denoted as R; extracting terminology from the selected requirements according to a preset text terminology extraction method, and representing the extracted terminology as RT; and constructing a mapping relationship between terminology and entities in the domain model based on the extracted terminology, denoted as S. The step of extracting terms from the selected requirements according to a preset text term extraction method includes: selecting multi-word terms (MWTs) from the corpus using the C-Value method. Specifically, first, based on a series of language filters for nested noun selection and a stop word filter, all candidate words are obtained. Then, based on the total frequency of candidate words, the frequency of candidate words as part of other candidate terms, the number of candidate terms exceeding a preset length threshold, and the length of candidate words, a term metric is assigned to the strings of all candidate words. Strings with a selection period exceeding a predefined threshold are selected as terms of the requirements.

2. The method according to claim 1, characterized in that, The preset text term extraction methods include: rule-based extraction method, statistical extraction method, and C-Value method, wherein the C-Value method is a term extraction method that combines rule extraction and statistical extraction.

3. The method according to claim 1, characterized in that, The process of de-synonyming terms and entities in the domain model that have synonym relationships includes: Based on word embedding similarity calculation, the synonym relationships between MWTs with the same head or tail components are detected, and the synonymous terms and entities in the domain model are removed. The word embedding similarity calculation includes: based on the extracted multi-word terminology MWT, using embedding vectors to represent terms, requirements, and domain models in the domain corpus, to calculate the similarity between terms, requirements, and domain models.

4. The method according to claim 3, characterized in that, If the term length between the head and tail exceeds a preset length threshold, the method further includes: Add an M element between the head and tail to expand the text and set multi-word terms. Where E is the head of MWT, M is the middle part of MWT, and T is the tail of MWT, and synonyms between MWTs are extracted based on preset rules; Let two terms be defined as follows: and ,and As two terms have a synonym relationship, it can be expressed by the formula: When two parts of two MWTs are equal, and one of them can be empty, the remaining part is a synonym, meaning the two terms are synonyms.

5. The method according to any one of claims 1-2, characterized in that, The entity types in the domain model are resource groups (Classes), attributes (Object Properties), data attributes (Data Properties), class instances (Named Individuals), and metadata (Annotation Properties).

Citation Information

Patent Citations

  • Personalized literature recommendation method based on domain knowledge atlas

    CN106960025A

  • Domain knowledge graph recommendation method for global comprehensive observation results

    CN113254630A