Automatic software product line feature model generation method based on use case enhancement

Through the automation method enhanced by use case, combined with formal modeling and large language model, the problem of manual dependence and modeling complexity in the generation of feature models of software products is solved, and efficient and accurate feature model construction is achieved, reducing costs and improving the integrity and adaptability of the model.

CN120234032AActive Publication Date: 2025-07-01SOUTH CHINA UNIV OF TECH
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510715849.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-07-01
Estimated Expiration
2045-05-30

AI Technical Summary

Technical Problem

The existing technology has insufficient experience in the field of manual modeling dependence in the generation of feature models of software product lines, high feature interaction complexity, difficulty in automated modeling, resulting in long construction cycles, high maintenance costs, and poor interaction modeling of sparse data and implicit features.

Method used

Using an automation method based on use case enhancement, combining formal modeling theory with generative artificial intelligence technology, by constructing feature constraints and use case constraint prompt frameworks, a standardized feature model is generated using large language models to achieve end-to-end intelligent transformation from unstructured requirements to formal feature models.

Benefits of technology

It realizes efficient and accurate modeling of complex software requirements, reduces manual maintenance costs, improves the integrity and correctness of feature models, and supports high-dimensional and dynamically evolved software product line requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120234032A_ABST
    Figure CN120234032A_ABST
Patent Text Reader

Abstract

The invention discloses an automatic software product line feature model generation method based on use case enhancement, and relates to the technical field of software product lines, a use case constraint prompt and a feature constraint prompt are obtained by constructing a feature constraint prompt framework and a use case constraint prompt framework and inputting a first software demand text, and then the feature constraint prompt and the feature constraint prompt are generated. Combining the two with a first software demand text, and respectively inputting a first large language model and a second large language model to obtain a first corpus and a second corpus; the first corpus and the second corpus are spliced into use case enhanced joint representation, the use case enhanced joint representation and the first software demand text form sample pairs, and a plurality of sample pairs form a training set; the training set guides the second large language model to be optimized into an SPL case enhancement large model; inputting the second software demand text into a feature constraint prompt framework to generate prompt words; and inputting the cue words into the SPL case enhanced large model to generate an SPL feature model. And finally, more efficient and accurate software product line modeling is carried out on complex software requirements.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of software product lines, and particularly to a method for generating a feature model of an automated software product line based on use case enhancement. Background Art

[0002] A software product line (SPL) refers to a collection of software systems constructed based on domain engineering methods. It derives software products that meet specific market or task requirements in a configurable manner by systematically reusing a predefined set of core assets. Multiple functions in a software product line can be abstracted into a series of features, and users can customize and generate software products that meet their needs by enabling or disabling these features.

[0003] A software product line feature model is a structured modeling method for functional features and their relationships in a software product line, used to systematically describe the commonalities and variability features of product families in a software product line. The feature model usually adopts a tree-like topological structure, where each node represents a feature, and the possible configuration space of product line members is formally defined through combination rules and constraint conditions between features. Compared with the traditional independent development mode, a software product line can achieve efficient development and personalized customization of software products by managing common features and variability features.

[0004] Although the software product line feature model can significantly improve the systematicness and reuse efficiency of software development, the complexity of its construction and maintenance is also significantly increased compared with traditional requirements modeling methods. Frequent operations of adding, deleting, and modifying features driven by software requirements often lead to continuous reconstruction of the model structure, and the time and labor costs consumed by manual maintenance methods are constantly increasing. In addition, the evolution process of the feature model involves multiple constraints. If the constraint relationships between features are not accurately modeled, it may cause configuration anomalies or runtime errors during the derivation stage. As software product lines develop towards large-scale and high-dimensional directions, the potential feature interaction combinations increase exponentially, and the complexity of feature model construction and maintenance costs also rise sharply. Therefore, how to efficiently and automatically generate a software product line feature model is of great significance.

[0005] The existing technologies for generating the feature model of a software product line have the following deficiencies: (1) The manual modeling method highly depends on the experience of domain experts, especially a large amount of manual intervention is required in the feature hierarchy division and constraint rule definition stages. As the scale of the SPL expands, the complexity of feature interactions increases sharply, and it is easy to have constraint omissions or logical contradictions in manual modeling.

[0006] (2) Although current machine learning-based generation methods can utilize historical configuration data, they perform poorly in modeling sparse data (such as product line in emerging fields) or implicit feature interactions (such as the implicit association between functional requirements and architectural features). Existing algorithms mostly rely on explicit feature labels and are difficult to automatically mine deep feature relationships from unstructured requirement texts, resulting in relatively low integrity and correctness of the generation model.

[0007] The above deficiencies directly lead to a long feature model construction cycle, high maintenance costs, and difficulty in supporting the requirements of high-dimensional and dynamically evolving industrial software product lines. Therefore, there is an urgent need for an automated software product line feature model generation method to solve the above technical problems. SUMMARY OF THE INVENTION

[0008] In view of the problems existing in the prior art, the present invention provides an automated software product line feature model generation method based on use case enhancement. By integrating formal modeling theory and generative artificial intelligence technology, and combining structured prompt engineering and domain adaptation optimization mechanism, it realizes more efficient and accurate software product line modeling for complex software requirements.

[0009] The technical solution of the present invention is realized as follows: An automated software product line feature model generation method based on use case enhancement, comprising the following steps: S1. Collect the feature model of the software product line, where the feature model includes multiple features; construct a feature constraint prompt framework according to the constraint relationships between the features; for the UML use case diagram, which includes multiple components, construct a use case constraint prompt framework according to the constraint relationships between the components; the feature model of the software product line contains several features, and its tree-like hierarchical structure reflects the constraint relationships between the features. The UML use case diagram contains use cases, actors, and relationships, where the relationships include the relationships between actors and use cases, and the relationships between use cases and use cases.

[0010] S2. Collect the first software requirement text, input the feature constraint prompt framework and the use case constraint prompt framework, and obtain use case constraint prompts and feature constraint prompts respectively; Input the use case constraint prompts and the first software requirement text into a preset first large language model to generate UML use case diagram structured data, denoted as the first corpus; Input the feature constraint prompts and the first software requirement text into a preset second large language model to generate SPL feature model structured data, denoted as the second corpus.

[0011] S3. Concatenate and process the first corpus and the second corpus to construct a use case-enhanced joint representation; a use case-enhanced joint representation and the corresponding first software requirement text form a sample pair.

[0012] S4. Repeat steps S2 and S3 to obtain multiple pairs of the samples, forming a training set; the second large language model is trained using the training set to maximize the correlation between the use case enhanced joint representation and the first software requirement text, obtaining a large model for SPL use case enhancement.

[0013] S5. Collect a second software requirement text, input it into the feature constraint prompt framework, and use the generated feature constraint prompt as a prompt word; input the prompt word into the large model for SPL use case enhancement for reasoning to generate an SPL feature model.

[0014] The "SPL feature model" here and the "feature model" collected in S1 have the same structure. Only for the purpose of distinguishing the former from the latter, the former is the result generated by the general large model of the method of the present invention.

[0015] The present invention combines the feature model of the software product line with the constraint relationship in the UML use case diagram, encodes the domain knowledge of software product line feature modeling and UML use case modeling as structured prompts, and guides the large language model to generate a standardized feature model corpus and UML use case diagram corpus. By splicing the SPL feature model corpus and the UML use case diagram corpus, the formed use case enhanced joint representation is input into the model as part of the training set data to drive the optimization of the parameters of the domain adaptive large language model. Finally, a specialized large model for SPL use case enhancement for software product line engineering is formed, realizing the end-to-end intelligent conversion from unstructured requirement descriptions to formal feature models, and achieving a leap in the efficiency and cost optimization of software product line feature model modeling through the automated modeling technology driven by generative artificial intelligence.

[0016] As a further optimization of the above solution, the constraint relationship between the features is a first constraint relationship, including mandatory, optional, or, exclusive or, requires, excludes; The constituent elements include actors and use cases; the constraint relationship between the constituent elements is a second constraint relationship, including the relationship between the actor and the use case, and the relationship between the use case and the use case; the second constraint relationship includes association, generalization, inclusion, extension.

[0017] As a further optimization of the above solution, in S2, the first software requirement text, the first corpus, and the second corpus are cleaned and standardized, including data desensitization, term standardization, and integrity verification; The data desensitization is: identifying sensitive entities in the first software requirement text and replacing the sensitive entities with placeholders; the sensitive entities include personal identity information and business sensitive information. Personal identity information includes name, ID number, phone number, email, address, etc. Business sensitive information includes customer order number, contract number, financial transaction amount, etc.

[0018] The above-mentioned terms are standardized as follows: creating a standard term library; identifying non-standard terms in the first software requirement text and replacing them with standard terms in the standard term library; for eliminating the ambiguity caused by synonyms, polysemes and abbreviations.

[0019] The integrity check is as follows: identifying missing fields in the first corpus and completing the missing fields according to the first software requirement text. It is used to ensure that each use case in the use case diagram has necessary fields, such as name, description, and actor. The missing fields are customized by users / developers. Identifying the features and the constraint relationships between the features in the second corpus, and correcting the features and the constraint relationships according to the first software requirement text.

[0020] As a further optimization of the above solution, in S2, the first corpus and the second corpus are encoded to obtain vectors and vector respectively; then the first software requirement text is encoded to obtain vector ; vectors and are concatenated to obtain the enhanced combined representation of the use case.

[0021] As a further optimization of the above solution, a delimiter is set and encoded to obtain vector ; the attention mechanism is used to dynamically calculate the positional relationship, and independent positional encodings are assigned to vector , vector , and vector , which are respectively; then the concatenation process is expressed as: ; where represents exclusive OR, and represents the enhanced combined representation of the use case.

[0022] As a further optimization of the above solution, the training is to calculate the SPL mutual information of the sample pair, that is: ; where represents the SPL mutual information; r represents vector ; j represents vector ; R is a set of multiple r's; J is a set of multiple j's; represents the joint distribution of r and j; and represent the marginal distributions of r and j respectively; Maximize the SPL mutual information to maximize the correlation between the enhanced joint representation of the use case and the first software requirement text.

[0023] Can be regarded as and The KL divergence between.

[0024] As a further optimization of the above solution, the training further includes using the MINE method and a neural network To calculate the SPL mutual information estimate, that is: ; Where Represents the SPL mutual information estimate; Is a function parameterized by the neural network parameters; Through gradient descent Maximize the SPL mutual information.

[0025] MINE stands for Mutual Information Neural Estimation. The SPL mutual information is transformed into the form of the KL divergence between the joint distribution and the marginal distribution. By means of a neural network, a function T that satisfies the integrability condition is approximated to obtain a lower bound estimate of the SPL mutual information.

[0026] As a further optimization of the above solution, the training further includes constructing the loss function of the second large language model, that is: ; Where Represents the loss function; Represents the semantic loss; Represents the SPL mutual information loss; Is the mutual information weight coefficient, and its value range is [0.1, 1.0]; N represents the number of sample pairs; K represents the number of features included in the second corpus in the i-th sample pair; Represents the ground truth value of the k-th feature in the i-th sample pair; Represents the predicted probability of the k-th feature in the i-th sample pair by the second large language model; After minimizing the loss function, the second large language model is the SPL use case enhanced large model.

[0027] By maximizing the SPL mutual information to strengthen the domain correlation between input and output, a specialized large model that not only retains the general language understanding ability but also has the characteristics of software product line domain structure perception is finally obtained.

[0028] As a further optimization of the above solution, in S5, the SPL feature model is XML-formatted data; the SPL feature model is encoded to generate a visualization model and display it.

[0029] Compared with the prior art, the present invention has the following beneficial effects: The present invention combines the feature model of the software product line and the constraint relationship in the UML use case diagram, encodes the domain knowledge of software product line feature modeling and UML use case modeling into structured prompts, and guides the large language model to generate standardized feature model corpus and UML use case diagram corpus. By splicing the feature model corpus and the UML use case diagram corpus, a use case-enhanced joint representation is formed as part of the training set data and input into the model to drive the optimization of the parameters of the domain-adaptive large language model. Finally, a specialized SPL use case-enhanced large model for software product line engineering is formed, realizing the end-to-end intelligent conversion from unstructured requirement description to formal feature model, and achieving the efficiency leap and cost optimization of software product line feature model modeling through the automated modeling technology driven by generative artificial intelligence, and realizing more efficient and accurate software product line modeling for complex software requirements. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] Figure 1 is a schematic flowchart of a method for generating an automated software product line feature model based on use case enhancement provided by an embodiment of the present invention; Figure 2 is a schematic flowchart of automatically generating a software product line feature model by using an SPL use case-enhanced large model provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0031] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0032] As Figure 1 shown, this embodiment provides a method for generating an automated software product line feature model based on use case enhancement, including the following steps: S1. Collect the feature model of the software product line, where the feature model includes multiple features; construct a feature constraint prompt framework according to the constraint relationship between the features; the constraint relationship between the features is the first constraint relationship, including mandatory, optional, or, exclusive or, requires, excludes.

[0033] For a UML use case diagram, it includes multiple constituent elements, which include actors and use cases.

[0034] Construct a use case constraint hint framework based on the constraint relationships between the constituent elements; the constraint relationships between the constituent elements are the second constraint relationships, including the relationships between actors and use cases, and the relationships between use cases and use cases; the second constraint relationships include association, generalization, inclusion, and extension. The feature model of a software product line contains several features, and its tree-like hierarchical structure reflects the constraint relationships between the features. The UML use case diagram contains use cases, actors, and relationships.

[0035] S2. Collect the first software requirement text, and input the first software requirement text into the feature constraint hint framework and the use case constraint hint framework respectively to obtain use case constraint hints and feature constraint hints.

[0036] For example, given the first software requirement text as: "The online translator system aims to provide users with convenient translation services, allowing users to directly input the text to be translated on the screen and supporting instant translation between multiple languages. When encountering text that cannot be recognized or translated, the system will promptly inform the user through a prompt message. Users can also save commonly used translation texts or phrases and view past translation contents through the history function. The system has a voice input function, can process and translate the text in uploaded pictures, and allows users to select different translation engines to obtain the best translation effect. In addition, the system provides a translation quality scoring function, supports the offline translation mode, allows users to customize the interface theme and font size, has multi-platform compatibility, and provides an API interface for third-party application integration. On the premise of protecting user privacy, the system regularly updates the language library to support newly emerging languages and dialects, and collects opinions through user feedback channels to continuously improve the product." Then the generated use case constraint hint is: "You are an expert in constructing UML use case diagrams. Given a software requirement text, you can generate the corresponding UML use case diagram, for example: " Actor 1: Actor 1 -> Use Case 1 Actor 1 -> Use Case 2 Use Case 1 -> Use Case 2 Actor 2: Actor 2 -> Use Case 3 ...... Association relationship: ...... Generalization relationship: ...... Inclusion relationship: ...... Extension relationship: ...... ” Next, I will give you a piece of software requirements text. Please generate a UML use case diagram based on the following software requirements text: {{The online translator system aims to provide users with convenient translation services, allowing users to directly input the text to be translated on the screen and supporting instant translation between multiple languages. When encountering text that cannot be recognized or translated, the system will promptly inform the user through a prompt message. Users can also save commonly used translation texts or phrases and view past translation content through the history function. The system has a voice input function, can process and translate the text in uploaded pictures, and allows users to select different translation engines to obtain the best translation effect. In addition, the system provides a translation quality scoring function, supports an offline translation mode, allows users to customize the interface theme and font size, has multi-platform compatibility, and provides an API interface for third-party application integration. On the premise of protecting user privacy, the system regularly updates the language library to support newly emerging languages and dialects, and collects opinions through user feedback channels to continuously improve the product.}} Note: The possible constraint relationships in the generated UML use case diagram include association relationship, generalization relationship, inclusion relationship, and extension relationship.

[0037] The generated result only needs to return the above format data. ”

[0038] The generated feature constraint prompt is: “You are a professional requirements analyst, good at analyzing software requirements to define the software product line (SPL) feature model. You know that there are six possible constraint types in the SPL feature model, including mandatory, optional, or, exclusive or, requires, and excludes. "Mandatory" applies to a single mandatory feature. "Optional" applies to a single optional feature. "Or" applies to a one-to-many constraint. When the parent feature exists, at least one child feature exists, and two or more child features can coexist. "Exclusive or" applies to a one-to-many constraint. When the parent feature exists, the child features coexist in a mutually exclusive manner. "Requires" applies to two unassociated but mutually dependent features. "Excludes" applies to two unassociated and mutually exclusive features.

[0039] Next, I will give you a piece of software requirements text. Please generate a software product line feature model based on the following software requirements text: {{The online translator system aims to provide users with convenient translation services, allowing users to directly input the text to be translated on the screen and supporting instant translation between multiple languages. When encountering text that cannot be recognized or translated, the system will promptly inform users through prompt messages. Users can also save commonly used translation texts or phrases and view past translation content through the history function. The system has a voice input function, can process and translate the text in uploaded pictures, and allows users to select different translation engines to obtain the best translation results. In addition, the system provides a translation quality scoring function, supports the offline translation mode, allows users to customize the interface theme and font size, has multi-platform compatibility, and provides an API interface for third-party application integration. On the premise of protecting user privacy, the system regularly updates the language library to support newly emerging languages and dialects, and collects opinions through user feedback channels to continuously improve the product.}} The generated result only needs to return data in XML format, i.e., {XML format code}.

[0040] Input the use case constraint prompt and the first software requirement text into the preset first large language model to generate the structured data of the UML use case diagram, denoted as the first corpus; in this embodiment, ChatGLM-9b is used as the first large language model. The generated first corpus is as follows: "User: User -> Input the text to be translated (input the text that needs to be translated on the screen) User -> Real-time translation (support instant translation between multiple languages) User -> Receive prompt messages (when the input text cannot be recognized or translated, the system will notify users through prompt messages) User -> Save commonly used translations (allow users to save texts or phrases that are often translated) User -> View translation history (provide a history function that allows users to view past translation records) User -> Voice input (support the voice input function, allowing users to input text by voice) User -> Image translation (able to process and translate the text in uploaded pictures) User -> Select translation engine (allow users to select different translation engines to obtain the best translation results) User -> Evaluate translation quality (provide a translation quality scoring function that allows users to evaluate the translation results) User -> Offline translation (support the offline translation mode to ensure that basic translation functions can still be used when the network is unavailable) User -> Customize interface (allow users to customize the interface theme and font size to meet the visual preferences of different users) Users -> Use multiple platforms (with multi-platform compatibility, supporting use on mobile phones, tablets, and computers) Users -> Provide feedback (Provide a user feedback channel to collect user opinions for continuous product improvement) System: System -> Protect privacy (Ensure that the text entered by users is not stored or shared without permission) System -> Update the language library (Regularly update the language library to support more emerging languages and dialects) System -> Provide API interfaces (Provide API interfaces to allow third-party applications to integrate the translation function) Association relationships: User enters text to be translated -> System performs real-time translation User saves frequently used translations -> System provides a translation history for viewing User voice input -> System image translation Generalization relationships: User real-time translation -> User offline translation User receives prompt messages -> User provides feedback Containment relationships: User real-time translation -> Select a translation engine + Evaluate translation quality User views translation history -> Save frequently used translations System protects privacy -> Update the language library + Provide API interfaces Extension relationships: User voice input -> User image translation User customizes the interface -> User uses multiple platforms System provides API interfaces -> System updates the language library.

[0041] Input the feature constraint prompt and the first software requirement text into a preset second large language model to generate structured data of the SPL feature model, denoted as the second corpus; in this embodiment, ChatGLM-9b is used as the second large language model. The generated second corpus is as follows: “<?xml version="1.0" encoding="UTF-8" standalone="no"?> <featuremodel> <struct> <and abstract="true" mandatory="true" name="OnlineTranslator"> <feature mandatory="true" name="TextEntry" / > <feature mandatory="true" name="InstantTranslation" / > <feature mandatory="true" name="PromptMessage" / > <feature mandatory="true" name="SaveFrequentText" / > <feature mandatory="true" name="HistoryFunction" / > <feature name="VoiceInput" / > <feature name="ImageTextTranslation" / > <feature name="SelectTranslationEngine" / > <feature name="TranslationQualityScoring" / > <feature name="OfflineTranslation" / > <feature name="CustomizeInterface" / > <feature name="MultiPlatformCompatibility" / > <feature name="APIIntegration" / > <feature mandatory="true" name="UserPrivacyProtection" / > <feature name="LanguageLibraryUpdate" / > <feature name="UserFeedbackChannel" / > < / and> < / struct> <constraints> <!-- Constraints can be added if there are specific dependencies or exclusions --> < / constraints> <comments / > < / featuremodel> ”。

[0042] Clean and standardize the first software requirement text, the first corpus, and the second corpus, including data desensitization, term standardization, and integrity verification.

[0043] Data desensitization is as follows: identify sensitive entities in the first software requirement text and replace the sensitive entities with placeholder; sensitive entities include personal identity information and business sensitive information. Personal identity information includes name, ID number, phone number, email, address, etc. Business sensitive information includes customer order number, contract number, financial transaction amount, etc. In this embodiment, the NER model of spaCy is used to identify sensitive entities in the collected first software requirement text.

[0044] Term standardization is as follows: create a standard term library; identify non-standard terms in the first software requirement text and replace them with standard terms in the standard term library; for eliminating the ambiguity caused by synonyms, polysemes and abbreviations. In this embodiment, the standard term library is created according to the "GB / T 8567-2006 Specification for Computer Software Documentation".

[0045] Integrity check is as follows: identify missing fields in the first corpus and complete the missing fields according to the first software requirement text. For ensuring that each use case in the use case diagram has necessary fields, such as name, description, actor. The missing fields are customized by users / developers. Identify the features in the second corpus and the constraint relationships between the features, and correct the features and the constraint relationships according to the first software requirement text.

[0046] In this embodiment, through the pre-trained language model BERT, the first corpus and the second corpus are encoded to obtain vectors and vector respectively; the formal representation of the UML use case diagram is a directed graph G=(V,E), where the node set V consists of actors and use cases, and the edge set E consists of the constraint relationships between actors and use cases, and between use cases. Extract use cases from the node set V, extract the constraint relationships between use cases from the edge set E, and encode them into UML use case diagram vectors using the pre-trained language model BERT, that is .

[0047] Then encode the first software requirement text to obtain vector ; set the separator [SEP] and encode it to obtain vector , which is used to explicitly distinguish the requirement text from the use case diagram features.

[0048] S3. Concatenate the first corpus and the second corpus to construct an enhanced joint representation of the use case; specifically, use the attention mechanism to dynamically calculate the position relationship, and assign independent position encodings to vector , vector , vector respectively, which are ; then the concatenation processing is expressed as: ; where Exclusive OR, which represents the use case enhanced union representation.

[0049] A use case enhanced union representation and the corresponding first software requirement text form a sample pair. The positional relationship is dynamically calculated through the attention mechanism, and independent positional encodings are assigned to the text, delimiters, and graph features, preserving the spatial order.

[0050] S4. Repeat steps S2 and S3 to obtain multiple sample pairs and form a training set; the second large language model is trained using the training set to maximize the correlation between the use case enhanced union representation and the first software requirement text, resulting in the SPL use case enhanced large model.

[0051] Specifically, in this embodiment, the training is to calculate the SPL mutual information of the sample pair, that is: ; where represents the SPL mutual information, which can be understood as a preliminary definition of the SPL mutual information between sample pairs; r represents the vector ; j represents the vector ; R is the set of multiple r; J is the set of multiple j; represents the joint distribution of r and j; and respectively represent the marginal distributions of r and j; can be regarded as and the KL divergence between. Maximizing the SPL mutual information maximizes the correlation between the use case enhanced union representation and the first software requirement text.

[0052] In this embodiment, the MINE method is adopted, and the neural network is used to calculate the SPL mutual information estimate, that is: ; where represents the SPL mutual information estimate; is a function parameterized by the neural network parameters; through gradient descent , the SPL mutual information is maximized.

[0053] MINE is short for Mutual Information Neural Estimation, that is, mutual information neural estimation. The SPL mutual information is transformed into the form of the KL divergence between the joint distribution and the marginal distribution, and the function T that satisfies the integrability condition is approximated by the neural network to obtain the lower bound estimate of the SPL mutual information.

[0054] In this embodiment, the training also includes constructing the loss function of the second large language model, that is: ; Among them, represents the loss function; represents the semantic loss; represents the SPL mutual information loss; is the mutual information weight coefficient, and its value range is [0.1, 1.0]; N represents the number of sample pairs; K represents the number of features included in the second corpus in the i-th sample pair; represents the ground truth value of the k-th feature in the i-th sample pair; represents the predicted probability of the k-th feature in the i-th sample pair by the second large language model; After minimizing the loss function, the second large language model is the SPL use case enhanced large model.

[0055] By maximizing the SPL mutual information to strengthen the domain correlation between input and output, a specialized large model that not only retains the general language understanding ability but also has the characteristics of software product line domain structure perception is finally obtained.

[0056] S5. As Figure 2 shown, collect the second software requirement text, input it into the feature constraint prompt framework, and the generated feature constraint prompt is used as the prompt word; input the prompt word into the SPL use case enhanced large model for reasoning to generate the SPL feature model.

[0057] In this embodiment, the SPL feature model is XML-formatted data.

[0058] Specifically, the second software requirement text is: "The online legal consultation service platform is committed to providing users with convenient and secure services, allowing users to register accounts and set login credentials, and at the same time providing a secure login mechanism to protect user privacy. Users can enter consultation questions or keywords on the home page, and the system will display relevant legal consultation articles, cases or lawyer information. If there is no matching result for the search request, the system will prompt the user to try other keywords or contact the customer service. Users can view the profiles and evaluations of legal experts in different fields, select suitable experts for online consultation, and communicate with the experts in real time through the instant messaging function. The system supports users to upload relevant files or evidence materials and ensures the security and confidentiality of all uploaded content and communications. The consultation history of users will be recorded for easy reference at any time, and users can evaluate and feedback on the consultation results, and the system will continuously optimize the service quality accordingly. In addition, the platform provides a self-service answer area for common legal questions, supports multiple payment methods, and provides users with electronic receipts. The customer service support helps to solve the problems encountered by users during the use process, and the system regularly updates the laws and regulations information to maintain the timeliness and accuracy of the content."

[0059] The output SPL feature model is: "<?xml version="1.0" encoding="UTF-8" standalone="no"?>" <featuremodel> <struct> <and abstract="true" mandatory="true" name="OnlineLegalConsultingService"> <feature mandatory="true" name="UserRegistration" / > <feature mandatory="true" name="SecureLogin" / > <feature mandatory="true" name="ConsultationSearch" / > <feature mandatory="true" name="DisplayRelatedContent" / > <feature mandatory="true" name="SearchSuggestion" / > <feature mandatory="true" name="ViewExpertProfiles" / > <feature mandatory="true" name="SelectExperts" / > <feature mandatory="true" name="InstantChat" / > <feature mandatory="true" name="UploadFiles" / > <feature mandatory="true" name="DataSecurity" / > <feature mandatory="true" name="ConsultationHistory" / > <feature mandatory="true" name="RateFeedback" / > <feature name="ServiceImprovement" / > <feature mandatory="true" name="SelfServiceArea" / > <feature mandatory="true" name="MultiplePaymentMethods" / > <feature mandatory="true" name="ElectronicReceipts" / > <feature mandatory="true" name="CustomerSupport" / > <feature mandatory="true" name="ContentUpdates" / > < / and> < / struct> <constraints> <rule> <imp> <var>Rate Feedback< / var> <var>Service Improvement< / var> < / imp> < / rule> < / constraints> <comments / > < / featuremodel> ".

[0060] Encode the SPL feature model, and use the FeatureIDE tool to generate and display a visual model. FeatureIDE is an open-source integrated development environment (IDE) based on Eclipse, focusing on feature-oriented software development (FOSD), aiming to support the development and management of software product lines (SPL).

[0061] This invention combines the feature model of the software product line with the constraint relationship in the UML use case diagram, encodes the domain knowledge of software product line feature modeling and UML use case modeling as structured prompts, and guides the large language model to generate standardized feature model corpus and UML use case diagram corpus. By splicing the feature model corpus and the UML use case diagram corpus, a use case-enhanced joint representation is formed as part of the training set data and input into the model to drive the optimization of the parameters of the domain-adaptive large language model. Finally, a specialized SPL use case-enhanced large model for software product line engineering is formed, realizing the end-to-end intelligent conversion from unstructured requirement description to formal feature model, and achieving the efficiency leap and cost optimization of software product line feature model modeling through the automated modeling technology driven by generative artificial intelligence.

[0062] According to the revelation and teaching of the above specification, those skilled in the art to which this invention pertains can also make changes and modifications to the above embodiments. Therefore, this invention is not limited to the specific embodiments disclosed and described above, and some modifications and changes to this invention should also fall within the protection scope of the claims of this invention. In addition, although some specific terms are used in this specification, these terms are only for convenience of description and do not constitute any limitation to this invention.

Claims

1. An automated software product line feature model generation method enhanced by use cases, characterized in that It includes the following steps: S1. Collect the feature model of the software product line, where the feature model includes multiple features; construct a feature constraint prompt framework according to the constraint relationships between the features; for the UML use case diagram, which includes multiple components, construct a use case constraint prompt framework according to the constraint relationships between the components; S2. Collect the first software requirement text, input it into the feature constraint prompt framework and the use case constraint prompt framework respectively, and obtain the use case constraint prompt and the feature constraint prompt; Input the use case constraint prompt and the first software requirement text into a preset first large language model to generate structured data of the UML use case diagram, denoted as the first corpus; Input the feature constraint prompt and the first software requirement text into a preset second large language model to generate structured data of the SPL feature model, denoted as the second corpus; S3. Concatenate and process the first corpus and the second corpus to construct a use case enhanced joint representation; one use case enhanced joint representation and the corresponding first software requirement text form a sample pair; S4. Repeat steps S2 and S3 to obtain multiple sample pairs, which constitute a training set; the second large language model is trained using the training set to maximize the correlation between the use case enhanced joint representation and the first software requirement text, and an SPL use case enhanced large model is obtained; S5. Collect the second software requirement text, input it into the feature constraint prompt framework, and use the generated feature constraint prompt as a prompt word; Input the prompt word into the SPL use case enhanced large model for inference to generate an SPL feature model.

2. The method for generating a feature model of an automated software product line enhanced by use cases according to claim 1, wherein The constraint relationship between the features is the first constraint relationship, including mandatory, optional, or, exclusive or, requires, excludes; The components include actors and use cases; the constraint relationship between the components is the second constraint relationship, including the relationship between the actor and the use case, and the relationship between the use cases; the second constraint relationship includes association, generalization, inclusion, extension.

3. The method for generating a feature model of an automated software product line based on use case enhancement according to claim 1, wherein, In S2, the first software requirement text, the first corpus, and the second corpus are cleaned and standardized, including data desensitization, term standardization, and integrity verification; The data desensitization is: identify sensitive entities in the first software requirement text and replace the sensitive entities with placeholders; The term standardization is: create a standard term library; identify non-standard terms in the first software requirement text and replace them with standard terms in the standard term library; The integrity verification is: identify missing fields in the first corpus and complete the missing fields according to the first software requirement text; identify the features and the constraint relationships between the features in the second corpus and correct the features and the constraint relationships according to the first software requirement text.

4. A method for generating a feature model of an automated software product line enhanced by use cases according to claim 1, wherein, In S2, the first corpus and the second corpus are encoded to obtain a vector and a vector ; Then, encode the first software requirement text to obtain a vector ; Concatenate the vector and to obtain the enhanced combined representation of the use case.

5. The method for generating a feature model of an automated software product line based on use case enhancement according to claim 4, wherein Set a delimiter and perform an encoding process to obtain a vector ; Dynamically calculate the positional relationship using the attention mechanism for the vector , the vector , the vector to allocate independent positional encodings, which are respectively ; Then the splicing process is expressed as: ; Among them, represents exclusive OR, represents the enhanced combined representation of the use case.

6. The method for generating a feature model of an automated software product line based on use case enhancement according to claim 5, wherein The training is to calculate the SPL mutual information of the sample pair, that is: ; Among them, represents the SPL mutual information; r represents the vector ; j represents the vector ; R is a set of multiple r; J is a set of multiple j; represents the joint distribution of r and j; and respectively represent the marginal distributions of r and j; Maximize the SPL mutual information to maximize the correlation between the use case enhanced joint representation and the first software requirement text.

7. A method for generating a feature model of an automated software product line enhanced by use cases according to claim 6, characterized in that , The training also includes using the MINE method and a neural network to calculate the SPL mutual information estimate, that is: ; Among them, represents the SPL mutual information estimation value; is a function parameterized by neural network parameters; By gradient descent , maximize the SPL mutual information.

8. A method for generating a feature model of an automated software product line enhanced by use cases according to claim 7, characterized in that, The training also includes constructing a loss function for the second large language model, that is: ; Among them, represents the loss function; represents the semantic loss; represents the SPL mutual information loss; is the mutual information weight coefficient, and its value range is [0.1, 1.0]; N represents the number of the sample pairs; K represents the number of the features included in the second corpus in the i-th sample pair; represents the ground truth value of the k-th feature in the i-th sample pair; represents the prediction probability of the k-th feature in the i-th sample pair by the second large language model; After minimizing the loss function, the second largest language model is the large model enhanced for the SPL use case.

9. The method for generating a feature model of an automated software product line based on use case enhancement according to claim 1, wherein In S5, the SPL feature model is XML-formatted data; the SPL feature model is encoded to generate a visualization model for display.

Citation Information

Patent Citations

  • Combination optimization automatic modeling-oriented training data generation method and system

    CN119294403A

  • Test case generation method and system based on large model and retrieval enhancement generation

    CN119537222A

  • UML model integration and refactoring method

    US20150100942A1