A method, medium, and system for obtaining an accurate prompt of a user

By extracting the semantic features of user problems and expanding them, building augmented semantic feature maps, generating candidate paths and propts, the problem of lack of coherence and difference in the existing technology is solved, and the propt generation is achieved that is closer to user needs is improved, and the application effect of large language models is improved.

CN119476500BActive Publication Date: 2025-05-30青岛网信信息科技有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510044709.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-13
Publication Date
2025-05-30
Estimated Expiration
2045-01-13

AI Technical Summary

Technical Problem

In the prior art, the generation of prompt focuses on phrase-level template filling, lacks consideration of the overall consistency and difference of prompt, which limits the effect of large language models in practical applications.

Method used

By obtaining the user's basic problems, extracting semantic features, and performing synonyms expansion, context association analysis, and domain knowledge supplementation, augmented semantic feature map is constructed, candidate paths are generated using the depth-first search algorithm, and a propt is generated based on the path, and a problem list is generated through the pre-adjusted propt generation problem model for users to select the final accurate propt.

Benefits of technology

The overall consistency and difference of propt is optimized, and the generated propt is closer to the actual needs of users and improves the effect of the large language model in practical applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119476500B_ABST
    Figure CN119476500B_ABST
Patent Text Reader

Abstract

The present invention provides a method, medium and system for obtaining an accurate prompt of a user, belonging to the technical field of large language models, including: obtaining a question raised by a user and extracting semantic features therefrom; expanding and supplementing these semantic features to form a richer semantic feature set; based on these expanded semantic features, establishing a semantic feature graph representing concepts and relationships; combining the user's context, analyzing their intentions and needs, and performing weighted ranking on the nodes in the semantic feature graph to form a weighted semantic feature graph; searching for multiple candidate paths from the weighted semantic feature graph and generating a prompt for each path; using a pre-trained question generation model to generate a series of questions according to these prompts and presenting them to the user; the user selects one of the questions as the final accurate prompt, solving the problem that the prior art lacks consideration of the overall coherence and difference of the prompt.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of large language models, and more particularly, relates to a method, medium, and system for obtaining accurate prompts from users. Background Art

[0002] With the wide application of large language models in the field of natural language processing, users can obtain the knowledge and reasoning ability of the models by asking various open-ended questions. However, the basic questions raised by users are often too simple or vague to well reflect their actual needs. To accurately capture the true intentions of users and generate more targeted answers, in-depth semantic analysis and expansion of the basic questions of users are required.

[0003] Currently, some researchers have attempted to enhance the pertinence and richness of user prompts through methods such as semantic feature extraction, knowledge graph construction, and path-based prompt generation. For example, semantic features such as keywords, named entities, and key phrases can be extracted from questions and expanded using dictionaries and knowledge bases to construct a knowledge graph that reflects semantic concepts and their associations. On this basis, by exploring the paths on the knowledge graph, prompts close to the user's needs can be generated. These methods have improved the quality of prompts to a certain extent and enhanced the user experience.

[0004] However, prompt generation in the prior art focuses on phrase-level template filling and lacks consideration of the overall coherence and differences of prompts. This problem limits the effectiveness of large language models in practical applications. Summary of the Invention

[0005] In view of this, the present invention provides a method, medium, and system for obtaining accurate prompts from users, which can solve the technical problem that prompt generation in the prior art focuses on phrase-level template filling and lacks consideration of the overall coherence and differences of prompts.

[0006] The present invention is implemented as follows:

[0007] In a first aspect of the present invention, there is provided a method for obtaining accurate prompts from users, which includes the following steps:

[0008] S10. Obtain the question raised by the user to the large language model as the basic question, and extract the semantic features of the basic question;

[0009] S20. According to the semantic features, perform synonym expansion, context association analysis, and domain knowledge supplementation to obtain an augmented semantic feature set;

[0010] S30. Establish an augmented semantic feature graph based on the augmented semantic feature set, where the nodes of the augmented semantic feature graph represent semantic concepts or keywords, and the edges represent the semantic relationships or logical connections between the nodes;

[0011] S40. According to the context input by the user, use the large language model to analyze the user's intentions and needs, identify key concepts and constraints, assign weights and prioritize the nodes in the augmented semantic feature graph to obtain a weighted semantic feature graph;

[0012] S50. For the weighted semantic feature graph, use the depth-first search algorithm to obtain multiple candidate paths of the weighted semantic feature graph;

[0013] S60. According to each candidate path, use the path-based prompt generation algorithm to generate the corresponding prompt, and combine each prompt to use the pre-fine-tuned prompt generation question model to generate a question, obtain a question list and output it to the user;

[0014] S70. Obtain the question selected by the user according to the question list as the selected question, and use the prompt corresponding to the selected question as the user's precise prompt.

[0015] On the basis of the above technical solution, a method for obtaining the user's precise prompt of the present invention can also be improved as follows:

[0016] Among them, the prompt generation question model is based on the large language model and is obtained by fine-tuning training with a large number of preset high-quality prompt-question pairs.

[0017] Specifically, the step S10 is a semantic feature extraction step, which specifically includes: performing word segmentation and part-of-speech tagging on the basic question proposed by the user to obtain the basic information of each word in the question; calculating the TF-IDF value of each word to reflect the importance of the word in the question; identifying the named entities in the question to discover important concepts such as person names, place names, and organization names; extracting the key phrases in the question to better capture the semantic information of the question. Combining the above three types of features to form a semantic feature vector of the question, which lays a foundation for subsequent semantic feature expansion and construction of the augmented semantic feature graph.

[0018] Among them, the step S20 is a semantic feature expansion step, which specifically includes: for each feature in the extracted semantic feature vector, look up its synonyms in a synonym dictionary or a word embedding model for synonym expansion; define a context window, analyze the adjacent words before and after each feature, and extract high-frequency co-occurring words as context-related words; map the features to a knowledge graph or a domain ontology library, and extract the nodes directly connected to these entities as supplementary knowledge. Through the above three expansion methods, the original semantic features are enriched, providing more comprehensive information for the construction of the subsequent augmented semantic feature graph.

[0019] Among them, the step S30 is an augmented semantic feature graph construction step, which specifically includes: using the expanded semantic features as the nodes of the graph, adding edges in the graph according to the condition that the similarity between nodes is greater than a certain threshold to represent the semantic association between nodes; assigning weights to each edge to represent the association strength between nodes. In this way, an augmented semantic feature graph reflecting semantic concepts and their relationships is constructed, providing a basis for the generation of the subsequent weighted semantic feature graph.

[0020] Among them, the step S40 is a weighted semantic feature graph generation step, which specifically includes: according to the context information input by the user, calculating the relevance score of each node to reflect the degree of relevance between the node and the user's needs; using the PageRank algorithm or centrality calculation to obtain the importance score of each node to reflect the status and influence of the node in the entire semantic system. The relevance score and the importance score act together on the edge weights of the augmented semantic feature graph to obtain a weighted semantic feature graph, highlighting the semantic concepts that are relevant and important to the user's needs.

[0021] Among them, the step S50 is a candidate path generation step, which specifically includes: determining the node with the highest weight in the weighted semantic feature graph as the starting node, and selecting the node most relevant to the user's intention as the ending node; using the depth-first search algorithm to explore multiple candidate paths from the starting node to the ending node on the weighted semantic feature graph, providing a basis for the subsequent prompt generation. In this way, multiple candidate paths reflecting the logical relationship and reasoning process between semantic concepts can be obtained.

[0022] Among them, the step S60 is a Prompt generation step, which specifically includes: defining a prompt template function, filling the nodes in the candidate path into a predefined template to generate a preliminary prompt; considering the coherence of the prompt, calculating the semantic similarity between adjacent nodes in the path to ensure the continuity and fluency of the internal semantics of the prompt; considering the diversity of the prompt, calculating the Jaccard distance between the current prompt and other prompts to ensure the difference between different prompts. Considering these factors comprehensively, multiple candidate prompts are generated for the user to choose from.

[0023] Among them, the step S70 is an accurate prompt acquisition step, specifically including: presenting multiple candidate prompts generated in step S60 to the user, and allowing the user to select according to their own needs; the prompt selected by the user is used as the final accurate prompt and serves as the input for subsequent processing by the large language model. Through this interactive method, it is ensured that the finally obtained prompt can accurately meet the actual needs of the user.

[0024] Optionally, in step S10, the semantic feature extraction algorithm may also include using a deep learning model to perform semantic representation learning on the question to obtain an embedded representation of the question, which is used as part of the semantic feature vector. This deep learning-based semantic feature extraction method can better capture the implicit semantic information in the question.

[0025] Optionally, in step S20, the synonym expansion may also combine context information to select synonyms that are more relevant to the current question, further improving the pertinence of the semantic features. At the same time, when supplementing domain knowledge, the keywords appearing in the question can also be considered to perform more accurate searches and associations on the knowledge graph or ontology library to obtain supplementary knowledge that is closer to the question.

[0026] Optionally, in step S40, in addition to using the PageRank algorithm to calculate the node importance score, other centrality metrics such as degree centrality and closeness centrality can also be used, and a suitable algorithm can be selected according to the specific application scenario.

[0027] Optionally, in step S50, the depth-first search algorithm may also combine a heuristic search strategy, such as the A* algorithm, to improve the efficiency and quality of path exploration. At the same time, multiple termination nodes can also be set to generate multiple sets of candidate paths for different user intents.

[0028] Among them, extracting the semantic features of the basic question is specifically expressed as:

[0029] The semantic feature extraction algorithm is specifically expressed as follows:

[0030] ;

[0031] In the formula, is the extracted semantic feature vector; is the basic question; is the th feature extraction function; is the th weight of the feature; is the number of feature extraction functions; is the embedded representation of the question; is the weight coefficient of the embedding representation.

[0032] Parameter acquisition method:

[0033] It can be achieved through different natural language processing techniques, such as TF-IDF, part-of-speech tagging, named entity recognition, etc. The specific steps are as follows:

[0034] Step 1: Segment and tag the part-of-speech of the question ;

[0035] Step 2: Calculate the TF-IDF value of each word;

[0036] Step 3: Identify named entities;

[0037] Step 4: Extract key phrases.

[0038] and default each are equal, take the mean value; or, and are obtained through machine learning model training, and the training data is a large number of labeled question-feature pairs.

[0039] Obtained through pre-trained language models (such as ChatGlm6B, Tongyi Qianwen).

[0040] According to the semantic features, perform synonym expansion, context correlation analysis, and domain knowledge supplementation to obtain an augmented semantic feature set, which is specifically represented as follows:

[0041] ;

[0042] In the formula, is the augmented semantic feature set; is the original semantic feature vector; is the synonym expansion function; is the context correlation analysis function; is the domain knowledge supplementation function; is the weight coefficient of each expansion method.

[0043] Parameter acquisition method:

[0044] Calculated through a synonym dictionary or a word embedding model, and the specific steps are as follows:

[0045] Step 1: For For each feature in, look up its synonyms in the thesaurus;

[0046] Step 2: Use the word embedding model to calculate the similarity, and select the words with similarity higher than the threshold as synonyms.

[0047] Obtained through context window analysis or topic model, the specific steps are as follows:

[0048] Step 1: Define the context window size ;

[0049] Step 2: For each feature in, analyze the words before and after it, and extract the high-frequency co-occurring words.

[0050] Obtained through the knowledge graph or domain ontology library, the specific steps are as follows:

[0051] Step 1: Map the features in to the entities in the knowledge graph;

[0052] Step 2: Extract the nodes directly connected to these entities as supplementary knowledge.

[0053] The default is 1 / 3 for all, or preferably determine the optimal value through the cross-validation method.

[0054] The algorithm for constructing the augmented semantic feature map is specifically expressed as follows:

[0055] ;

[0056] ;

[0057] ;

[0058] ;

[0059] In the formula, is the augmented semantic feature map; is the node set, corresponding to the extended semantic features; is the edge set; is the edge weight matrix; is the and similarity function between nodes; is the similarity threshold.

[0060] Parameter acquisition method:

[0061] It can be calculated by cosine similarity or Jaccard similarity:

[0062] ;

[0063] ;

[0064] The default value is 0.75, or it can be determined through experiments. It can start from 0.5 and be gradually adjusted to obtain the best performance.

[0065] The weighted semantic feature map generation algorithm is specifically expressed as follows:

[0066] ;

[0067] ;

[0068] ;

[0069] In the formula, is the weighted semantic feature map; is the new edge weight matrix; is the matrix elementwise multiplication; is the adjustment matrix; is the sigmoid function; is the node 's relevance score; is the node 's importance score.

[0070] Parameter acquisition method:

[0071] It is calculated through the relevance with the user input context:

[0072] ;

[0073] Among them, is the word set in the user input context.

[0074] It is calculated through the PageRank algorithm or centrality:

[0075] ;

[0076] Among them, is the damping factor, usually taken as 0.85; is the neighbor node set of the node .

[0077] The depth-first search algorithm is specifically expressed as follows:

[0078] ;

[0079] ;

[0080] In the formula, is the candidate path set; is the starting node; is the ending node; is the maximum path length; is the neighbor node set of node .

[0081] Parameter acquisition method:

[0082] Select as the node with the highest weight in the weighted semantic feature map;

[0083] Select as the node most relevant to the user intention;

[0084] Determined through experiments, usually the value range is 3 - 7.

[0085] The path-based prompt generation algorithm is specifically expressed as follows:

[0086] ;

[0087] In the formula, is the prompt generated by the th candidate path; is the th candidate path; is the template filling function; is the coherence scoring function; is the diversity scoring function; is the weight coefficient.

[0088] Parameter acquisition method:

[0089] Implemented through a predefined prompt template, filling the nodes in the path into the placeholders in the template.

[0090] Obtained by calculating the semantic similarity of adjacent nodes in the path:

[0091] ;

[0092] Obtained by calculating the average Jaccard distance between the current path and other paths:

[0093] ;

[0094] Default Both are 0.5. Additionally, The optimal value can also be determined by grid search and cross-validation.

[0095] The second aspect of the present invention provides a computer-readable storage medium, wherein program instructions are stored in the computer-readable storage medium, and when the program instructions run on a computer, they are used to execute the method for obtaining an accurate prompt of a user as described above.

[0096] The third aspect of the present invention provides a system for obtaining an accurate prompt of a user, which includes the computer-readable storage medium described above.

[0097] Compared with the prior art, the beneficial effects of the method, medium, and system for obtaining an accurate prompt of a user provided by the present invention are as follows:

[0098] 1. The extraction and expansion of semantic features are more comprehensive and dynamic. In addition to using static dictionaries and knowledge bases, this method also combines pre-trained language models to perform in-depth semantic representation learning on questions and capture implicit semantic information. At the same time, in aspects such as synonym expansion, context correlation analysis, and domain knowledge supplementation, context information and prior knowledge are fully utilized to make the semantic features closer to the actual semantics of the questions.

[0099] 2. The construction and weighting of the augmented semantic feature graph are more intelligent. This method not only considers the similarity relationship between nodes but also introduces the relevance and importance information of nodes. Through weighted adjustment, concepts that are more relevant to user needs and more important in the entire semantic system receive more attention. This modeling and analysis method based on the semantic feature graph can better reflect the semantic structure and internal logic involved in the questions.

[0100] 3. The prompt generation not only considers template filling at the phrase level but also optimizes the overall coherence and difference of the prompt. By calculating the semantic similarity of adjacent nodes in the candidate path, the internal semantics of the generated prompt is ensured to be fluent; at the same time, by analyzing the Jaccard distance between different prompts, the diversity of prompt selection is improved, meeting the user's need for differential selection.

[0101] Generally speaking, the method proposed by the present invention makes full use of technical means such as semantic analysis, knowledge graph, and graph algorithms, comprehensively considers various factors such as user needs, semantic associations, and concept importance, and can better extract and generate accurate prompts that meet the actual needs of users from the basic questions of users. It solves the technical problem that prompt generation in the prior art focuses on template filling at the phrase level and lacks consideration of the overall coherence and difference of prompts. BRIEF DESCRIPTION OF THE DRAWINGS

[0102] Figure 1 It is a flowchart of the method provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0103] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention.

[0104] As Figure 1 shown, it is a flowchart of a method for obtaining an accurate prompt of a user provided by the present invention. The method includes the following steps:

[0105] S10. Obtain the question raised by the user to the large language model as the basic question, and extract the semantic features of the basic question;

[0106] S20. According to the semantic features, perform synonym expansion, context association analysis, and domain knowledge supplementation to obtain an augmented semantic feature set;

[0107] S30. Establish an augmented semantic feature graph according to the augmented semantic feature set, where the nodes of the augmented semantic feature graph represent semantic concepts or keywords, and the edges represent the semantic relationships or logical connections between the nodes;

[0108] S40. According to the context input by the user, use the large language model to analyze the user's intentions and needs, identify key concepts and constraints, assign weights and priorities to the nodes in the augmented semantic feature graph to obtain a weighted semantic feature graph;

[0109] S50. For the weighted semantic feature graph, use the depth-first search algorithm to obtain multiple candidate paths of the weighted semantic feature graph;

[0110] S60. Generate corresponding prompts according to each candidate path using a path-based prompt generation algorithm, and combine each prompt to generate a question using a pre-fine-tuned prompt generation question model to obtain a question list and output it to the user;

[0111] S70. Obtain the question selected by the user from the question list as the selected question, and use the prompt corresponding to the selected question as the user's precise prompt.

[0112] The following is a detailed description of the specific implementation of the above steps:

[0113] Step S10: First, the purpose of this step is to extract semantic features from the basic question proposed by the user. The specific implementation is as follows:

[0114] First, perform word segmentation and part-of-speech tagging on the basic question Q. For word segmentation, common word segmentation algorithms can be used, such as HanLP, jieba, etc.; for part-of-speech tagging, corresponding natural language processing tools can be used, such as StanfordCoreNLP. Through word segmentation and part-of-speech tagging, the basic information of each word in the question can be obtained.

[0115] Next, calculate the TF-IDF value of each word. TF-IDF is a commonly used text feature representation method that can reflect the importance of a word in the text. The specific calculation formula is:

[0116] ;

[0117] where TF represents term frequency and IDF represents inverse document frequency. The higher the TF-IDF value, the more important the word is in the question.

[0118] In addition, it is also necessary to identify the named entities in the question. Named entity recognition is an important task in natural language processing, which can discover important concepts such as person names, place names, and organization names in the question. Pretrained named entity recognition models can be used, such as the models provided by tools like NLTK and spaCy. This feature can be represented as:

[0119] ;

[0120] Finally, extract the key phrases in the question. Key phrases usually contain multiple words and can better capture the semantic information of the question. Key phrase extraction algorithms based on grammar rules or machine learning can be used, such as TextRank, RAKE, etc. This feature can be represented as:

[0121] ;

[0122] Combining the above three types of features, the semantic feature vector F of the question Q can be obtained:

[0123] ;

[0124] where are the weight coefficients of each feature, is the weight coefficient of the embedding representation, is the embedding representation of the question, which can be obtained through pre-trained language models (such as ChatGlm6B, Tongyi Qianwen).

[0125] These features can comprehensively capture the semantic information of the question, laying a foundation for subsequent semantic feature expansion and construction of the augmented semantic feature map. The weight coefficient and can be obtained through machine learning model training, and the training data is a large number of labeled question-feature pairs.

[0126] Step S20: The purpose of this step is to expand on the basis of the basic semantic features extracted in step S10 to obtain a richer semantic feature set. The specific implementation is as follows:

[0127] First, perform synonym expansion. For each feature in the semantic feature vector F obtained in step S10, look up its synonyms in a synonym dictionary or a word embedding model. The synonym dictionary can use existing resources such as HowNet, WordNet, etc.; the word embedding model can use pre-trained models such as Word2Vec, GloVe, etc., calculate the similarity and select words with similarity higher than a certain threshold as synonyms. This expansion step can be expressed as:

[0128] ;

[0129] Secondly, perform context correlation analysis. Define a context window size k. For each feature in F, analyze the k words before and after it, and extract the high-frequency co-occurring words as context-related words. This expansion step can be expressed as:

[0130] ;

[0131] Finally, supplement domain knowledge. Map the features in F to the entities in the knowledge graph or domain ontology library, and extract the nodes directly connected to these entities as supplementary knowledge. This expansion step can be expressed as:

[0132] ;

[0133] Combining the above three expansion methods, the expanded semantic feature set :

[0134] ;

[0135] Among them, are the weight coefficients of each expansion method, and the optimal values can be determined through cross-validation methods.

[0136] The purpose of this step is to enrich semantic features and provide more comprehensive information for the construction of the subsequent augmented semantic feature map. Synonym expansion can increase the concept coverage, context analysis can capture semantic associations, and domain knowledge supplementation can increase professionalism. By combining these three expansion methods with weights, a more comprehensive semantic feature set can be obtained.

[0137] Step S30: The purpose of this step is to construct an augmented semantic feature map based on the expanded semantic feature set obtained in step S20 to better reflect the relationships and structures between semantics. The specific implementation is as follows:

[0138] First, each feature in the expanded semantic feature set is used as a node of the graph , forming a node set V:

[0139] ;

[0140] Next, determine the edges E between the nodes. For any two nodes and , if their similarity is greater than a certain threshold , then an edge is added to the graph. The similarity can be calculated by cosine similarity or Jaccard similarity:

[0141] ;

[0142] ;

[0143] is an adjustable parameter, usually taking a value around 0.5 and can be adjusted according to the actual situation.

[0144] Finally, assign a weight to each edge, indicating the association strength between nodes and :

[0145] ;

[0146] In summary, an augmented semantic feature map G can be obtained:

[0147] ;

[0148] where V is the node set, E is the edge set, and W is the edge weight matrix.

[0149] The purpose of this step is to construct a semantic feature map that reflects semantic concepts and their relationships. Nodes in the map represent semantic concepts or keywords, and edges represent semantic relationships or logical connections between nodes. By calculating node similarities and constructing edges, the connections between semantic features can be captured, providing a basis for generating a weighted semantic feature map in the subsequent steps. The threshold needs to be adjusted according to the specific application scenario to obtain the best semantic feature associations.

[0150] Step S40: The purpose of this step is to weight the augmented semantic feature map constructed in step S30 based on the user's context information to reflect the importance of nodes and their relevance to the user's needs. The specific implementation is as follows:

[0151] First, calculate the relevance score for each node . The relevance score reflects the degree of relevance of the node to the user input context and can be calculated using the following formula:

[0152] ;

[0153] where is the set of words in the user input context, is the similarity between node and context word , which can be calculated using cosine similarity or Jaccard similarity.

[0154] Next, calculate the importance score for each node . The importance score reflects the status and influence of the node in the entire semantic feature map and can be obtained using the PageRank algorithm or centrality calculation:

[0155] ;

[0156] where is the damping factor, usually taken as 0.85; is the set of neighbor nodes of node .

[0157] Finally, adjust the edge weights W of the augmented semantic feature map G according to the relevance score and importance score of the nodes to obtain the weighted semantic feature map :

[0158] ;

[0159] ;

[0160] where Denotes element - wise multiplication of the matrix, which is the sigmoid function.

[0161] The purpose of this step is to weight the semantic feature map according to the context information input by the user, so as to highlight the semantic concepts and keywords related to the user's needs and reflect their status in the entire semantic system. The relevance score and importance score jointly act on the adjustment of the edge weights, so that more relevant and important semantic features receive more attention and are given priority in the subsequent candidate path generation.

[0162] Step S50: The purpose of this step is to generate multiple candidate paths based on the weighted semantic feature map to provide choices for subsequent prompt generation. The specific implementation is as follows:

[0163] First, determine the starting node and the ending node . is selected as the node with the highest weight in the weighted semantic feature map, reflecting the core concept of the entire semantic system; is selected as the node most relevant to the user's intention, which can be determined according to the context information input by the user.

[0164] Next, use the depth - first search (DFS) algorithm to generate candidate paths on the weighted semantic feature map as follows:

[0165] ;

[0166] Among them, is the maximum path length, usually with a value range of 3 - 7, which can be adjusted according to the actual situation. The specific implementation of the DFS algorithm is as follows:

[0167] ;

[0168] This algorithm starts from the starting node , performs a depth - first search along the edges in the graph until it encounters the ending node or reaches the maximum path length , and records the searched path.

[0169] In this way, multiple candidate paths from to can be obtained to provide a basis for subsequent prompt generation.

[0170] The purpose of this step is to generate multiple candidate paths from the core concept to the concepts related to the user's intention based on the weighted semantic feature map. These candidate paths reflect the logical relationships and reasoning processes between semantic concepts, providing diverse choices for subsequent prompt generation. The DFS algorithm can effectively explore the paths in the graph, and the determination of the starting and ending nodes reflects the attention to the user's needs.

[0171] Step S60: The purpose of this step is to generate multiple candidate prompts for the user using a path-based prompt generation algorithm based on the candidate paths generated in step S50. The specific implementation is as follows:

[0172] First, define a prompt template function , and fill the nodes in the candidate path into the predefined prompt template to obtain a preliminary prompt:

[0173] ;

[0174] Next, consider the coherence and diversity of the prompt. Coherence reflects the continuity and fluency of the semantics within the prompt and can be evaluated by calculating the semantic similarity between adjacent nodes in the path:

[0175] ;

[0176] Diversity, on the other hand, reflects the differences between different prompts and can be evaluated by calculating the Jaccard distance between the current prompt and other prompts:

[0177] ;

[0178] Finally, the i-th candidate prompt can be obtained:

[0179] ;

[0180] Among them, and are the weight coefficients for coherence and diversity, and the optimal values can be determined through grid search and cross-validation.

[0181] The purpose of this step is to generate multiple prompts for the user to choose from based on the candidate paths. The prompt template function fills the semantic concepts in the path into the predefined template to form a complete prompt; the coherence scoring function ensures the continuity and fluency of the semantics within the prompt, and the diversity scoring function This ensures the difference between different prompts, thus providing users with richer choices.

[0182] By adding coherence and diversity scores, it can be ensured that the generated prompts can not only cover various semantic concepts involved in user needs, but also reflect the logical relationships between semantics, and at the same time can provide differentiated choices to meet the diverse needs of users.

[0183] Step S70: The purpose of this step is to obtain the precise prompt finally selected by the user. The specific implementation method is as follows:

[0184] First, present the list of prompts generated in step S60 to the user and let the user make a choice according to their own needs. The prompt selected by the user is used as the final precise prompt.

[0185] Next, use the prompt selected by the user as the output and as the input for subsequent processing by the large language model. This prompt has been optimized through steps such as semantic feature extraction, expansion and weighting, and candidate path generation, and can better reflect the actual needs of users.

[0186] The core of this step is to allow the user to independently select the prompt that best suits their needs. The previous steps provide users with diverse prompt choices, and users can make a choice according to their own understanding and preferences. This interactive method can ensure that the finally obtained prompt more precisely meets the needs of users.

[0187] The second aspect of the present invention provides a computer-readable storage medium, wherein program instructions are stored in the computer-readable storage medium, and when the program instructions run on a computer, they are used to execute the method for obtaining the precise prompt of the user as described above.

[0188] The third aspect of the present invention provides a system for obtaining the precise prompt of the user, which includes the above-mentioned computer-readable storage medium.

[0189] Specifically, the principle of the present invention: The method for obtaining the precise prompt proposed by the present invention mainly includes the following key steps:

[0190] 1. Semantic feature extraction

[0191] The core idea of this step is to lay a foundation for subsequent semantic expansion and knowledge modeling by deeply analyzing the basic questions raised by users and extracting multi-dimensional semantic features including word frequency, named entities, key phrases, etc. Specifically, first, word segmentation and part-of-speech tagging are performed to obtain the basic information of each word in the question; then, the TF-IDF value of each word is calculated to reflect the importance of the word in the question; at the same time, named entities in the question are identified to discover important concepts such as people's names and place names; finally, key phrases are extracted to better capture the semantic connotation of the question. These features are combined to form the semantic feature vector of the question.

[0192] The reason for adopting this way of multi-dimensional feature extraction is that a single feature is difficult to comprehensively reflect the semantics of the question. Word frequency can reflect the topic keywords of the question, named entities reveal the important entities involved in the question, and key phrases better capture the internal logic between semantics. By integrating these features, the semantic features of the question can be more comprehensively characterized, providing richer information for subsequent expansion and analysis.

[0193] 2. Semantic Feature Expansion

[0194] The purpose of this step is to expand on the basis semantic feature vector extracted in Step 1 to obtain a richer set of semantic features. It specifically includes three aspects:

[0195] Synonym Expansion: For each word in the basic features, look up its synonyms in a synonym dictionary or a word embedding model to increase the concept coverage. This can make up for the possible missing semantic information in the basic features.

[0196] Context Association Analysis: Define a context window, analyze the co-occurring words before and after each feature word, extract the frequently occurring related words, and capture the associations between semantics. This is conducive to reflecting the context semantics of the question.

[0197] Domain Knowledge Supplement: Map the basic features to a knowledge graph or an ontology library, extract the nodes directly related to these concepts, and add professional knowledge. This can make up for the lack of domain information in the basic features.

[0198] By organically combining the above three expansion methods, a more comprehensive, dynamic, and professional set of semantic features can be constructed, providing a richer basis for subsequent knowledge modeling and prompt generation.

[0199] 3. Construction of Augmented Semantic Feature Graph

[0200] Based on the semantic feature set expanded in Step 2, the purpose of this step is to construct a knowledge graph that reflects semantic concepts and their relationships. Specifically, first, each semantic feature is used as a node in the graph, and then edges are constructed according to the similarity relationships between the nodes to form an augmented semantic feature graph. The weight of the edge reflects the strength of the semantic association between the nodes.

[0201] This knowledge representation method based on the semantic feature graph can better capture the semantic structure and internal logic involved in the problem. Compared with the traditional triple-based knowledge base, the semantic feature graph can depict various associations between concepts more delicately, providing richer information for subsequent reasoning and analysis.

[0202] 4. Generation of Weighted Semantic Feature Graph

[0203] Based on the augmented semantic feature graph constructed in Step 3, the purpose of this step is to weight the nodes in the graph according to the context information input by the user, so as to highlight the concepts that are more relevant to the user's needs and have a more important position in the entire semantic system.

[0204] Specifically, first calculate the relevance score of each node, which reflects the degree of relevance between the node and the user's context. At the same time, calculate the importance score of each node, which reveals the position and influence of the node in the entire semantic system. Combining these two scoring factors, weight adjustment is performed on the edges of the augmented semantic feature graph, so that more relevant and important semantic concepts receive more attention in the subsequent candidate path generation.

[0205] This weighting method based on node relevance and importance can effectively incorporate user demand information into knowledge modeling, making the knowledge graph better fit the actual application scenario. Compared with the simple similarity-based construction method, this method can more intelligently discover and highlight semantic concepts related to user needs.

[0206] 5. Generation of Candidate Paths

[0207] Based on the weighted semantic feature graph generated in Step 4, the purpose of this step is to explore multiple candidate paths from the core semantic concept to the concepts related to the user's intention, providing a basis for subsequent prompt generation.

[0208] Specifically, first determine the starting node as the core concept of the entire semantic system, and the ending node as the concept most relevant to the user's intention. Then use the depth-first search algorithm to explore paths on the weighted semantic feature graph to generate multiple candidate paths. These paths reflect the logical relationships and reasoning processes between semantic concepts, providing rich choices for subsequent prompt generation.

[0209] Compared with simple shortest path algorithms, depth-first search can better discover the hidden semantic associations in the graph and generate more diverse candidate paths. At the same time, by reasonably determining the starting and ending nodes, it also reflects the emphasis on user needs, making the generated paths closer to the actual application scenarios.

[0210] 6.Prompt Generation

[0211] The purpose of this step is to use the path-based prompt generation algorithm to provide multiple candidate prompts for the user based on the candidate paths generated in step 5.

[0212] Specifically, first define a prompt template function, fill the semantic concepts in the candidate paths into the predefined template to form a preliminary prompt. Next, consider the coherence and diversity of the prompt. Coherence reflects the fluency of the internal semantics of the prompt and can be evaluated by calculating the semantic similarity of adjacent nodes in the path; diversity reflects the difference between different prompts and can be measured by analyzing the Jaccard distance between prompts. Considering these two factors comprehensively, the final prompt is generated.

[0213] This path-based prompt generation method can make full use of the semantic knowledge graph constructed in the previous steps and is closer to the actual semantic connotation of the problem. Compared with simple template filling, this method not only considers the coherence within the prompt but also improves the diversity of prompt selection, and can better meet the personalized needs of users.

[0214] To better understand and implement the present invention, a specific embodiment 1 of the method of the present invention is provided below. The relevant steps of this embodiment 1 are specifically described as follows:

[0215] Step S10: Semantic Feature Extraction

[0216] The purpose of this step is to extract semantic feature vectors from the basic question posed by the user. The specific implementation is as follows:

[0217] First, perform word segmentation and part-of-speech tagging on the question to obtain the basic information of each word in the question. Word segmentation can use common word segmentation algorithms such as HanLP and jieba, and part-of-speech tagging can use natural language processing tools such as StanfordCoreNLP. This step can obtain the word set of the question and the corresponding part-of-speech set .

[0218] Next, calculate the TF-IDF value of each word. TF-IDF is a commonly used method for representing text features and can reflect the importance of words in the text. The specific calculation formula is:

[0219] ;

[0220] where, is the total number of documents, is the number of documents containing the word . The higher the TF-IDF value, the more important the word is in the question .

[0221] In addition, it is also necessary to identify the named entities in the question. Named entity recognition is an important task in natural language processing and can discover important concepts such as person names, place names, and organization names. Pretrained named entity recognition models can be used, such as the models provided by tools like NLTK and spaCy. This feature can be represented as:

[0222] ;

[0223] where, is the number of named entities in the question , is the type of the i-th named entity, such as person name, place name, etc.

[0224] Finally, extract the key phrases in the question. Key phrases usually contain multiple words and can better capture the semantic information of the question. Key phrase extraction algorithms based on grammar rules or machine learning can be used, such as TextRank and RAKE. This feature can be represented as:

[0225] ;

[0226] where, is the number of key phrases extracted from the question , and each represents a key phrase.

[0227] Combining the above three types of features, the semantic feature vector of the question can be obtained :

[0228] ;

[0229] where, is the weight coefficient of each feature, is the weight coefficient of the embedding representation, is the embedding representation of the question, which can be obtained through a pre-trained language model (such as ChatGlm6B, Tongyi Qianwen).

[0230] These features can comprehensively capture the semantic information of the problem, laying a foundation for subsequent semantic feature expansion and the construction of an augmented semantic feature map. The weight coefficients and can be obtained through machine learning model training, and the training data is a large number of labeled problem-feature pairs.

[0231] Step S20: Semantic feature expansion

[0232] The purpose of this step is to expand on the basis of the basic semantic feature vectors extracted in step S10 to obtain a richer set of semantic features . The specific implementation is as follows:

[0233] First, perform synonym expansion. For each feature in , look up its set of synonyms in a synonym dictionary or a word embedding model . The synonym dictionary can use existing resources such as HowNet and WordNet; for the word embedding model, a pre-trained model such as Word2Vec or GloVe can be used to calculate similarities and select words with similarities higher than a certain threshold as synonyms. This expansion step can be expressed as:

[0234] ;

[0235] Secondly, perform context correlation analysis. Define a context window size , and for each feature in , analyze the words before and after it, and extract the high-frequency co-occurring words as the context-related word set . This expansion step can be expressed as:

[0236] ;

[0237] Finally, supplement domain knowledge. Map the features in to the entities in the knowledge graph or domain ontology library, and extract the set of nodes directly connected to these entities as supplementary knowledge. This expansion step can be expressed as:

[0238] ;

[0239] Combining the above three expansion methods, the expanded set of semantic features can be obtained:

[0240] ;

[0241] Among them, is the weight coefficient of each expansion method, and the optimal value can be determined by the cross-validation method.

[0242] The purpose of this step is to enrich the semantic features and provide more comprehensive information for the construction of the subsequent augmented semantic feature map. Synonym expansion can increase the concept coverage, context analysis can capture semantic associations, and domain knowledge supplementation can increase professionalism. By combining these three expansion methods with weights, a more comprehensive semantic feature set can be obtained.

[0243] Step S30: Construction of the augmented semantic feature map

[0244] The purpose of this step is to construct an augmented semantic feature map based on the extended semantic feature set obtained in step S20 to better reflect the relationships and structures between semantics. The specific implementation method is as follows:

[0245] First, each feature in the extended semantic feature set is used as a node of the graph to form a node set :

[0246] ;

[0247] Next, determine the edges between the nodes . For any two nodes and , if their similarity is greater than a certain threshold , then add an edge in the graph. The similarity can be calculated by cosine similarity or Jaccard similarity:

[0248] ;

[0249] ;

[0250] is an adjustable parameter, usually taking a value of about 0.5 and can be adjusted according to the actual situation.

[0251] Finally, assign a weight to each edge , representing the association strength between nodes and :

[0252] ;

[0253] In summary, an augmented semantic feature map can be obtained. :

[0254] ;

[0255] Among them, is the node set, is the edge set, is the edge weight matrix.

[0256] The purpose of this step is to construct a semantic feature map that reflects semantic concepts and their relationships. Nodes in the graph represent semantic concepts or keywords, and edges represent semantic relationships or logical connections between nodes. By calculating node similarities and constructing edges, the connections between semantic features can be captured, providing a basis for generating the subsequent weighted semantic feature map. The threshold needs to be adjusted according to the specific application scenario to obtain the best semantic feature association.

[0257] Step S40: Generation of weighted semantic feature map

[0258] The purpose of this step is to weight the augmented semantic feature map constructed in step S30 according to the user's context information to reflect the importance of nodes and their relevance to the user's needs. The specific implementation is as follows: First, calculate the relevance score

[0259] of each node . The relevance score reflects the degree of relevance between the node and the user input context and can be calculated by the following formula: ;

[0260] ;

[0261] Among them, is the set of words in the user input context, is the similarity between the node and the context word and can be calculated using cosine similarity or Jaccard similarity.

[0262] Next, calculate the importance score of each node . The importance score reflects the status and influence of the node in the entire semantic feature map and can be obtained using the PageRank algorithm or centrality calculation:

[0263] ;

[0264] Among them, is the damping factor, usually taken as 0.85; is the node The set of neighbor nodes of

[0265] Finally, according to the relevance score and importance score of the nodes, the edge weights of the augmented semantic feature map are adjusted to obtain the weighted semantic feature map : :

[0266] ;

[0267] ;

[0268] where, represents element-wise multiplication of matrices, is the sigmoid function.

[0269] The purpose of this step is to weight the semantic feature map according to the context information input by the user, so as to highlight the semantic concepts and keywords related to the user's needs and reflect their status in the entire semantic system. The relevance score and importance score jointly act on the adjustment of the edge weights, so that more relevant and important semantic features receive more attention and are given priority in the subsequent candidate path generation.

[0270] Step S50: Candidate path generation

[0271] The purpose of this step is to generate multiple candidate paths based on the weighted semantic feature map for subsequent prompt generation. The specific implementation is as follows:

[0272] First, determine the starting node and the ending node . is selected as the node with the highest weight in the weighted semantic feature map, reflecting the core concept of the entire semantic system; is then selected as the node most relevant to the user's intention, which can be determined according to the context information input by the user.

[0273] Next, use the depth-first search (DFS) algorithm to generate candidate paths on the weighted semantic feature map :

[0274] ;

[0275] where, is the maximum path length, usually with a value range of 3 - 7, which can be adjusted according to the actual situation.

[0276] The algorithm starts from the starting node Start from a node and perform a depth - first search along the edges in the graph until a termination node is encountered or the maximum path length is reached , and record the searched path

[0277] In this way, multiple candidate paths from to can be obtained , providing a basis for subsequent prompt generation

[0278] The purpose of this step is to generate multiple candidate paths from the core concept to the concepts related to the user's intention according to the weighted semantic feature graph. These candidate paths reflect the logical relationships and reasoning processes between semantic concepts, providing diverse choices for subsequent prompt generation. The DFS algorithm can effectively explore the paths in the graph, and the determination of the starting and termination nodes reflects the attention to the user's needs

[0279] Step S60: Prompt Generation

[0280] The purpose of this step is to generate multiple candidate prompts for the user using a path - based prompt generation algorithm based on the candidate paths generated in step S50 . The specific implementation is as follows: .

[0281] First, define a prompt template function , fill the nodes in the candidate path into the predefined prompt template to obtain a preliminary prompt :

[0282] ;

[0283] Next, consider the coherence and diversity of the prompt

[0284] ;

[0285] Diversity reflects the differences between different prompts and can be evaluated by calculating the Jaccard distance between the current prompt and other prompts:

[0286] ;

[0287] Finally, the th candidate prompt can be obtained:

[0288] ;

[0289] Among them, and are the weight coefficients for coherence and diversity, and the optimal values can be determined through grid search and cross-validation.

[0290] The purpose of this step is to generate multiple prompts for the user to choose from according to the candidate paths. The prompt template function fills the semantic concepts in the path into a predefined template to form a complete prompt; the coherence scoring function ensures the continuity and fluency of the internal semantics of the prompt, and the diversity scoring function ensures the differences between different prompts, thus providing the user with richer choices.

[0291] By adding the scoring of coherence and diversity, it can be ensured that the generated prompts can not only cover various semantic concepts involved in the user's needs, but also reflect the logical relationships between the semantics, and at the same time can provide differentiated choices to meet the diverse needs of the user.

[0292] Step S70: Obtaining the precise prompt

[0293] The purpose of this step is to obtain the precise prompt finally selected by the user. The specific implementation is as follows:

[0294] First, present the prompt list generated in step S60 to the user and let the user choose according to their own needs. The prompt selected by the user is used as the final precise prompt.

[0295] Next, use the prompt selected by the user as the output and as the input for the subsequent processing of the large language model. This prompt has been optimized through steps such as semantic feature extraction, expansion and weighting, and candidate path generation, and can better reflect the actual needs of the user.

[0296] To further better understand and implement the present invention, an embodiment 2 of a specific application scenario of the present invention is provided below: A technical team developed a dialogue assistant based on a large language model. Users can ask various questions through natural language, and the system gives corresponding answers. In the actual application process, the technical team found that the basic questions raised by users were often too simple or ambiguous to accurately reflect their real needs, resulting in certain deviations in the answers given by the system and poor user experience. Therefore, the technical team decided to adopt the method for obtaining precise prompts proposed by the present invention to improve the service quality of the system.

[0297] The dialogue system of this technical team mainly includes three modules: a question understanding module, a knowledge base construction module, and an answer generation module. Taking the question "How to prevent water pipes from bursting" raised by a user named Zhang San as an example, the implementation process of each module will be described in detail below.

[0298] 1. Question Understanding Module

[0299] The main task of this module is to extract an accurate prompt using the method of the present invention based on the basic question raised by the user. The specific steps are as follows:

[0300] Step S10: Semantic Feature Extraction

[0301] The basic question raised by the user is "How to prevent water pipes from bursting". First, the question is segmented and part-of-speech tagged, and the following results are obtained, as shown in Table 1:

[0302] Table 1 Part-of-Speech Tagging Table

[0303] Word Part of speech How Conjunction Prevent Verb Water pipe Noun Burst Verb

[0304] Next, calculate the TF-IDF value of each word. Taking "water pipe" as an example, it appears 10 times in the question corpus, and there are a total of 100 words, so the TF-IDF value of "water pipe" is:

[0305] ;

[0306] Similarly, calculate the TF-IDF values of other words and organize them as shown in Table 2 below:

[0307] Table 2 TF-IDF Value Table

[0308] Word TF-IDF value How 0.8 Prevent 1.5 Water pipe 2.0 Burst 1.8

[0309] In addition, through named entity recognition, it is found that "water pipe" is an important geographical object concept.

[0310] Finally, use the TextRank algorithm to extract the key phrase, and get "prevent water pipes from bursting".

[0311] Based on the above features, form the semantic feature vector of the question :

[0312] ;

[0313] Step S20: Semantic Feature Expansion

[0314] Next, expand the semantic feature vector :

[0315] 1) Synonym Expansion:

[0316] Search for synonyms of "water pipe" in HowNet, and obtain "pipe", "water pipe", "water supply pipe", etc.

[0317] Calculate the similarity in the Word2Vec model, and select words with a similarity greater than 0.7, such as "drain pipe", "water supply pipe", etc.

[0318] Comprehensively obtain the synonym set Pipe, water pipe, water supply pipe, drain pipe, water supply pipe 。

[0319] 2) Context correlation analysis:

[0320] Define the context window size , analyze the first 3 words and the last 3 words of "water pipe", and find that the frequently occurring words include "household", "installation", "material", etc.

[0321] Obtain the context-related word set Household, installation, material 。

[0322] 3) Domain knowledge supplementation:

[0323] Map "water pipe" to the knowledge graph, and find that the concepts directly connected to it include "pipe material", "water pressure", "anti-corrosion", etc.

[0324] Obtain the domain knowledge supplementation set Pipe material, water pressure, anti-corrosion 。

[0325] Integrate the above three expansion methods to obtain the expanded semantic feature set :

[0326] ;

[0327] ;

[0328] Step S30: Construction of the augmented semantic feature graph

[0329] Based on the expanded semantic feature set , construct the augmented semantic feature graph 。

[0330] First, take each semantic feature as a node of the graph, "How", "prevent", "water pipe", "burst", "pipe", "water pipe", "water supply pipe", "drain pipe", "water supply pipe", "household", "installation", "material", "pipe material", "water pressure", "anti-corrosion" 。

[0331] Next, calculate the cosine similarity between any two nodes. If the similarity is greater than 0.6, add an edge in the graph. The weight of the edge is the similarity value divided by the sum of the total similarities of the nodes.

[0332] The finally constructed augmented semantic feature graph is shown in Table 3:

[0333] Table 3 Augmented Semantic Feature Graph Adjacency list representation

[0334] Node Edge and weight How ("Prevent", 0.72) Prevent ("How", 0.72), ("Water pipe", 0.68), ("Burst", 0.74) Water pipe ("Prevent", 0.68), ("Pipeline", 0.78), ("Water supply pipe", 0.71), ("Drain pipe", 0.65), ("Water pipe", 0.62), ("Pipe material", 0.67), ("Water pressure", 0.59) Burst ("Prevent", 0.74) Pipeline ("Water pipe", 0.78), ("Water supply pipe", 0.75), ("Drain pipe", 0.68) Water pipeline ("Water supply pipe", 0.73), ("Water pipe", 0.68) Water supply pipe ("Water pipe", 0.71), ("Pipeline", 0.75), ("Water pipeline", 0.73), ("Material", 0.62) Drain pipe ("Water pipe", 0.65), ("Pipeline", 0.68) Water pipe ("Water pipe", 0.62), ("Water pipeline", 0.68) Household ("Water pipe", 0.58), ("Material", 0.63) Install ("Water pipe", 0.61), ("Material", 0.67) Material ("Water supply pipe", 0.62), ("Household", 0.63), ("Install", 0.67) Pipe material ("Water pipe", 0.67), ("Anticorrosion", 0.72) Water pressure ("Water pipe", 0.59) Anticorrosion ("Pipe material", 0.72)

[0335] Step S40: Generation of weighted semantic feature graph

[0336] According to the question "How to prevent water pipes from bursting" raised by the user, calculate the relevance score and importance score of each node, and weight the augmented semantic feature graph to obtain a weighted semantic feature graph .

[0337] 1) Calculation of node relevance score:

[0338] The keywords in the user's question are "water pipe", "burst", and "prevent". Calculate the similarity between each node and these keywords to obtain the node relevance scores as shown in Table 4.

[0339] Table 4 Node Relevance Scores

[0340] Node Relevance score How 0.62 Prevent 0.78 Water pipe 0.88 Burst 0.82 Pipeline 0.72 Water pipeline 0.65 Water supply pipe 0.75 Drain pipe 0.61 Water pipe 0.58 Household 0.53 Install 0.55 Material 0.62 Pipe material 0.68 Water pressure 0.59 Anticorrosion 0.51

[0341] 2) Calculation of node importance score:

[0342] Use the PageRank algorithm to calculate the importance scores of each node, and the results are shown in Table 5.

[0343] Table 5 Node Importance Scores

[0344] Node Importance score How 0.068 Prevent 0.129 Water pipe 0.186 Burst 0.112 Pipeline 0.152 Water pipeline 0.093 Water supply pipe 0.127 Drain pipe 0.084 Water pipe 0.072 Household 0.045 Install 0.052 Material 0.091 Pipe material 0.108 Water pressure 0.038 Anticorrosion 0.043

[0345] 3) Generation of weighted semantic feature graph :

[0346] According to the node relevance scores and importance scores, adjust the edge weights of the augmented semantic feature graph to obtain a weighted semantic feature graph , as shown in Table 6.

[0347] Table 6 Adjacency list representation of weighted semantic feature graph

[0348] Node Edge and weight How ("Prevent", 0.78) Prevent ("How", 0.78), ("Water pipe", 0.84), ("Burst", 0.88) Water pipe ("Prevent", 0.84), ("Pipeline", 0.92), ("Water supply pipe", 0.85), ("Drain pipe", 0.72), ("Water pipe", 0.65), ("Pipe material", 0.81), ("Water pressure", 0.62) Burst ("Prevent", 0.88) Pipeline ("Water pipe", 0.92), ("Water supply pipe", 0.88), ("Water pipeline", 0.82) Water pipeline ("Water supply pipe", 0.75), ("Water pipe", 0.71) Water supply pipe ("Water pipe", 0.85), ("Pipeline", 0.88), ("Water pipeline", 0.75), ("Material", 0.72) Drain pipe ("Water pipe", 0.72), ("Pipeline", 0.75) Water pipe ("Water pipe", 0.65), ("Water pipeline", 0.71) Household ("Water pipe", 0.48), ("Material", 0.58) Install ("Water pipe", 0.52), ("Material", 0.62) Material ("Water supply pipe", 0.72), ("Household", 0.58), ("Installation", 0.62) Pipe material ("Water pipe", 0.81), ("Anticorrosion", 0.84) Water pressure ("Water pipe", 0.62) Anticorrosion ("Pipe material", 0.84) ​

[0349] It can be seen that through weighted adjustment, the edge weights of nodes with higher relevance to the user's question and more important positions in the entire semantic system, such as "water pipe", "prevent", "rupture", etc., have been increased and will receive more attention in the subsequent generation of candidate paths.

[0350] Step S50: Generation of candidate paths

[0351] Based on the weighted semantic feature graph , the depth-first search algorithm is used to generate candidate paths from the starting node "water pipe" to the ending node "rupture". The maximum path length is set to 5, and the following 3 candidate paths are explored:

[0352] = {"water pipe" -> "pipe" -> "water supply pipe" -> "material" -> "rupture"};

[0353] = {"water pipe" -> "pipe material" -> "anti-corrosion" -> "rupture"};

[0354] = {"water pipe" -> "water pressure" -> "rupture"};

[0355] These candidate paths reflect different semantic logics and reasoning processes, providing a basis for the subsequent generation of prompts.

[0356] Step S60: Generation of prompts

[0357] According to the 3 candidate paths generated in step S50, using the path-based prompt generation algorithm, the following 3 candidate prompts are generated:

[0358] = "To prevent the rupture of water pipes, first, it is necessary to understand the material properties of water pipes and select appropriate pipe materials and anti-corrosion measures."

[0359] = "The root cause of water pipe rupture is often excessive water pressure, so it is necessary to reasonably control the water pressure and select pipe materials with good pressure resistance."

[0360] = "Common causes of water pipe rupture include material deterioration, improper installation, excessive water pressure, etc., and prevention needs to be carried out from multiple aspects."

[0361] Among them, emphasizes material and anti-corrosion, highlights water pressure control, and proposes preventive measures from a comprehensive perspective.

[0362] To further optimize the prompt, the following two factors were considered:

[0363] Coherence: Calculate the cosine similarity between adjacent nodes in the candidate paths to obtain the coherence scores of each prompt as follows:

[0364] ;

[0365] ;

[0366] ;

[0367] Diversity: Calculate the average Jaccard distance between three prompts to obtain a diversity score of 0.71.

[0368] Add the coherence and diversity scores to the prompt generation formula to finally obtain three optimized candidate prompts:

[0369] To prevent the water pipe from bursting, first understand the material properties of the water pipe and select appropriate pipe materials and anti-corrosion measures;

[0370] The root cause of water pipe bursting is often excessive water pressure Therefore, it is necessary to reasonably control the water pressure and select pipe materials with good pressure resistance;

[0371] Common causes of water pipe bursting include material deterioration, improper installation, excessive water pressure, etc., and prevention needs to be carried out from multiple aspects;

[0372] Step S70: Precise prompt acquisition

[0373] Present the three optimized candidate prompts to user Zhang San and let him select according to his own needs.

[0374] User Zhang San selected as the final precise prompt: "To prevent the water pipe from bursting, first understand the material properties of the water pipe and select appropriate pipe materials and anti-corrosion measures."

[0375] This prompt has been optimized through the foregoing steps and can better meet Zhang San's actual needs. Next, the dialogue system uses this prompt as input and provides corresponding answers to the user through the answer generation module.

[0376] The answer generation module mainly includes the following steps:

[0377] 1. Input the prompt selected by the user, "To prevent water pipe rupture, first understand the material properties of the water pipes, select appropriate pipe materials and anti-corrosion measures", into the question-answering generation model based on the large language model.

[0378] 2. The model has been pre-trained for prompt-answer pairs and can generate smooth and coherent answers according to the input prompt.

[0379] 3. The generated answer is as follows:

[0380] "The water pipe material is the key factor in preventing water pipe rupture. Common water pipe materials include PVC, copper, stainless steel, etc., and each material has its own characteristics. PVC pipes have good corrosion resistance but low strength and are prone to rupture under high water pressure; copper pipes have high strength but are easily corroded; stainless steel pipes have better comprehensive performance and are an ideal choice. In addition to selecting the appropriate pipe material, it is also important to do daily anti-corrosion maintenance. Measures such as anti-corrosion coatings and anti-corrosion boxes can be adopted to effectively extend the service life of the water pipes and reduce water leakage and rupture problems caused by corrosion. In short, starting from the two aspects of water pipe material properties and anti-corrosion measures can better prevent the occurrence of water pipe rupture."

[0381] 4. Return the generated answer to user Zhang San to complete the entire question-answering interaction process.

[0382] Through the above implementation process, it can be seen how the method for obtaining accurate prompts proposed by the present invention works in practical applications:

[0383] First, the question understanding module can comprehensively extract the semantic features of the user's question and generate prompts that are close to the user's actual needs through means such as expansion and knowledge modeling. This provides a good foundation for subsequent answer generation.

[0384] Secondly, the prompt generation process takes into account two factors, coherence and diversity, to ensure that the generated prompts are not only semantically coherent but also can meet the user's needs for differentiated choices. This improves the user experience to a certain extent.

[0385] Finally, the dialogue system can generate more targeted answer content according to the accurate prompt selected by the user. Compared with the answers generated directly based on the original question, this method can better understand the user's true intention and give more in-line solutions to the needs.

[0386] Generally speaking, through the application of the method of the present invention, the dialogue system has made remarkable progress in improving service quality and enhancing user satisfaction. This technical path based on semantic analysis, knowledge modeling, and prompt optimization provides a valuable reference solution for similar dialogue system applications.

[0387] Of course, in actual applications, further optimization and adjustment are required for different scenarios. For example, it is possible to consider combining domain knowledge in the knowledge graph to perform more professional and detailed processing on the generation of prompts and the content of answers; it is also possible to try methods such as reinforcement learning to continuously optimize the system performance through interaction feedback with users.

[0388] The above is only a specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered within the protection scope of the present invention.

Claims

1. A method for obtaining a user's precise prompt, characterized in that: The following steps are involved: S10, obtaining questions raised by the user to the large language model as basic questions, and extracting semantic features of the basic questions; S20, performing synonym expansion, context association analysis, and domain knowledge supplementation according to the semantic features to obtain an augmented semantic feature set; S30, establishing an augmented semantic feature graph according to the augmented semantic feature set, wherein the nodes of the augmented semantic feature graph represent semantic concepts or keywords, and the edges represent semantic relationships or logical connections between nodes; S40, according to the context input by the user, using the large language model to analyze the user's intention and needs, identify key concepts and constraints, assign weights and prioritize nodes in the augmented semantic feature graph, and obtain a weighted semantic feature graph; S50, using a depth-first search algorithm for the weighted semantic feature graph to obtain multiple candidate paths of the weighted semantic feature graph; S60, generating a corresponding prompt according to each candidate path using a prompt generation algorithm based on the path, and combining each prompt with a pre-adjusted prompt generation question model to generate a question, obtain a question list and output it to the user; S70, obtaining the question selected by the user according to the question list as the selected question, and using the prompt corresponding to the selected question as the user's precise prompt; The path-based prompt generation algorithm is specifically expressed as follows: ; In the formula, For the A prompt for candidate paths to be generated; For the candidate paths; Fill in the template with functions; is the coherence scoring function; is the diversity scoring function; is the weight coefficient; This is achieved through a predefined prompt template, where the nodes in the path are filled into the placeholders in the template. It is obtained by calculating the semantic similarity of adjacent nodes in the path; It is obtained by calculating the average Jaccard distance between the current path and other paths.

2. A method for obtaining a user's precise prompt according to claim 1, characterized in that: The prompt question generation model is based on the large language model and is obtained by fine-tuning and training a large number of preset high-quality prompt-question pairs.

3. A method for obtaining a user's precise prompt according to claim 1, characterized in that: The semantic features of the basic question are extracted, which are specifically expressed as follows: ; In the formula, is the extracted semantic feature vector; For basic questions; For the feature extraction function; For the The weight of each feature; is the number of feature extraction functions; Embedding representation for the problem; is the weight coefficient of the embedding representation.

4. A method for obtaining a user's precise prompt according to claim 3, characterized in that: The steps of obtaining the augmented semantic feature set are specifically expressed as follows: ; In the formula, is the expanded semantic feature set; is the original semantic feature vector; Extension functions for synonyms; It is the context association analysis function; Supplement functions for domain knowledge; is the weight coefficient of each expansion method.

5. A method for obtaining a user's precise prompt according to claim 4, characterized in that: The augmented semantic feature graph is specifically expressed as follows: ; ; ; ; In the formula, To augment the semantic feature map; is a node set, corresponding to the expanded semantic features; is the edge set; is the edge weight matrix; For Node and The similarity function between is the similarity threshold, Represents an index.

6. A method for obtaining a user's precise prompt according to claim 5, characterized in that: The weighted semantic feature map is specifically expressed as follows: ; ; ; In the formula, is the weighted semantic feature map; is the new edge weight matrix; is matrix element-wise multiplication; To adjust the matrix; is the sigmoid function; For Node 's relevance score; For Node 's relevance score; For Node Importance score.

7. A method for obtaining a user's precise prompt according to claim 6, characterized in that: The depth-first search algorithm is specifically expressed as follows: ; ; In the formula, is a set of candidate paths; is the starting node; is the termination node; is the maximum path length; For Node The set of neighbor nodes.

8. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores program instructions, and when the program instructions are executed in a computer, the method for obtaining a precise prompt from a user as described in any one of claims 1 to 7 is used to execute the method.

9. A system for obtaining accurate prompts from users, characterized in that: Contains the computer-readable storage medium of claim 8.

Citation Information

Patent Citations

  • Method, device and equipment for generating chat reply based on dynamic Prompt and medium

    CN115599896A

  • Method and device for training generative large language model based on knowledge base feedback

    CN117009490A