Automatic prompt construction method based on human-computer dialogue history and semantic retrieval

Through the automatic construction method of prompt based on human-computer dialogue history and semantic retrieval, the pre-trained language model and convex optimization method are used to extract relevant context fragments and construct dynamic Prompt templates, which solves the problem of lack of context adaptability of generated replies in multiple rounds of dialogue, and achieves high semantic correlation and logical coherence of generated content.

CN119441443BActive Publication Date: 2025-05-16BEIJING QIBU TIANXIA TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510037633.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-10
Publication Date
2025-05-16
Estimated Expiration
2045-01-10

AI Technical Summary

Technical Problem

In multiple rounds of dialogue scenarios, it is difficult for the prior art to accurately extract relevant context information, resulting in the generated reply lacking context adaptability and even appearing stiff and unnatural.

Method used

Through the propt automatic construction method based on human-computer dialogue history and semantic retrieval, the pre-trained language model is used to map historical dialogue and user requests into semantic vectors, and match them with keyword weights to build a collection of historical vectors and user request vectors. Then, the context fragments that are most relevant to user requests are extracted based on the convex optimization method, and the context is filtered through sparse optimization problems, and finally a dynamic Prompt template is constructed and a generative language model is input to generate a reply.

Benefits of technology

It significantly improves the semantic correlation and context adaptability of the generated content, makes the reply more natural and reasonable, and can smoothly connect historical dialogue content, improving semantic accuracy and logical coherence in multiple rounds of dialogue scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119441443B_ABST
    Figure CN119441443B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of artificial intelligence technology, and discloses a prompt automatic construction method based on human-computer dialogue history and semantic retrieval, comprising the following steps: S1, text preprocessing of historical dialogue data and user input, including dynamically generating query keywords, verifying and parsing user uploaded files, and extracting core content; S2, mapping historical dialogues and user requests into semantic vectors using a pre-trained language model, matching them in combination with keyword weights, and constructing a historical vector set and a user request vector; S3, extracting the context fragment most relevant to the user request from the historical dialogue vector based on a convex optimization method. Context information highly relevant to the user request is accurately extracted through sparse optimization technology, and combined with dynamic prompt template design, background information, user request and generated instructions are organically integrated, thereby improving the semantic relevance and context coherence of generated replies and avoiding interference from redundant and irrelevant information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a method for automatically constructing prompts based on human-computer dialogue history and semantic retrieval. Background Art

[0002] With the continuous development of artificial intelligence, multi-turn dialogue systems have gradually become an important application core in the fields of intelligent customer service, virtual assistants, etc. The main goal of these systems is to interact efficiently with users through natural language understanding and generation. However, in multi-turn dialogue scenarios, how to use the dialogue history to generate high-quality natural language responses has always been a technical problem that needs to be solved urgently. Although existing methods have made progress to a certain extent, they still have many limitations in dealing with complex contexts and generating accurate and coherent responses.

[0003] Traditional multi-round dialogue systems usually rely on rule templates or semantic vector retrieval technology. Although rule templates are clearly structured, they are static and cannot adapt to the diverse needs of users, especially when dealing with dynamic contexts and complex dialogue scenarios. Although semantic vector retrieval can screen historical dialogues to a certain extent, it is not accurate enough in extracting the relevance of information in high-dimensional semantic space and is prone to introducing redundant or irrelevant content. These problems often make the generated responses lack context adaptability and may even appear stiff and unnatural.

[0004] In addition, in multi-turn dialogue systems, the dynamics and diversity of historical dialogues pose significant challenges to context processing. The system needs to quickly filter out the most relevant information to the current request in the dialogue history while avoiding redundancy and noise interference. However, simple linear combination or retrieval methods often lead to excessive or insufficient information screening, affecting the coherence and focus of the generated content. More importantly, the existing technology lacks a dynamic control mechanism for the generation quality, and cannot optimize the accuracy and adaptability of the generated content in real time.

[0005] Therefore, in order to solve the above problems, the present invention proposes a prompt automatic construction method based on human-computer dialogue history and semantic retrieval. Summary of the invention

[0006] In view of the shortcomings of the prior art, the present invention provides a method for automatically constructing prompts based on human-computer dialogue history and semantic retrieval, which solves the problem of how to accurately extract relevant context information and dynamically construct the optimal prompt in a multi-round dialogue system to generate semantically coherent, content-related and natural and reasonable responses.

[0007] To achieve the above objectives, the present invention is implemented through the following technical solutions: a prompt automatic construction method based on human-computer dialogue history and semantic retrieval, comprising the following steps:

[0008] S1. Perform text preprocessing on historical conversation data and user input, including dynamically generating query keywords, verifying and parsing user uploaded files, and extracting core content;

[0009] S2. Use the pre-trained language model to map historical conversations and user requests into semantic vectors, match them with keyword weights, and construct a historical vector set and user request vector;

[0010] S3, extract the most relevant context fragments from the historical conversation vector based on the convex optimization method, combine the highly relevant content captured by the web crawler, and filter the context through the sparse optimization problem;

[0011] S4, build a dynamic prompt template, integrate the extracted context fragments, user requests and generation instructions to generate prompts;

[0012] S5, inputting the generated Prompt into the generative language model for processing to generate a natural language reply text;

[0013] S6. Optimize the sparse optimization parameters and prompt template design according to the generation quality evaluation results.

[0014] Preferably, the step S1 includes the following sub-steps:

[0015] S1.1. Perform sentence and word segmentation on historical conversation data and user input, filter stop words, and perform part-of-speech tagging and dependency analysis;

[0016] S1.2. The system dynamically generates query keywords based on user input, and selects highly relevant keywords based on semantic similarity to ensure accurate retrieval;

[0017] S1.3. Verify the file formats uploaded by users, including PDF, DOCX, and Markdown, clean up irrelevant characters, and extract the core content of the file for subsequent knowledge base retrieval.

[0018] Preferably, the step S2 includes the following sub-steps:

[0019] S2.1, map the historical conversation data into vector representation through the sentence embedding model to form a historical vector set V = {,,…,};

[0020] S2.2, map the user request text into a semantic vector, combine it with the dynamically generated keyword weights, and calculate the relevance score based on the word frequency and position weights to optimize the vectorization results;

[0021] S2.3. Construct the historical conversation semantic vector set V and the user request vector.

[0022] Preferably, the process of extracting the context fragment most relevant to the user request in step S3 includes the following contents:

[0023] S3.1. The system simulates the behavior of search engines through web crawlers, crawls web page content in real time, and extracts highly relevant text data;

[0024] S3.2, construct a sparse optimization problem, whose objective function is:

[0025] Among them, is the user request vector, is the historical conversation vector, = is the sparse coefficient vector, and λ is the sparse regularization parameter;

[0026] S3.3. The system has a built-in exception handling mechanism that supports automatic retry of network requests and records exception logs.

[0027] Preferably, the sparse optimization problem in step S3.2 is solved by an alternating direction multiplier method, which specifically includes:

[0028] Update the sparse coefficient α:

[0029] Among them, is the sparse coefficient vector optimized in the k+1th iteration, α is the sparse coefficient vector to be optimized in the current iteration, is the semantic vector representation of the user request, is the semantic vector representation of the historical dialogue, N is the total number of rounds of historical dialogue, is the linear combination of the historical dialogue vectors, fits the user request vector, is the reconstruction error term, ρ is the penalty factor, is the constraint term, is the sparse variable updated in the kth iteration, and is the Lagrange multiplier in the kth iteration;

[0030] Update the sparsity constraint variable z:

[0031] Where, is the sparse variable updated in the k+1th iteration, (·,τ) is the soft threshold function, is the sparse coefficient vector optimized in the k+1th iteration, is the Lagrange multiplier in the kth iteration, λ is the regularization parameter, and ρ is the penalty factor;

[0032] Update the Lagrange multiplier μ:

[0033] Among them, is the Lagrange multiplier updated in the k+1th iteration, is the Lagrange multiplier in the kth iteration, ρ is the penalty factor, is the sparse coefficient vector optimized in the k+1th iteration, and is the sparse variable updated in the k+1th iteration;

[0034] Repeat the iteration until the sparse coefficients meet the preset convergence conditions.

[0035] Preferably, the sparse optimization result in step S3.2 is a sparse coefficient vector, from which a set of historical conversation segments corresponding to non-zero weights is screened out, and its expression is:

[0036] The most relevant context snippet for the user request.

[0037] Preferably, the construction of the dynamic prompt template in step S4 includes:

[0038] Extracting background information: sorting the extracted historical conversation fragments by chronological order or semantic importance;

[0039] User request content: Integrate the request text currently entered by the user into the template;

[0040] Instruction design: Clearly generate instruction content, including tone, style, and length limits.

[0041] Preferably, the step S5 includes the following specific steps:

[0042] S5.1. Input the generative language model into the generated Prompt and trigger the generation task;

[0043] S5.2. Based on the historical dialogue fragments and semantic retrieval results, the output of the generated model is verified for contextual consistency, and the best output is selected by calculating the semantic relevance score between the generated text and the input prompt;

[0044] S5.3. Processing the output text of the generated model according to a preset format, including automatic sentence segmentation, punctuation adjustment, and weighted sorting of key content;

[0045] S5.4. Evaluate the generation results, adjust the prompt content according to the real-time feedback of the generation quality, and re-execute the generation steps until the output quality requirements are met.

[0046] Preferably, the evaluation of the generation quality in step S6 includes the following indicators:

[0047] Semantic relevance: Calculate the cosine similarity between the user request vector and the generated reply vector;

[0048] Contextual coherence: The matching degree between the generated text and the reference text is evaluated by BLEU and ROUGE-L indicators;

[0049] User satisfaction: Subjective evaluation of the generated results based on user ratings.

[0050] Preferably, the calculation formula of cosine similarity in step S6 is as follows:

[0051] Among them, represents the semantic similarity between the user request vector and the generated reply vector, is the semantic vector representation of the user request, is the semantic vector representation of the generated reply, · represents the dot product of and , is the Euclidean norm of the user request vector, and is the Euclidean norm of the generated reply vector.

[0052] The present invention provides a prompt automatic construction method based on human-computer dialogue history and semantic retrieval. It has the following beneficial effects:

[0053] 1. The present invention constructs a sparse optimization problem to filter out the context information most relevant to the user request from the historical conversations and dynamically constructs the Prompt input. The sparse representation method can effectively select the conversation segments with the greatest semantic contribution to the generation task in the high-dimensional vector space, while eliminating the interference of irrelevant or redundant information. Compared with traditional rule matching or simple vector retrieval methods, the semantic relevance and context adaptability of the generated content are significantly improved. Combined with context management technology and semantic consistency evaluation, the generated replies are not only more natural and reasonable, but also can smoothly connect the historical conversation content, so that it can show higher semantic accuracy and logical coherence in multi-round conversation scenarios.

[0054] 2. The present invention integrates the screened context fragments, user requests and clear generation instructions into a standardized input structure through dynamic Prompt template design, and combines convex optimization theory to ensure the global optimality of context fragment selection. The output results of the generation model are accurately controlled through the dynamic parameters of the generation instructions in the Prompt template, so that the generated content can be dynamically adjusted according to different task requirements and dialogue scenarios. This method significantly improves the coherence and logic of the generation results, and enhances the adaptability of the generation model to complex scenarios. In terms of semantic coherence, content relevance and generation efficiency, the present invention achieves multi-dimensional optimization, so that the generated replies are more in line with user needs and have a higher interactive experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] Figure 1 Flow chart of the method of the present invention. DETAILED DESCRIPTION

[0056] The following will be combined with the drawings in the specification of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0057] Please refer to the attached Figure 1 The embodiment of the present invention provides a prompt automatic construction method based on human-computer dialogue history and semantic retrieval, comprising the following steps:

[0058] S1. Perform text preprocessing on historical conversation data and user input, including dynamically generating query keywords, verifying and parsing user uploaded files, and extracting core content;

[0059] The S1 step includes the following sub-steps:

[0060] S1.1. Perform sentence and word segmentation on historical conversation data and user input, filter stop words, and perform part-of-speech tagging and dependency analysis;

[0061] S1.2. The system dynamically generates query keywords based on user input, and selects highly relevant keywords based on semantic similarity to ensure accurate retrieval;

[0062] S1.3. Verify the file formats uploaded by users, including PDF, DOCX, and Markdown, clean up irrelevant characters, and extract the core content of the file for subsequent knowledge base retrieval.

[0063] Specifically, in this embodiment, the historical conversation data and user input are firstly processed by sentence and word segmentation to ensure the accuracy of subsequent language processing operations. The specific steps are as follows:

[0064] The sentence segmentation algorithm is used to segment the input text sentence by sentence. A rule-based segmentation method based on punctuation marks such as periods, question marks, and commas is adopted, and the segmentation strategy is dynamically adjusted in combination with an adaptive language model. The word segmentation model is used to segment each sentence after text segmentation, and part-of-speech tagging technology is used to assign grammatical labels to each word. Dependency parsing technology is introduced to construct a syntactic tree to analyze the grammatical structure within the sentence, extract the subject-verb-object relationship and modifying components in the sentence, and provide structured support for subsequent keyword generation.

[0065] Extract the core words in the user input text and build a candidate keyword set K based on historical conversation data;

[0066] Calculate the semantic similarity of candidate keywords. The formula is as follows:

[0067] Among them, represents the semantic vector of the keyword, represents the semantic vector of the user input text, and is the cosine similarity between the two;

[0068] Filter highly relevant keywords based on the semantic similarity scores of keywords and build a high-priority keyword set;

[0069] The filtered keyword set is applied to the semantic retrieval module to obtain contextual information that is highly relevant to the user request.

[0070] Verify that the format of the file uploaded by the user is a supported type, including PDF, DOCX, and Markdown;

[0071] For PDF files, use optical character recognition (OCR) technology to extract file content and recognize text in embedded images;

[0072] For DOCX files, parse the text content, tables, charts, and comments of the file, and remove formatting marks and redundant characters;

[0073] For Markdown files, parse their structured content and extract information such as titles, paragraphs, and code blocks;

[0074] All parsed text content is preprocessed, including removing irrelevant characters, repeated content, and meaningless stop words; the preprocessed core text content is vectorized, and the formula is as follows:

[0075] Among them, is the vector representation of the core content of the file, is the semantic vector of the iii-th important word in the file, is the importance weight of the word, and m is the number of important words in the file.

[0076] Through the preprocessing of historical conversation data and user input in this embodiment, not only can context information highly related to user requests be effectively extracted, but also external resources provided by users can be fully utilized during the file parsing process to provide high-quality input data for subsequent Prompt construction. At the same time, the combination of word segmentation, keyword generation and file parsing ensures the versatility and adaptability of the method, which is applicable to a variety of input data types and scenario requirements.

[0077] S2. Use the pre-trained language model to map historical conversations and user requests into semantic vectors, match them with keyword weights, and construct a historical vector set and user request vector;

[0078] Step S2 includes the following sub-steps:

[0079] S2.1, map the historical conversation data into vector representation through the sentence embedding model to form a historical vector set V = {,,…,};

[0080] S2.2, map the user request text into a semantic vector, combine it with the dynamically generated keyword weights, and calculate the relevance score based on the word frequency and position weights to optimize the vectorization results;

[0081] S2.3. Construct the historical conversation semantic vector set V and the user request vector.

[0082] Specifically, in this embodiment, the construction of the historical vector set and the user request vector is completed by using the pre-trained language model combined with the keyword optimization technology, which specifically includes the following steps:

[0083] In this embodiment, the historical conversation data is semantically vectorized by using a sentence embedding model, which specifically includes the following steps: the system calls a pre-trained language model based on the Transformer architecture (such as BERT or Sentence-BERT), processes each conversation in the historical conversation data, and generates a corresponding sentence embedding vector;

[0084] Define a historical vector set, where N represents the semantic vector of the i-th historical conversation and N is the number of historical conversations;

[0085] The system normalizes the generated historical conversation semantic vectors. The normalization formula is:

[0086] Among them, represents the modulus of the vector, and normalization ensures the scale consistency of the semantic vector;

[0087] The normalized vector set V is stored for use in semantic matching of subsequent user request vectors.

[0088] In this embodiment, the system maps the user request text into a semantic vector and optimizes the semantic vectorization result in combination with the keyword weight. The specific steps are as follows:

[0089] The system calls the generate_query_keywords function to dynamically generate a preliminary keyword set K= according to the user request text u; generates a corresponding semantic vector for each keyword, and calculates the semantic similarity of the keyword in combination with the semantic vector of the user request text. The semantic similarity formula is:

[0090] Among them, represents the semantic vector of the user request text, and represents the semantic similarity between the keyword and the user request;

[0091] The system selects keywords with semantic similarity higher than the threshold to form a set of highly relevant keywords;

[0092] The system calls the weighted_keyword_matching function to set a weight for each keyword in the keyword set, and calculates the comprehensive score based on the word frequency and position weight. The weight calculation formula is as follows:

[0093] Among them, is the frequency value of the keyword in the user request text, is the relative position weight of the keyword, and max(p) is the maximum value of the frequency and position weight, which is used for normalization;

[0094] The semantic vector optimization formula for optimizing user request text according to keyword weight is:

[0095] Among them, is the optimized user request semantic vector, is the weight of the keyword, and is the semantic vector of the keyword.

[0096] In this embodiment, by combining the historical conversation semantic vector set V and the optimized user request vector, the semantic vector set is constructed. The specific steps are as follows:

[0097] Ensure that the history vector set and the user request vector are in the same semantic vector space;

[0098] The system calculates the semantic similarity between the user request vector and each vector in the historical conversation vector set V. The similarity calculation formula is:

[0099] Among them, represents the semantic similarity between the optimized user request vector and the semantic vector of the i-th historical conversation;

[0100] The historical conversation vectors are sorted from high to low according to the semantic similarity score, and several vectors V with high relevance are screened out; the system passes the optimized user request vector and the screened high-relevance historical vector set to the subsequent context extraction and prompt generation module.

[0101] Dynamic keyword generation and optimization are achieved by calling the generate_query_keywords and weighted_keyword_matching functions, and the semantic vector optimization method of the user request text is combined to build a historical conversation semantic vector set and an optimized user request vector. With the dynamic adjustment of keyword weights and the semantic similarity calculation mechanism, the accuracy and precision of semantic matching are significantly improved, laying a high-quality semantic foundation for subsequent prompt generation and context extraction.

[0102] S3, extract the most relevant context fragments from the historical conversation vector based on the convex optimization method, combine the highly relevant content captured by the web crawler, and filter the context through the sparse optimization problem;

[0103] The process of extracting the most relevant context fragments for the user request in the S3 step includes the following:

[0104] S3.1. The system simulates the behavior of search engines through web crawlers, crawls web page content in real time, and extracts highly relevant text data;

[0105] S3.2, construct a sparse optimization problem, whose objective function is:

[0106] Among them, is the user request vector, is the historical conversation vector, = is the sparse coefficient vector, and λ is the sparse regularization parameter;

[0107] S3.3. The system has a built-in exception handling mechanism that supports automatic retry of network requests and records exception logs.

[0108] The sparse optimization problem in step S3.2 is solved by the alternating direction multiplier method, which includes:

[0109] Update the sparse coefficient α:

[0110] Among them, is the sparse coefficient vector optimized in the k+1th iteration, α is the sparse coefficient vector to be optimized in the current iteration, is the semantic vector representation of the user request, is the semantic vector representation of the historical dialogue, N is the total number of rounds of historical dialogue, is the linear combination of the historical dialogue vectors, fits the user request vector, is the reconstruction error term, ρ is the penalty factor, is the constraint term, is the sparse variable updated in the kth iteration, and is the Lagrange multiplier in the kth iteration;

[0111] Update the sparsity constraint variable z:

[0112] Where, is the sparse variable updated in the k+1th iteration, (·,τ) is the soft threshold function, is the sparse coefficient vector optimized in the k+1th iteration, is the Lagrange multiplier in the kth iteration, λ is the regularization parameter, and ρ is the penalty factor;

[0113] Update the Lagrange multiplier μ:

[0114] Among them, is the Lagrange multiplier updated in the k+1th iteration, is the Lagrange multiplier in the kth iteration, ρ is the penalty factor, is the sparse coefficient vector optimized in the k+1th iteration, and is the sparse variable updated in the k+1th iteration;

[0115] Repeat the iteration until the sparse coefficients meet the preset convergence conditions.

[0116] The sparse optimization result in step S3.2 is a sparse coefficient vector, from which the set of historical dialogue segments corresponding to non-zero weights is screened out, and its expression is:

[0117] The most relevant context snippet for the user request.

[0118] Specifically, in this embodiment, the system uses the Crawler class to implement web crawling to crawl web page content that is highly relevant to the user's request, which specifically includes the following steps:

[0119] Call the Crawler class to simulate the search engine behavior, build a query link through the dynamic URL splicing mechanism, and generate the user's request keyword and the search engine's complete query address;

[0120] Use delay mechanism and random request header strategy to avoid being blocked by search engines. The system simulates real user behavior by setting dynamic interval time to reduce the probability of crawling failure.

[0121] Parse the HTML pages returned by the search engine and extract the main content, including the title, abstract and body fragment, and use these contents as external supplementary information for subsequent context expansion;

[0122] To improve the stability of network requests, the system introduces an exponential backoff mechanism in the _send_request method. When a network request fails, the system will extend the retry interval exponentially until the preset maximum number of retries is reached.

[0123] The system records exception information in each crawling task and uses a logging mechanism to save exceptions in the crawling process (such as timeouts, link failures, etc.) to log files to facilitate subsequent debugging and problem location.

[0124] In this embodiment, the system extracts the context fragment most relevant to the user request from the historical conversation vector by a sparse optimization method, which specifically includes the following steps:

[0125] Define the objective function of the sparse optimization problem as follows:

[0126] Among them, is the user request vector, is the historical conversation vector, = is the sparse coefficient vector, and λ is the sparse regularization parameter;

[0127] The reconstruction error term is introduced to ensure that the optimized historical conversation vector combination can fit the user request vector to the greatest extent;

[0128] Through the sparse regularization term, the minimum number of historical conversation vectors are selected to participate in the reconstruction of the user request vector, thereby optimizing the efficiency of context extraction;

[0129] Combine the historical conversation vector set V and the web page content semantic vector obtained by the crawler to jointly construct a context vector set for subsequent screening.

[0130] In this embodiment, the sparse optimization problem is solved by the alternating direction multiplier method, which specifically includes the following steps:

[0131] Update the sparse coefficient α:

[0132] Among them, is the sparse coefficient vector optimized in the k+1th iteration, α is the sparse coefficient vector to be optimized in the current iteration, is the semantic vector representation of the user request, is the semantic vector representation of the historical dialogue, N is the total number of rounds of historical dialogue, is the linear combination of the historical dialogue vectors, fits the user request vector, is the reconstruction error term, ρ is the penalty factor, is the constraint term, is the sparse variable updated in the kth iteration, and is the Lagrange multiplier in the kth iteration.

[0133] Update the sparsity constraint variable z:

[0134] Among them, is the sparse variable updated in the k+1th iteration, (·,τ) is the soft threshold function, is the sparse coefficient vector optimized in the k+1th iteration, is the Lagrange multiplier in the kth iteration, λ is the regularization parameter, and ρ is the penalty factor.

[0135] Update the Lagrange multiplier μ:

[0136] Among them, is the Lagrange multiplier updated in the k+1th iteration, is the Lagrange multiplier in the kth iteration, ρ is the penalty factor, is the sparse coefficient vector optimized in the k+1th iteration, and is the sparse variable updated in the k+1th iteration.

[0137] Iteration convergence conditions:

[0138] The above optimization process is repeated iteratively until the sparse coefficient α meets the convergence condition:

[0139] Among them, ∈ is the preset convergence threshold.

[0140] In this embodiment, the system extracts the context fragment most relevant to the user request through the sparse optimization result, which specifically includes the following steps:

[0141] Get the sparse coefficient vector from the result of the sparse optimization problem, filter the historical dialogue segments corresponding to non-zero weights, and extract the formula as follows:

[0142] Among them, represents the set of historical conversation segments most relevant to the user request, and is the index corresponding to the non-zero weight in the sparse coefficient; the extracted historical conversation segments and the highly relevant text content obtained by the crawler are spliced ​​into a complete context vector and passed to the subsequent generation module;

[0143] Use ContextManager to dynamically expand the context, stitching together historical conversation fragments, crawler results, and the user's current input into a complete prompt to ensure the coherence and accuracy of the generated content.

[0144] S4, build a dynamic prompt template, integrate the extracted context fragments, user requests and generation instructions to generate prompts;

[0145] The construction of the dynamic prompt template in step S4 includes:

[0146] Extracting background information: sorting the extracted historical conversation fragments by chronological order or semantic importance;

[0147] User request content: Integrate the request text currently entered by the user into the template;

[0148] Instruction design: Clearly generate instruction content, including tone, style, and length limits.

[0149] Specifically, in this embodiment, the system constructs the background information part of the Prompt template according to the extracted context fragment set, which specifically includes the following steps:

[0150] According to the context fragment set, the fragments are first sorted in chronological order to ensure the coherence of context semantics;

[0151] If the user request involves specific keywords or topics, the system will filter the fragment collection based on semantic similarity and only retain the fragment collection that is highly relevant to the user request. The calculation formula for semantic similarity is as follows:

[0152] Among them, is the semantic vector of the user request, and is the semantic vector of the context segment.

[0153] The system further sorts the fragments according to their importance. The importance of a fragment can be assessed based on its frequency of occurrence in context or its degree of match with a keyword, ultimately forming a collection of highly relevant background information.

[0154] The filtered background information is used as the front part of the Prompt template to provide a clear semantic context for the generation model.

[0155] In this embodiment, the system integrates the request text currently input by the user into the Prompt template to ensure that the generated content meets the actual needs of the user, which specifically includes the following steps:

[0156] Extract the user input text and use it as the core part of the Prompt template, placing it after the background information;

[0157] If the user request content contains clear generation intentions or task instructions (such as "generate a summary" or "answer questions"), the system extracts these task instructions and displays them as an important part of the template;

[0158] The combination logic of user requests and background information must ensure semantic consistency. The system will set a clear separation in the template to ensure that the generated model can effectively distinguish between background and request content.

[0159] In this embodiment, the system designs clear generation instructions for the Prompt template to guide the tone, style, output length, etc. of the generation model, which specifically includes the following steps:

[0160] Dynamically adjust the tone and style of generated instructions based on the scenario requirements of the user's request. Instruction design rules include:

[0161] If the user's request involves formal content (such as academic papers or business reports), the system will generate instructions designed to "use a formal and professional tone";

[0162] If the user's request is biased towards daily communication or interaction scenarios, the instructions will prompt the generation model to adopt a "concise and friendly tone"; set a length limit for the generated content based on the context of the user's input text, and specify the constraints in the instructions. For example:

[0163] "Please limit your output to 300 words or less."

[0164] “The output should include the following points: 1. Background of the problem; 2. Detailed analysis; 3. Conclusion.”

[0165] The content of the instruction is adapted to the capabilities of the generation model according to the user's task requirements (e.g., whether a detailed explanation or a list of specific steps is required), ensuring that the prompt can effectively guide the generation task.

[0166] In this embodiment, the system adjusts the structure and content of the Prompt template in real time according to the dynamic changes of user input and context content, which specifically includes the following steps:

[0167] Dynamically expand context content: Combined with the multi-turn dialogue manager, it extracts the most relevant fragments in the current dialogue context and splices the historical dialogue content with the latest user input;

[0168] Dynamically integrate external search results: If the system obtains external content related to the user's request through web crawlers or semantic search, the search results will be added to the background information part of the prompt template according to priority;

[0169] During multiple runs of the generative model, the Prompt template is optimized based on the quality of the generated results, such as adjusting the number of background clips or modifying the tone and style cues in the generated instructions.

[0170] S5, inputting the generated Prompt into the generative language model for processing to generate a natural language reply text;

[0171] Step S5 includes the following specific steps:

[0172] S5.1. Input the generative language model into the generated Prompt and trigger the generation task;

[0173] S5.2. Based on the historical dialogue fragments and semantic retrieval results, the output of the generated model is verified for contextual consistency, and the best output is selected by calculating the semantic relevance score between the generated text and the input prompt;

[0174] S5.3. Processing the output text of the generated model according to a preset format, including automatic sentence segmentation, punctuation adjustment, and weighted sorting of key content;

[0175] S5.4. Evaluate the generation results, adjust the prompt content according to the real-time feedback of the generation quality, and re-execute the generation steps until the output quality requirements are met.

[0176] S6. Optimize the sparse optimization parameters and prompt template design according to the generation quality evaluation results.

[0177] Specifically, in this embodiment, the system first inputs the dynamically generated Prompt into the generative language model, which specifically includes the following steps:

[0178] Prompt input and calling generation tasks: The system directly uses the built Prompt template as input and calls the generative language model to perform the generation task;

[0179] Dynamic context expansion: Use ContextManager to dynamically manage the user's historical conversations, search results, and current request content to ensure that the prompt contains complete context information. The logic of context splicing is:

[0180] Among them, represents the historical conversation fragment, represents the retrieved related content, and represents the user's current request;

[0181] Triggering the generation task: After the generative language model receives the prompt, it starts the generation task and generates the corresponding natural language response based on the prompt content.

[0182] In this embodiment, the system verifies the semantic consistency of the output of the generative language model, and selects the optimal output result by calculating the semantic relevance between the generated result and the prompt content, which specifically includes the following steps:

[0183] Semantic relevance calculation: Score the semantic relevance between the output text TgT_gTg of the generation model and the prompt content PPP. The calculation formula for the semantic relevance score is:

[0184] Among them, and represent the semantic vectors of generated text and prompt content respectively;

[0185] Context consistency verification: Further verify the context consistency of the content in the generated text to ensure its semantic match with the historical conversation fragments and user request content;

[0186] Filter the best output: Filter the generated results from high to low according to the semantic relevance score, and select the best output text as the final generated result.

[0187] In this embodiment, the system processes the output text of the generated model according to a preset format to ensure the semantic clarity and readability of the text, which specifically includes the following steps:

[0188] Automatic sentence segmentation: segment the sentences of the generated text to ensure clear paragraph structure and easy reading;

[0189] Punctuation adjustment: Correct punctuation errors in generated text to ensure grammatical standardization;

[0190] Weighted sorting of key content: Sort the key information in the generated text according to semantic importance, and put the key content at the front of the text. The formula for weighted sorting is as follows:

[0191] Among them, is the weight of the content, indicating the semantic relevance between the content and the user request, indicates the position of the content in the generated text, and α and β are weight parameters.

[0192] In this embodiment, the system dynamically optimizes the Prompt content based on the generated quality feedback, specifically including the following steps:

[0193] Generation quality assessment: The system evaluates the quality of the generated text, including the following indicators:

[0194] Semantic relevance: evaluates the semantic consistency of the generated text with the user request;

[0195] Contextual coherence: evaluates the semantic coherence between the generated text and the historical dialogue fragments;

[0196] User satisfaction: based on subjective feedback ratings from users;

[0197] Prompt content optimization: Dynamically adjust the prompt template content based on the generation quality assessment results, for example:

[0198] Add or subtract background information snippets;

[0199] Modify the tone or style requirements of the generated instructions;

[0200] Adjust the length limit of generated content;

[0201] Re-execute the generation task: If the generation quality does not meet the expected standard, the system re-optimizes the prompt and triggers the generation task again until the generation result meets the quality requirements.

[0202] In this embodiment, the system realizes the gradual return of the generated results through a streaming response mechanism to reduce the user's waiting time, which specifically includes the following steps:

[0203] Streaming data chunking: During the generation process, the generative model divides the output text into chunks and returns them step by step;

[0204] Gradual output and dynamic update: The generated task returns the generated content in a streaming form, and users can view part of the results before the generation is completed, which improves the response efficiency;

[0205] Context extension and update: In multi-turn dialogue scenarios, the system dynamically expands the context based on real-time feedback from streaming responses and adjusts the subsequent parts of the generated task based on new user input.

[0206] In this embodiment, the sparse optimization parameters and the prompt template design are further optimized according to the generation quality evaluation result, which specifically includes the following steps:

[0207] Sparse optimization parameter adjustment: Dynamically adjust the regularization parameter λ and penalty factor ρ in sparse optimization based on the generation quality feedback to optimize the selection and sorting of context fragments;

[0208] Prompt template structure optimization: Based on the semantic analysis results of the generated text quality, the background information order, user request expression method and generated instruction content in the Prompt template are adjusted;

[0209] Multi-round optimization and iteration: In multiple rounds of generation tasks, the system gradually optimizes the context screening mechanism and Prompt template structure based on the generation results to improve the overall quality of the generation task.

[0210] The evaluation of the generated quality in step S6 includes the following indicators:

[0211] Semantic relevance: Calculate the cosine similarity between the user request vector and the generated reply vector;

[0212] Contextual coherence: The matching degree between the generated text and the reference text is evaluated by BLEU and ROUGE-L indicators;

[0213] User satisfaction: Subjective evaluation of the generated results based on user ratings.

[0214] The calculation formula of cosine similarity in step S6 is as follows:

[0215] Among them, represents the semantic similarity between the user request vector and the generated reply vector, is the semantic vector representation of the user request, is the semantic vector representation of the generated reply, · represents the dot product of and , is the Euclidean norm of the user request vector, and is the Euclidean norm of the generated reply vector.

[0216] Specifically, in this embodiment, the system uses cosine similarity to quantify the semantic correlation between the user request vector and the generated reply vector, which specifically includes the following steps:

[0217] Generate a semantic vector to represent the user request, which represents the semantic vector of the generated reply text;

[0218] The semantic relevance between user requests and generated responses is calculated using the cosine similarity formula, as follows:

[0219] Where, represents the semantic similarity between the user request vector and the generated reply vector, is the semantic vector representation of the user request, is the semantic vector representation of the generated reply, represents the dot product of and, is the Euclidean norm of the user request vector, is the Euclidean norm of the generated reply vector

[0220] The system determines whether the semantic matching degree between the generated reply and the user request meets the expectation according to the set similarity threshold. If >, the semantic relevance is considered to meet the requirements.

[0221] In this embodiment, the system evaluates the contextual matching degree between the generated text and the reference text through the BLEU and ROUGE-L indicators, which specifically includes the following steps:

[0222] The system calculates the BLEU score based on the generated text and the reference text to evaluate the consistency between the two at the vocabulary and phrase levels. The formula is as follows:

[0223] Among them: BP is the penalty factor used to control the length of the generated text; is the weight of different n-grams; is the precision of n-gram.

[0224] The longest common subsequence (LCS) matching degree between the generated text and the reference text is evaluated by the ROUGE-L indicator, and the formula is as follows:

[0225] Where: is the ratio of the length of LCS that matches the reference text in the generated text to the length of the generated text; is the ratio of the length of LCS to the length of the reference text;

[0226] β is a tuning parameter used to balance the weights of Precision and Recall.

[0227] The system combines BLEU and ROUGE-L scores to evaluate whether the contextual coherence between the generated text and the reference text meets the quality requirements.

[0228] In this embodiment, the system performs a subjective evaluation on the generated results based on the user ratings, specifically including the following steps:

[0229] User rating mechanism: The system collects user satisfaction ratings on the generated results through interface interaction, with a rating range of 0 to 5 points; Subjective evaluation weight: User ratings are given a higher weight to assist in adjusting the optimization strategy of the generated model. For example, when the user rating is less than 3 points, the system will optimize the prompt template or context fragment.

[0230] Real-time feedback from user ratings is used together with semantic relevance and contextual coherence indicators to optimize generation quality.

[0231] In this embodiment, the system comprehensively calculates the scores of semantic relevance, contextual coherence, and user satisfaction to form a comprehensive score of generation quality, and optimizes the prompt template and sparse optimization parameters accordingly, specifically including the following steps:

[0232] The comprehensive score Q of the generated quality is defined as:

[0233] Among them: α, β, γ are weight coefficients; is the semantic relevance score; BLEU and ROUGE-L are context coherence indicators; UserScore is the user score.

[0234] Generation quality feedback and optimization: According to the changing trend of the comprehensive score Q, the regularization parameter λ and penalty factor ρ in sparse optimization, as well as the context fragment sorting and generation instruction content in the Prompt template are dynamically adjusted to ensure continuous improvement of generation quality.

[0235] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A prompt automatic construction method based on human-computer dialogue history and semantic retrieval, characterized in that: The following steps are involved: S1. Perform text preprocessing on historical conversation data and user input, including dynamically generating query keywords, verifying and parsing user uploaded files, and extracting core content; S2. Use the pre-trained language model to map historical conversations and user requests into semantic vectors, match them with keyword weights, and construct a historical vector set and user request vector; S3, extract the most relevant context fragments from the historical conversation vector based on the convex optimization method, combine the highly relevant content captured by the web crawler, and filter the context through the sparse optimization problem; The process of extracting the most relevant context fragment to the user request in the S3 step includes the following: S3.

1. The system simulates the behavior of search engines through web crawlers, crawls web page content in real time, and extracts highly relevant text data; S3.2, construct a sparse optimization problem, whose objective function is: Among them, v u is the semantic vector representation of the user request, v i is the semantic vector representation of the historical dialogue, α=[α1,α2,…,α N ] is the sparse coefficient vector, λ is the sparse regularization parameter; S3.

3. The system has a built-in exception handling mechanism, which supports automatic retry of network requests and records exception logs; The sparse optimization problem in step S3.2 is solved by the alternating direction multiplier method, which specifically includes: The sparse coefficient vector α to be optimized in the current iteration: Among them, α k+1 is the sparse coefficient vector optimized in the k+1th iteration, α is the sparse coefficient vector to be optimized in the current iteration, v u is the semantic vector representation of the user request, v i is the semantic vector representation of the historical dialogue, N is the total number of rounds of the historical dialogue, is a linear combination of historical conversation vectors, fitting the semantic vector representation v of the user request u , is the reconstruction error term, ρ is the penalty factor, is the constraint term, z k is the sparse variable updated in the kth iteration, μ k is the Lagrange multiplier in the kth iteration; Update the sparsity constraint variable z: z k+1 =soft-thresholding(a k+1 +m k ,l / r); Among them, z k+1 is the sparse variable updated in the k+1th iteration, soft-thresholding(·,τ) is the soft threshold function, α k+1 is the sparse coefficient vector optimized in the k+1th iteration, μ k is the Lagrange multiplier in the kth iteration, λ is the regularization parameter, and ρ is the penalty factor; Update the Lagrange multiplier μ: m k+1 =μ k +p(a k+1 -z k+1 ); Among them, μ k+1 is the Lagrange multiplier updated in the k+1th iteration, μ k is the Lagrange multiplier in the kth iteration, ρ is the penalty factor, α k+1 is the sparse coefficient vector obtained by the k+1th iteration optimization, z k+1 is the sparse variable updated in the k+1th iteration; Repeat the iteration until the sparse coefficient meets the preset convergence condition; The sparse optimization result in step S3.2 is a sparse coefficient vector α * , from which the most relevant context fragment D′ corresponding to the non-zero weight and the user request is filtered out, and its expression is: Among them, D′ is the context segment most relevant to the user request; S4, build a dynamic prompt template, integrate the extracted context fragments, user requests and generation instructions to generate prompts; S5, inputting the generated Prompt into the generative language model for processing to generate a natural language reply text; S6. Optimize the sparse optimization parameters and prompt template design according to the generation quality evaluation results.

2. The method for automatically constructing prompts based on human-computer dialogue history and semantic retrieval according to claim 1 is characterized in that: The S1 step includes the following sub-steps: S1.

1. Perform sentence and word segmentation on historical conversation data and user input, filter stop words, and perform part-of-speech tagging and dependency analysis; S1.

2. The system dynamically generates query keywords based on user input, and selects highly relevant keywords based on semantic similarity to ensure accurate retrieval; S1.

3. Verify the file formats uploaded by users, including PDF, DOCX, and Markdown, clean up irrelevant characters, and extract the core content of the file for subsequent knowledge base retrieval.

3. The method for automatically constructing prompts based on human-computer dialogue history and semantic retrieval according to claim 1 is characterized in that: The S2 step includes the following sub-steps: S2.

1. Map the historical conversation data into vector representations through the sentence embedding model to form a historical vector set V = {v1, v2, ..., v N }; S2.

2. Map the user request text into a semantic vector v u , combined with dynamically generated keyword weights, the relevance score is calculated based on word frequency and position weights to optimize the vectorization results; S2.

3. Constructing the historical conversation semantic vector set V and the user request vector v u .

4. The method for automatically constructing prompts based on human-computer dialogue history and semantic retrieval according to claim 1, characterized in that: The construction of the dynamic prompt template in step S4 includes: Extracting background information: sorting the extracted historical conversation fragments by chronological order or semantic importance; User request content: Integrate the request text currently entered by the user into the template; Instruction design: Clearly generate instruction content, including tone, style, and length limits.

5. The method for automatically constructing prompts based on human-computer dialogue history and semantic retrieval according to claim 1, characterized in that: The S5 step includes the following specific steps: S5.

1. Input the generative language model into the generated Prompt and trigger the generation task; S5.

2. Based on the historical dialogue fragments and semantic retrieval results, the output of the generated model is verified for contextual consistency, and the best output is selected by calculating the semantic relevance score between the generated text and the input prompt; S5.

3. Processing the output text of the generated model according to a preset format, including automatic sentence segmentation, punctuation adjustment, and weighted sorting of key content; S5.

4. Evaluate the generation results, adjust the prompt content according to the real-time feedback of the generation quality, and re-execute the generation steps until the output quality requirements are met.

6. The method for automatically constructing prompts based on human-computer dialogue history and semantic retrieval according to claim 1 is characterized in that: The evaluation of the generated quality in step S6 includes the following indicators: Semantic relevance: Calculate the user request vector v u and generate the reply vector v r The cosine similarity of Contextual coherence: The matching degree between the generated text and the reference text is evaluated by BLEU and ROUGE-L indicators; User satisfaction: Subjective evaluation of the generated results based on user ratings.

7. The method for automatically constructing prompts based on human-computer dialogue history and semantic retrieval according to claim 6 is characterized in that: The calculation formula of cosine similarity in step S6 is as follows: Among them, Similarity(v u , v r ) represents the user request vector v u and generate the reply vector v r The semantic similarity between u is the semantic vector representation of the user request, v r To generate the semantic vector representation of the reply, v u ·v r Indicates v u and v r The dot product of ||v u || is the Euclidean norm of the user-requested vector, ||v r || is the Euclidean norm of the generated reply vector.

Citation Information

Patent Citations

  • Multi-round dialogue method and device, equipment and storage medium

    CN118093796A

  • Intelligent customer service question and answer method based on large language model technology

    CN118364084A