Index extraction method, electronic device and computer program product
By matching similar indicators of user input problems in the intelligent question-answer system, using the index prompt words and generation models of the language model to determine the target indicators, the problem of inaccurate extraction caused by inaccurate indicator description is solved, and the system's analysis accuracy is improved.
Patent Information
- Application Number
- CN202510413936.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-08-26
- Estimated Expiration
- 2045-04-03
AI Technical Summary
When the existing intelligent question-and-answer system describes the indicators inaccurately in the user input questions, it leads to inaccurate indicator extraction, affecting the accuracy of the system analysis.
The first similar indicator is determined by matching user input problems, and the indicator prompt words and generation models of the pre-trained language model are used to extract indicators, and the target indicators are determined in combination with the pending problems and the first indicator, so as to improve the accuracy of indicator extraction.
It improves the accuracy of indicator extraction, avoids output results that deviate from user intentions, and enhances the analysis capabilities of the intelligent question-and-answer system.
Smart Images

Figure CN119938870B_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the technical field of indicator extraction, and in particular relates to an indicator extraction method, device, electronic equipment and computer program product. Background Art
[0002] With the development of artificial intelligence (AI) technology, demand for intelligent question-answering (Q&A) systems is growing. By simulating human intelligence, these systems can quickly analyze and answer various user-entered questions, improving the efficiency of information acquisition. In the field of intelligent Q&A, indicator extraction refers to the process of extracting indicators from user-entered questions or matching relevant indicators to user-entered questions. The accuracy of indicator extraction is crucial to the accuracy of analysis performed by intelligent Q&A systems.
[0003] At present, since the description of indicators in the questions input by users may not be accurate enough, it is difficult for the intelligent question-answering system to extract accurate indicators, which affects the accuracy of the intelligent question-answering system. Summary of the Invention
[0004] The embodiments of the present application provide an indicator extraction method, device, electronic device, and computer program product, which can improve the accuracy of indicator extraction.
[0005] In a first aspect, an embodiment of the present application provides an indicator extraction method, comprising:
[0006] Get pending questions input by the user;
[0007] Determine a first similarity index based on the problem matched to the problem to be processed; wherein the first similarity index is an index corresponding to the problem matched to the problem to be processed;
[0008] Determining an indicator prompt word for a pre-trained language model based on the first similarity indicator, and extracting an indicator from the problem to be processed using the indicator prompt word and the pre-trained language generation model to obtain a first indicator; wherein the indicator prompt word is used to guide the pre-trained language generation model to extract an indicator using the first similarity indicator as a positive example;
[0009] A target indicator is determined according to the problem to be processed and the first indicator, wherein the target indicator is an indicator similar to both the problem to be processed and the first indicator.
[0010] In a second aspect, an embodiment of the present application provides an indicator extraction device, comprising:
[0011] The acquisition module is used to obtain the pending questions input by the user;
[0012] An indicator matching module, configured to determine a first similarity indicator based on the problem matched to the problem to be processed; wherein the first similarity indicator is an indicator corresponding to the problem matched to the problem to be processed;
[0013] an indicator extraction module, configured to determine an indicator prompt word for a pre-trained language model based on the first similarity indicator, and to extract an indicator for the problem to be processed using the indicator prompt word and the pre-trained language generation model to obtain a first indicator; wherein the indicator prompt word is used to guide the pre-trained language generation model to perform indicator extraction using the first similarity indicator as a positive example;
[0014] An indicator determination module is used to determine a target indicator based on the problem to be processed and the first indicator, wherein the target indicator is an indicator similar to both the problem to be processed and the first indicator.
[0015] In a third aspect, an embodiment of the present application provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the method described in the first aspect when executing the computer program.
[0016] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, wherein the computer storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the method described in the first aspect are implemented.
[0017] In a fifth aspect, an embodiment of the present application provides a computer program product, which, when executed on an electronic device, enables the electronic device to execute the method described in the first aspect above.
[0018] Compared with the prior art, the embodiments of the present application have the following beneficial effects:
[0019] This application determines a first similarity index by matching the problem to be processed, then determines the index prompt word of the pre-trained language model based on the first similarity index, and extracts indicators for the problem to be processed based on the indicator prompt word and the pre-trained language generation model to obtain the first indicator, and finally determines the target indicator based on the problem to be processed and the above-mentioned first indicator, which can improve the accuracy of indicator extraction. Specifically, since the first similarity indicator is the indicator corresponding to the problem matched by the problem to be processed, it means that the above-mentioned first similarity indicator is the relevant indicator matched by the problem to be processed. Then, using the first similarity indicator as a positive example and determining the indicator prompt word of the pre-trained language model, the indicator extraction process of the pre-trained language generation model can be positively guided, thereby improving the accuracy of the pre-trained language generation model in extracting indicators for the problem to be processed; in addition, the target indicator is determined based on the first indicator extracted from the problem to be processed and the pre-trained language generation model. Since the target indicator is an indicator similar to the problem to be processed and the first indicator, it means that the problem to be processed input by the user and the output result of the pre-trained language model can be comprehensively considered, thereby avoiding deviation from the problem to be processed during indicator extraction, and the problem of inaccurate output results caused by extracting indicators only based on the problem to be processed. Therefore, this method can improve the accuracy of indicator extraction. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments or descriptions of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0021] Figure 1 This is a flow chart of an indicator extraction method provided in an embodiment of the present application;
[0022] Figure 2 This is a schematic diagram of a prompt word template provided in an embodiment of the present application;
[0023] Figure 3 This is a flow chart of the indicator extraction process provided by the embodiment of the present application;
[0024] Figure 4 It is a structural diagram of the index extraction device provided in an embodiment of the present application;
[0025] Figure 5 It is a structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0026] In the following description, specific details such as specific system structures and techniques are provided for purposes of illustration rather than limitation to facilitate a thorough understanding of the embodiments of the present application. However, it will be apparent to those skilled in the art that the present application may be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid obscuring the description of the present application with unnecessary detail.
[0027] It should be understood that when used in the present specification and the appended claims, the term "comprising" indicates the presence of described features, integers, steps, operations, elements and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or collections thereof.
[0028] It will also be understood that the term "and / or" used in this specification and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.
[0029] As used in this specification and the appended claims, the term "if" can be interpreted as "when" or "upon" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrase "if it is determined" or "if [described condition or event] is detected" can be interpreted as meaning "upon determination" or "in response to determining" or "upon detection of [described condition or event]" or "in response to detecting [described condition or event]," depending on the context.
[0030] In addition, in the description of the present application specification and the appended claims, the terms "first", "second", "third", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance.
[0031] References to "one embodiment" or "some embodiments" in this specification mean that a particular feature, structure, or characteristic described in conjunction with that embodiment is included in one or more embodiments of the present application. Thus, phrases such as "in one embodiment," "in some embodiments," "in other embodiments," and "in other embodiments" appearing in various places in this specification do not necessarily refer to the same embodiment, but rather mean "one or more but not all embodiments," unless otherwise specifically emphasized. The terms "including," "comprising," "having," and variations thereof all mean "including but not limited to," unless otherwise specifically emphasized.
[0032] In the intelligent question-and-answer system interface provided by merchants or platforms, users can enter questions. The intelligent question-and-answer system will first extract the indicators included in the input question or related indicators, and then answer the extracted questions based on the user's question format. For example, if a user enters "Help me check the daily community member trends of XX store for the past 10 weeks", the intelligent question-and-answer system will first extract the indicators included and then provide a corresponding answer based on the extracted indicators and the question.
[0033] Currently, metrics extraction can be performed primarily through the following methods: 1. Entity recognition technology: This primarily utilizes deep neural networks to identify and label metrics in user-entered questions. 2. Language generation technology: This utilizes a pre-trained large language generation model to understand the meaning of user-entered questions and generate corresponding metrics.
[0034] However, when extracting indicators using entity recognition technology, the user's input question may not include accurate indicator description information, resulting in inaccurate identification of indicators and other entities. Furthermore, the pre-trained large language model may misunderstand the user's input question, leading to incorrect indicator extraction. Therefore, the above method will lead to inaccurate indicator extraction.
[0035] In order to improve the accuracy of indicator extraction, the present application provides an indicator extraction method, in which a first similarity indicator is determined based on the problem matched to the problem to be processed; wherein the above-mentioned first similarity indicator is the indicator corresponding to the problem matched to the problem to be processed; based on the above-mentioned first similarity indicator, the indicator prompt word of the pre-trained language model is determined, and the indicator prompt word and the above-mentioned pre-trained language generation model are used to extract indicators for the problem to be processed to obtain the first indicator; wherein the above-mentioned indicator prompt word is used to guide the above-mentioned pre-trained language generation model to extract indicators using the above-mentioned first similarity indicator as a positive example; and a target indicator is determined based on the above-mentioned problem to be processed and the above-mentioned first indicator, wherein the above-mentioned target indicator is an indicator that is similar to both the above-mentioned problem to be processed and the above-mentioned first indicator.
[0036] Figure 1 A flow chart of an indicator extraction method provided in an embodiment of the present application is shown, and is described in detail as follows:
[0037] S11. Obtain the pending questions input by the user.
[0038] Specifically, after a user enters a question through the intelligent question-and-answer interface provided by the intelligent question-and-answer system, the intelligent question-and-answer system will obtain the question in the intelligent question-and-answer interface and analyze and answer it based on the question. The above intelligent question-and-answer interface can be provided by an e-commerce platform, an intelligent assistant, an after-sales robot, etc.
[0039] S12. Determine a first similarity index based on the problem matched with the problem to be processed; wherein the first similarity index is an index corresponding to the problem matched with the problem to be processed.
[0040] Specifically, after the intelligent question-and-answer system obtains the pending questions input by the user, it can match questions based on the pending questions or the keywords contained in the pending questions, and then sort the matched questions based on the degree of similarity between the matched questions and the pending questions or the keywords contained in the pending questions, determine a preset number of questions from the sorted questions as candidate questions, and determine the indicators corresponding to the candidate questions as the above-mentioned first similarity indicator.
[0041] The indicators corresponding to the candidate questions are indicators included in the candidate questions, or indicators pre-associated with the candidate questions.
[0042] For example, if the user inputs "Help me check the daily community member trend of XX store in the past 10 weeks", the top N (preset number, which can be set according to actual situation) questions can be selected from the matched questions as candidate questions based on the degree of similarity, and then the indicators included in the candidate questions or the indicators associated with the candidate questions can be extracted as the above-mentioned first similarity indicators.
[0043] It should also be noted that in order to improve the accuracy of matching the questions to be processed, the questions to be processed can be preprocessed before matching based on the questions to be processed or the keywords included in the questions to be processed. The above preprocessing includes at least one of the following: deleting time point or date information, deleting stop words, deleting repeated text, deleting punctuation, etc.
[0044] For example, since the indicators do not cover specific time points or date information, the above-mentioned deletion of time points or date information may include: for the text to be processed, the open source jionlp tool can be used to identify the time and date text in the question and delete it; the above-mentioned deletion of stop words includes: deleting "yeah, ah" and other information that carries almost no meaning in the text to be processed; the above-mentioned deletion of duplicate text includes: deleting duplicate text in the question to be processed.
[0045] In the embodiment of the present application, the first similarity index is determined by matching the problem to be processed, which can improve the accuracy of determining the relevant index (ie, the first similarity index).
[0046] S13. Determine the indicator prompt word of the pre-trained language model based on the above-mentioned first similarity indicator, and use the above-mentioned indicator prompt word and the above-mentioned pre-trained language generation model to extract indicators for the above-mentioned problem to be processed to obtain the first indicator; wherein the above-mentioned indicator prompt word is used to guide the above-mentioned pre-trained language generation model to extract indicators using the above-mentioned first similarity indicator as a positive example.
[0047] The pre-trained language model is a deep learning model trained on large amounts of text data. It can understand the meaning of input text and generate corresponding natural language text. This pre-trained language model can handle a variety of natural language tasks, such as text classification, question-answering, and conversation. The positive examples are correct examples that guide the pre-trained language model in extracting metrics, specifically guiding the pre-trained language model to extract metrics based on the content, style, and format of the correct examples.
[0048] Specifically, a question-indicator pair can be constructed based on the above-mentioned first similarity indicator and the corresponding candidate question, and then the above-mentioned question-indicator pair and the question to be processed are used as inputs of the pre-trained language model. The above-mentioned pre-trained language model will use the question-indicator pair as a positive example to extract indicators for the question to be processed and obtain the above-mentioned first indicator.
[0049] For example, if the determined first similarity index includes: index 1, index 2 and index 3, then the question-indicator pairs can be constructed according to the candidate questions corresponding to the first similarity index: candidate question 1-indicator 1, candidate question 2-indicator 2 and candidate question 3-indicator 3, and together with the question to be processed, they are used as the input of the pre-trained language model.
[0050] In an embodiment of the present application, an indicator prompt word is constructed by using the first similarity indicator and its corresponding candidate question. The question-indicator pair can be used as a positive example to guide the indicator extraction process of the pre-trained language model and improve the accuracy of indicator extraction.
[0051] S14. Determine a target indicator based on the problem to be processed and the first indicator, wherein the target indicator is an indicator similar to both the problem to be processed and the first indicator.
[0052] Specifically, after obtaining the above-mentioned first indicator, a similarity indicator can be retrieved based on the above-mentioned problem to be processed and the above-mentioned first indicator, and then the similarity between the retrieved similarity indicator and the problem to be processed and the similarity between the retrieved similarity indicator and the first indicator can be determined, and the target indicator can be determined based on these two similarities. For example, the similarity indicator with the largest sum of the two similarities can be determined as the above-mentioned target indicator. Of course, the above-mentioned target indicator can also be determined by setting weights, which is not limited here. For example, corresponding similarity weights are set for the two similarities respectively and multiplied, and the two similarities multiplied by the corresponding weights are added, and the above-mentioned target indicator is determined based on the addition result. Since the original problem (i.e., the problem to be processed) and the output result of the language model (i.e., the first indicator) are taken into account at the same time when determining the target indicator, the accuracy of the target indicator determination can be improved.
[0053] In an embodiment of the present application, a first similarity index is determined by matching the problem to be processed, and then the index prompt word of the pre-trained language model is determined based on the first similarity index, and the index of the problem to be processed is extracted based on the indicator prompt word and the pre-trained language generation model to obtain the first index. Finally, the target index is determined based on the problem to be processed and the above-mentioned first index, which can improve the accuracy of the index extraction. Specifically, since the first similarity indicator is the indicator corresponding to the problem matched by the problem to be processed, it means that the above-mentioned first similarity indicator is the relevant indicator matched by the problem to be processed. Then, using the first similarity indicator as a positive example and determining the indicator prompt word of the pre-trained language model, the indicator extraction process of the pre-trained language generation model can be positively guided, thereby improving the accuracy of the pre-trained language generation model in extracting indicators for the problem to be processed; in addition, the target indicator is determined based on the first indicator extracted from the problem to be processed and the pre-trained language generation model. Since the target indicator is an indicator similar to the problem to be processed and the first indicator, it means that the problem to be processed input by the user and the output result of the pre-trained language model can be comprehensively considered, thereby avoiding deviation from the problem to be processed during indicator extraction, and the problem of inaccurate output results caused by extracting indicators only based on the problem to be processed. Therefore, this method can improve the accuracy of indicator extraction.
[0054] In some embodiments, determining the first similarity index based on the questions matched to the question to be processed includes:
[0055] Determine one or more questions matched from a question knowledge base as first candidate questions; wherein the question knowledge base includes a plurality of preset questions and question indicators corresponding to the questions;
[0056] The question indicator corresponding to the first candidate question is determined as the first similarity indicator.
[0057] The question knowledge base refers to a high-quality question-and-answer database maintained by operations and maintenance personnel based on business scenarios. For example, this database can be constructed and adjusted based on historical conversation records of the intelligent question-and-answer system, manually constructed standard question-and-answer scripts, and open-source question-answer databases. The data storage format for each question in the question knowledge base can be: <question vector (1*1024 dimensions), question, question metric>. To improve the accuracy of question storage and retrieval, the question knowledge base can also include user information (such as user ID). For example, the data storage format for each question in the question knowledge base can be: <question vector (1*1024 dimensions), question, question metric, user ID>.
[0058] Specifically, the pre-processed questions to be processed can be vectorized using a preset vector encoding model to obtain a query question vector, and then the similarity between the query question vector and the question vector stored in the question knowledge base can be calculated using a vector retrieval method. Then, the top M (a preset number, which can be set according to actual conditions) questions with the highest similarity are selected as the first candidate questions, and the question index corresponding to the first candidate question is used as the first similarity index.
[0059] It should be understood that the above-mentioned preset vector encoding model is a model for converting text into a high-dimensional vector representation. The above-mentioned preset vector encoding model may include one of the BGE-Large model, the text2vec-large-chinese model, the BERT model, etc.; the above-mentioned vector retrieval method may be a retrieval augmented generation (RAG) method; the above-mentioned similarity calculation may include one of the cosine similarity calculation method, the Euclidean distance calculation method, the Manhattan distance calculation method, etc.
[0060] For example, if a user inputs "Help me check the daily community member trend of XX store in the past 10 weeks", it can be converted into a 1*1024-dimensional query question vector: [0.3, 0.4, 0.5.....] through the BGE-Large model. Then, the RAG method is used to retrieve M first candidate questions with the highest similarity from the question knowledge base, and the corresponding question indicator is determined to be the above-mentioned first similarity indicator.
[0061] In an optional embodiment of the present application, the above-mentioned question knowledge base may be a database stored in milvus.
[0062] In the embodiment of the present application, by searching questions in the question knowledge base through the pending questions, the first candidate question and its corresponding first indicator can be determined quickly and accurately.
[0063] In some embodiments, determining the indicator prompt word of the pre-trained language model based on the first similarity indicator includes:
[0064] Constructing a problem indicator sample according to the first candidate problem and the first similarity indicator;
[0065] Fill the above-mentioned problem indicator samples into the pre-built prompt word template to obtain the above-mentioned indicator prompt words; wherein the above-mentioned pre-built prompt word template includes instructions or rules for guiding the above-mentioned pre-trained language model to extract indicators.
[0066] The instructions or rules in the pre-built prompt word template may include at least one of the following: role instructions, indicator extraction rules, output rules, and related knowledge.
[0067] Specifically, the question-indicator pair consisting of each first candidate question and its corresponding question indicator can be used as a question indicator sample. The question indicator sample can be used as a correct example to guide the pre-trained language model to extract indicators according to the content, style, format, etc. corresponding to the question indicator sample. Then all the question indicator samples are used as supplements to fill in the above-mentioned pre-built prompt word template to obtain the above-mentioned indicator prompt word.
[0068] For example, refer to Figure 2 As shown, it is a schematic diagram of a prompt word template provided by an embodiment of the present application. The above-mentioned pre-built prompt word template may include: Role instructions: "You are an intelligent indicator extraction assistant, specializing in identifying and extracting user intention indicators based on the user's historical conversation records."; Related knowledge: "**Indicator**: An indicator is a combination of concepts and numerical values used to illustrate overall quantitative characteristics. Common indicators include **amount, **amount, **expenditure, **actual collection, **expense, **fee, **amortization, **number, **number of people, **number of times, **number of orders, **ROI, **price, etc."; Indicator extraction rules: "**Indicator priority**: Based on the order of user conversations, give priority to the indicators mentioned in the user's current question or supplement as the indicators in the user's current needs. **Rigorousness**: Ensure that the improved data requirements are consistent with User requests are precisely matched, demands are not spread, and unnecessary data retrieval and analysis demands are eliminated. **Focus**: Focus on the core content of user requests to ensure that data requirements directly reflect user needs. **Clear indicators**: Data requirements must contain clear indicators and only one indicator, and cannot be date and time related, otherwise it is an empty string. **Indicator inheritance**: If the current user question or supplementary information does not have a clear indicator, but there is a clear indicator in the previous round, the current question should inherit the indicator mentioned by the user last. "; Output rules: "Please directly output the user's intention indicator without any thinking steps"; Other related rules: "Current user conversation record: {Questions}. Please ensure that your processing flow strictly follows the above rules to efficiently, accurately and concisely parse user intent indicators. ".
[0069] If the above-mentioned question indicator samples include: "Help me check the daily community member trend of XX store in the past week" and "Daily community member number", then the above-mentioned indicator prompt words can be obtained by adding the question indicator sample to the above-mentioned pre-built prompt word template.
[0070] In the embodiment of the present application, by adding problem indicator samples to the above-mentioned pre-built prompt word template, the indicator prompt words can include positive examples and guiding rules for indicator extraction, thereby improving the richness of the indicator prompt words.
[0071] It should be understood that after filling the above-mentioned problem indicator samples into the pre-built prompt word template and obtaining the above-mentioned indicator prompt word, the prompt word template and the above-mentioned problem to be processed can be used as input of the pre-trained language model to guide the above-mentioned pre-trained language model to output the first indicator, thereby improving the accuracy of the generation of the first indicator.
[0072] In some embodiments, determining the target indicator based on the problem to be addressed and the first indicator includes:
[0073] Match indicators similar to the problem to be processed from a pre-built indicator knowledge base to obtain a second similar indicator;
[0074] Matching an indicator similar to the first indicator from the pre-built indicator knowledge base to obtain a third similar indicator;
[0075] The target indicator is determined based on the second similarity indicator and the third similarity indicator.
[0076] The pre-built indicator knowledge base is a high-quality indicator database maintained by operations and maintenance personnel based on business scenarios. For example, this database can be constructed and adjusted based on historical conversation records from intelligent question-and-answer systems, manually constructed standard indicators, and open-source indicator databases. The data storage format for each indicator in the indicator knowledge base can be: <indicator vector (1*1024 dimensions), indicator (indicator name)>.
[0077] Specifically, a second similarity index can be obtained by recalling multiple similar indicators from a pre-built indicator knowledge base through the query question vector corresponding to the problem to be processed; the first indicator can be vectorized to obtain a query indicator vector, and based on the query indicator vector, multiple similar indicators can be recalled from the pre-built indicator knowledge base to obtain a third similarity index; for each indicator included in the second similarity index and the third similarity index, the similarity between the indicator and the problem to be processed and the first indicator is determined respectively, and the two similarities are added or multiplied by the corresponding weights and the added similarity is used as the final similarity of the indicator, and the indicator with the largest final similarity is determined as the above-mentioned target indicator.
[0078] It should be noted that the above-mentioned vectorization method and recall process can refer to the relevant description of the problem knowledge base matching in the above-mentioned embodiment, and will not be repeated here. The number of indicators included in the above-mentioned second similarity index and the number of indicators included in the above-mentioned third similarity index can be the same or different, and are not limited here.
[0079] For example, assuming the number of indicators recalled by the pending question is the same as the number of indicators recalled by the first indicator, after calculating the cosine similarity between the query vector corresponding to the pending question and the indicator vector in the indicator knowledge base, the first n indicators with the highest similarity are obtained, namely the second similarity index (assuming it is labeled A), where the similarity corresponding to the i-th indicator in the second similarity index is A_score_i; after calculating the cosine similarity between the query indicator vector quantized by the first indicator and the indicator vector in the indicator knowledge base, the first n indicators with the highest similarity are obtained, namely the third similarity index (assuming it is labeled B), where the similarity corresponding to the j-th indicator in the third similarity index is B_score_j, where 0 < i ≤ n, 0 < j ≤ n. Then, based on the similarity between each indicator and the pending question and the first indicator, the final similarity is determined, and the indicator with the highest final similarity is determined as the target indicator. It should be noted that when determining the final similarity, for the indicators in the second similarity index, since its similarity with the problem to be processed is already known (for example, A_score_i), it is only necessary to calculate the similarity between this indicator and the first indicator; for the indicators in the third similarity index, since its similarity with the first indicator is already known (for example, B_score_i), it is only necessary to calculate the similarity between this indicator and the problem to be processed.
[0080] In an embodiment of the present application, multiple similar indicators (including second similar indicators and third similar indicators) are recalled from a pre-built indicator knowledge base by respectively using the problem to be processed and the first indicator, and the indicator with the greatest final similarity is determined as the target indicator from the similar indicators, thereby improving the accuracy of the target indicator determination.
[0081] In some embodiments, in order to comprehensively consider the semantic information of the problem to be processed and the user's preferences and ensure the accuracy of the target indicator ultimately produced, the target indicator is determined based on the second similarity indicator and the third similarity indicator, including:
[0082] Determining a set of candidate indicators based on the second similarity index and the third similarity index; wherein the candidate indicators in the set of candidate indicators are indicators that have appeared in the second similarity index and / or the third similarity index, and the first similarity corresponding to the candidate indicators is determined based on the number of times they appear in the second similarity index and / or the third similarity index;
[0083] For any candidate indicator in the candidate indicator set, determining a character similarity between the candidate indicator and the question to be processed to obtain a second similarity, and determining a historical preference similarity between the candidate indicator and the historical questions recorded by the user to obtain a third similarity;
[0084] Determining a comprehensive similarity of each candidate indicator in the candidate indicator set according to the first similarity, the second similarity, and the third similarity;
[0085] Determine the candidate indicator with the highest comprehensive similarity in the candidate indicator set as the target indicator.
[0086] Among them, the above-mentioned first similarity can reflect the degree of vector matching between the candidate indicator and the question to be processed and / or the first indicator; the above-mentioned second similarity can reflect the degree of character matching between the candidate indicator and the question to be processed; the above-mentioned third similarity can reflect the degree of preference matching between the candidate indicator and the historical questions recorded by the user.
[0087] Specifically, for any one of the second and third similarity indicators, if it only appears in the second similarity indicator, a first similarity is determined based on its similarity corresponding to the second similarity indicator and a preset first weight. If it only appears in the third similarity indicator, a first similarity is determined based on its similarity corresponding to the third similarity indicator and a preset second weight. If it appears in both the second and third similarity indicators, a first similarity is determined based on its similarity corresponding to the second and third similarity indicators and a preset third weight. Then, based on the first similarity, a preset number of indicators are determined as the candidate indicators to obtain a candidate indicator set. For any candidate indicator in the candidate indicator set, a character similarity is determined based on the number of characters that the candidate indicator and the question to be processed share, thereby obtaining a second similarity. Simultaneously, a similarity between the candidate indicator and all historical questions recorded by the user can be calculated to obtain a historical preference similarity for each historical question, with the largest historical preference similarity being used as the third similarity. The first, second, and third similarities of each candidate indicator are then summed to determine a comprehensive similarity for each candidate indicator in the candidate indicator set. Finally, the candidate indicator with the highest comprehensive similarity is selected as the target indicator. It should be noted that the above-mentioned preset first weight, preset second weight and preset third weight can be set according to actual conditions and are not limited here.
[0088] For example, assuming that the preset first weight, the preset second weight, and the preset third weight are the same, then for the second similarity index (assuming it is marked as A), the similarity corresponding to the i-th index is A_score_i. If this index only appears in the second similarity index, its first similarity can be determined to be A_score_i / 2; for the third similarity index (assuming it is marked as B), the similarity corresponding to the j-th index is B_score_j. If this index only appears in the third similarity index, its first similarity can be determined to be B_score_j / 2; if the index appears at the same time, If it appears in the second similarity index (assuming it is marked as A) and the third similarity index (assuming it is marked as B), the corresponding first similarity is (A_score_i+B_score_j) / 2. Finally, sort according to the first similarity and select the first m (set according to actual situation) indicators as candidate indicators to obtain a candidate indicator set (assuming it is marked as C). Then, for any candidate indicator in the candidate indicator set, calculate the corresponding second similarity and third similarity respectively. Finally, obtain the comprehensive similarity of each candidate indicator based on the first similarity, second similarity and third similarity.
[0089] In some embodiments, the comprehensive similarity can be determined by the following formula: comprehensive similarity = a * first similarity + b * second similarity + c * third similarity, where a, b, and c are preset weight coefficients, a+b+c=1, and the values of a, b, and c are all greater than or equal to 0 and less than or equal to 1.
[0090] In an embodiment of the present application, a set of candidate indicators is determined by the above-mentioned second similarity indicator and the above-mentioned third similarity indicator, and the comprehensive similarity of each candidate indicator is determined. The comprehensive similarity corresponding to each candidate indicator includes a first similarity reflecting the degree of vector matching between the candidate indicator and the problem to be processed and / or the first indicator, a second similarity reflecting the degree of character matching between the candidate indicator and the problem to be processed, and a third similarity reflecting the degree of matching between the candidate indicator and the historical question preferences recorded by the user. Therefore, the degree of matching between each candidate indicator and the problem to be processed and the first indicator can be comprehensively considered to improve the accuracy of determining the target indicator.
[0091] In some embodiments, for any candidate indicator in the candidate indicator set, determining the character similarity between the candidate indicator and the problem to be processed to obtain the second similarity includes:
[0092] Splitting the above-mentioned problem to be processed into single characters to obtain a problem character set, and splitting the above-mentioned candidate indicators into single characters to obtain an indicator character set;
[0093] Determine the intersection of the problem character set and the indicator character set;
[0094] The second similarity is determined according to the number of characters in the intersection.
[0095] Specifically, the problem to be processed can be split into a set consisting of only single characters, namely, the problem character set, and the length of the problem character set (i.e., the number of characters in the problem character set) can be determined. Furthermore, the kth candidate indicator in the candidate indicator set can be split into a set consisting of only single characters, namely, the indicator character set, and the length of the indicator character set (i.e., the number of characters in the indicator character set) can be determined. The number of characters in the intersection of the problem character set and the indicator character set can be determined. The second similarity can be determined based on the number of characters, the length of the problem character set, and the length of the indicator character set. It should be noted that, to avoid the influence of repeated characters, the characters in the problem character set and the indicator character set can be deduplicated before determining the second similarity.
[0096] For example, the problem to be processed is broken up into a set of single characters (i.e., the problem character set) and duplicates are removed to obtain a first character set Q, and the kth candidate indicator in the candidate indicator set (assuming it is marked as C) is broken up into a set of single characters (i.e., the indicator character set corresponding to the kth candidate indicator) and duplicates are removed to obtain a second character set Ck; the intersection of the first character set Q and the second character set Ck is determined, and the number of characters in the intersection is calculated to obtain Lk_inte, then the second similarity of the kth candidate indicator is: Lk_inte / len(Q) + Lk_inte / len(Ck), where the above len(Q) is the length of the first character set, and the above len(Ck) is the length of the second character set.
[0097] In the embodiment of the present application, the accuracy of the second similarity calculation can be improved according to the intersection of the question character set and the indicator character set.
[0098] In some embodiments, for any candidate indicator in the candidate indicator set, determining the historical preference similarity between the candidate indicator and the historical question recorded by the user to obtain the third similarity includes:
[0099] According to the problem to be processed and the user information of the user, the problem knowledge base is matched to obtain a second candidate question;
[0100] Determine the problem indicator corresponding to the second candidate problem as the second similarity indicator;
[0101] When the second similarity index is the same as the candidate index, the matching score of the second candidate question corresponding to the second similarity index in the matching process is used as the historical preference similarity of the candidate index, and the largest historical preference similarity is selected to obtain the third similarity.
[0102] Specifically, according to the problem knowledge base matched with the user information of the problem to be processed and the user, a preset number of problems (which can be set according to actual conditions) with the highest similarity matching scores are selected as the second candidate problems, the problem indicators corresponding to the second candidate problems are determined as the second similarity indicators, and the similarity matching scores of the second candidate problems and the problem to be processed during the matching process are determined as the similarity scores of the second similarity indicators corresponding to the second candidate problems. Only the second similarity indicators and their corresponding similarity scores existing in the candidate indicator set are retained, wherein the similarity score can reflect the historical preference similarity of the candidate indicators that are the same as the second similarity indicators; for any candidate indicator, if the same second similarity indicator exists, the largest similarity score among the same second similarity indicators is selected as the third similarity of the candidate indicator. If the same second similarity indicator does not exist, the preset value can be determined as the third similarity of the candidate indicator.
[0103] For example, suppose that three second candidate questions are matched: candidate question 1, candidate question 2, and candidate question 3, and their corresponding second similarity indicators are indicator 1, indicator 2, and indicator 1 respectively. The similarity matching scores of each second candidate question are similarity A, similarity B, and similarity C, then the similarity scores of the second similarity indicators are similarity A (corresponding to indicator 1), similarity B (corresponding to indicator 2), and similarity C (corresponding to indicator 1), respectively; if the candidate indicator set includes three candidate indicators: indicator 1, indicator 2, and indicator 3, then indicator 1 and indicator 2 (that is, the second similarity indicators existing in the candidate indicator set) are retained. For the three candidate indicators included in the indicator set, the similarity score corresponding to indicator 1 includes similarity A and similarity C, the similarity score corresponding to indicator 2 includes similarity B, and indicator 3 does not have the same second similarity indicator. Then the third similarity of indicator 1 is max(similarity A, similarity C), the third similarity of indicator 2 is similarity B, and the third similarity corresponding to indicator 3 is 0 (that is, the preset value).
[0104] In an embodiment of the present application, when the second similarity index is the same as the above-mentioned candidate index, the maximum similarity matching score of the second candidate question corresponding to the second similarity index in the matching process is used as the third similarity of the candidate index, thereby improving the accuracy of determining the third similarity.
[0105] In another optional embodiment of the present application, after determining the target indicator based on the problem to be processed and the first indicator, the following steps are further included:
[0106] Store the above target indicators that are correct based on user feedback.
[0107] Specifically, when the user feedback target indicator is correct, the target indicator can be stored in the above-mentioned question knowledge base and / or indicator knowledge base according to the preset time interval. It should be noted that during the storage process, it needs to be stored in accordance with the storage format of the question knowledge base and / or indicator knowledge base.
[0108] For example, if a user reports a correct target indicator, the target indicator, the pending question currently entered by the user, and the user's user information can be stored in the question knowledge base in a regular manner in the format of <question vector (1*1024 dimensions), question, target indicator, user ID>. The process of storing the indicator knowledge base is similar and will not be repeated here.
[0109] In an embodiment of the present application, by regularly maintaining the correct target indicators fed back by users and storing them in the problem knowledge base and / or indicator knowledge base, the accuracy, richness and timeliness of the problem knowledge base and / or indicator knowledge base can be ensured, forming a dynamic closed loop of indicator extraction, and further improving the accuracy of indicator extraction.
[0110] In order to better illustrate the indicator extraction process, the following Figure 3 For explanation, refer to Figure 3 Shown, including:
[0111] Step S301: Obtain a pending question input by a user, wherein the pending question may be a question input by the user in the intelligent question-answering system interface;
[0112] Step S302: performing text preprocessing on the problem to be processed, wherein the text preprocessing may include one or more of deleting time point or date information, deleting stop words, deleting repeated text, deleting punctuation, etc.;
[0113] Step S303: vectorize the question to be processed after the input text preprocessing using a vector encoding model to obtain a query question vector, wherein the vector encoding model may be a BGE-Large model;
[0114] Step S304: determining a first similarity index in the question knowledge base by querying the question vector, wherein the first similarity index is an index corresponding to the question matched by the query question vector;
[0115] Step S305: construct an indicator prompt word based on the first similarity index and the question to be processed after text preprocessing, and use the indicator prompt word and the question to be processed after preprocessing as inputs of the pre-trained language model;
[0116] Step S306: Output a first indicator using the pre-trained language model;
[0117] Step S307: vectorize the first indicator using a vector coding model to obtain a query indicator vector;
[0118] Step S308: Recall similar indices from the indicator knowledge base using the input query question vector and query indicator vector to obtain a second similar indicator (i.e., an indicator similar to the question to be processed) and a third similar indicator (i.e., an indicator similar to the first indicator).
[0119] Step S309: construct a candidate indicator set based on the second similarity index and the third similarity index, and determine the first similarity, the second similarity, and the third similarity of each candidate indicator in the candidate indicator set;
[0120] Step S310: perform similarity fusion on each candidate indicator to obtain a comprehensive similarity, and select the candidate indicator with the highest comprehensive similarity as the target indicator. Finally, perform regular maintenance, that is, store the target indicators with correct user feedback into the question knowledge base according to the preset time interval.
[0121] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0122] Corresponding to the indicator extraction method described in the above embodiment, Figure 4 A structural schematic diagram of the indicator extraction device provided in an embodiment of the present application is shown. For the sake of convenience, only the parts related to the embodiment of the present application are shown.
[0123] Reference Figure 4 The device may be an indicator extraction device 41 , and the indicator extraction device 41 may include an acquisition module 411 , an indicator matching module 412 , an indicator extraction module 413 and an indicator determination module 414 .
[0124] Reference Figure 4 , the above-mentioned indicator extraction device 41 includes:
[0125] The acquisition module 411 is used to obtain the pending question input by the user;
[0126] The indicator matching module 412 is used to determine a first similarity indicator based on the problem matched to the problem to be processed; wherein the first similarity indicator is an indicator corresponding to the problem matched to the problem to be processed;
[0127] An indicator extraction module 413 is configured to determine an indicator prompt word for a pre-trained language model based on the first similarity indicator, and to extract an indicator for the problem to be processed using the indicator prompt word and the pre-trained language generation model to obtain a first indicator; wherein the indicator prompt word is used to guide the pre-trained language generation model to perform indicator extraction using the first similarity indicator as a positive example;
[0128] The indicator determination module 414 is configured to determine a target indicator based on the problem to be processed and the first indicator, wherein the target indicator is an indicator similar to both the problem to be processed and the first indicator.
[0129] In some embodiments, when the indicator matching module 412 determines the first similarity indicator based on the problem matched to the problem to be processed, it includes:
[0130] Determine one or more questions matched from a question knowledge base as first candidate questions; wherein the question knowledge base includes a plurality of preset questions and question indicators corresponding to the questions;
[0131] The question indicator corresponding to the first candidate question is determined as the first similarity indicator.
[0132] In some embodiments, when determining the indicator prompt word of the pre-trained language model based on the first similarity indicator, the indicator extraction module 413 includes:
[0133] Constructing a problem indicator sample according to the first candidate problem and the first similarity indicator;
[0134] Fill the above-mentioned problem indicator sample into the pre-built prompt word template to obtain the above-mentioned indicator prompt word; wherein the above-mentioned pre-built prompt word template includes instructions or rules for guiding the above-mentioned pre-trained language model to extract indicators.
[0135] In some embodiments, when determining the target indicator based on the problem to be processed and the first indicator, the indicator determination module 414 includes:
[0136] Match indicators similar to the problem to be processed from a pre-built indicator knowledge base to obtain a second similar indicator;
[0137] Matching an indicator similar to the first indicator from the pre-built indicator knowledge base to obtain a third similar indicator;
[0138] The target indicator is determined based on the second similarity indicator and the third similarity indicator.
[0139] In some embodiments, in order to comprehensively consider the semantic information of the problem to be processed and the user's preferences and ensure the accuracy of the target indicator ultimately produced, the indicator determination module 414, when determining the target indicator based on the second similarity indicator and the third similarity indicator, includes:
[0140] Determining a set of candidate indicators based on the second similarity index and the third similarity index; wherein the candidate indicators in the set of candidate indicators are indicators that have appeared in the second similarity index and / or the third similarity index, and the first similarity corresponding to the candidate indicators is determined based on the number of times they appear in the second similarity index and / or the third similarity index;
[0141] For any candidate indicator in the candidate indicator set, determining a character similarity between the candidate indicator and the question to be processed to obtain a second similarity, and determining a historical preference similarity between the candidate indicator and the historical questions recorded by the user to obtain a third similarity;
[0142] Determining a comprehensive similarity of each candidate indicator in the candidate indicator set according to the first similarity, the second similarity, and the third similarity;
[0143] Determine the candidate indicator with the highest comprehensive similarity in the candidate indicator set as the target indicator.
[0144] In some embodiments, for any candidate indicator in the candidate indicator set, when the indicator determination module 414 determines the character similarity between the candidate indicator and the question to be processed to obtain the second similarity, the following steps are performed:
[0145] Splitting the above-mentioned problem to be processed into single characters to obtain a problem character set, and splitting the above-mentioned candidate indicators into single characters to obtain an indicator character set;
[0146] Determine the intersection of the problem character set and the indicator character set;
[0147] The second similarity is determined according to the number of characters in the intersection.
[0148] In some embodiments, for any candidate indicator in the candidate indicator set, when the indicator determination module 414 determines the historical preference similarity between the candidate indicator and the historical question recorded by the user to obtain the third similarity, the following steps may be performed:
[0149] According to the problem to be processed and the user information of the user, the problem knowledge base is matched to obtain a second candidate question;
[0150] Determine the problem indicator corresponding to the second candidate problem as the second similarity indicator;
[0151] When the second similarity index is the same as the candidate index, the matching score of the second candidate question corresponding to the second similarity index in the matching process is used as the historical preference similarity of the candidate index, and the largest historical preference similarity is selected to obtain the third similarity.
[0152] In another optional embodiment of the present application, the indicator extraction device 41 further includes a maintenance module. The maintenance module is configured to, after determining the target indicator based on the problem to be processed and the first indicator, include:
[0153] Store the above target indicators that are correct based on user feedback.
[0154] It should be noted that the information interaction, execution process, etc. between the devices / units are based on the same concept as the method embodiments of this application. Their specific functions and technical effects can be found in the method embodiment section and will not be repeated here.
[0155] Figure 5 This is a schematic diagram of the structure of an electronic device provided in one embodiment of the present application. Figure 5 As shown, the electronic device 5 of this embodiment includes: at least one processor 50 ( Figure 5 The computer program 52 is stored in the memory 51 and can be run on the at least one processor 50. When the processor 50 executes the computer program 52, the steps of any of the method embodiments are implemented.
[0156] The electronic device 5 can be a computing device such as a desktop computer, a notebook, a PDA, a cloud server, etc. The electronic device can include, but is not limited to, a processor 50 and a memory 51. It can be understood by those skilled in the art that Figure 5 It is only an example of the electronic device 5 and does not constitute a limitation of the electronic device 5. It may include more or fewer components than shown in the figure, or a combination of certain components, or different components. For example, the electronic device may also include an input sending device, a network access device, a bus, etc.
[0157] The processor 50 may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.
[0158] In some embodiments, the memory 51 may be an internal storage unit of the electronic device 5, such as a hard drive or memory of the electronic device 5. The memory 51 may also be an external storage device of the electronic device 5, such as a plug-in hard drive, a Smart Media Card (SMC), a Secure Digital (SD) card, a flash memory card, etc. Furthermore, the memory 51 may include both an internal storage unit of the electronic device 5 and an external storage device. The memory 51 is used to store an operating system, application programs, a boot loader, data, and other programs, such as the program code of the computer program. The memory 51 may also be used to temporarily store data that has been sent or is about to be sent.
[0159] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the functional units and modules is used as an example for illustration. In actual applications, the function allocation can be completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the system can refer to the corresponding process in the aforementioned method embodiment, which will not be repeated here.
[0160] An embodiment of the present application also provides a network device, which includes: at least one processor, a memory, and a computer program stored in the memory and executable on the at least one processor, wherein the processor implements the steps of any of the method embodiments when executing the computer program.
[0161] An embodiment of the present application further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in the various method embodiments can be implemented.
[0162] An embodiment of the present application provides a computer program product. When the computer program product is run on an electronic device, the electronic device can implement the steps in the various method embodiments when executing the computer program product.
[0163] If the integrated unit is implemented as a software functional unit and sold or used as a standalone product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the process steps in the method embodiments described herein can be implemented by a computer program instructing the relevant hardware. The computer program can be stored in a computer-readable storage medium. When executed by a processor, the computer program can implement the steps of each method embodiment described. The computer program includes computer program code, which can be in source code form, object code form, executable file, or some intermediate form. The computer-readable medium can include at least: any entity or device capable of carrying computer program code to a camera / electronic device, recording medium, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunications signals, and software distribution media. Examples include USB flash drives, removable hard drives, magnetic disks, or optical disks. In some jurisdictions, based on legislation and patent practice, computer-readable media cannot be electrical carrier signals or telecommunications signals.
[0164] In the embodiments, the description of each embodiment has its own focus. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0165] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0166] In the embodiments provided in this application, it should be understood that the disclosed devices / network equipment and methods can be implemented in other ways. For example, the device / network equipment embodiments described above are merely illustrative. For example, the division of the modules or units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0167] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0168] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the scope of protection of the present application.
Claims
1. An index extraction method, characterized in that: include: Get pending questions input by the user; Determine a first similarity index based on the problem matched to the problem to be processed; wherein the first similarity index is an index corresponding to the problem matched to the problem to be processed; Determining an indicator prompt word for a pre-trained language model based on the first similarity indicator, and extracting an indicator from the problem to be processed using the indicator prompt word and the pre-trained language generation model to obtain a first indicator; wherein the indicator prompt word includes an instruction or rule for guiding the pre-trained language model to extract the indicator, and the indicator prompt word is used to use the first similarity indicator as a positive example and guide the pre-trained language generation model to extract the indicator according to the instruction or rule; Matching the similarity index according to the question to be processed to obtain a second similarity index, and matching the similarity index according to the first index to obtain a third similarity index, determining a candidate index set according to the second similarity index and the third similarity index, determining, for any candidate index in the candidate index set, a historical preference similarity between the candidate index and a historical question recorded by the user to obtain a third similarity, and determining a character similarity between the candidate index and the question to be processed to obtain a second similarity; Determining a target indicator from the candidate indicator set according to the first similarity, the second similarity, and the third similarity, wherein the first similarity is used to reflect the degree of vector matching between the candidate indicator and the problem to be processed and / or the first indicator, and the target indicator is an indicator in the candidate indicator set that is similar to both the problem to be processed and the first indicator; Determining the historical preference similarity between the candidate indicator and the historical questions recorded by the user to obtain a third similarity includes: Matching the question to be processed with the user information of the user to a question knowledge base to obtain a second candidate question; Determining the question indicator corresponding to the second candidate question as a second similarity indicator; When the second similarity index is the same as the candidate index, the matching score of the second candidate question corresponding to the second similarity index in the matching process is used as the historical preference similarity of the candidate index, and the largest historical preference similarity is selected to obtain the third similarity.
2. The index extraction method according to claim 1, characterized in that: The determining of a first similarity index according to the problem matched with the problem to be processed includes: Determining one or more questions matched to the problem to be processed from the problem knowledge base as first candidate questions; wherein the problem knowledge base includes a plurality of preset questions and problem indicators corresponding to the questions; The question indicator corresponding to the first candidate question is determined as the first similarity indicator.
3. The index extraction method according to claim 2, characterized in that: The step of determining the indicator prompt word of the pre-trained language model according to the first similarity indicator includes: Constructing a problem indicator sample based on the first candidate problem and the first similarity indicator; Filling the problem indicator sample into a pre-constructed prompt word template to obtain the indicator prompt word; wherein the pre-constructed prompt word template includes the instructions or rules that guide the pre-trained language model to extract indicators.
4. The index extraction method according to any one of claims 1 to 3, characterized in that: The determining of a candidate indicator set according to the second similarity indicator and the third similarity indicator includes: Matching an indicator similar to the problem to be processed from a pre-built indicator knowledge base to obtain the second similar indicator; Matching an indicator similar to the first indicator from the pre-built indicator knowledge base to obtain the third similar indicator; The candidate indicator set is determined according to the second similarity indicator and the third similarity indicator, and a candidate indicator that is similar to the problem to be processed and the first indicator is selected from the candidate indicator set to obtain the target indicator.
5. The index extraction method according to claim 4, characterized in that: The determining the target indicator from the candidate indicator set according to the first similarity, the second similarity, and the third similarity includes: Determining the candidate indicator set based on the second similarity index and the third similarity index; wherein the candidate indicators in the candidate indicator set are indicators that have appeared in the second similarity index and / or the third similarity index, and the first similarity corresponding to the candidate indicator is determined based on the number of times it appears in the second similarity index and / or the third similarity index; Determining a comprehensive similarity of each candidate indicator in the candidate indicator set according to the first similarity, the second similarity, and the third similarity; Determine the candidate indicator with the highest comprehensive similarity in the candidate indicator set as the target indicator.
6. The index extraction method according to claim 5, characterized in that: Determining the character similarity between the candidate indicator and the problem to be processed to obtain a second similarity includes: Splitting the problem to be processed into single characters to obtain a problem character set, and splitting the candidate indicators into single characters to obtain an indicator character set; Determining the intersection of the question character set and the indicator character set; The second similarity is determined according to the number of characters in the intersection.
7. The index extraction method according to any one of claims 1 to 3, characterized in that: After determining the target indicator from the candidate indicator set according to the first similarity, the second similarity, and the third similarity, the method further includes: The target indicator that the user feedback is correct is stored.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the method according to any one of claims 1 to 7 is implemented.
9. A computer program product, characterized in that The invention comprises a computer program, which implements the method according to any one of claims 1 to 7 when the computer program is executed.
Citation Information
Patent Citations
Deep learning-based intelligent customer service semantic retrieval method, device and system
CN114817461A
Intention recognition method, system and equipment
CN119358563A
Intelligent number asking method and device and readable storage medium
CN119739730A