Index extraction method, electronic equipment and computer program product
By matching similar indicators of user input problems in the intelligent question-answer system and using pre-trained language models for indicator extraction, the problem of inaccurate indicator description in user input problems is solved, and the accuracy of indicator extraction is improved.
Patent Information
- Application Number
- CN202510413936.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-04-03
AI Technical Summary
When the indicator descriptions of the existing intelligent question-and-answer system are not accurate enough in user input questions, it is difficult to accurately extract the indicators, which affects the system's analysis accuracy.
By obtaining the pending problems input by the user, matching the similar problems to determine the first similar indicator, using the pre-trained language model and indicator prompt words to perform indicator extraction, and finally determining the target indicator based on the pending problems and the extracted indicators.
The accuracy of indicator extraction is improved, and the generation model of the forward guide language is generated and the comprehensive consideration of user input and language model output is avoided by the deviation of indicator extraction and inaccurate output caused by extraction based on the problem only.
Smart Images

Figure CN119938870A_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the technical field of indicator extraction, and in particular, relates to an indicator extraction method, device, electronic equipment and computer program product. Background Art
[0002] With the development of artificial intelligence technology, people's demand for intelligent question-answering systems is growing. Intelligent question-answering systems can quickly analyze and answer various questions input by users by simulating human intelligence, thereby improving the efficiency of users in obtaining information. In the field of intelligent question-answering, indicator extraction refers to the process of extracting indicators from questions input by users or matching relevant indicators according to questions input by users. The accuracy of indicator extraction is crucial to the accuracy of analysis by intelligent question-answering systems.
[0003] At present, since the description of indicators in the questions input by users may not be accurate enough, it is difficult for the intelligent question-answering system to extract accurate indicators, which affects the accuracy of the intelligent question-answering system. Summary of the invention
[0004] The embodiments of the present application provide an indicator extraction method, device, electronic device and computer program product, which can improve the accuracy of indicator extraction.
[0005] In a first aspect, an embodiment of the present application provides an indicator extraction method, comprising: Get pending questions input by the user; Determine a first similarity index according to the problem matched by the problem to be processed; wherein the first similarity index is an index corresponding to the problem matched by the problem to be processed; Determining an indicator prompt word of a pre-trained language model according to the first similarity indicator, and extracting an indicator from the problem to be processed using the indicator prompt word and the pre-trained language generation model to obtain a first indicator; wherein the indicator prompt word is used to guide the pre-trained language generation model to extract an indicator using the first similarity indicator as a positive example; A target indicator is determined according to the problem to be processed and the first indicator, wherein the target indicator is an indicator similar to both the problem to be processed and the first indicator.
[0006] In a second aspect, an embodiment of the present application provides an indicator extraction device, comprising: The acquisition module is used to obtain pending issues input by the user; An indicator matching module, used to determine a first similarity indicator according to the problem matched by the problem to be processed; wherein the first similarity indicator is an indicator corresponding to the problem matched by the problem to be processed; An indicator extraction module, used to determine an indicator prompt word of a pre-trained language model according to the first similarity indicator, and use the indicator prompt word and the pre-trained language generation model to extract indicators from the problem to be processed to obtain a first indicator; wherein the indicator prompt word is used to guide the pre-trained language generation model to extract indicators using the first similarity indicator as a positive example; An indicator determination module is used to determine a target indicator based on the problem to be processed and the first indicator, wherein the target indicator is an indicator similar to both the problem to be processed and the first indicator.
[0007] In a third aspect, an embodiment of the present application provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the method described in the first aspect when executing the computer program.
[0008] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, wherein the computer storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the method described in the first aspect are implemented.
[0009] In a fifth aspect, an embodiment of the present application provides a computer program product. When the computer program product is run on an electronic device, the electronic device executes the method described in the first aspect above.
[0010] Compared with the prior art, the embodiments of the present invention have the following beneficial effects: The present application determines a first similarity index by matching the problem to be processed, then determines the index prompt word of the pre-trained language model based on the first similarity index, and extracts the index of the problem to be processed based on the index prompt word and the pre-trained language generation model to obtain the first index, and finally determines the target index based on the problem to be processed and the above-mentioned first index, which can improve the accuracy of index extraction. Specifically, since the first similarity indicator is the indicator corresponding to the problem matched by the problem to be processed, it means that the above-mentioned first similarity indicator is the relevant indicator matched by the problem to be processed. Then, using the first similarity indicator as a positive example and determining the indicator prompt word of the pre-trained language model, the indicator extraction process of the pre-trained language generation model can be positively guided, thereby improving the accuracy of the pre-trained language generation model in extracting indicators for the problem to be processed; in addition, the target indicator is determined according to the first indicator extracted from the problem to be processed and the pre-trained language generation model. Since the target indicator is an indicator similar to the problem to be processed and the first indicator, it means that the problem to be processed input by the user and the output result of the pre-trained language model can be comprehensively considered, thereby avoiding deviation from the problem to be processed during indicator extraction, and the problem of inaccurate output results caused by extracting indicators only based on the problem to be processed. Therefore, the method can improve the accuracy of indicator extraction. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0012] Figure 1 It is a flowchart of an indicator extraction method provided in an embodiment of the present application; Figure 2 This is a schematic diagram of a prompt word template provided in an embodiment of the present application; Figure 3 It is a flowchart of the indicator extraction process provided by the embodiment of the present application; Figure 4 is a schematic diagram of the structure of the index extraction device provided in the embodiment of the present application; Figure 5 It is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0013] In the following description, specific details such as specific system structures, technologies, etc. are provided for the purpose of illustration rather than limitation, so as to provide a thorough understanding of the embodiments of the present application. However, it should be clear to those skilled in the art that the present application may also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to prevent unnecessary details from obstructing the description of the present application.
[0014] It should be understood that when used in the present specification and the appended claims, the term "comprising" indicates the presence of described features, wholes, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or combinations thereof.
[0015] It should also be understood that the term “and / or” used in the specification and appended claims refers to any and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0016] As used in the specification and appended claims of this application, the term "if" can be interpreted as "when" or "uponce" or "in response to determining" or "in response to detecting", depending on the context. Similarly, the phrase "if it is determined" or "if [described condition or event] is detected" can be interpreted as meaning "uponce it is determined" or "in response to determining" or "uponce [described condition or event] is detected" or "in response to detecting [described condition or event]", depending on the context.
[0017] In addition, in the description of the present application specification and the appended claims, the terms "first", "second", "third", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance.
[0018] References to "one embodiment" or "some embodiments" etc. described in the specification of this application mean that one or more embodiments of the present application include specific features, structures or characteristics described in conjunction with the embodiment. Therefore, the statements "in one embodiment", "in some embodiments", "in some other embodiments", "in some other embodiments", etc. that appear in different places in this specification do not necessarily refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "including", "comprising", "having" and their variations all mean "including but not limited to", unless otherwise specifically emphasized in other ways.
[0019] In the intelligent question-and-answer system interface provided by merchants or platforms, users can ask questions by inputting questions. The intelligent question-and-answer system will first extract the indicators included in or related to the input questions, and answer the extracted questions according to the way the user asks the questions. For example, if a user inputs "Help me check the daily community number trend of XX store in the past 10 weeks", the intelligent question-and-answer system will first extract the indicators included, and then answer the questions accordingly based on the extracted indicators and questions.
[0020] At present, the following methods can be used to extract indicators: 1. Entity recognition technology: mainly using deep neural networks to identify and label indicators in user input questions. 2. Language generation technology: for user input questions, use a pre-trained large language generation model to understand the meaning of user input questions and generate corresponding indicators.
[0021] However, when extracting indicators through entity recognition technology, the questions input by users may not include accurate indicator description information, which will lead to inaccurate identification of indicators and other entities; at the same time, the pre-trained large language model may misunderstand the questions input by users, resulting in incorrect indicator extraction. Therefore, the above method will lead to inaccurate indicator extraction.
[0022] In order to improve the accuracy of indicator extraction, the present application provides an indicator extraction method, in which a first similarity indicator is determined based on a problem matched to a problem to be processed; wherein the first similarity indicator is an indicator corresponding to the problem matched to the problem to be processed; an indicator prompt word of a pre-trained language model is determined based on the first similarity indicator, and the indicator prompt word and the pre-trained language generation model are used to extract indicators for the problem to be processed to obtain a first indicator; wherein the indicator prompt word is used to guide the pre-trained language generation model to extract indicators using the first similarity indicator as a positive example; a target indicator is determined based on the problem to be processed and the first indicator, wherein the target indicator is an indicator that is similar to both the problem to be processed and the first indicator.
[0023] Figure 1 A flow chart of an indicator extraction method provided in an embodiment of the present application is shown, which is described in detail as follows: S11. Obtain the pending questions input by the user.
[0024] Specifically, after the user enters the pending question through the intelligent question-and-answer interface provided by the intelligent question-and-answer system, the intelligent question-and-answer system will obtain the pending question in the intelligent question-and-answer interface and analyze and answer it based on the pending question. The above intelligent question-and-answer interface can be an interface provided by an e-commerce platform, an intelligent assistant, an after-sales robot, etc.
[0025] S12. Determine a first similarity index according to the problem matched by the problem to be processed; wherein the first similarity index is an index corresponding to the problem matched by the problem to be processed.
[0026] Specifically, after the intelligent question-and-answer system obtains the pending questions input by the user, it can match questions according to the pending questions or the keywords contained in the pending questions, and then sort the matched questions according to the degree of similarity between the matched questions and the pending questions or the keywords contained in the pending questions, determine a preset number of questions from the sorted questions as candidate questions, and determine the indicators corresponding to the candidate questions as the above-mentioned first similarity indicator.
[0027] The indicators corresponding to the above-mentioned candidate questions refer to the indicators included in the candidate questions, or the indicators pre-associated with the candidate questions.
[0028] For example, if the user inputs "Help me check the daily community member trend of XX store in the past 10 weeks", the top N (preset number, which can be set according to actual conditions) questions can be selected from the matched questions as candidate questions based on the degree of similarity, and then the indicators included in the candidate questions or the indicators associated with the candidate questions can be extracted as the above-mentioned first similarity indicators.
[0029] It should also be noted that in order to improve the accuracy of matching the questions to be processed, before matching according to the questions to be processed or the keywords included in the questions to be processed, the questions to be processed can be preprocessed. The above preprocessing includes at least one of the following: deleting time point or date information, deleting stop words, deleting repeated text, deleting punctuation, etc.
[0030] For example, since the indicators do not cover specific time points or date information, the above-mentioned deletion of time points or date information may include: for the text to be processed, the open source jionlp tool can be used to identify the time and date text in the question and delete it; the above-mentioned deletion of stop words includes: deleting information that carries almost no meaning, such as "ah, ah" in the text to be processed; the above-mentioned deletion of repeated text includes: deleting repeated text in the question to be processed.
[0031] In the embodiment of the present application, the first similarity index is determined by matching the problem to be processed, so that the accuracy of determining the relevant index (ie, the first similarity index) can be improved.
[0032] S13. Determine the indicator prompt words of the pre-trained language model according to the above-mentioned first similarity indicator, and use the above-mentioned indicator prompt words and the above-mentioned pre-trained language generation model to extract indicators for the above-mentioned problem to be processed to obtain the first indicator; wherein the above-mentioned indicator prompt words are used to guide the above-mentioned pre-trained language generation model to extract indicators with the above-mentioned first similarity indicator as a positive example.
[0033] The above-mentioned pre-trained language model refers to a deep learning model trained with a large amount of text data, which can understand the meaning of the input language text and generate the corresponding natural language text. The above-mentioned pre-trained language model can handle a variety of natural language tasks, such as text classification, question answering, dialogue, etc. The above-mentioned positive examples refer to the correct examples that guide the above-mentioned pre-trained language model to extract indicators, that is, guide the pre-trained language model to extract indicators according to the content, style, format, etc. corresponding to the correct examples.
[0034] Specifically, a question-indicator pair can be constructed based on the above-mentioned first similarity indicator and the corresponding candidate question, and then the above-mentioned question-indicator pair and the problem to be processed are used as inputs of the pre-trained language model. The above-mentioned pre-trained language model will use the question-indicator pair as a positive example, extract indicators for the problem to be processed, and obtain the above-mentioned first indicator.
[0035] For example, if the determined first similarity index includes: index 1, index 2 and index 3, then the question-indicator pairs can be constructed according to the candidate questions corresponding to the first similarity index: candidate question 1-indicator 1, candidate question 2-indicator 2 and candidate question 3-indicator 3, and used together with the question to be processed as the input of the pre-trained language model.
[0036] In the embodiment of the present application, the indicator prompt word is constructed by the first similar indicator and its corresponding candidate question, and the question-indicator pair can be used as a positive example to guide the indicator extraction process of the pre-trained language model and improve the accuracy of indicator extraction.
[0037] S14. Determine a target indicator based on the above-mentioned problem to be processed and the above-mentioned first indicator, wherein the above-mentioned target indicator is an indicator similar to the above-mentioned problem to be processed and the above-mentioned first indicator.
[0038] Specifically, after obtaining the above-mentioned first indicator, a similarity indicator can be retrieved based on the above-mentioned problem to be processed and the above-mentioned first indicator, and then the similarity between the retrieved similarity indicator and the problem to be processed and the similarity between the retrieved similarity indicator and the first indicator can be determined, and the target indicator can be determined based on the two similarities. For example, the similarity indicator with the largest sum of two similarities can be determined as the above-mentioned target indicator. Of course, the above-mentioned target indicator can also be determined by setting weights, which is not limited here. For example, corresponding similarity weights are set for the two similarities respectively and multiplied, and the two similarities multiplied by the corresponding weights are added, and the above-mentioned target indicator is determined based on the addition result. Since the original problem (i.e., the problem to be processed) and the output result of the language model (i.e., the first indicator) are considered at the same time when determining the target indicator, the accuracy of the determination of the target indicator can be improved.
[0039] In an embodiment of the present application, a first similarity index is determined by matching the problem to be processed, and then the index prompt word of the pre-trained language model is determined based on the first similarity index, and indicators are extracted for the problem to be processed based on the indicator prompt word and the pre-trained language generation model to obtain the first indicator, and finally the target indicator is determined based on the problem to be processed and the above-mentioned first indicator, which can improve the accuracy of indicator extraction. Specifically, since the first similarity indicator is the indicator corresponding to the problem matched by the problem to be processed, it means that the above-mentioned first similarity indicator is the relevant indicator matched by the problem to be processed. Then, using the first similarity indicator as a positive example and determining the indicator prompt word of the pre-trained language model, the indicator extraction process of the pre-trained language generation model can be positively guided, thereby improving the accuracy of the pre-trained language generation model in extracting indicators for the problem to be processed; in addition, the target indicator is determined according to the first indicator extracted from the problem to be processed and the pre-trained language generation model. Since the target indicator is an indicator similar to the problem to be processed and the first indicator, it means that the problem to be processed input by the user and the output result of the pre-trained language model can be comprehensively considered, thereby avoiding deviation from the problem to be processed during indicator extraction, and the problem of inaccurate output results caused by extracting indicators only based on the problem to be processed. Therefore, the method can improve the accuracy of indicator extraction.
[0040] In some embodiments, determining the first similarity index according to the problem matched to the problem to be processed includes: Determine one or more questions matched from the question knowledge base as the first candidate question; wherein the question knowledge base includes a plurality of preset questions and question indicators corresponding to the questions; The question indicator corresponding to the first candidate question is determined as the first similarity indicator.
[0041] The above-mentioned question knowledge base refers to a high-quality question-and-answer database maintained by operation and maintenance personnel based on business scenarios. For example, the high-quality question-and-answer database adjusted and constructed by operation and maintenance personnel based on historical conversation records of the intelligent question-and-answer system, manually constructed standard question-and-answer scripts, open source question-and-answer databases, etc. The data storage format of each question in the above-mentioned question knowledge base can be: <question vector (1*1024 dimensions), question, question index>; of course, in order to improve the accuracy of question storage and retrieval, the above-mentioned question knowledge base can also include user information (such as user ID). For example, the data storage format of each question in the question knowledge base can be: <question vector (1*1024 dimensions), question, question index, user ID>.
[0042] Specifically, the pre-processed problem to be processed can be vectorized using a preset vector encoding model to obtain a query question vector, and then the query question vector and the question vector stored in the question knowledge base can be similarly calculated using a vector retrieval method, and then the top M (a preset number, which can be set according to actual conditions) questions with the highest similarity are selected as the first candidate question, and the question indicator corresponding to the first candidate question is used as the first similarity indicator.
[0043] It should be understood that the above-mentioned preset vector encoding model is a model for converting text into a high-dimensional vector representation, and the above-mentioned preset vector encoding model may include one of the BGE-Large model, the text2vec-large-chinese model, the BERT model, etc.; the above-mentioned vector retrieval method may be a retrieval augmented generation (RAG) method; the above-mentioned similarity calculation may include one of the cosine similarity calculation method, the Euclidean distance calculation method, the Manhattan distance calculation method, etc.
[0044] For example, if the user inputs "Help me check the daily community member trend of XX store in the past 10 weeks", it can be converted into a 1*1024-dimensional query question vector through the BGE-Large model: [0.3, 0.4, 0.5.....], and then the RAG method is used to retrieve M first candidate questions with the highest similarity from the question knowledge base, and the corresponding question indicator is determined to be the above-mentioned first similarity indicator.
[0045] In an optional embodiment of the present application, the above-mentioned question knowledge base may be a database stored in milvus.
[0046] In the embodiment of the present application, by searching the questions in the question knowledge base through the questions to be processed, the first candidate question and its corresponding first indicator can be determined quickly and accurately.
[0047] In some embodiments, determining the indicator prompt word of the pre-trained language model according to the first similarity indicator includes: Constructing a problem indicator sample according to the first candidate problem and the first similarity indicator; Fill the above-mentioned problem indicator sample into the pre-constructed prompt word template to obtain the above-mentioned indicator prompt word; wherein the above-mentioned pre-constructed prompt word template includes instructions or rules for guiding the above-mentioned pre-trained language model to extract indicators.
[0048] The instructions or rules in the pre-built prompt word template may include at least one of the following: role instructions, indicator extraction rules, output rules, and related knowledge.
[0049] Specifically, a question-indicator pair consisting of each first candidate question and its corresponding question indicator can be used as a question indicator sample. The question indicator sample can be used as a correct example to guide the pre-trained language model to extract indicators according to the content, style, format, etc. corresponding to the question indicator sample. Then all the question indicator samples are used as supplements to fill in the above-mentioned pre-built prompt word template to obtain the above-mentioned indicator prompt word.
[0050] For example, refer to Figure 2 As shown, it is a schematic diagram of a prompt word template provided by an embodiment of the present application. The above-mentioned pre-built prompt word template may include: Role instruction: "You are an intelligent indicator extraction assistant, specializing in identifying and extracting user intention indicators based on the user's historical conversation records."; Related knowledge: "**Indicator**: An indicator is a combination of concepts and numerical values used to illustrate overall quantitative characteristics. Common indicators include **amount, **amount, **expenditure, **actual collection, **expense, **fee, **amortization, **number, **number of people, **number of times, **number of orders, **ROI, **price, etc."; Indicator extraction rules: "**Indicator priority**: According to the order of user conversations, the indicators mentioned in the user's current question or supplement are given priority as the indicators in the user's current needs. **Rigorousness**: Ensure that the improved data requirements are consistent with User requests are precisely matched, demands are not spread, and unnecessary data retrieval and analysis demands are eliminated. **Focus**: Focus on the core content of user requests to ensure that data requirements directly reflect user needs. **Clear indicators**: Data requirements must contain clear indicators and only one indicator, and cannot be related to date and time, otherwise it is an empty string. **Indicator inheritance**: If the indicator is not clear in the current user question or supplementary information, but there is a clear indicator in the previous round, the current question should inherit the indicator mentioned by the user last. "; Output rules: "Please output the user's intention indicator directly without any thinking steps"; Other related rules: "Current user conversation record: {Questions}. Please make sure that your processing flow strictly follows the above rules to efficiently, accurately and concisely parse user intention indicators. ".
[0051] If the above-mentioned problem indicator samples include: "Help me check the daily community member number trend of XX store in the past week" and "Daily community member number", then the above-mentioned indicator prompt words can be obtained by adding the problem indicator sample to the above-mentioned pre-built prompt word template.
[0052] In the embodiment of the present application, by adding problem indicator samples to the above-mentioned pre-constructed prompt word template, the indicator prompt words can include positive samples and guiding rules for indicator extraction, thereby improving the richness of the indicator prompt words.
[0053] It should be understood that after filling the above-mentioned problem indicator samples into the pre-built prompt word template to obtain the above-mentioned indicator prompt word, the prompt word template and the above-mentioned problem to be processed can be used as input of the pre-trained language model to guide the above-mentioned pre-trained language model to output the first indicator, thereby improving the accuracy of the generation of the first indicator.
[0054] In some embodiments, determining the target indicator according to the problem to be processed and the first indicator includes: Matching indicators similar to the above-mentioned problem to be processed from a pre-built indicator knowledge base to obtain a second similar indicator; Matching an indicator similar to the first indicator from the pre-built indicator knowledge base to obtain a third similar indicator; The target indicator is determined according to the second similarity indicator and the third similarity indicator.
[0055] The above-mentioned pre-built indicator knowledge base refers to a high-quality indicator database maintained by operation and maintenance personnel based on business scenarios. For example, the high-quality indicator database adjusted and constructed by operation and maintenance personnel based on historical conversation records of the intelligent question-answering system, manually constructed standard indicators, open source indicator databases, etc. The data storage format of each indicator in the above-mentioned indicator knowledge base can be: <indicator vector (1*1024 dimension), indicator (indicator name)>.
[0056] Specifically, a second similarity indicator can be obtained by recalling multiple similar indicators from a pre-built indicator knowledge base through a query problem vector corresponding to the problem to be processed; the first indicator can be vectorized to obtain a query indicator vector, and multiple similar indicators can be recalled from a pre-built indicator knowledge base based on the query indicator vector to obtain a third similarity indicator; for each indicator included in the second similarity indicator and the third similarity indicator, the similarity between the indicator and the problem to be processed and the first indicator is determined respectively, and the two similarities are added or multiplied by corresponding weights to obtain the added similarity as the final similarity of the indicator, and the indicator with the largest final similarity is determined as the above-mentioned target indicator.
[0057] It should be noted that the above-mentioned vectorization method and recall process can refer to the relevant description of the problem knowledge base matching in the above-mentioned embodiment, and will not be repeated here. The number of indicators included in the above-mentioned second similarity index and the number of indicators included in the above-mentioned third similarity index can be the same or different, and are not limited here.
[0058] For example, assuming that the number of indicators recalled through the pending problem is the same as the number of indicators recalled through the first indicator, after calculating the cosine similarity between the query problem vector corresponding to the pending problem and the indicator vector in the indicator knowledge base, the first n indicators with the highest similarity are obtained, that is, the second similarity indicator (assuming it is marked as A), where the similarity corresponding to the i-th indicator in the second similarity indicator is A_score_i; after calculating the cosine similarity between the query indicator vector quantized by the first indicator and the indicator vector in the indicator knowledge base, the first n indicators with the highest similarity are obtained, that is, the third similarity indicator (assuming it is marked as B), where the similarity corresponding to the j-th indicator in the third similarity indicator is B_score_j, where 0<i≤n, 0<j≤n. Then, the final similarity is determined based on the similarity between each indicator and the pending problem and the first indicator, and the indicator with the highest final similarity is determined as the above target indicator. It should be noted that when determining the final similarity, for the indicators in the second similarity index, since its similarity with the problem to be processed is already known (for example, A_score_i), it is only necessary to calculate the similarity between the indicator and the first indicator; for the indicators in the third similarity index, since its similarity with the first indicator is already known (for example, B_score_i), it is only necessary to calculate the similarity between the indicator and the problem to be processed.
[0059] In an embodiment of the present application, by respectively recalling multiple similar indicators (including a second similar indicator and a third similar indicator) from a pre-built indicator knowledge base for the problem to be processed and the first indicator, and determining the indicator with the largest final similarity from the similar indicators as the target indicator, the accuracy of determining the target indicator can be improved.
[0060] In some embodiments, in order to comprehensively consider the semantic information of the problem to be processed and the user's preferences and ensure the accuracy of the target indicator finally outputted, the above-mentioned target indicator is determined according to the above-mentioned second similarity indicator and the above-mentioned third similarity indicator, including: Determine a candidate indicator set according to the second similarity indicator and the third similarity indicator; wherein the candidate indicators in the candidate indicator set are indicators that have appeared in the second similarity indicator and / or the third similarity indicator, and the first similarity corresponding to the candidate indicators is determined according to the number of times they appear in the second similarity indicator and / or the third similarity indicator; For any one of the candidate indicators in the candidate indicator set, determine the character similarity between the candidate indicator and the question to be processed to obtain a second similarity, and determine the historical preference similarity between the candidate indicator and the historical question recorded by the user to obtain a third similarity; Determine the comprehensive similarity of each candidate indicator in the candidate indicator set according to the first similarity, the second similarity and the third similarity; Determine the candidate indicator with the highest comprehensive similarity in the candidate indicator set as the target indicator.
[0061] Among them, the above-mentioned first similarity can reflect the degree of vector matching between the candidate indicator and the question to be processed and / or the first indicator; the above-mentioned second similarity can reflect the degree of character matching between the candidate indicator and the question to be processed; the above-mentioned third similarity can reflect the degree of preference matching between the candidate indicator and the historical questions recorded by the user.
[0062] Specifically, for any one of the second similarity index and the third similarity index, if it only appears in the second similarity index, the first similarity is determined according to the similarity corresponding to the second similarity index and the preset first weight; if it only appears in the third similarity index, the first similarity is determined according to the similarity corresponding to the third similarity index and the preset second weight; if it appears in both the second similarity index and the third similarity index, the first similarity is determined according to the similarity corresponding to the second similarity and the third similarity index and the preset third weight; then, according to the first similarity, a preset number of indicators are determined as the above-mentioned candidate indicators to obtain a candidate indicator set; for any candidate indicator in the candidate indicator set, the character similarity can be determined according to the number of characters that are the same as the candidate indicator and the problem to be processed, and then the second similarity can be obtained. At the same time, the similarity between the above-mentioned candidate indicator and all historical questions recorded by the user can be calculated to obtain the historical preference similarity of each historical question, and the largest historical preference similarity among them is used as the above-mentioned third similarity. Then, the first similarity, the second similarity and the third similarity of each candidate indicator are added together to determine the comprehensive similarity of each candidate indicator in the candidate indicator set, and finally the candidate indicator with the highest comprehensive similarity is selected as the above-mentioned target indicator. It should be noted that the above-mentioned preset first weight, preset second weight and preset third weight can be set according to actual conditions and are not limited here.
[0063] For example, assuming that the preset first weight, the preset second weight, and the preset third weight are the same, then for the second similarity index (assuming it is marked as A), the similarity corresponding to the i-th index is A_score_i. If the index only appears in the second similarity index, its first similarity can be determined to be A_score_i / 2; for the third similarity index (assuming it is marked as B), the similarity corresponding to the j-th index is B_score_j. If the index only appears in the third similarity index, its first similarity can be determined to be B_score_j / 2; if the index is both If it appears in the second similarity index (assuming it is marked as A) and the third similarity index (assuming it is marked as B), the corresponding first similarity is (A_score_i+B_score_j) / 2. Finally, sort according to the first similarity, and select the first m (set according to actual conditions) indicators as candidate indicators to obtain a candidate indicator set (assuming it is marked as C); then for any candidate indicator in the candidate indicator set, calculate the corresponding second similarity and third similarity respectively, and finally obtain the comprehensive similarity of each candidate indicator according to the first similarity, the second similarity and the third similarity.
[0064] In some embodiments, the above-mentioned comprehensive similarity can be determined by the following formula: comprehensive similarity = a * first similarity + b * second similarity + c * third similarity, wherein a, b, c are preset weight coefficients, a+b+c=1, and the values of a, b, c are all greater than or equal to 0 and less than or equal to 1.
[0065] In an embodiment of the present application, a set of candidate indicators is determined by the second similarity indicator and the third similarity indicator, and a comprehensive similarity of each candidate indicator is determined. The comprehensive similarity corresponding to each candidate indicator includes a first similarity reflecting the degree of vector matching between the candidate indicator and the problem to be processed and / or the first indicator, a second similarity reflecting the degree of character matching between the candidate indicator and the problem to be processed, and a third similarity reflecting the degree of matching between the candidate indicator and the historical question preferences recorded by the user. Therefore, the degree of matching between each candidate indicator and the problem to be processed and the first indicator can be comprehensively considered to improve the accuracy of determining the target indicator.
[0066] In some embodiments, for any candidate indicator in the candidate indicator set, determining the character similarity between the candidate indicator and the problem to be processed to obtain the second similarity includes: Splitting the above-mentioned problem to be processed into single characters to obtain a problem character set, and splitting the above-mentioned candidate indicators into single characters to obtain an indicator character set; Determine the intersection of the problem character set and the indicator character set; The second similarity is determined according to the number of characters in the intersection.
[0067] Specifically, the problem to be processed can be split into a set including only single characters, namely, the problem character set, and the length of the problem character set (i.e., the number of characters in the problem character set) can be determined, and the kth candidate indicator in the candidate indicator set can be split into a set including only single characters, namely, the indicator character set, and the length of the indicator character set (i.e., the number of characters in the indicator character set) can be determined; the number of characters in the intersection of the problem character set and the indicator character set can be determined; and the second similarity can be determined based on the number of the above characters, the length of the above problem character set, and the length of the above indicator character set. It should be noted that, in order to avoid the influence of repeated characters, before determining the second similarity, the characters in the problem character set and the indicator character set can be deduplicated.
[0068] For example, the problem to be processed is broken up into a set of single characters (i.e., the problem character set) and duplicates are removed to obtain a first character set Q, and the kth candidate indicator in the candidate indicator set (assuming it is marked as C) is broken up into a set of single characters (i.e., the indicator character set corresponding to the kth candidate indicator) and duplicates are removed to obtain a second character set Ck; the intersection of the first character set Q and the second character set Ck is determined, and the number of characters in the intersection is calculated to obtain Lk_inte, then the second similarity of the kth candidate indicator is: Lk_inte / len(Q) + Lk_inte / len(Ck), wherein the above len(Q) is the length of the first character set, and the above len(Ck) is the length of the second character set.
[0069] In the embodiment of the present application, the accuracy of the second similarity calculation can be improved according to the intersection of the question character set and the indicator character set.
[0070] In some embodiments, for any candidate indicator in the candidate indicator set, determining the historical preference similarity between the candidate indicator and the historical question recorded by the user to obtain the third similarity includes: According to the problem to be processed and the user information of the user, the problem knowledge base is matched to obtain a second candidate question; Determine the problem indicator corresponding to the second candidate problem as the second similarity indicator; When the second similarity index is the same as the candidate index, the matching score of the second candidate question corresponding to the second similarity index in the matching process is used as the historical preference similarity of the candidate index, and the largest historical preference similarity is selected to obtain the third similarity.
[0071] Specifically, according to the problem knowledge base matched with the user information of the problem to be processed and the user, a preset number of problems (which can be set according to actual conditions) with the highest similarity matching scores are taken as second candidate problems, the problem indicator corresponding to the second candidate problem is determined as the second similarity indicator, and the similarity matching score between the second candidate problem and the problem to be processed during the matching process is determined as the similarity score of the second similarity indicator corresponding to the second candidate problem, and only the second similarity indicators existing in the candidate indicator set and their corresponding similarity scores are retained, wherein the similarity score can reflect the historical preference similarity of the candidate indicators that are the same as the second similarity indicator; for any candidate indicator, if the same second similarity indicator exists, the largest similarity score among the same second similarity indicators is selected as the third similarity of the candidate indicator, and if the same second similarity indicator does not exist, the preset value can be determined as the third similarity of the candidate indicator.
[0072] For example, suppose that three second candidate questions are matched: candidate question 1, candidate question 2 and candidate question 3, and their corresponding second similarity indicators are indicator 1, indicator 2 and indicator 1 respectively, and the similarity matching scores of each second candidate question are similarity A, similarity B and similarity C, then the similarity scores of the second similarity indicators are similarity A (corresponding to indicator 1), similarity B (corresponding to indicator 2) and similarity C (corresponding to indicator 1) respectively; if the candidate indicator set includes three candidate indicators: indicator 1, indicator 2 and indicator 3, then indicator 1 and indicator 2 (that is, the second similarity indicators existing in the candidate indicator set) are retained, for the three candidate indicators included in the indicator set, the similarity score corresponding to indicator 1 includes similarity A and similarity C, the similarity score corresponding to indicator 2 includes similarity B, and indicator 3 does not have the same second similarity indicator, then the third similarity of indicator 1 is max(similarity A, similarity C), the third similarity of indicator 2 is similarity B, and the third similarity corresponding to indicator 3 is 0 (that is, the preset value).
[0073] In an embodiment of the present application, when the second similarity index and the above-mentioned candidate index are the same, the maximum similarity matching score of the second candidate question corresponding to the second similarity index in the matching process is used as the third similarity of the candidate index, so as to improve the accuracy of determining the third similarity.
[0074] In another optional embodiment of the present application, after determining the target indicator according to the above-mentioned problem to be processed and the above-mentioned first indicator, the following further comprises: Store the above target indicators that are correct based on user feedback.
[0075] Specifically, when the user feedback target indicator is correct, the target indicator can be stored in the above-mentioned question knowledge base and / or indicator knowledge base at a preset time interval. It should be noted that the storage process needs to be stored in the storage format of the question knowledge base and / or indicator knowledge base.
[0076] For example, for a target indicator that is correctly fed back by a user, the target indicator, the pending question currently input by the user, and the user information of the user can be stored in the question knowledge base in a regular manner in the format of <question vector (1*1024 dimension), question, target indicator, user ID>. The process of storing in the indicator knowledge base is similar and will not be repeated here.
[0077] In an embodiment of the present application, by regularly maintaining the correct target indicators fed back by the user and storing them in the problem knowledge base and / or the indicator knowledge base, the accuracy, richness and timeliness of the problem knowledge base and / or the indicator knowledge base can be ensured, forming a dynamic closed loop of indicator extraction, and further improving the accuracy of indicator extraction.
[0078] In order to better illustrate the index extraction process, the following Figure 3 For explanation, refer to Figure 3 As shown, including: Step S301: obtaining a pending question input by a user, wherein the pending question may be a question input by the user in an intelligent question-answering system interface; Step S302: performing text preprocessing on the problem to be processed, wherein the text preprocessing may include one or more of deleting time point or date information, deleting stop words, deleting repeated text, deleting punctuation marks, etc.; Step S303: vectorize the question to be processed after the input text preprocessing using a vector coding model to obtain a query question vector, wherein the vector coding model may be a BGE-Large model; Step S304: determining a first similarity index in the question knowledge base by querying the question vector, wherein the first similarity index is an index corresponding to the question matched by the query question vector; Step S305: construct an indicator prompt word through the first similarity indicator and the problem to be processed after the text preprocessing, and use the indicator prompt word and the problem to be processed after the preprocessing as the input of the pre-trained language model; Step S306: Output a first indicator using a pre-trained language model; Step S307: vectorize the first indicator using a vector coding model to obtain a query indicator vector; Step S308: Recall similar indices from the indicator knowledge base using the input query question vector and query indicator vector to obtain a second similar indicator (i.e., an indicator similar to the problem to be processed) and a third similar indicator (i.e., an indicator similar to the first indicator); Step S309: construct a candidate indicator set according to the second similarity indicator and the third similarity indicator, and determine the first similarity, the second similarity and the third similarity of each candidate indicator in the candidate indicator set; Step S310: perform similarity fusion on each candidate indicator to obtain comprehensive similarity, and take the candidate indicator with the highest comprehensive similarity as the target indicator. Finally, perform regular maintenance, that is, store the target indicators with correct user feedback into the question knowledge base according to the preset time interval.
[0079] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0080] Corresponding to the index extraction method described in the above embodiment, Figure 4 A schematic diagram of the structure of the indicator extraction device provided in an embodiment of the present application is shown. For the sake of convenience of explanation, only the parts related to the embodiment of the present application are shown.
[0081] Reference Figure 4 The device may be an indicator extraction device 41 , and the indicator extraction device 41 may include an acquisition module 411 , an indicator matching module 412 , an indicator extraction module 413 and an indicator determination module 414 .
[0082] Reference Figure 4 , the above-mentioned index extraction device 41 comprises: The acquisition module 411 is used to acquire the pending question input by the user; The indicator matching module 412 is used to determine a first similarity indicator according to the problem matched by the problem to be processed; wherein the first similarity indicator is an indicator corresponding to the problem matched by the problem to be processed; The indicator extraction module 413 is used to determine the indicator prompt word of the pre-trained language model according to the first similarity indicator, and use the indicator prompt word and the pre-trained language generation model to extract the indicator of the problem to be processed to obtain the first indicator; wherein the indicator prompt word is used to guide the pre-trained language generation model to extract the indicator with the first similarity indicator as a positive example; The indicator determination module 414 is used to determine the target indicator according to the problem to be processed and the first indicator, wherein the target indicator is an indicator similar to the problem to be processed and the first indicator.
[0083] In some embodiments, when the indicator matching module 412 determines the first similarity indicator according to the problem matched to the problem to be processed, it includes: Determine one or more questions matched from the question knowledge base as the first candidate question; wherein the question knowledge base includes a plurality of preset questions and question indicators corresponding to the questions; The question indicator corresponding to the first candidate question is determined as the first similarity indicator.
[0084] In some embodiments, when the indicator extraction module 413 determines the indicator prompt word of the pre-trained language model according to the first similarity indicator, it includes: Constructing a problem indicator sample according to the first candidate problem and the first similarity indicator; Fill the above-mentioned problem indicator sample into the pre-constructed prompt word template to obtain the above-mentioned indicator prompt word; wherein the above-mentioned pre-constructed prompt word template includes instructions or rules for guiding the above-mentioned pre-trained language model to extract indicators.
[0085] In some embodiments, when the indicator determination module 414 determines the target indicator according to the problem to be processed and the first indicator, it includes: Matching indicators similar to the above-mentioned problem to be processed from a pre-built indicator knowledge base to obtain a second similar indicator; Matching an indicator similar to the first indicator from the pre-built indicator knowledge base to obtain a third similar indicator; The target indicator is determined according to the second similarity indicator and the third similarity indicator.
[0086] In some embodiments, in order to comprehensively consider the semantic information of the problem to be processed and the user's preferences and ensure the accuracy of the target indicator finally outputted, the indicator determination module 414, when determining the target indicator according to the second similarity indicator and the third similarity indicator, includes: Determine a candidate indicator set according to the second similarity indicator and the third similarity indicator; wherein the candidate indicators in the candidate indicator set are indicators that have appeared in the second similarity indicator and / or the third similarity indicator, and the first similarity corresponding to the candidate indicators is determined according to the number of times they appear in the second similarity indicator and / or the third similarity indicator; For any one of the candidate indicators in the candidate indicator set, determine the character similarity between the candidate indicator and the question to be processed to obtain a second similarity, and determine the historical preference similarity between the candidate indicator and the historical question recorded by the user to obtain a third similarity; Determine the comprehensive similarity of each candidate indicator in the candidate indicator set according to the first similarity, the second similarity and the third similarity; Determine the candidate indicator with the highest comprehensive similarity in the candidate indicator set as the target indicator.
[0087] In some embodiments, for any candidate indicator in the candidate indicator set, when the indicator determination module 414 determines the character similarity between the candidate indicator and the problem to be processed to obtain the second similarity, it includes: Splitting the above-mentioned problem to be processed into single characters to obtain a problem character set, and splitting the above-mentioned candidate indicators into single characters to obtain an indicator character set; Determine the intersection of the problem character set and the indicator character set; The second similarity is determined according to the number of characters in the intersection.
[0088] In some embodiments, for any candidate indicator in the candidate indicator set, when the indicator determination module 414 determines the historical preference similarity between the candidate indicator and the historical question recorded by the user to obtain the third similarity, the following steps are performed: According to the problem to be processed and the user information of the user, the problem knowledge base is matched to obtain a second candidate question; Determine the problem indicator corresponding to the second candidate problem as the second similarity indicator; When the second similarity index is the same as the candidate index, the matching score of the second candidate question corresponding to the second similarity index in the matching process is used as the historical preference similarity of the candidate index, and the largest historical preference similarity is selected to obtain the third similarity.
[0089] In another optional embodiment of the present application, the indicator extraction device 41 further includes a maintenance module, and the maintenance module is used to, after determining the target indicator according to the problem to be processed and the first indicator, include: Store the above target indicators that are correct based on user feedback.
[0090] It should be noted that the information interaction, execution process, etc. between the devices / units are based on the same concept as the method embodiments of the present application. Their specific functions and technical effects can be found in the method embodiment section and will not be repeated here.
[0091] Figure 5 This is a schematic diagram of the structure of an electronic device provided by an embodiment of the present application. Figure 5 As shown, the electronic device 5 of this embodiment includes: at least one processor 50 ( Figure 5 Only one is shown in the figure), a memory 51, and a computer program 52 stored in the memory 51 and executable on the at least one processor 50. When the processor 50 executes the computer program 52, the steps in any of the method embodiments are implemented. The electronic device 5 may be a computing device such as a desktop computer, a notebook, a PDA, or a cloud server. The electronic device may include, but is not limited to, a processor 50 and a memory 51. Those skilled in the art will appreciate that Figure 5 It is only an example of the electronic device 5 and does not constitute a limitation of the electronic device 5. It may include more or fewer components than shown in the figure, or a combination of certain components, or different components. For example, the electronic device may also include an input sending device, a network access device, a bus, etc.
[0092] The processor 50 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc.
[0093] In some embodiments, the memory 51 may be an internal storage unit of the electronic device 5, such as a hard disk or memory of the electronic device 5. The memory 51 may also be an external storage device of the electronic device 5, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the electronic device 5. Further, the memory 51 may also include both an internal storage unit of the electronic device 5 and an external storage device. The memory 51 is used to store an operating system, an application program, a boot loader (BootLoader), data, and other programs, such as the program code of the computer program. The memory 51 may also be used to temporarily store data that has been sent or is to be sent.
[0094] The technicians in the relevant field can clearly understand that for the convenience and simplicity of description, only the division of the functional units and modules is used as an example. In practical applications, the function allocation can be completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated into a processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the system can refer to the corresponding process in the aforementioned method embodiment, which will not be repeated here.
[0095] An embodiment of the present application also provides a network device, which includes: at least one processor, a memory, and a computer program stored in the memory and executable on the at least one processor, wherein the processor implements the steps in any of the method embodiments when executing the computer program.
[0096] An embodiment of the present application further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in the various method embodiments can be implemented.
[0097] An embodiment of the present application provides a computer program product. When the computer program product runs on an electronic device, the electronic device can implement the steps in the various method embodiments when executing the computer program product.
[0098] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application implements all or part of the process in the embodiment method, which can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a computer-readable storage medium. When the computer program is executed by the processor, the steps of the various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium can at least include: any entity or device that can carry the computer program code to the camera / electronic device, recording medium, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, RandomAccess Memory), electric carrier signal, telecommunication signal and software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk or an optical disk. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electric carrier signals and telecommunication signals.
[0099] In the embodiments, the description of each embodiment has its own emphasis. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0100] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0101] In the embodiments provided in the present application, it should be understood that the disclosed devices / network equipment and methods can be implemented in other ways. For example, the device / network equipment embodiments described above are merely schematic. For example, the division of the modules or units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0102] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0103] The embodiments described above are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, a person skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. Such modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.
Claims
1. An index extraction method, characterized in that: include: Get pending questions input by the user; Determine a first similarity index according to the problem matched by the problem to be processed; wherein the first similarity index is an index corresponding to the problem matched by the problem to be processed; Determining an indicator prompt word of a pre-trained language model according to the first similarity indicator, and extracting an indicator from the problem to be processed using the indicator prompt word and the pre-trained language generation model to obtain a first indicator; wherein the indicator prompt word is used to guide the pre-trained language generation model to extract an indicator using the first similarity indicator as a positive example; A target indicator is determined according to the problem to be processed and the first indicator, wherein the target indicator is an indicator similar to both the problem to be processed and the first indicator.
2. The index extraction method according to claim 1, characterized in that: The determining of the first similarity index according to the problem matched with the problem to be processed includes: Determine one or more questions matched from a question knowledge base as the first candidate question; wherein the question knowledge base includes a plurality of preset questions and question indicators corresponding to the questions; The question indicator corresponding to the first candidate question is determined as the first similarity indicator.
3. The index extraction method according to claim 2, characterized in that: The step of determining the indicator prompt word of the pre-trained language model according to the first similarity indicator includes: Constructing a problem indicator sample according to the first candidate problem and the first similarity indicator; Fill the problem indicator sample into a pre-constructed prompt word template to obtain the indicator prompt word; wherein the pre-constructed prompt word template includes instructions or rules for guiding the pre-trained language model to extract indicators.
4. The index extraction method according to any one of claims 1 to 3, characterized in that: The step of determining the target indicator according to the problem to be processed and the first indicator includes: Matching an indicator similar to the problem to be processed from a pre-built indicator knowledge base to obtain a second similar indicator; Matching an indicator similar to the first indicator from the pre-built indicator knowledge base to obtain a third similar indicator; The target indicator is determined according to the second similarity indicator and the third similarity indicator.
5. The index extraction method according to claim 4, characterized in that: The determining the target indicator according to the second similarity indicator and the third similarity indicator comprises: Determine a candidate indicator set according to the second similarity indicator and the third similarity indicator; wherein the candidate indicators in the candidate indicator set are indicators that have appeared in the second similarity indicator and / or the third similarity indicator, and the first similarity corresponding to the candidate indicator is determined according to the number of times it appears in the second similarity indicator and / or the third similarity indicator; For any candidate indicator in the candidate indicator set, determine the character similarity between the candidate indicator and the question to be processed to obtain a second similarity, and determine the historical preference similarity between the candidate indicator and the historical questions recorded by the user to obtain a third similarity; Determine a comprehensive similarity of each candidate indicator in the candidate indicator set according to the first similarity, the second similarity, and the third similarity; Determine the candidate indicator with the highest comprehensive similarity in the candidate indicator set as the target indicator.
6. The index extraction method according to claim 5, characterized in that: The step of determining the character similarity between the candidate indicator and the problem to be processed to obtain a second similarity includes: Splitting the problem to be processed into single characters to obtain a problem character set, and splitting the candidate indicators into single characters to obtain an indicator character set; Determine the intersection of the question character set and the indicator character set; The second similarity is determined according to the number of characters in the intersection.
7. The index extraction method according to claim 5, characterized in that: The determining of the historical preference similarity between the candidate indicator and the historical question recorded by the user to obtain a third similarity includes: Matching the question knowledge base according to the question to be processed and the user information of the user to obtain a second candidate question; Determining the problem indicator corresponding to the second candidate problem as a second similarity indicator; When the second similarity index is the same as the candidate index, the matching score of the second candidate question corresponding to the second similarity index in the matching process is used as the historical preference similarity of the candidate index, and the largest historical preference similarity is selected to obtain the third similarity.
8. The index extraction method according to any one of claims 1 to 3, characterized in that: After determining the target indicator according to the problem to be processed and the first indicator, the method further includes: The target indicator that the user feedback is correct is stored.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the method according to any one of claims 1 to 8 is implemented.
10. A computer program product, characterized in that The invention comprises a computer program, which implements the method according to any one of claims 1 to 8 when being executed.
Citation Information
Patent Citations
Deep learning-based intelligent customer service semantic retrieval method, device and system
CN114817461A
Intelligent number asking and reading method, system and equipment based on large language model and medium
CN117874181A
Intention recognition method, system and equipment
CN119358563A
Intelligent number asking method and device and readable storage medium
CN119739730A
Merchandise retrieval device having function for presenting reference keyword and merchandise retrieval method
JP2009151734A