Method for determining recall content, device for determining recall content, and electronic device
By calculating the similarity between the paragraph vector and the target vector and weighted calculations based on the number of keyword matching, the recall content is determined, which solves the problem of low matching accuracy in traditional question-and-answer systems and improves the accuracy and relevance of the recall content.
Patent Information
- Application Number
- CN202311616680.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-29
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2043-11-29
AI Technical Summary
When traditional document question and answer systems deal with a large number of documents, it is difficult to accurately match problems and documents, resulting in more noise in recalls and low accuracy.
By obtaining the target question and the corresponding question and answer documents, the similarity between the paragraph vector and the target vector is calculated, and the weighted calculation is performed based on the number of matches of the target keywords and the semantic similarity, and the recall score is obtained to determine the recall content.
Improves matching accuracy between questions and documents, reduces noise, and provides more accurate recall content, allowing users to obtain the most relevant and high-quality answers.
Smart Images

Figure CN117708283B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computers, and particularly relates to a method for determining recalled content, a device for determining recalled content, and an electronic device. Background Art
[0002] In a traditional document question-and-answer system, the method for answer recall is usually based on the semantic similarity between the question and each paragraph of the document, which has certain limitations. For example, when the document content is too large, it is difficult to accurately find the answer paragraph related to the question, resulting in a low matching accuracy between the question and the document, a lot of recall noise, and thus affecting the accuracy of recall. Summary of the Invention
[0003] The embodiments of the present application provide a method for determining recalled content, a device for determining recalled content, and an electronic device, which can improve the accuracy of recall.
[0004] In a first aspect, the embodiments of the present application provide a method for determining recalled content, the method including: obtaining a target question sentence and a plurality of Q&A documents corresponding to the target question sentence; respectively calculating the similarity between each paragraph vector corresponding to each Q&A document and a target vector corresponding to the target question sentence, where each paragraph vector is a feature vector of each paragraph content, the target vector includes a sentence vector and a plurality of word vectors, and the similarity includes a first semantic similarity between each paragraph vector and each word vector, and a second semantic similarity between each paragraph vector and the sentence vector; performing a weighted calculation on the number of target keyword matches corresponding to each paragraph content, a plurality of the first semantic similarities, and a plurality of the second semantic similarities to obtain a recall score for each paragraph content, where the target keyword is obtained by expanding an initial keyword corresponding to the target question sentence through a target expansion model; determining recalled content according to the recall score of each paragraph content.
[0005] Second aspect, an embodiment of the present application provides a device for determining recalled content. The device includes: an acquisition module, configured to acquire a target question sentence and a plurality of Q&A documents corresponding to the target question sentence; a first calculation module, configured to calculate the similarity between each paragraph vector corresponding to each Q&A document and a target vector corresponding to the target question sentence respectively, where each of the paragraph vectors is a feature vector of each paragraph content, the target vector includes a sentence vector and a plurality of word vectors, and the similarity includes a first semantic similarity between each paragraph vector and each word vector, and a second semantic similarity between each paragraph vector and the sentence vector; a second calculation module, configured to perform weighted calculation on the number of target keyword matches corresponding to each paragraph content, a plurality of the first semantic similarities, and a plurality of the second semantic similarities to obtain a recall score for each paragraph content, where the plurality of target keywords are obtained by expanding the initial keywords corresponding to the target question sentence through a target expansion model; a determination module, configured to determine the recalled content according to the recall score of each paragraph content.
[0006] Third aspect, an embodiment of the present application provides an electronic device, which includes a processor, a memory, and a program or instruction stored on the memory and executable on the processor. When the program or instruction is executed by the processor, the steps of the method described in the first aspect are implemented.
[0007] Fourth aspect, an embodiment of the present application provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the steps of the method described in the first aspect are implemented.
[0008] Fifth aspect, an embodiment of the present application provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor, and the processor is configured to run a program or instruction to implement the method described in the first aspect.
[0009] In the embodiments of the present application, after obtaining the target question sentence and multiple Q&A documents corresponding to the target question sentence, the similarity between each paragraph vector corresponding to each Q&A document and the target vector corresponding to the target question sentence is calculated respectively, where each paragraph vector is the feature vector of each paragraph content, the target vector includes a sentence vector and multiple word vectors, and the similarity includes a first semantic similarity between each paragraph vector and each word vector, and a second semantic similarity between each paragraph vector and the sentence vector; then the number of target keyword matches corresponding to each paragraph content, multiple first semantic similarities, and multiple second semantic similarities are weighted and calculated to obtain the recall score of each paragraph content, where the target keyword is obtained by expanding the initial keyword corresponding to the target question sentence through the target expansion model. Finally, according to the recall score of each paragraph content, the recalled content is determined. In this way, not only the semantic similarities of the word vectors and sentence vectors between the target statement and the Q&A documents are calculated, but also the matching quantity of the keywords and expanded words in the paragraph content slices is considered. This multi-dimensional matching method improves the matching accuracy between the question and the document. At the same time, based on the weighted score result, more accurate recalled content can be provided, enabling the user to obtain the most relevant and high-quality answers. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] Figure 1 FIG. is a schematic flowchart of a method for determining recalled content provided by an embodiment of the present application;
[0011] Figure 2 FIG. is another schematic flowchart of a method for determining recalled content provided by an embodiment of the present application;
[0012] Figure 3 FIG. is a schematic structural diagram of a device for determining recalled content provided by an embodiment of the present application;
[0013] Figure 4 FIG. is a schematic structural diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0014] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0015] Next, a method for determining recalled content, a device for determining recalled content, and an electronic device provided by an embodiment of the present application will be described in detail with reference to the accompanying drawings through specific embodiments and their application scenarios.
[0016] Figure 1Disclosed is a method for determining recall content provided by an embodiment of the present application. This method can be executed by an electronic device, which may include: a server and / or a terminal device. In other words, this method can be executed by software or hardware installed in the electronic device. The method includes the following steps:
[0017] S110: Obtain a target question sentence and multiple Q&A documents corresponding to the target question sentence.
[0018] Among them, the target question sentence can be the key content input by the user according to their query purpose. Optionally, the target question sentence can be a complete sentence, a phrase, or a simple description. The multiple Q&A documents include information documents of multiple types.
[0019] S120: Calculate the similarity between each paragraph vector corresponding to each Q&A document and the target vector corresponding to the target question sentence respectively.
[0020] Among them, each paragraph vector is a feature vector of each paragraph content. The target vector includes a sentence vector and multiple word vectors. The similarity includes a first semantic similarity between each paragraph vector and each word vector, and a second semantic similarity between each paragraph vector and the sentence vector.
[0021] Among them, before calculating the similarity between each paragraph vector corresponding to each Q&A document and the target vector corresponding to the target question sentence, it may further include: segmenting the target question sentence, for example, dividing the target question sentence into multiple single characters or words, and then deleting some words irrelevant to the question itself. For example, in the sentence "How is the execution process of the XX language model?", words such as "of", "is", and "how" are non-keywords and can be deleted. After segmenting, perform vectorization processing on each segment to obtain a word vector corresponding to each word. Optionally, an n-gram model can be used to segment the target question sentence of the question content. Further, it may further include: performing vectorization processing on the target question sentence to obtain a sentence vector corresponding to the target question sentence.
[0022] Further, before calculating the similarity between each paragraph vector corresponding to each Q&A document and the target vector corresponding to the target question sentence, it may further include: slicing each Q&A document to obtain the paragraph content corresponding to each Q&A document, and then performing vectorization processing on each paragraph content to obtain a paragraph vector corresponding to each paragraph content. Optionally, a document parsing model can be used to slice the document content into paragraph content slices according to paragraphs. Optionally, the multiple Q&A documents can be merged into one document.
[0023] After determining the sentence vector corresponding to the target sentence, the word vectors corresponding to each word segment, and the paragraph vectors corresponding to the content of each paragraph, the first semantic similarity between each paragraph vector and each word vector, and the second semantic similarity between each paragraph vector and the sentence vector can be calculated respectively. For example, the content of multiple paragraphs includes a first paragraph and a second paragraph, the paragraph vector corresponding to the first paragraph is D1, the paragraph vector corresponding to the second paragraph is D2, the multiple word vectors include W1, W2, W3, and the sentence vector is S1. Among them, this example takes the first semantic similarity and the second semantic similarity as cosine similarities for illustration. Then, calculate the first semantic similarity between the paragraph vector corresponding to the first paragraph and each word vector: cos(D1, W1), cos(D1, W2), cos(D1, W3); calculate the second semantic similarity between the paragraph vector corresponding to the first paragraph and the sentence vector: cos(D1, S1); calculate the first semantic similarity between the paragraph vector corresponding to the second paragraph and each word vector: cos(D2, W1), cos(D2, W2), cos(D2, W3); calculate the second semantic similarity between the paragraph vector corresponding to the second paragraph and the sentence vector: cos(D2, S1). That is to say, each paragraph content corresponds to multiple first semantic similarities and one second semantic similarity.
[0024] S130: Weightedly calculate the number of target keyword matches corresponding to each of the paragraph contents, the multiple first semantic similarities, and the multiple second semantic similarities to obtain the recall score for each of the paragraph contents.
[0025] Among them, the target keyword is obtained by expanding the initial keyword corresponding to the target question through a target expansion model.
[0026] It can be understood that the number of target keyword matches corresponding to each of the paragraph contents refers to the number of target keywords hit by each paragraph content. That is to say, the number of target keywords hit by each paragraph content is weightedly calculated with the corresponding first semantic similarity and second semantic similarity, so that the recall score corresponding to each paragraph content can be obtained.
[0027] S140: Determine the recalled content according to the recall score of each of the paragraph contents.
[0028] Among them, in one implementation manner, the recall scores can be sorted from largest to smallest, and the paragraph content corresponding to the recall scores ranked in the top preset positions can be used as the recalled content. In another implementation manner, the determining the recalled content according to the recall score of each of the paragraph contents includes: determining the paragraph content with a recall score greater than a preset threshold as the recalled content corresponding to the target question. In this way, the most relevant and high-quality answers can be presented to the user, thereby increasing user satisfaction.
[0029] In an embodiment of the present application, after obtaining a target question sentence and a plurality of Q&A documents corresponding to the target question sentence, the similarity between each paragraph vector corresponding to each Q&A document and the target vector corresponding to the target question sentence is calculated respectively, where each paragraph vector is a feature vector of each paragraph content, the target vector includes a sentence vector and a plurality of word vectors, and the similarity includes a first semantic similarity between each paragraph vector and each word vector, and a second semantic similarity between each paragraph vector and the sentence vector; then, the number of target keyword matches corresponding to each paragraph content, a plurality of first semantic similarities, and a plurality of second semantic similarities are weighted and calculated to obtain a recall score for each paragraph content, where the target keyword is obtained by expanding the initial keyword corresponding to the target question sentence through a target expansion model. Finally, according to the recall score of each paragraph content, the recalled content is determined. In this way, not only the semantic similarities of the word vectors and sentence vectors between the target statement and the Q&A documents are calculated, but also the matching quantity of the keywords and the expanded words in the paragraph content slices is considered. This multi-dimensional matching method improves the matching accuracy between the question and the document. At the same time, based on the weighted score result, more accurate recalled content can be provided, enabling the user to obtain the most relevant and high-quality answers.
[0030] In one implementation, before calculating the weighted sum of the number of target keyword matches corresponding to each paragraph content, a plurality of first semantic similarities, and a plurality of second semantic similarities to obtain the recall score for each paragraph content, the method further includes: inputting the target question sentence into a target keyword extraction model to obtain a plurality of the initial keywords of the target question sentence. In another implementation, expanding the initial keyword corresponding to the target question sentence through a target expansion model includes: inputting a plurality of the initial keywords into the target expansion model to obtain expanded words semantically related to the plurality of the initial keywords; and combining the plurality of the initial keywords and the expanded words semantically related to the plurality of the initial keywords to obtain a plurality of the target keywords.
[0031] It can be understood that the target keyword extraction model is used to automatically extract keywords or key phrases from the target question, and the target expansion model is used to expand based on the extracted keywords or key phrases to generate semantically related expansion words. It should be noted that when the target expansion model performs semantic expansion, it can expand the semantic relevance of the initial keywords, not limited to synonyms. For example, for the target statement "What is the net profit of xx company this year?", the target keyword extraction model is used to extract keywords, and the obtained keywords are: xx, company, this year, net profit. Then, the target expansion model is used to expand the above keywords. Then, the expansion words corresponding to "company" can include operation, employees, development, shareholders, products, etc.; the expansion words corresponding to "this year" can include the whole year of 2023, annual, etc.; the expansion words corresponding to "net profit" can include finance, revenue, US dollars, hundreds of millions of yuan, losses, etc. Therefore, in this implementation method, by using the target keyword extraction model, the keywords of the target statement can be extracted more accurately, which helps to better understand the meaning of the target statement, thereby improving the quality of the question-answering system. At the same time, it also avoids the user's question being too colloquial, causing difficulties in semantic understanding. By using the target expansion model, the semantic matching degree between the document and the question can be improved, and the user's intention can be better understood when matching answers.
[0032] Among them, in another implementation method, before inputting the target question into the target keyword extraction model to obtain multiple initial keywords of the target question, the method further includes: obtaining a training instruction set, where the training instruction set contains multiple sample questions and the question keywords in each sample question; inputting the training instruction set into the initial keyword extraction model for keyword extraction; and iteratively training the initial keyword extraction model a target number of times based on the keyword extraction result and the question keyword to obtain the target keyword extraction model.
[0033] Exemplarily, the training instruction set can be expressed as:
[0034]
[0035] It should be noted that the above instruction set is only an example, and the present application does not make specific limitations.
[0036] During the training process of the initial keyword extraction model, the instruction set is input into the initial keyword extraction model for keyword extraction to obtain a keyword extraction result. Then, the keyword extraction result is compared with the corresponding keyword in the instruction set. Based on the comparison result, the initial keyword extraction model is fine-tuned and then continued to be trained. After iteratively training a target number of times, the trained initial keyword extraction model, that is, the target keyword extraction model, can be obtained.
[0037] Optionally, the initial keyword extraction model can be constructed based on a large language model (LLM).
[0038] In one implementation, before calculating the weighted sum of the number of target keyword matches corresponding to each paragraph content, the multiple first semantic similarities, and the multiple second semantic similarities to obtain the recall score for each paragraph content, the method further includes: determining the number of target keyword matches corresponding to each paragraph content by respectively matching each paragraph content with multiple target keywords. It can be understood that respectively matching each paragraph content with multiple target keywords means hitting multiple target keywords with each paragraph content, and the number of hit target keywords is the number of target keyword matches corresponding to the paragraph content. It should be noted that matching or hitting not only refers to exact identity but also includes semantic correspondence. Exemplarily, assuming the target statement is "What are the ancient poems about rain in spring", the determined initial keywords are "spring", "rain", "spring rain", and "ancient poem", and the extended words determined based on the initial keywords include "green, gentle rain, moisten, regulated verse, quatrain", etc.; assuming multiple target documents include the following content:
[0039] Document 1: Good rain knows its season, arriving when spring comes. With wind it steals in by night, moistening everything gently without a sound.
[0040] Document 2: The fine rain that wets the clothes but not the apricot blossoms, the gentle breeze that caresses the face without a chill.
[0041] Document 3: The gentle rain on the imperial street is as soft as butter, and the grass color can be faintly seen but seems to disappear when approached closely.
[0042] Document 4: Spring rain melts the remaining frost, and the warm wind reaches the cold ashes.
[0043] Document 5: In the east wind, the gentle rain that skims across my face, like stars, is the fluff of spring.
[0044] Then the keywords that Document 1 can match include: spring, rain, spring rain, gentle rain, moisten, regulated verse, that is, the number of target keyword matches is 6; the keywords that Document 2 can match include: spring, rain, spring rain, gentle rain, quatrain, that is, the number of target keyword matches is 5; the keywords that Document 3 can match include: green, spring, rain, spring rain, gentle rain, quatrain, that is, the number of target keyword matches is 6. The keywords that Document 4 can match include: spring, rain, spring rain, regulated verse, that is, the number of target keyword matches is 4; the keywords that Document 5 can match include: spring, rain, spring rain, that is, the number of target keyword matches is 3. In this implementation, not limited to exact matches, semantic correspondence is also considered, which can improve the breadth, relevance, and diversity of the recalled content.
[0045] In one implementation, the weighted calculation of the number of target keyword matches corresponding to each of the paragraph contents, the multiple first semantic similarities, and the multiple second semantic similarities includes: a first weight can be set for the number of target keyword matches corresponding to each of the paragraph contents, a second weight can be set for each of the first semantic similarities, and a third weight can be set for each of the second semantic similarities; according to the first weight, the second weight, and the third weight, the weighted sum of the number of matches, the first semantic similarity, and the second semantic similarity corresponding to each of the paragraph contents is calculated to determine the recall score corresponding to each paragraph content.
[0046] Among them, the formula for weighted calculation can be:
[0047] score1 = w2S w,p + w3S s,p + w1N w,o ;
[0048] Among them, w1 is the first weight, N w,p is the number of target keyword matches corresponding to the first paragraph content, w2 is the second weight, S w,p is the first semantic similarity, w3 is the third weight, S s,p is the second semantic similarity. In this implementation, by calculating the weighted scores of each dimension, more accurate recall segments can be provided, enabling users to find relevant information faster.
[0049] Furthermore, based on the above various implementations, as Figure 2 shown, the embodiments of the present application further provide a schematic diagram of a method for determining recall content, and this method may include the following steps:
[0050] S201: Obtain word vectors.
[0051] The n-gram model can be used to segment the question content. N-Gram is an algorithm based on a statistical language model. It can perform a sliding window operation of size N on the content in the text according to bytes, forming a sequence of byte segments of length N, and then encode each segmented word to obtain word vectors.
[0052] S202: Obtain sentence vectors.
[0053] Encode the entire question sentence to obtain sentence vectors.
[0054] S203: Obtain keywords.
[0055] Extract the keywords of the question through a keyword extraction model.
[0056] S204: Expand keywords.
[0057] Semantic expansion can be performed based on LSI.
[0058] S205: Paragraph slicing.
[0059] A document parsing model can be used to slice the content of each document into paragraph content slices according to paragraphs.
[0060] S206: Obtain paragraph vectors.
[0061] Calculate the paragraph vectors corresponding to the paragraph content of all slices of the document.
[0062] S207: Obtain the number of keyword matches.
[0063] S208: Calculate the cosine similarity.
[0064] Among them, the cosine similarity includes the cosine similarity between each paragraph content and each word vector, which can be denoted as , and the cosine similarity between each paragraph content and each word vector, which can be denoted as , .
[0065] S209: Calculate the weighted score.
[0066] Among them, the number of keyword matches corresponding to each paragraph content can be denoted as N w,p . Perform weighted calculation on S w,p , S s,p and N w,p , and the calculation formula can be:
[0067] score = w1S w,p + w2S s,p + w3N w,p ;
[0068] Among them, score is the calculation result of the single-paragraph content slice and words, sentences, and keywords. For multiple paragraph content slices, calculate score respectively. Finally, obtain the score set;
[0069] S210: Determine the recall segments.
[0070] The segment content ranked top k in the score set can be used as the recall segments of the question.
[0071] In this implementation method, multiple semantic processing methods are combined, including the n-gram model, word vector embedding, sentence vector embedding, keyword extraction, and LSI latent semantic expansion. This comprehensive processing can better capture the semantic information of the text and improve the matching accuracy between the question and the document.
[0072] Figure 3 A schematic structural diagram of a recall content determination device provided by an embodiment of the present application is shown, as Figure 3 shown. The recall content determination device 300 may include: an acquisition module 310, a first calculation module 320, a second calculation module 330, and a determination module 340.
[0073] In this embodiment, the acquisition module 310 is configured to acquire a target question sentence and a plurality of Q&A documents corresponding to the target question sentence; the first calculation module 320 is configured to calculate the similarity between each paragraph vector corresponding to each Q&A document and a target vector corresponding to the target question sentence respectively, where each paragraph vector is a feature vector of each paragraph content, the target vector includes a sentence vector and a plurality of word vectors, and the similarity includes a first semantic similarity between each paragraph vector and each word vector, and a second semantic similarity between each paragraph vector and the sentence vector; the second calculation module 330 is configured to perform weighted calculation on the number of target keyword matches corresponding to each paragraph content, a plurality of the first semantic similarities, and a plurality of the second semantic similarities to obtain a recall score for each paragraph content, where the plurality of target keywords are obtained by expanding the initial keywords corresponding to the target question sentence through a target expansion model; the determination module 340 is configured to determine the recall content according to the recall score of each paragraph content.
[0074] In one implementation manner, the device further includes: a first processing module, configured to, before performing weighted calculation on the number of target keyword matches corresponding to each paragraph content, a plurality of the first semantic similarities, and a plurality of the second semantic similarities to obtain a recall score for each paragraph content, obtain a plurality of the initial keywords of the target question sentence by inputting the target question sentence into a target keyword extraction model.
[0075] In one implementation manner, the device further includes: a training module, configured to, before obtaining a plurality of the initial keywords of the target question sentence by inputting the target question sentence into a target keyword extraction model, acquire a training instruction set, where the training instruction set includes a plurality of sample questions and the query keywords in each sample question; input the training instruction set into an initial keyword extraction model for keyword extraction; perform iterative training on the initial keyword extraction model for a target number of times based on the keyword extraction result and the query keywords to obtain the target keyword extraction model.
[0076] In one implementation, expanding the initial keywords corresponding to the target question by the target expansion model includes: inputting a plurality of the initial keywords into the target expansion model to obtain expansion words semantically related to the plurality of the initial keywords; and merging the plurality of the initial keywords and the expansion words semantically related to the plurality of the initial keywords to obtain a plurality of the target keywords.
[0077] In one implementation, the apparatus further includes: a second processing module, configured to determine the number of target keyword matches corresponding to each paragraph content by respectively matching each paragraph content with a plurality of the target keywords before calculating a weighted sum of the number of target keyword matches corresponding to each paragraph content, a plurality of the first semantic similarities, and a plurality of the second semantic similarities to obtain a recall score for each paragraph content.
[0078] In one implementation, calculating a weighted sum of the number of target keyword matches corresponding to each paragraph content, a plurality of the first semantic similarities, and a plurality of the second semantic similarities includes: setting a first weight for the number of target keyword matches corresponding to each paragraph content, setting a second weight for each of the first semantic similarities, and setting a third weight for each of the second semantic similarities; and performing a weighted sum of the number of matches, the first semantic similarity, and the second semantic similarity corresponding to each paragraph content according to the first weight, the second weight, and the third weight to determine a recall score for each paragraph content.
[0079] In one implementation, determining recall content according to the recall score of each paragraph content includes: determining the paragraph content with a recall score greater than a preset threshold as the recall content corresponding to the target question.
[0080] The apparatus for determining recall content provided by the embodiments of the present application can implement Figure 1 - Figure 2 each process implemented in the method embodiments shown. To avoid repetition, details are not described herein again.
[0081] The apparatus for determining recall content in the embodiments of the present application may be a device, or a component, an integrated circuit, or a chip in an electronic device. The embodiments of the present application do not make specific limitations.
[0082] The apparatus for determining recall content in the embodiments of the present application may be a device having an operating system. The operating system may be an Android operating system, an iOS operating system, or other possible operating systems. The embodiments of the present application do not make specific limitations.
[0083] Optionally, as Figure 4As shown in the figure, an embodiment of the present application further provides an electronic device 400, which includes a processor 410, a memory 420, and a program or instruction stored on the memory 420 and executable on the processor 410. When the program or instruction is executed by the processor 410, it implements the above Figure 1 - Figure 2 each process of the method embodiment shown, and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.
[0084] An embodiment of the present application further provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, it implements the above Figure 1 - Figure 2 each process of the method embodiment shown, and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.
[0085] Among them, the processor is the processor in the electronic device described in the above embodiment. The readable storage medium includes a computer-readable storage medium, such as a computer read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc, etc.
[0086] Another embodiment of the present application provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor, and the processor is used to run a program or instruction to implement the above Figure 1 - Figure 2 each process of the method embodiment shown. To avoid repetition, it will not be elaborated here.
[0087] It should be understood that the chip mentioned in the embodiment of the present application may also be referred to as a system-on-chip, a system chip, a chip system, or a system-on-chip.
[0088] It should be noted that in this article, the term "including", "comprising" or any other variant thereof is intended to cover a non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the phrase "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including that element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in a reverse order according to the functions involved. For example, the described methods may be performed in an order different from that described, and various steps may be added, omitted, or combined. In addition, the features described with reference to certain examples may be combined in other examples.
[0089] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present application, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in various embodiments of the present application.
[0090] The embodiments of the present application have been described above in conjunction with the accompanying drawings. However, the present application is not limited to the above specific implementation manners. The above specific implementation manners are merely illustrative and not restrictive. Under the inspiration of the present application, those of ordinary skill in the art can also make many forms without departing from the purpose of the present application and the scope protected by the claims, and all of them belong to the protection scope of the present application.
Claims
1. A method for determining recalled content, characterized in that, Including: Obtain a target question sentence and multiple Q&A documents corresponding to the target question sentence; Calculate the similarity between each paragraph vector corresponding to each Q&A document and the target vector corresponding to the target question sentence respectively. Wherein, each paragraph vector is a feature vector of each paragraph content, the target vector includes a sentence vector and multiple word vectors, the similarity includes a first semantic similarity between each paragraph vector and each word vector, and a second semantic similarity between each paragraph vector and the sentence vector. One paragraph content corresponds to multiple first semantic similarities and one second semantic similarity; Perform weighted calculation on the number of target keyword matches corresponding to each paragraph content, multiple first semantic similarities, and multiple second semantic similarities to obtain a recall score for each paragraph content. Wherein, the target keyword is obtained by performing semantic relevance and synonym expansion on the initial keyword corresponding to the target question sentence through a target expansion model; Determine the paragraph content with a recall score greater than a preset threshold as the recall content corresponding to the target question sentence. The recall content is the answer paragraph with the highest relevance to the target question sentence among the multiple Q&A documents; 2. The method according to claim 1, wherein Before performing the weighted calculation on the number of target keyword matches corresponding to each paragraph content, multiple first semantic similarities, and multiple second semantic similarities to obtain a recall score for each paragraph content, the method further includes: Input the target question sentence into a target keyword extraction model to obtain multiple initial keywords of the target question sentence; 3. The method according to claim 2, wherein Before inputting the target question sentence into the target keyword extraction model to obtain multiple initial keywords of the target question sentence, the method further includes: Obtain a training instruction set, wherein the training instruction set contains multiple sample questions and the query keywords in each sample question; Input the training instruction set into an initial keyword extraction model for keyword extraction; Through iterative training of the initial keyword extraction model for a target number of times based on the keyword extraction result and the query keyword, obtain the target keyword extraction model; 4. The method according to claim 2, wherein Expanding the initial keyword corresponding to the target question sentence through a target expansion model includes: Input multiple initial keywords into the target expansion model to obtain expansion words semantically related to the multiple initial keywords; Merge multiple initial keywords and expansion words semantically related to the multiple initial keywords to obtain multiple target keywords; 5. The method according to claim 1, wherein Before performing the weighted calculation on the number of target keyword matches corresponding to each paragraph content, multiple first semantic similarities, and multiple second semantic similarities to obtain a recall score for each paragraph content, the method further includes: Determine the number of target keyword matches corresponding to each paragraph content by respectively matching each paragraph content with multiple target keywords; 6. The method according to any one of claims 1-5, characterized in that, The weighted calculation of the number of target keyword matches corresponding to each paragraph content, multiple first semantic similarities, and multiple second semantic similarities includes: Set a first weight for the number of matches of the target keywords corresponding to each of the paragraph contents, a second weight for each of the first semantic similarities, and a third weight for each of the second semantic similarities; Based on the first weight, the second weight, and the third weight, perform a weighted sum of the number of matches, the first semantic similarity, and the second semantic similarity corresponding to each of the paragraph contents to determine the recall score corresponding to each paragraph content.
7. A device for determining recalled content, characterized in that, Includes: An acquisition module for acquiring a target question sentence and a plurality of Q&A documents corresponding to the target question sentence; A first calculation module for respectively calculating the similarity between each paragraph vector corresponding to each of the Q&A documents and a target vector corresponding to the target question sentence, where each of the paragraph vectors is a feature vector of each paragraph content, the target vector includes a sentence vector and a plurality of word vectors, and the similarity includes a first semantic similarity between each paragraph vector and each word vector, and a second semantic similarity between each paragraph vector and the sentence vector, and one paragraph content corresponds to a plurality of first semantic similarities and one second semantic similarity; A second calculation module for performing a weighted calculation on the number of matches of the target keywords corresponding to each of the paragraph contents, a plurality of the first semantic similarities, and a plurality of the second semantic similarities to obtain the recall score corresponding to each of the paragraph contents, where the plurality of target keywords are obtained by performing semantic relevance and synonym expansion on the initial keywords corresponding to the target question sentence through a target expansion model; A determination module for determining the paragraph content with a recall score greater than a preset threshold as the recall content corresponding to the target question sentence, where the recall content is the answer paragraph with the highest relevance to the target question sentence among the plurality of Q&A documents.
8. An electronic device, characterized in that, Includes a processor, a memory, and a program or instruction stored on the memory and executable on the processor, and when the program or instruction is executed by the processor, it implements the steps of the method for determining the recall content according to any one of claims 1-6.
9. A readable storage medium, characterized in that, A program or instruction is stored on the readable storage medium, and when the program or instruction is executed by the processor, it implements the steps of the method for determining the recall content according to any one of claims 1-6.
Citation Information
Patent Citations
Question and answer processing method and device, computer equipment and readable storage medium
CN112948562A