Title matching method and apparatus

By acquiring feature information from the search page and similar candidate pages, matching questions, and performing merging operations in response to anomalies, the problem of low accuracy in question search in existing technologies is solved, resulting in more accurate matching results and a better user experience.

CN115705731BActive Publication Date: 2026-08-25BEIJING YUANLI WEILAI SCI & TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202110921109.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-08-11
Publication Date
2026-08-25
Estimated Expiration
2041-08-11

AI Technical Summary

Technical Problem

Existing technologies have low accuracy when searching for questions with rich text content or graphic elements, especially in primary and preschool stages. Furthermore, full-page search technology cannot handle abnormal matching results, leading to inaccurate matching.

Method used

By acquiring feature information from the search page and similar candidate pages, and responding to anomalies after question matching, some questions are merged to correct the matching results, including duplicate or unmatched questions, and then re-matching to obtain accurate similar pages.

Benefits of technology

It improves the accuracy of question search, corrects abnormal matching results, enhances user experience, and ensures that the questions and solutions users want are matched.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115705731B_ABST
    Figure CN115705731B_ABST
Patent Text Reader

Abstract

The application relates to the field of natural language processing, and discloses a title matching method and device. The method comprises the following steps: obtaining a search page and a similar candidate page of the search page, wherein the search page comprises a first title set, the first title set comprises a title, and the candidate page comprises a second title set and title analysis of each candidate title in the second title set; performing title matching on the first title set and the second title set to obtain a title matching result; in response to the title matching result being abnormal, obtaining an abnormal type of the title matching result; according to the abnormal type, determining a title set that needs to be combined from the first title set and the second title set; performing title combination operation on at least part of the title set; and re-performing the title matching. The application can correct the abnormal title matching result, obtain an accurate similar page, and match the title and analysis that a user wants to search.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of natural language processing, and in particular to a topic matching method and apparatus thereof. Background Technology

[0002] Most common question search scenarios are based on a single question. The accuracy rate of searching for questions with rich text content is relatively high, while the effect of searching for questions with blurry or incomplete images is not good. In addition, there are a lot of graphic questions and questions with the same question stem in the primary school and preschool stages. In this case, searching for a single question is less effective.

[0003] Currently, while full-page search technology can improve the accuracy of single-question searches, it cannot handle abnormal matching results to obtain more accurate results. Summary of the Invention

[0004] This application aims to at least partially address one of the technical problems in the related art.

[0005] Therefore, one objective of this application is to propose a topic matching method.

[0006] The second objective of this application is to propose a question matching device.

[0007] The third objective of this application is to propose an electronic device.

[0008] The fourth objective of this application is to provide a non-transitory computer-readable storage medium.

[0009] The fifth objective of this application is to provide a computer program product.

[0010] To achieve the above objectives, a first aspect of this application proposes a question matching method, comprising: obtaining a search page and candidate pages similar to the search page, wherein the search page includes a first question set, the first question set includes target questions, and the candidate pages include a second question set and question explanations for each candidate question in the second question set; performing question matching on the first question set and the second question set to obtain question matching results; in response to an anomaly in the question matching results, obtaining the anomaly type of the question matching results; determining a target question set from the first question set and the second question set that needs to be merged based on the anomaly type; performing a question merging operation on at least a portion of the questions in the target question set, and re-performing the question matching. This application can correct abnormal question matching results during the question matching process, obtaining accurate similar pages, thereby matching the questions and explanations that the user wants to search for.

[0011] According to one embodiment of this application, the question merging operation on at least a portion of the questions in the target question set includes: determining the at least a portion of the questions from the target question set based on the question matching result; and performing a question merging operation on the at least a portion of the questions.

[0012] According to one embodiment of this application, determining the target question set to be merged from the first question set and the second question set based on the exception type includes: in response to the exception type indicating that there is a first target question with duplicate matching in the first question set, then taking the first question set as the target question set.

[0013] According to one embodiment of this application, determining the at least part of the questions includes: determining the first target questions with duplicate matching as the at least part of the questions.

[0014] According to one embodiment of this application, before performing the question merging operation on at least a portion of the questions in the target question set, the method further includes: obtaining location information of the at least a portion of the questions; determining the area of ​​the question region formed by the question merging operation based on the location information; in response to the question region area being greater than or equal to an area threshold, not performing the question merging operation on at least a portion of the questions in the target question set; obtaining the similarity between the first target question and the matched first candidate question; selecting the first target question with the highest similarity to the first candidate question, and establishing a matching relationship between the target question and the first candidate question.

[0015] According to one embodiment of this application, determining the target question set to be merged from the first question set and the second question set according to the exception type includes: in response to the exception type indicating that there is a second target question in the first question set that has not been successfully matched, then taking the second question set as the target question set.

[0016] According to one embodiment of this application, determining the at least part of the topics includes: obtaining the word count of the candidate topics, and selecting a second candidate topic with a word count less than the first target topic, and determining it as the at least part of the topics.

[0017] According to one embodiment of this application, after determining the at least part of the questions, the method further includes: performing a question merging operation on the at least part of the questions to generate merged candidate questions; obtaining the similarity between the second target question and the merged candidate questions; and establishing a matching relationship between the second target question and the merged candidate questions in response to the similarity being greater than or equal to the similarity between the second target question and the at least part of the questions before merging.

[0018] According to one embodiment of this application, after re-matching the questions, the method further includes: in response to the fact that there are still unmatched third target questions in the first question set, obtaining a fourth target question adjacent to the third target question; and obtaining a third candidate question matched by the third target question based on the fourth candidate question matched by the fourth target question.

[0019] According to one embodiment of this application, obtaining a third candidate question matching the third target question based on a fourth candidate question matching the fourth target question includes: obtaining adjacent candidate questions adjacent to the fourth candidate question; obtaining a first positional relationship between the third target question and the fourth target question; and determining the third candidate question for the third target question from the adjacent candidate questions based on the first positional relationship.

[0020] According to one embodiment of this application, determining the third candidate question from the adjacent candidate questions based on the first positional relationship includes: obtaining a second positional relationship between the fourth candidate question and the adjacent candidate questions; determining a target second positional relationship from the second positional relationship, wherein the target second positional relationship is the same as the positional relationship of the first positional relationship; and determining the adjacent candidate question corresponding to the target second positional relationship as the third candidate question.

[0021] To achieve the above objectives, a second aspect of this application provides a question matching apparatus, comprising: a first acquisition module, configured to acquire a search page and a candidate page similar to the search page, wherein the search page includes a first question set, the first question set includes target questions, and the candidate page includes a second question set and question parsing for each candidate question in the second question set; a matching module, configured to perform question matching on the first question set and the second question set to obtain a question matching result; a second acquisition module, configured to acquire the exception type of the question matching result in response to an exception type; a determination module, configured to determine a target question set that needs to be merged from the first question set and the second question set based on the exception type; and a merging module, configured to perform a question merging operation on at least a portion of the questions in the target question set and re-perform the question matching.

[0022] According to one embodiment of this application, the merging module is further configured to: determine the at least part of the questions from the target question set based on the question matching result; and perform a question merging operation on the at least part of the questions.

[0023] According to one embodiment of this application, the determining module is further configured to: in response to the exception type indicating that there is a duplicate matching first target question in the first question set, then take the first question set as the target question set.

[0024] According to one embodiment of this application, the determining module is further configured to: determine the first target question with duplicate matching as the at least part of the questions.

[0025] According to one embodiment of this application, before the merging module, the method further includes: obtaining location information of at least some of the questions; determining the area of ​​the question region formed by the question merging operation based on the location information; in response to the question region area being greater than or equal to an area threshold, not performing the question merging operation on at least some of the questions in the target question set; obtaining the similarity between the first target question and the matched first candidate question; selecting the first target question with the highest similarity to the first candidate question, and establishing a matching relationship between the target question and the first candidate question.

[0026] According to one embodiment of this application, the determining module is further configured to: in response to the exception type indicating that there is a second target question in the first question set that has not been successfully matched, then take the second question set as the target question set.

[0027] According to one embodiment of this application, the determining module is further configured to: obtain the word count of the candidate questions, and select a second candidate question with a word count less than the second target question, and determine it as the at least part of the questions.

[0028] According to one embodiment of this application, after the merging module, the method further includes: performing a question merging operation on the at least some questions to generate merged candidate questions; obtaining the similarity between the second target question and the merged candidate questions; and establishing a matching relationship between the second target question and the merged candidate questions in response to the similarity being greater than or equal to the similarity of the second target question before merging.

[0029] According to one embodiment of this application, after the matching module, the method further includes: in response to the fact that there is still a third target question that has not been matched in the first question set, obtaining a fourth target question that is adjacent to the third target question; and obtaining a third candidate question that matches the third target question based on the fourth candidate question that matches the fourth target question.

[0030] According to one embodiment of this application, the second acquisition module is further configured to: acquire adjacent candidate questions that are adjacent to the fourth candidate question; acquire a first positional relationship between the third target question and the fourth target question; and determine the third candidate question from the adjacent candidate questions based on the first positional relationship.

[0031] According to one embodiment of this application, the second acquisition module is further configured to: acquire a second positional relationship between the fourth candidate question and the adjacent candidate questions; determine a target second positional relationship from the second positional relationship, wherein the target second positional relationship is the same as the positional relationship of the first positional relationship; and determine the adjacent candidate question corresponding to the target second positional relationship as the third candidate question.

[0032] To achieve the above objectives, a third aspect of this application provides an electronic device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor to implement the question matching method as described in the first aspect of this application.

[0033] To achieve the above objectives, a fourth aspect of this application provides a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to implement the question matching method as described in the first aspect of this application.

[0034] To achieve the above objectives, a fifth aspect of this application provides a computer program product, including a computer program that, when executed by a processor, implements the question matching method as described in the first aspect of this application. Attached Figure Description

[0035] Figure 1 This is a flowchart illustrating a question matching method according to one embodiment of this application;

[0036] Figure 2 This is a specific example diagram of a title matching method according to one embodiment of this application;

[0037] Figure 3 This is a specific example diagram illustrating another title matching method according to one implementation of this application;

[0038] Figure 4 This is a flowchart illustrating another question matching method according to one embodiment of this application;

[0039] Figure 5 This is a flowchart illustrating another question matching method according to one embodiment of this application;

[0040] Figure 6 This is a flowchart illustrating another question matching method according to one embodiment of this application;

[0041] Figure 7 This is a flowchart illustrating another question matching method according to one embodiment of this application;

[0042] Figure 8 This is a schematic diagram of a question matching device according to one embodiment of this application;

[0043] Figure 9 This is a schematic diagram of an electronic device according to one embodiment of this application. Detailed Implementation

[0044] The embodiments of this application are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain this application, and should not be construed as limiting this application.

[0045] Figure 1 This is a schematic diagram of a question matching method proposed in an embodiment of this application, as shown below. Figure 1 As shown, the question matching method includes the following steps:

[0046] S101, obtain the search page and similar candidate pages, wherein the search page includes a first set of questions, the first set of questions includes the target question, and the candidate pages include a second set of questions and question analysis for each candidate question in the second set of questions.

[0047] In this embodiment of the application, a user can take a picture of a page on a test / workbook using a terminal device (e.g., a mobile phone, a tablet computer, etc.) as the search page to be searched. The user can then upload the search page to the server through a client on the terminal device. The server can identify the search page and extract its feature information. For example, the feature information may include at least one of text information, graphic information, and spatial coordinate information.

[0048] Before searching, a question bank can be pre-built, storing all pages of test questions / workbooks as candidate pages in the question bank. Feature information of each candidate page can be extracted, and after establishing a relationship with that candidate page, it can be stored in the question bank. The candidate pages also include question explanations for each question.

[0049] In some implementations, after obtaining the feature information of the search page, the server can search for candidate pages similar to the search page in the question bank based on this feature information, and obtain the question and question analysis that the user wants to search for from the second question set of the candidate pages. There can be one or more similar candidate pages for the search page.

[0050] One possible approach is to compare the feature information of the search page with the feature information of each candidate page in the question bank, and then determine the similar candidate pages of the search page from the candidate pages in the question bank based on the comparison similarity.

[0051] As another possible implementation, the question bank can store sub-question banks for each subject. When searching a search page, the corresponding sub-question bank can be determined first based on the subject type of the search page. Then, based on the feature information of the search page, similar candidate pages can be searched from the candidate pages of the sub-question banks to narrow the search scope and improve search efficiency. For example, the subject type can be determined based on the title or text information of the search page.

[0052] S102, perform question matching on the first question set and the second question set to obtain the question matching results.

[0053] In this embodiment of the application, after obtaining the search page and the candidate pages similar to the search page, the server can perform question matching on the first question set and the second question set by setting a strategy (e.g., text matching, image matching, question position coordinate matching, etc.) and obtain the matching results. The matching results include the correspondence between the target questions in the first question set and the candidate questions in the second question set.

[0054] In some implementations, the location coordinates of the target question in the first question set can be used to match the corresponding candidate question in the second question set, thus determining the correspondence between the target question and the candidate question.

[0055] In other implementations, the correspondence between the target question and the candidate question can be determined by matching the text information of the target question in the first question set with the text information of the candidate questions in the second question set based on the similarity.

[0056] In other implementations, the similarity can be matched between the image information of the target question in the first question set and the image information of the candidate questions in the second question set, and the correspondence between the target question and the candidate questions can be determined based on the similarity.

[0057] In other embodiments, the similarity between the target question and the candidate question is calculated based on the text information and the image information, and the similarity is weighted. The correspondence between the target question and the candidate question is determined based on the weighted similarity.

[0058] In other embodiments, the location coordinates of the target question in the first question set are used to match the candidate question at the corresponding position in the second question set, and the similarity of the two questions at the position is matched based on the text information and / or image information. The correspondence between the target question and the candidate question is determined based on the matched similarity.

[0059] S103, in response to an anomaly in the question matching result, the anomaly type of the question matching result is obtained. In this embodiment of the application, when the server performs question matching on the first question set and the second question set, if each question in the first question set matches the only question with high similarity in the second question set, that is, a one-to-one match, the matching result is normal; otherwise, the matching result is abnormal. The abnormal situation may be many-to-one, one-to-many, many-to-many, or no similar question is matched. When the question matching result has the above-mentioned abnormal situation, the anomaly type of the question matching result can be obtained.

[0060] S104. Based on the exception type, determine the target question set that needs to be merged from the first question set and the second question set.

[0061] In this embodiment, when anomalies occur in the matching results, different processing methods can be applied to different anomaly types. Specifically, the anomaly-matching questions in the first or second question set are merged to obtain normal matching results. These anomaly types include two categories: First, the first question set contains duplicate target questions, i.e., multiple questions match the same similar question (many-to-one), in which case the multiple questions in the first question set need to be merged; second, the first question set contains unmatched target questions, i.e., one question matches multiple similar questions (one-to-many), in which case the multiple similar questions in the second question set need to be merged. Based on the processing of the above anomaly types, the target question set requiring question merging can be determined from the first and second question sets.

[0062] S105, perform a question merging operation on at least a portion of the questions in the target question set, and then re-match the questions.

[0063] In this embodiment of the application, if the first question set is determined to be the target question set, it means that at least some questions in the first question set are mismatched, and a merging operation can be performed on the at least some questions; if the second question set is determined to be the target question set, it means that at least some questions in the second question set are mismatched, and a merging operation can be performed on the at least some questions.

[0064] Optionally, after merging at least some of the questions in the target question set, the server can re-match the questions to correct abnormal matching results and obtain normal matching results.

[0065] This application proposes a question matching method. First, a search page and candidate pages similar to the search page are obtained. The search page includes a first question set containing a target question, while the candidate pages include a second question set and question explanations for each candidate question in the second question set. Question matching is performed between the first and second question sets to obtain matching results. Then, in response to anomalies in the matching results, the anomaly type is determined. Based on the anomaly type, a target question set requiring question merging is identified from the first and second question sets. Finally, at least a portion of the questions in the target question set are merged, and question matching is performed again. This application can correct abnormal question matching results during the question matching process, obtaining accurate similar pages, thereby matching the question and explanation the user wants to search for.

[0066] To clearly illustrate the previous embodiment, in one embodiment of this application, performing a question merging operation on at least a portion of the questions in the target question set may include: determining at least a portion of the questions from the target question set based on the matching results, and performing a merging operation on at least a portion of the questions.

[0067] Optionally, a target question set can be determined from the first question set and the second question set based on the matching results, and at least some questions can be determined from the target question set. If the first question set is determined as the target question set, and the matching result is that multiple target questions match the same similar question, that is, there are duplicate matching target questions, then the duplicate matching target questions can be determined as at least some questions. If the second question set is determined as the target question set, and the matching result is that one target question matches multiple similar questions, then the multiple similar questions can be determined as at least some questions.

[0068] In some implementations, when the server searches the search page, if the granularity of the target questions in the first question set of the search page is smaller than the granularity of the candidate questions entered in the question bank, there will be abnormal matching results where multiple target questions match the same candidate question. In this case, there are duplicate matching target questions in the first question set. The duplicate matching target questions can be identified as at least some questions and merged.

[0069] For example, see Figure 2 In actual searching, the search page is divided into three single-question areas. That is, a large question in the first question set of the search page is split into three sub-questions. These three sub-questions are entered into the question bank as a large question, so that each of the three sub-questions matches the same large question during the search. At this time, the above three sub-questions are duplicate matches. The above three sub-questions can be merged into a large question to avoid duplicate matches.

[0070] In some other implementations, when the server searches the search page, if the granularity of the target questions in the first question set is larger than the granularity of the candidate questions entered in the question bank, an abnormal matching situation will occur where a target question is found to have multiple candidate questions. In this case, the multiple candidate questions can be identified as at least some questions and merged.

[0071] For example, see Figure 3 During the actual search, the first set of questions on the search page is split into two target questions. These two target questions are recorded as three candidate questions when they are entered into the question bank. One of the target questions is split into two corresponding candidate questions in the question bank, so that the target question can be matched with two candidate questions. At this time, the two candidate questions can be identified as at least some questions and merged into one candidate question.

[0072] This application can avoid duplicate matching of questions and correct abnormal matching results during the question matching process to obtain normal matching results, thereby improving the user experience.

[0073] Before merging at least some of the above questions, it is necessary to consider the area of ​​the merged region and the similarity of the matching, such as... Figure 4 As shown, before merging at least a portion of the questions in the target question set, the following steps are also included:

[0074] S401, Obtain the location information of at least some of the questions, and determine the area of ​​the question region formed by the question merging operation based on the location information.

[0075] In this embodiment of the application, if the area of ​​the merged region of at least some of the questions is large, for example, if the merged region exceeds the range that the server can recognize, it will be detrimental to question search. Therefore, before merging, it is necessary to obtain the location information of at least some of the questions and determine the area of ​​the merged question region based on the location information.

[0076] S402, in response to the question area being greater than or equal to the area threshold, then at least some questions in the target question set will not be merged.

[0077] The area threshold can be preset in the storage space of the terminal device according to the actual situation and needs, and no restrictions are imposed here.

[0078] In this embodiment of the application, it can be determined whether the area of ​​the question region is greater than or equal to the area threshold. If so, at least some questions in the target question set will not be merged. If not, a merging operation will be performed on at least some questions in the target question set.

[0079] This application can avoid affecting the question search due to the large area of ​​the question region after the merging operation.

[0080] Regarding one type of abnormal matching situation in question matching, in one embodiment of this application, such as Figure 5 As shown, the question matching method also includes the following steps:

[0081] S501, in response to the exception type indicating that there is a duplicate matching first target question in the first question set, the first question set is taken as the target question.

[0082] In this embodiment of the application, if the question matching result is abnormal, and the abnormality type of the question matching result indicates that there is a duplicate matching of the first target question in the first question set, the server can respond to the abnormality type and use the first question set as the target question set.

[0083] S502, obtain the similarity between the first target question and the first matched candidate question.

[0084] The similarity can be determined by a combination of factors, including the similarity of text, the similarity of graphic features, and the difference in question length between the first target question and the first matched candidate question. Appropriate weights can be set for text similarity, graphic feature similarity, and the difference in question length, and the similarity between the first target question and the first matched candidate question can be calculated based on the set weights.

[0085] S503: Select the first target question that has the highest similarity to the first candidate question, and establish a matching relationship between it and the first candidate question.

[0086] In this embodiment, if the question merging operation is not performed because the area of ​​the merged question region in the target question set is too large, the first target question with the highest similarity to the first candidate question can be selected from multiple repeatedly matched first target questions, and a matching relationship can be established between the first candidate question and the first target question for re-matching. This application can select the most suitable question from the repeatedly matched search questions during the question matching process, thereby avoiding abnormal matching results.

[0087] For another possible question matching anomaly, such as Figure 6 As shown, in one embodiment of this application, the question matching method further includes the following steps:

[0088] S601, in response to the exception type indicating that there is a second target question in the first question set that has not been successfully matched, the second target set is taken as the target question set.

[0089] In this embodiment of the application, if the question matching result is abnormal, and the abnormality type of the question matching result indicates that there is a second target question in the first question set that has not been successfully matched, the server may respond to the abnormality type indication and use the second question set as the target question set.

[0090] It should be noted that the aforementioned situation where a second target question failed to match means that a target question was found to be split into multiple questions, but no complete similar question was matched. See [link to relevant documentation]. Figure 3 One of the target questions is matched with two questions that are split from the target question. The aforementioned target question is called the second target question.

[0091] S602, obtain the word count of the candidate topics, and select the second candidate topics with a word count less than the second target topic, and determine them as at least some topics.

[0092] In this embodiment of the application, if the second target question is partially similar to the matched second candidate question, but the second candidate question has more words than the second target question, that is, there is a redundant part in the second candidate question besides the similar part, then there will be a redundant part during merging, which is not conducive to merging a candidate question with higher similarity and affects the matching result. Therefore, the second candidate question with fewer words than the second target question can be selected as at least part of the question to avoid the redundant part during merging.

[0093] S603, perform a question merging operation on at least some of the questions to generate candidate questions for merging.

[0094] At least some of the questions identified above can be merged into a single candidate question to generate merged candidate questions.

[0095] S604, obtain the similarity between the second target question and the merged candidate questions.

[0096] The similarity can be determined by a combination of factors, including the similarity of the text, the similarity of the graphic features, and the difference in question length between the second target question and the merged candidate question. Appropriate weights can be set for the text similarity, the similarity of the graphic features, and the difference in question length, and the similarity between the first target question and the merged candidate question can be calculated based on the set weights.

[0097] S605, in response to a similarity greater than or equal to the similarity between the second target question and at least some of the questions before merging, a matching relationship is established between the second target question and the candidate questions to be merged.

[0098] In this embodiment, if the similarity between the second target question and the candidate question to be merged is greater than or equal to the similarity between the second target question and at least some questions before merging, a matching relationship can be established between the second target question and the candidate question to be merged, and re-matching can be performed. This application can select suitable questions from multiple similar questions found in the search and merge them into a complete similar question, thereby improving the similarity, making the matching results more accurate, and improving the user experience.

[0099] After the above merging operation, a corresponding question matching relationship was established, and question matching was performed again. In one embodiment of this application, after re-matching the questions, the following steps are also included:

[0100] S701, in response to the fact that there are still unmatched third target questions in the first question set.

[0101] After re-matching the questions, there may still be third target questions in the first question set that have not been matched successfully, that is, the third target questions have not been matched with similar candidate questions.

[0102] S702, obtain the fourth objective question that is adjacent to the third objective question.

[0103] Among them, adjacent can be left and right adjacent or top and bottom adjacent.

[0104] S703, based on the fourth candidate question that matches the fourth target question, obtain the third candidate question that matches the third target question.

[0105] In this embodiment of the application, adjacent candidate questions that are adjacent to the fourth candidate question can be obtained, and a first positional relationship between the third question and the fourth question can be obtained. Based on the first positional relationship, a third candidate question is determined from the adjacent candidate questions for the third target question.

[0106] Optionally, the second positional relationship between the fourth candidate question and the adjacent questions can be obtained first, and the target second positional relationship can be determined from the second positional relationship, wherein the target second positional relationship is the same as the positional relationship of the first positional relationship. Then, the adjacent candidate questions corresponding to the target second relationship are determined as the third candidate questions.

[0107] In some implementations, if the third target question and the fourth target question are adjacent vertically, and the fourth question is located below the third question, then the candidate question above the fourth candidate question can be determined as the third candidate question that matches the third target question.

[0108] In some other implementations, if the first and third target questions in the first question set of the search page match the first and third candidate questions in the second question set of the candidate page, but the second target question (i.e., the third target question) does not match a similar candidate question, then the second target question in the first question set of the search page and the second candidate question in the second question set of the candidate page can be considered as a matching result. That is, the second candidate question is determined as the third candidate question that matches the third target question.

[0109] This application can match similar questions and solutions to unmatched questions by using adjacent questions, thereby enhancing the question matching capability and further improving the user experience.

[0110] The question matching method proposed in this application first obtains a search page and a candidate page similar to the search page. The search page includes a first question set containing a target question, while the candidate pages include a second question set and question parsing for each candidate question in the second question set. Question matching is then performed on the first and second question sets to obtain matching results. Next, in response to anomalies in the matching results, the anomaly type is determined. Based on the anomaly type, a target question set requiring question merging is identified from the first and second question sets. Finally, at least a portion of the questions in the target question set are merged, and question matching is performed again. This application can correct abnormal question matching results during the question matching process, obtaining accurate similar pages and thus matching the question the user wants to search for.

[0111] Figure 8 This is a schematic diagram of a title matching device according to one embodiment of this application, as shown in the diagram. Figure 8 As shown, the question matching device 800 includes: a first acquisition module 810, a matching module 820, a second acquisition module 830, a determination module 840, and a merging module 850.

[0112] The first acquisition module 810 is used to acquire the search page and similar candidate pages of the search page. The search page includes a first question set, which includes the target question. The candidate pages include a second question set and question analysis for each candidate question in the second question set.

[0113] The matching module 820 is used to perform question matching on the first question set and the second question set to obtain question matching results;

[0114] The second acquisition module 830 is used to acquire the exception type of the question matching result in response to an exception in the question matching result;

[0115] The determination module 840 is used to determine the target question set that needs to be merged from the first question set and the second question set based on the exception type.

[0116] The merging module 850 is used to merge at least a portion of the questions in the target question set and to re-match the questions.

[0117] The question matching device proposed in this application acquires a search page and a candidate page similar to the search page through a first acquisition module. The search page includes a first question set containing a target question, and the candidate pages include a second question set and question explanations for each candidate question in the second question set. A matching module performs question matching on the first and second question sets to obtain matching results. Then, in response to an anomaly in the question matching results, the second acquisition module acquires the anomaly type of the matching results. Based on the anomaly type, a determination module determines a target question set from the first and second question sets that needs to be merged. Finally, a merging module performs a question merging operation on at least a portion of the questions in the target question set and performs question matching again. This application can correct abnormal question matching results during the question matching process, obtaining accurate similar pages, thereby matching the question and explanation the user wants to search for.

[0118] Furthermore, the merging module 850 is also used to: determine at least a portion of the questions from the target question set based on the question matching results; and perform a question merging operation on the at least a portion of the questions.

[0119] Furthermore, the determining module 840 is also configured to: in response to an exception type indicating that there is a duplicate matching first target question in the first question set, use the first question set as the target question set.

[0120] Furthermore, the determination module 840 is also used to: identify the first target question with duplicate matching as at least a portion of the questions.

[0121] Furthermore, before the merging module 850, the method further includes: obtaining the location information of at least some of the questions; determining the area of ​​the question region formed by the question merging operation based on the location information; in response to the question region area being greater than or equal to an area threshold, not performing the question merging operation on at least some of the questions in the target question set; obtaining the similarity between the first target question and the matched first candidate question; selecting the first target question with the highest similarity to the first candidate question, and establishing a matching relationship between the first target question and the first candidate question.

[0122] Furthermore, the determination module 840 is also configured to: in response to an exception type indication that there is a second target question in the first question set that has not been successfully matched, use the second question set as the target question set.

[0123] Furthermore, the determination module 840 is also used to: obtain the word count of the candidate topics, and select the second candidate topics with a word count less than the second target topic, and determine them as at least some topics.

[0124] Furthermore, after the merging module 850, the method further includes: performing a question merging operation on at least some of the questions to generate candidate questions for merging; obtaining the similarity between the second target question and the candidate questions for merging; and establishing a matching relationship between the second target question and the candidate questions for merging in response to the similarity being greater than or equal to the similarity of the second target question before merging.

[0125] Furthermore, after the matching module 820, the module further includes: in response to the fact that there are still unmatched third target questions in the first question set, obtaining a fourth target question adjacent to the third target question; and obtaining a third candidate question matched by the third target question based on the fourth candidate question matched by the fourth target question.

[0126] Furthermore, the second acquisition module 830 is also used to: acquire adjacent candidate questions that are adjacent to the fourth candidate question; acquire a first positional relationship between the third target question and the fourth target question; and determine a third candidate question for the third target question from the adjacent candidate questions based on the first positional relationship.

[0127] Furthermore, the second acquisition module 830 is also used to: acquire the second positional relationship between the fourth candidate question and the adjacent candidate questions; determine the target second positional relationship from the second positional relationship, wherein the target second positional relationship is the same as the positional relationship of the first positional relationship; and determine the adjacent candidate question corresponding to the target second positional relationship as the third candidate question.

[0128] To achieve the above embodiments, this application also proposes an electronic device 900, such as... Figure 9 As shown, the electronic device 900 includes a processor 910 and a memory 920 communicatively connected to the processor. The memory 920 stores instructions executable by at least one processor. The instructions are executed by at least one processor 910 to implement the title matching method as described in the first aspect embodiment of this application.

[0129] To implement the above embodiments, this application also proposes a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to implement the title matching method as described in the first aspect of this application.

[0130] To implement the above embodiments, this application also proposes a computer program product, including a computer program that, when executed by a processor, implements the title matching method as described in the first aspect of this application.

[0131] In the description of this application, it should be understood that the terms "center", "longitudinal", "lateral", "length", "width", "thickness", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", "clockwise", "counterclockwise", "axial", "radial", "circumferential", etc., indicating the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings, are only for the convenience of describing this application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of this application.

[0132] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this application, "multiple" means two or more, unless otherwise explicitly specified.

[0133] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0134] Although embodiments of this application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting this application. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of this application.

Claims

1. A question matching method, characterized in that, include: Obtain a search page and similar candidate pages of the search page, wherein the search page includes a first set of questions, the first set of questions includes the target question, and the similar candidate pages include a second set of questions and question analysis for each candidate question in the second set of questions; Perform question matching on the first question set and the second question set to obtain question matching results; If an anomaly is found in the question matching result, the anomaly type of the question matching result is obtained. Based on the anomaly type, determine the target question set that needs to be merged from the first question set and the second question set; At least a portion of the questions in the target question set are merged, and the question matching is performed again. The step of determining the target question set to be merged from the first question set and the second question set according to the exception type includes: In response to the exception type indicating that there is a duplicate matching first target question in the first question set, the first question set is taken as the target question set; Determine at least some of the topics, including: The first target question with duplicate matching is identified as at least a portion of the questions; The step of determining the target question set to be merged from the first question set and the second question set according to the exception type includes: In response to the exception type indicating that there is a second target question in the first question set that has not been successfully matched, the second question set is taken as the target question set; Determine at least some of the topics, including: Obtain the word count of the candidate questions, and select a second candidate question with a word count less than the second target question, and determine it as the at least part of the questions; After re-matching the questions, the process also includes: In response to the fact that there is still a third target question that has not been matched in the first question set, a fourth target question that is adjacent to the third target question is obtained; Based on the fourth candidate question that matches the fourth target question, obtain the third candidate question that matches the third target question.

2. The method according to claim 1, characterized in that, The process of merging at least a portion of the questions in the target question set includes: Based on the question matching results, at least a portion of the questions are determined from the target question set; Perform a question merging operation on at least some of the questions.

3. The method according to claim 2, characterized in that, Before performing the question merging operation on at least a portion of the questions in the target question set, the method further includes: Obtain the location information of at least some of the questions, and determine the area of ​​the question region formed by the question merging operation based on the location information; If the area of ​​the question region is greater than or equal to the area threshold, then the question merging operation will not be performed on at least a portion of the questions in the target question set; Obtain the similarity between the first target question and the matched first candidate question; Select the first target question that has the highest similarity to the first candidate question, and establish a matching relationship between it and the first candidate question.

4. The method according to claim 3, characterized in that, After determining that the questions are at least a portion of the total questions, the method further includes: Perform a question merging operation on at least some of the questions to generate candidate merged questions; Obtain the similarity between the second target question and the merged candidate questions; In response to the similarity being greater than or equal to the similarity between the second target question and at least some of the questions before merging, a matching relationship is established between the second target question and the candidate questions to be merged.

5. The method according to claim 4, characterized in that, The step of obtaining the third candidate question matching the third target question based on the fourth candidate question matching the fourth target question includes: Obtain the adjacent candidate questions that are adjacent to the fourth candidate question; Obtain the first positional relationship between the third target question and the fourth target question; Based on the first positional relationship, the third candidate question is determined from the adjacent candidate questions for the third target question.

6. The method according to claim 5, characterized in that, The step of determining the third candidate question from the adjacent candidate questions based on the first positional relationship includes: Obtain the second positional relationship between the fourth candidate question and the adjacent candidate questions; Determine a target second positional relationship from the second positional relationship, wherein the target second positional relationship is the same as the positional relationship of the first positional relationship; The adjacent candidate questions corresponding to the second positional relationship of the target are determined as the third candidate questions.

7. A question matching device, characterized in that, include: The first acquisition module is used to acquire a search page and a candidate page similar to the search page, wherein the search page includes a first question set, the first question set includes target questions, and the candidate page includes a second question set and question analysis for each candidate question in the second question set; The matching module is used to match questions between the first question set and the second question set to obtain question matching results; The second acquisition module is used to acquire the type of the exception in the question matching result if there is an exception. The determination module is used to determine the target question set that needs to be merged from the first question set and the second question set based on the exception type. The merging module is used to perform a question merging operation on at least a portion of the questions in the target question set, and to re-match the questions; The determining module is further configured to: In response to the exception type indicating that there is a duplicate matching first target question in the first question set, the first question set is taken as the target question set; The determining module is further configured to: The first target question with duplicate matching is identified as at least a portion of the questions; The determining module is further configured to: In response to the exception type indicating that there is a second target question in the first question set that has not been successfully matched, the second question set is taken as the target question set; The determining module is further configured to: Obtain the word count of the candidate questions, and select a second candidate question with a word count less than the second target question, and determine it as the at least part of the questions; In response to the fact that there is still a third target question that has not been matched in the first question set, a fourth target question that is adjacent to the third target question is obtained; Based on the fourth candidate question that matches the fourth target question, obtain the third candidate question that matches the third target question.

8. The apparatus according to claim 7, characterized in that, The merging module is also used for: Based on the question matching results, at least a portion of the questions are determined from the target question set; Perform a question merging operation on at least some of the questions.

9. The apparatus according to claim 7, characterized in that, Prior to the merging module, the following also includes: Obtain the location information of at least some of the questions, and determine the area of ​​the question region formed by the question merging operation based on the location information; If the area of ​​the question region is greater than or equal to the area threshold, then the question merging operation will not be performed on at least a portion of the questions in the target question set; Obtain the similarity between the first target question and the matched first candidate question; Select the first target question that has the highest similarity to the first candidate question, and establish a matching relationship between it and the first candidate question.

10. The apparatus according to claim 7, characterized in that, Following the merging module, the system further includes: Perform a question merging operation on at least some of the questions to generate candidate merged questions; Obtain the similarity between the second target question and the merged candidate questions; If the similarity is greater than or equal to the similarity of the second target question before merging, a matching relationship is established between the second target question and the candidate questions to be merged.

11. The apparatus according to claim 7, characterized in that, The second acquisition module is further configured to: Obtain the adjacent candidate questions that are adjacent to the fourth candidate question; Obtain the first positional relationship between the third target question and the fourth target question; Based on the first positional relationship, the third candidate question is determined from the adjacent candidate questions for the third target question.

12. The apparatus according to claim 11, characterized in that, The second acquisition module is further configured to: Obtain the second positional relationship between the fourth candidate question and the adjacent candidate questions; Determine a target second positional relationship from the second positional relationship, wherein the target second positional relationship is the same as the positional relationship of the first positional relationship; The adjacent candidate questions corresponding to the second positional relationship of the target are determined as the third candidate questions.

13. An electronic device, comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-6.

14. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-6.

15. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1-6.

Citation Information

Patent Citations

  • Question production method and device and electronic equipment

    CN113204617A