A text sorting method, device, and electronic device for a retrieval system
By performing multi-dimensional feature analysis of user input text and combining the sorting method of knowledge base documents, the problem of neglecting semantic information in the existing technology is solved, and more accurate text sorting and higher quality search results are achieved.
Patent Information
- Application Number
- CN202411304177.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-19
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2044-09-19
AI Technical Summary
The prior art ignores the rich semantic information contained in structures such as global context, keywords and topics in generative search engines, resulting in deviations from the search results from user intentions.
By performing multi-dimensional feature analysis on user input text, and combining pre-constructed knowledge base documents for rough sorting and precise sorting, we will enhance the understanding of text subjects and topics and strengthen multi-grained semantic features.
It realizes the accurate sorting of candidate texts during the search process, improves the quality of candidate data and document analysis effect, and significantly improves the overall performance and user satisfaction of the search system.
Smart Images

Figure CN118885570B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of natural language processing. Specifically, it relates to a method, apparatus, electronic device, and storage medium for text sorting in a retrieval system. Background Art
[0002] In the existing generative search engines, the rough recall strategy mainly focuses on the local semantic characteristics of paragraph texts, ignoring the rich semantic information contained in the global context, keywords, and themes. This bias may lead to insufficient understanding of the deep meaning and context of the text, resulting in a deviation between the retrieval results and the user's query intention. The fine-ranking method mainly uses a model based on the cross-encoding structure. This model can effectively achieve direct interaction between the user's query and the candidate set, thus achieving a better semantic matching effect. However, when the subject of the user's query is inconsistent with the subject of the candidate set while other descriptive contents are the same, it is often impossible to achieve an ideal matching result in this case.
[0003] In addition, in the generative retrieval system, due to the text chunk vectorization processing method, the quality of the candidate set data is not high, which directly affects the effects of preliminary recall and precise ranking. The low readability and poor quality of the text will have an adverse impact on answer generation. Summary of the Invention
[0004] The purpose of this application is to provide a method, apparatus, electronic device, and storage medium for text sorting in a retrieval system, which can achieve precise sorting of candidate texts during the retrieval process, improve the quality of candidate data and the document parsing effect, use two methods of rough sorting and precise sorting to enhance the understanding of the text subject and theme, strengthen multi-granularity semantic features, and significantly improve the overall performance of the retrieval system and user satisfaction.
[0005] In a first aspect, an embodiment of this application provides a method for text sorting in a retrieval system, and the method includes:
[0006] Obtain the preprocessed user input text;
[0007] Perform multi-dimensional feature parsing on the user input text to obtain the data to be retrieved;
[0008] Roughly sort the data to be retrieved according to the pre-constructed knowledge base document to obtain candidate data;
[0009] Precisely sort the candidate data according to the pre-constructed precise ranking model to obtain a ranking result.
[0010] In the above implementation process, by performing multi-dimensional parsing on the text input by the user and then combining it with the knowledge base documents for rough sorting and precise sorting, it is possible to achieve precise sorting of candidate texts during the retrieval process, improve the quality of candidate data and the effect of document parsing. By using two methods of rough sorting and precise sorting, the understanding of the text subject and theme is enhanced, multi-granularity semantic features are strengthened, and the overall performance of the retrieval system and user satisfaction are significantly improved.
[0011] Further, the steps of constructing the knowledge base documents include:
[0012] Extract the table information, picture information and text data in the initial documents used to construct the knowledge base documents;
[0013] Vectorize the table information, picture information and text data to obtain segmented texts;
[0014] Optimize the segmented texts to obtain optimized segmented texts;
[0015] Construct the knowledge base documents according to the optimized segmented texts. The knowledge base documents are a document set containing multiple sub-documents, and each sub-document contains at least one of the segmented texts.
[0016] In the above implementation process, by processing the table information, picture information and text data separately and then performing vectorization, the degree of fusion of various types of data in the vectorization process can be improved, the loss of key information in the vectorization process can be reduced, and the accuracy can be improved.
[0017] Further, the step of vectorizing the table information, picture information and text data to obtain segmented texts includes:
[0018] Obtain the context information of the table information;
[0019] Generate a summary text of the table information according to the context information;
[0020] Obtain a summary text of the picture information according to the picture information;
[0021] Vectorize the summary text of the table information, the summary text of the picture information and the text data, and retain the corresponding hierarchical structure relationship to obtain the segmented texts.
[0022] In the above implementation process, by vectorizing the summary text of the table information, the summary text of the picture information and the text data and retaining the hierarchical structure relationship of the text data, the semantic features of the obtained segmented texts are more explicit, and the information entropy of the segmented texts can be increased.
[0023] Further, the step of roughly sorting the data to be retrieved according to the pre-constructed knowledge base document to obtain candidate data includes:
[0024] Query the knowledge base document to obtain the topic information of each sub-document in the knowledge base and the topic information of the segmented text in each sub-document;
[0025] Match the topic information of the data to be retrieved with the topic information of each sub-document and the topic information of the segmented text in each sub-document respectively;
[0026] If the topic information of the data to be retrieved is consistent with the topic information of each sub-document and the topic information of the segmented text in each sub-document, determine the sub-document as the first candidate data;
[0027] Obtain the candidate data according to the first candidate data.
[0028] In the above implementation process, retrieving the knowledge base document according to the segmented text, selecting candidate data for rough sorting, can quickly and accurately screen out the data in the knowledge base document that conforms to the topic of the segmented text, reduce the error probability, and reduce the error.
[0029] Further, the step of obtaining the candidate data according to the first candidate data includes:
[0030] Extract the semantic features with different fine-grained levels in the first candidate data;
[0031] Match the semantic features of the data to be retrieved with the semantic features of the first candidate data according to the fine-grained level;
[0032] Remove the first candidate data in the first candidate data whose fine-grained level of semantic information does not match the fine-grained level of the semantic information of the data to be retrieved, to obtain the candidate data.
[0033] In the above implementation process, matching for semantic features with different fine-grained levels, selecting the data with the most matching semantic features in the first candidate data, improves the usability and accuracy of the candidate data, and ensures the effective progress of the retrieval process.
[0034] Further, the step of precisely sorting the candidate data according to the pre-constructed precise sorting model to obtain the sorting result includes:
[0035] Evaluate the candidate data according to the precise sorting model to obtain the estimated score of the candidate data;
[0036] Determine the candidate data whose estimated score meets the evaluation threshold as the second candidate data;
[0037] Perform a secondary verification on the second candidate data to obtain third candidate data;
[0038] Perform data filling on the third candidate data to obtain the sorting result.
[0039] In the above implementation process, after evaluating the candidate data according to the precise sorting model and then performing secondary verification and data filling, the candidate data can be calibrated from multiple dimensions, improving the result of precise sorting, and the selection of candidate data can be refined, making the sorting result closer to the user's intention.
[0040] Further, the step of performing a secondary verification on the second candidate data to obtain third candidate data includes:
[0041] Obtain the subject information of the data to be retrieved;
[0042] Match the subject information of the data to be retrieved with the subject information of the second candidate data;
[0043] Filter out the data in the second candidate data whose subject information does not match the subject information of the data to be retrieved to obtain the third candidate data.
[0044] In the above implementation process, by matching according to the subject information, the secondary filtering and calibration of the second candidate data are realized, which can improve the data accuracy and perfect the sorting process.
[0045] In a second aspect, an embodiment of the present application further provides a text sorting device for a retrieval system, and the device includes:
[0046] An acquisition module, configured to acquire a pre-processed user input text;
[0047] A multi-dimensional feature analysis module, configured to perform multi-dimensional feature analysis on the user input text to obtain data to be retrieved;
[0048] A rough sorting module, configured to perform rough sorting on the data to be retrieved according to a pre-constructed knowledge base document to obtain candidate data;
[0049] A precise sorting module, configured to perform precise sorting on the candidate data according to a pre-constructed precise sorting model to obtain a sorting result.
[0050] In the above implementation process, by performing multi-dimensional parsing on the text input by the user and then combining it with the knowledge base document for rough sorting and precise sorting, it is possible to achieve precise sorting of candidate texts during the retrieval process, improve the quality of candidate data and the effect of document parsing. By using two methods of rough sorting and precise sorting, the understanding of the text subject and theme is enhanced, multi-granularity semantic features are strengthened, and the overall performance of the retrieval system and user satisfaction are significantly improved.
[0051] In a third aspect, an electronic device provided by an embodiment of the present application includes: a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the method according to any one of the first aspects are implemented.
[0052] In a fourth aspect, a computer-readable storage medium provided by an embodiment of the present application has instructions stored thereon. When the instructions are run on a computer, the computer is made to execute the method according to any one of the first aspects.
[0053] In a fifth aspect, a computer program product provided by an embodiment of the present application, when run on a computer, causes the computer to execute the method according to any one of the first aspects.
[0054] Other features and advantages of the present disclosure will be described in the subsequent description, or some features and advantages can be inferred from the description or determined without doubt, or can be learned by implementing the above technologies of the present disclosure.
[0055] And it can be implemented according to the content of the description. The following will be described in detail with reference to the preferred embodiments of the present application and the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] To more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required to be used in the embodiments of the present application. It should be understood that the following drawings only show some embodiments of the present application, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.
[0057] Figure 1 It is a flowchart of the text sorting method of the retrieval system provided by the embodiment of the present application;
[0058] Figure 2 It is a schematic structural diagram of the text sorting device of the retrieval system provided by the embodiment of the present application;
[0059] Figure 3 It is a schematic structural diagram of the electronic device provided by the embodiment of the present application. Detailed implementation manners
[0060] Next, the technical solutions in the embodiments of the present application will be described with reference to the accompanying drawings in the embodiments of the present application.
[0061] It should be noted that similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. At the same time, in the description of the present application, terms such as "first" and "second" are only used for differential description and cannot be construed as indicating or implying relative importance.
[0062] Next, with reference to the drawings and embodiments, the specific implementation manners of the present application will be further described in detail. The following embodiments are used to illustrate the present application but are not used to limit the scope value of the present application.
[0063] Embodiment 1
[0064] Figure 1 is a schematic flowchart of a text sorting method for a retrieval system provided by an embodiment of the present application. As Figure 1 shown, the method includes:[[]]
[0065] S1. Obtain the preprocessed user input text;
[0066] S2. Perform multi-dimensional feature parsing on the user input text to obtain the data to be retrieved;
[0067] S3. Roughly sort the data to be retrieved according to the pre-constructed knowledge base document to obtain candidate data;
[0068] S4. Accurately sort the candidate data according to the pre-constructed accurate sorting model to obtain a sorting result.
[0069] In the above implementation process, by performing multi-dimensional parsing on the text input by the user and then performing rough sorting and accurate sorting in combination with the knowledge base document, accurate sorting of candidate texts during the retrieval process can be achieved, improving the quality of candidate data and the document parsing effect. By using two methods of rough sorting and accurate sorting, the understanding of the text subject and theme is enhanced, and multi-granularity semantic features are strengthened, significantly improving the overall performance of the retrieval system and user satisfaction.
[0070] The present application effectively solves the limitations of the prior art in text processing, data quality optimization, and model improvement by introducing a multi-feature fusion retrieval strategy, optimizing the quality of the knowledge base document, and improving the cross-encoder model. The performance of generative search is significantly improved, providing users with more accurate and comprehensive retrieval results and enhancing the user experience.
[0071] In S1 and S2, the present application preprocesses and analyzes multi-dimensional features of the input text to be queried entered by the user, including user semantic analysis, keyword analysis, topic recognition, subject recognition, etc. When retrieving the user text, both the global semantic information and local keyword information of the text are considered.
[0072] Further, the steps of constructing the knowledge base document include:
[0073] Extract the table information, picture information, and text data in the initial document used to construct the knowledge base document;
[0074] Vectorize the table information, picture information, and text data to obtain segmented text;
[0075] Optimize the segmented text to obtain the optimized segmented text;
[0076] Construct the knowledge base document according to the optimized segmented text. The knowledge base document is a document set containing multiple sub-documents, and each sub-document contains at least one segmented text.
[0077] In the above implementation process, the table information, picture information, and text data are processed separately and then vectorized, which can improve the integration degree of various types of data in the vectorization process, reduce the loss of key information in the vectorization process, and improve the accuracy.
[0078] The processing of the knowledge base document in the present application involves multiple levels, such as OCR recognition, document parsing, text optimization, and topic recognition.
[0079] In the initial document, accurately extracting the picture information, table information, and text content from the document is the basis for the subsequent construction of the question-answering system. Since different elements in the initial document carry unique information respectively, it is crucial for comprehensive data understanding and analysis. The present application uses a fine-tuned OCR technology to accurately identify and extract the picture information, table information, and text data from the initial document.
[0080] Readable text is crucial for the recall of user questions and the generation of answers. Therefore, it is necessary to optimize the segmented text. The optimization method of the present application is implemented based on a fine-tuned large language model, rewriting and polishing the segmented text, combining the content of the current paragraph to be polished and the adjacent paragraphs for rewriting, and ensuring the integrity and smoothness of the sentences, so as to achieve the polishing of the segmented text.
[0081] The present application can achieve the effect of enhanced retrieval from two aspects. The polished paragraphs have more complete sentences and clearer semantics, which helps to improve data recall and solves the problem of difficult data recall caused by paragraph truncation or incompleteness and low text quality. Secondly, the optimization of segmented text improves the understanding of the context in the prompts by large models. The entire process combines automated processing and a small amount of manual verification to ensure the construction of high-quality knowledge base documents.
[0082] Further, the steps of vectorizing table information, picture information, and text data to obtain segmented text include:
[0083] Obtain the context information of the table information;
[0084] Generate summary text of the table information based on the context information;
[0085] Obtain summary text of the picture information based on the picture information;
[0086] Vectorize the summary text of the table information, the summary text of the picture information, and the text data, and retain the corresponding hierarchical structure relationship to obtain segmented text.
[0087] In the above implementation process, vectorize the summary text of the table information, the summary text of the picture information, and the text data, retain the hierarchical structure relationship of the text data, make the semantic features of the obtained segmented text more explicit, and can increase the information entropy of the segmented text.
[0088] Deeply process the table information, picture information, and text data in the initial document. Specifically, in the process of processing table information, combine the context information of the table information (such as title information and table description information) to generate a description and summary text of the table, and realize the vectorization of the table information from the text level to ensure more accurate retrieval of the table information.
[0089] For the picture information extracted from the initial document, generate corresponding summary information for each picture information, and convert the summary text of the picture information into a vector representation that can be processed at the semantic level by a vectorization method. Achieve effective recall of image data in the semantic dimension.
[0090] In the process of vectorizing the summary text of the table information, the summary text of the picture information, and the text data for text paragraph division and vectorization, while retaining the hierarchical structure relationship, realize the association between parent and child documents. Ensure that semantic recall can be achieved at different granularities during the user query process, thereby improving the accuracy of data recall.
[0091] Further, S3 includes:
[0092] Query the knowledge base documents to obtain the topic information of each sub-document in the knowledge base and the topic information of the segmented text in each sub-document;
[0093] Match the topic information of the data to be retrieved with the topic information of each sub-document and the topic information of the segmented text in each sub-document respectively;
[0094] If the topic information of the data to be retrieved is consistent with the topic information of each sub-document and the topic information of the segmented text in each sub-document, determine the sub-document as the first candidate data;
[0095] Obtain candidate data based on the first candidate data.
[0096] In the above implementation process, retrieving the knowledge base documents according to the segmented text and selecting candidate data for rough sorting can quickly and accurately screen the data in the knowledge base documents that conform to the topic of the segmented text, reduce the probability of errors, and reduce errors.
[0097] In the embodiments of the present application, topic recognition focuses on extracting the topic information of the data to be retrieved and the topic information of each sub-document and the segmented text, considering from the perspectives of the whole and the local, and achieving a more accurate match with the user's query.
[0098] If the topic information of the data to be retrieved, the sub-document, and the segmented text are all the same, it means that the best paragraph has been recalled. If the topics are not the same, it means that a paragraph with similar semantics but unable to answer the user's question has been recalled, and such segmented text cannot improve the retrieved answer.
[0099] The topic recognition method is mainly used to filter out the relevant documents in the knowledge base documents that are related to the semantic features of the user's query but do not match in the topic information pair.
[0100] Further, the step of obtaining candidate data based on the first candidate data includes:
[0101] Extract the semantic features with different fine-grained levels in the first candidate data;
[0102] Match the semantic features of the data to be retrieved with the semantic features of the first candidate data according to the fine-grained level;
[0103] Remove the first candidate data in the first candidate data whose fine-grained level of semantic information does not match the fine-grained level of the semantic information of the data to be retrieved to obtain candidate data.
[0104] In the above implementation process, matching for semantic features with different fine-grained levels, and selecting the data with the most matching semantic features in the first candidate data can improve the usability and accuracy of the candidate data, and ensure the effective progress of the retrieval process.
[0105] By integrating semantic features at different fine-grained levels, this application improves the precision, recall rate, and MRR in the initial recall process to a certain extent. The integration of semantic features at different fine-grained levels can be fine-tuned based on the vertical domain, which can improve the accuracy of document recall to a certain extent.
[0106] Further, S4 includes:
[0107] Evaluating the candidate data according to the exact ranking model to obtain the predicted scores of the candidate data;
[0108] Determining the candidate data with predicted scores meeting the evaluation threshold as the second candidate data;
[0109] Performing a secondary verification on the second candidate data to obtain the third candidate data;
[0110] Performing data filling on the third candidate data to obtain the ranking result.
[0111] In the above implementation process, after evaluating the candidate data according to the exact ranking model and then performing secondary verification and data filling, the candidate data can be calibrated in multiple dimensions, improving the result of the exact ranking, and the selection of candidate data can be refined, making the ranking result closer to the user's intention.
[0112] When performing an exact ranking on the candidate data, this application uses a cross-encoding model fine-tuned for the vertical domain to comprehensively score the candidate data, and filters out candidate data with a high semantic match according to the threshold (evaluation threshold) determined by comprehensive analysis.
[0113] Further, the step of performing a secondary verification on the second candidate data to obtain the third candidate data includes:
[0114] Obtaining the subject information of the data to be retrieved;
[0115] Matching the subject information of the data to be retrieved with the subject information of the second candidate data;
[0116] Filtering out the data in the second candidate data whose subject information does not match the subject information of the data to be retrieved to obtain the third candidate data.
[0117] In the above implementation process, by matching according to the subject information, the secondary filtering and calibration of the second candidate data are realized, which can improve the data accuracy and perfect the ranking process.
[0118] In the process of subject matching, by comparing and analyzing the subject information of the user query with the subject information of the candidate data, the secondary verification of the candidate data is realized, and most of the candidate data with a small difference from the user query can be filtered out.
[0119] Subject recognition is mainly used to solve the problem that when the subject information of the user query is inconsistent with the subject information of the candidate data during the precise sorting process, but other descriptions are the same, a good sorting cannot be achieved.
[0120] When constructing the prompt, based on the filling method of adaptive context length, efficient and lossless retrieval content compression can be achieved. It solves the problem of information loss caused by the context length limitation of large models, improves the accuracy and integrity of retrieval. In addition, it also enhances the interaction quality between the user query and the large model, thus significantly improving the overall user retrieval experience.
[0121] The embodiment of this application proposes a retrieval strategy that fuses multiple features, comprehensively considers the global and local information of the text, and also introduces a theme recognition technology, which can effectively filter out candidate data that is not relevant to the theme information and improve the relevance between the candidate data and the user query.
[0122] The method of this application optimizes the knowledge base and text quality at the data bottom layer, greatly enhancing the retrieval potential of the data. At the same time, aiming at the limitations of the cross-encoder model in vertical field applications, while performing fine-tuning, a subject recognition mechanism is added to correct the scoring deviation caused by the small differences between the user query and the candidate data, and improve the relevance of the candidate data.
[0123] Embodiment 2
[0124] In order to execute the method corresponding to the above Embodiment 1 to achieve the corresponding functions and technical effects, a text sorting device for a retrieval system is provided below, as Figure 2 shown, the device includes:
[0125] An acquisition module 1, configured to acquire the pre-processed user input text;
[0126] A multi-dimensional feature parsing module 2, configured to perform multi-dimensional feature parsing on the user input text to obtain the data to be retrieved;
[0127] A rough sorting module 3, configured to perform rough sorting on the data to be retrieved according to the pre-constructed knowledge base documents to obtain candidate data;
[0128] An accurate sorting module 4, configured to perform accurate sorting on the candidate data according to the pre-constructed accurate sorting model to obtain a sorting result.
[0129] In the above implementation process, by performing multi-dimensional parsing on the text input by the user and then combining it with the knowledge base documents for rough sorting and precise sorting, it is possible to achieve precise sorting of candidate texts during the retrieval process, improve the quality of candidate data and the effect of document parsing. By using the two methods of rough sorting and precise sorting, the understanding of the text subject and theme is enhanced, multi-granularity semantic features are strengthened, and the overall performance of the retrieval system and user satisfaction are significantly improved.
[0130] Furthermore, the device further includes a construction module for constructing the knowledge base documents, which includes:
[0131] Extract the table information, picture information, and text data in the initial document used to construct the knowledge base documents;
[0132] Vectorize the table information, picture information, and text data to obtain segmented text;
[0133] Optimize the segmented text to obtain the optimized segmented text;
[0134] Construct the knowledge base documents according to the optimized segmented text. The knowledge base documents are a document set containing multiple sub-documents, and each sub-document contains at least one segmented text.
[0135] In the above implementation process, by processing the table information, picture information, and text data separately and then performing vectorization, the integration degree of various types of data in the vectorization process can be improved, the loss of key information in the vectorization process can be reduced, and the accuracy can be improved.
[0136] Furthermore, the construction module is also used for:
[0137] Obtain the context information of the table information;
[0138] Generate a summary text of the table information according to the context information;
[0139] Obtain the summary text of the picture information according to the picture information;
[0140] Vectorize the summary text of the table information, the summary text of the picture information, and the text data, and retain the corresponding hierarchical structure relationship to obtain segmented text.
[0141] In the above implementation process, by vectorizing the summary text of the table information, the summary text of the picture information, and the text data and retaining the hierarchical structure relationship of the text data, the semantic features of the obtained segmented text are more explicit, and the information entropy of the segmented text can be increased.
[0142] Furthermore, the rough sorting module 3 is also used for:
[0143] Query the knowledge base document to obtain the topic information of each sub-document in the knowledge base and the topic information of the segmented text in each sub-document;
[0144] Match the topic information of the data to be retrieved with the topic information of each sub-document and the topic information of the segmented text in each sub-document respectively;
[0145] If the topic information of the data to be retrieved is consistent with the topic information of each sub-document and the topic information of the segmented text in each sub-document, determine the sub-document as the first candidate data;
[0146] Obtain candidate data based on the first candidate data.
[0147] In the above implementation process, retrieve the knowledge base document according to the segmented text, select candidate data for rough sorting, and can quickly and accurately screen the data in the knowledge base document that conforms to the topic of the segmented text, reduce the probability of errors, and reduce errors.
[0148] Furthermore, the rough sorting module 3 is also used for:
[0149] Extract semantic features with different fine-grained levels in the first candidate data;
[0150] Match the semantic features of the data to be retrieved with the semantic features of the first candidate data according to the fine-grained level;
[0151] Remove the first candidate data in the first candidate data whose fine-grained level of semantic information does not match the fine-grained level of the semantic information of the data to be retrieved to obtain candidate data.
[0152] In the above implementation process, match the semantic features with different fine-grained levels, select the data with the most matching semantic features in the first candidate data, improve the usability and accuracy of the candidate data, and ensure the effective progress of the retrieval process.
[0153] Furthermore, the precise sorting module 4 is also used for:
[0154] Evaluate the candidate data according to the precise sorting model to obtain the estimated score of the candidate data;
[0155] Determine the candidate data whose estimated score meets the evaluation threshold as the second candidate data;
[0156] Perform secondary verification on the second candidate data to obtain the third candidate data;
[0157] Perform data filling on the third candidate data to obtain the sorting result.
[0158] In the above implementation process, after evaluating the candidate data according to the precise sorting model and then performing secondary verification and data filling, the candidate data can be calibrated from multiple dimensions, improving the result of precise sorting, and the selection of candidate data can be refined, making the sorting result closer to the user's intention.
[0159] Furthermore, the precise sorting module 4 is further configured to:
[0160] Obtain the subject information of the data to be retrieved;
[0161] Match the subject information of the data to be retrieved with the subject information of the second candidate data;
[0162] Filter out the data in the second candidate data whose subject information does not match the subject information of the data to be retrieved, obtaining the third candidate data.
[0163] In the above implementation process, by matching according to the subject information, the secondary filtering and calibration of the second candidate data are realized, which can improve the data accuracy and perfect the sorting process.
[0164] The text sorting device of the above retrieval system can implement the method of Embodiment 1. The optional items in Embodiment 1 above are also applicable to this embodiment and will not be elaborated here.
[0165] The remaining content of the embodiments of this application can refer to the content of Embodiment 1 above and will not be repeated in this embodiment.
[0166] Embodiment 3
[0167] The embodiment of this application provides an electronic device, including a memory and a processor. The memory is used to store a computer program, and the processor runs the computer program to enable the electronic device to execute the text sorting method of the retrieval system in Embodiment 1.
[0168] Optionally, the above electronic device can be a server.
[0169] Please refer to Figure 3 , Figure 3 which is a schematic diagram of the structural composition of the electronic device provided by the embodiment of this application. The electronic device can include a processor 31, a communication interface 32, a memory 33, and at least one communication bus 34. Among them, the communication bus 34 is used to realize the direct connection communication between these components. Among them, the communication interface 32 of the device in the embodiment of this application is used to communicate with other node devices for signaling or data. The processor 31 can be an integrated circuit chip with signal processing capabilities.
[0170] The above-mentioned processor 31 may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it may also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or the processor 31 may also be any conventional processor, etc.
[0171] The memory 33 may be, but is not limited to, a Random Access Memory (RAM), a Read Only Memory (ROM), a Programmable Read-Only Memory (PROM), an Erasable Programmable Read-Only Memory (EPROM), an Electric Erasable Programmable Read-Only Memory (EEPROM), etc. The memory 33 stores computer-readable instructions. When the computer-readable instructions are executed by the processor 31, the device can execute the Figure 1 various steps involved in the method embodiments.
[0172] Optionally, the electronic device may further include a storage controller and an input / output unit. The memory 33, the storage controller, the processor 31, the peripheral interface, and the input / output unit are electrically connected directly or indirectly to each other to achieve data transmission or interaction. For example, these components may be electrically connected to each other through one or more communication buses 34. The processor 31 is used to execute the executable modules stored in the memory 33, such as software function modules or computer programs included in the device.
[0173] The input / output unit is used to provide the user with the creation of tasks and the creation of a start optional period or a preset execution time for the task to achieve the interaction between the user and the server. The input / output unit may be, but is not limited to, a mouse and a keyboard, etc.
[0174] It can be understood that Figure 3 the structure shown is only schematic, and the electronic device may further include more or fewer components than those Figure 3 shown, or have a configuration different from that Figure 3 shown. Figure 3 The components shown in may be implemented using hardware, software, or a combination thereof.
[0175] In addition, an embodiment of the present application further provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the text sorting method of the retrieval system in the first embodiment.
[0176] The embodiment of the present application further provides a computer program product, which, when running on a computer, causes the computer to execute the method described in the method embodiment.
[0177] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions, and operations of devices, methods, and computer program products according to multiple embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order from that marked in the accompanying drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based device for performing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.
[0178] In addition, in each embodiment of the present application, the functional modules may be integrated together to form an independent part, or each module may exist alone, or two or more modules may be integrated to form an independent part.
[0179] If the above functions are implemented in the form of software function modules and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present application. The aforementioned storage medium includes various media such as USB flash drives, mobile hard disks, ROM, RAM, magnetic disks, or optical discs that can store program codes.
[0180] The above are only embodiments of the present application and are not intended to limit the protection scope of the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application. It should be noted that similar reference numerals and letters indicate similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.
[0181] As described above, this is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or replacements within the technical scope disclosed by the present application, and all should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
[0182] It should be noted that in this text, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or device including the said element.
Claims
1. A text sorting method for a retrieval system, characterized in that: The method comprises: Get the preprocessed user input text; Performing multi-dimensional feature analysis on the user input text to obtain data to be retrieved; Roughly sorting the data to be retrieved according to the pre-built knowledge base documents to obtain candidate data; Accurately sorting the candidate data according to a pre-built accurate sorting model to obtain a sorting result; The step of accurately sorting the candidate data according to the pre-built accurate sorting model to obtain a sorting result includes: Evaluate the candidate data according to the precise ranking model to obtain an estimated score of the candidate data; Determine the candidate data whose estimated score meets the evaluation threshold as the second candidate data; Performing a second verification on the second candidate data to obtain third candidate data; Performing data filling on the third candidate data to obtain the sorting result; The step of performing a secondary check on the second candidate data to obtain the third candidate data comprises: Acquire subject information of the data to be retrieved; Matching the subject information of the data to be retrieved with the subject information of the second candidate data; The data whose subject information in the second candidate data does not match the subject information of the data to be retrieved is filtered to obtain the third candidate data.
2. The text sorting method of the retrieval system according to claim 1, characterized in that: The steps to build a knowledge base document include: Extracting table information, image information and text data from an initial document used to construct the knowledge base document; Vectorizing the table information, picture information and text data to obtain segmented text; Optimizing the segmented text to obtain optimized segmented text; The knowledge base document is constructed according to the optimized segmented text, wherein the knowledge base document is a document collection including a plurality of sub-documents, and each of the sub-documents includes at least one segmented text.
3. The text sorting method of the retrieval system according to claim 2, characterized in that: The step of vectorizing the table information, picture information and text data to obtain segmented text includes: Obtaining context information of the table information; generating a summary text of the table information according to the context information; Obtaining a summary text of the picture information according to the picture information; The summary text of the table information, the summary text of the picture information and the text data are vectorized, and the corresponding hierarchical structure relationship is retained to obtain the segmented text.
4. The text sorting method of the retrieval system according to claim 1, characterized in that: The step of roughly sorting the data to be retrieved according to the pre-built knowledge base documents to obtain candidate data includes: Querying the knowledge base document to obtain subject information of each sub-document in the knowledge base and subject information of the segmented text in each sub-document; Respectively matching the subject information of the data to be retrieved with the subject information of each of the sub-documents and the subject information of the segmented text in each of the sub-documents; If the subject information of the data to be retrieved is consistent with the subject information of each of the sub-documents and the subject information of the segmented text in each of the sub-documents, determining the sub-document as the first candidate data; The candidate data is obtained according to the first candidate data.
5. The text sorting method of the retrieval system according to claim 4, characterized in that: The step of obtaining the candidate data according to the first candidate data comprises: Extracting semantic features with different fine-grainedness from the first candidate data; Matching the semantic features of the data to be retrieved with the semantic features of the first candidate data according to fine granularity; The first candidate data whose semantic information granularity does not match the semantic information granularity of the data to be retrieved are removed from the first candidate data to obtain the candidate data.
6. A text sorting device for a retrieval system, characterized in that: The device comprises: The acquisition module is used to obtain the pre-processed user input text; A multi-dimensional feature analysis module, used to perform multi-dimensional feature analysis on the user input text to obtain data to be retrieved; A rough sorting module is used to roughly sort the data to be retrieved according to the pre-built knowledge base documents to obtain candidate data; An accurate sorting module is used to accurately sort the candidate data according to a pre-built accurate sorting model to obtain a sorting result; The precise sorting module is also used for: Evaluate the candidate data according to the precise ranking model to obtain an estimated score of the candidate data; Determine the candidate data whose estimated score meets the evaluation threshold as the second candidate data; Performing a second verification on the second candidate data to obtain third candidate data; Performing data filling on the third candidate data to obtain the sorting result; Acquire subject information of the data to be retrieved; Matching the subject information of the data to be retrieved with the subject information of the second candidate data; The data whose subject information in the second candidate data does not match the subject information of the data to be retrieved is filtered to obtain the third candidate data.
7. An electronic device, characterized in that: The electronic device comprises a memory and a processor, wherein the memory is used to store a computer program, and the processor runs the computer program to enable the electronic device to execute the text sorting method of the retrieval system according to any one of claims 1 to 5.
8. A storage medium, characterized in that: It stores a computer program, which, when executed by a processor, implements the text sorting method of the retrieval system as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Retrieval method and device based on multistage semantic matching, computer equipment and storage medium
CN114298055A