Document matching method and device, equipment and storage medium
By continuously combining multiple terms in the query request, combined with the enhanced BM25 algorithm and diversity weighting logic, the problem of insufficient sorting of document search results in the prior art is solved, and more accurate and relevant search results are achieved.
Patent Information
- Application Number
- CN202510308302.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-06-13
AI Technical Summary
In the existing search system, the sorting accuracy of document search results is insufficient, especially the literal matching algorithms such as the BM25 algorithm are insufficient in dealing with the content and diversity of continuous matching phrases, resulting in the search results that do not match the user's intention.
A document matching method is proposed, which generates continuous content by continuously combining multiple terms in the query request, and uses the enhanced BM25 algorithm to combine continuous content scores and diversity weighting logic to calculate more accurate matching results.
Improve the accuracy and relevance of search results, ensure that the results with exact matches or high hit rate can rank at the forefront, and improve the user experience and overall search sorting effect of the search system.
Smart Images

Figure CN120144741A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of artificial intelligence technology, and in particular to technical fields such as deep learning, intelligent search, and large models. Background Art
[0002] In a search system, the sorting accuracy of document retrieval results affects the overall search effect and user experience. Among them, the literal matching score is an important factor affecting the sorting of document retrieval results and may determine the final sorting effect. The Best Matching (BM) 25 algorithm is a classic literal matching algorithm, which can be used for information retrieval and text mining, and is widely used in search engines and related fields, with advantages such as simplicity and efficiency. Summary of the Invention
[0003] This disclosure provides a document matching method, apparatus, device, and storage medium.
[0004] According to one aspect of this disclosure, a document matching method is provided, including:
[0005] Obtaining continuous content based on multiple terms in a query request;
[0006] Obtaining a first matching result based on the multiple terms and a document;
[0007] Obtaining a second matching result based on the continuous content and the document;
[0008] Obtaining a third matching result based on the first matching result and the second matching result.
[0009] According to another aspect of this disclosure, a document matching apparatus is provided, including:
[0010] A continuous content module for obtaining continuous content based on multiple terms in a query request;
[0011] A first matching module for obtaining a first matching result based on the multiple terms and a document;
[0012] A second matching module for obtaining a second matching result based on the continuous content and the document;
[0013] A third matching module for obtaining a third matching result based on the first matching result and the second matching result.
[0014] According to another aspect of this disclosure, an electronic device is provided, including:
[0015] At least one processor; and
[0016] A memory communicatively connected to the at least one processor; wherein,
[0017] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute any method in the embodiments of the present disclosure.
[0018] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute any method in the embodiments of the present disclosure.
[0019] According to another aspect of the present disclosure, there is provided a computer program product including a computer program, and the computer program implements any method in the embodiments of the present disclosure when executed by a processor.
[0020] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:
[0022] Figure 1 is a flowchart of a document matching method according to an embodiment of the present disclosure;
[0023] Figure 2 is a flowchart of a document matching method according to another embodiment of the present disclosure;
[0024] Figure 3 is a flowchart of a document matching method according to another embodiment of the present disclosure;
[0025] Figure 4 is a flowchart of a document matching method according to another embodiment of the present disclosure;
[0026] Figure 5 is a flowchart of a document matching method according to another embodiment of the present disclosure;
[0027] Figure 6 is a structural diagram of a document matching device according to an embodiment of the present disclosure;
[0028] Figure 7 is a structural diagram of a document matching device according to another embodiment of the present disclosure;
[0029] Figure 8 is a flowchart of extracting consecutive matching phrases of a query request;
[0030] Figure 9It is a block diagram of an electronic device for implementing the document matching method of the embodiments of the present disclosure. Detailed implementation manners
[0031] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0032] Figure 1 It is a schematic flowchart of a document matching method 100 according to an embodiment of the present disclosure. The method may include:
[0033] S110. Obtain continuous content based on multiple terms in a query request;
[0034] S120. Obtain a first matching result based on the multiple terms and a document;
[0035] S130. Obtain a second matching result based on the continuous content and the document;
[0036] S140. Obtain a third matching result based on the first matching result and the second matching result.
[0037] In the embodiments of the present disclosure, a query request may include a request sent by a user or a system to a data source such as a server or a database to obtain information, and may also be referred to as a query statement, a query, etc. By performing word segmentation processing on the query request, one or more terms included in the query request can be obtained. A term can also be referred to as a word segment, a word, etc. For example, the terms included in query request Q include q1, q2, and q3. Multiple continuous contents of the query request can be obtained by performing continuous combination (or continuous splicing) according to the multiple terms in the query request. The continuous content can be referred to as continuous terms, continuous phrases, continuous matching phrases, continuous matching contents, etc. For example, based on the terms q1, q2, and q3 of query request Q, continuous contents q1q2, q1q2q3, and q2q3 can be obtained. Another example is that by performing word segmentation on the query statement "historical data details list", the word segmentation result can be "historical", "data", "details", "list", and the continuous matching phrases obtained after disassembling the terms can be "historical data", "historical data details", "historical data details list", "data details", "data details list", "details list", etc.
[0038] In the embodiments of the present disclosure, the documents to be retrieved can form a document collection or a document database. The number of documents, the content of the documents, etc. in the document collection or document database for different application scenarios may vary. For example, the document collection within an enterprise can include various content documents of a certain enterprise. Based on multiple terms in the query request and the documents in the document collection, a matching algorithm such as the BM25 algorithm can be used to calculate the first matching result. For example, the first matching result can include the first matching score, and can also include the document ranking or screening result obtained according to the first matching score. For example, the first matching result can include the first matching scores of all documents in the document collection, and can also include the document identifiers of the top N (e.g., N = 1000) ranked documents screened from the document collection. The BM25 algorithm is a ranking algorithm for information retrieval that can calculate the relevance score between a query statement and a document. The BM25 algorithm comprehensively evaluates the relevance between a document and a query by calculating factors such as the term frequency of the document, the document length, and the inverse document frequency of the query term in the entire document collection.
[0039] In the embodiments of the present disclosure, by continuously merging or concatenating multiple terms of a query statement, the continuous content of the query statement can be obtained. Based on multiple continuous contents in the query request and the documents in the document collection, a matching algorithm can be used to calculate the second matching result. For example, the second matching result can include the second matching score, and can also include the document ranking or screening result obtained according to the second matching score. For example, the second matching result can include the second matching scores of all documents in the document collection, and can also include the second matching scores of the top N documents screened based on the above first matching score. In addition, based on the second matching score, it is also possible to further screen out the top M (e.g., M = 100) ranked documents from the top N ranked documents.
[0040] In the embodiments of the present disclosure, after obtaining the first matching result between a document and a term and the second matching result between a document and continuous content, further processing can be performed according to the first matching result and the second matching result. For example, the first matching score and the second matching score are directly summed or weighted summed to obtain a third matching result such as a third matching score.
[0041] According to the embodiments of the present disclosure, the matching result comprehensively obtained from the terms and continuous contents in the query request can improve the accuracy and relevance of the search results. For example, when searching the document collection within an enterprise, the relevance between the query request and the search results can be improved, thereby improving the accuracy of the search results.
[0042] Figure 2FIG. 200 is a schematic flowchart of a document matching method 200 according to another embodiment of the present disclosure. The method 200 can be used to implement step S110 in the document matching method 100. In one implementation, the method 200 includes: obtaining continuous content based on multiple terms in a query request, further including:
[0043] S210. Perform word segmentation processing on the query request to obtain the multiple terms;
[0044] S220. Determine the sequence numbers of the first item and the last item in the continuous content based on the sequence numbers of the multiple terms;
[0045] S230. Generate the continuous content based on the multiple terms and the sequence numbers of the first item and the last item.
[0046] In the embodiments of the present disclosure, the text in the query request can be segmented into words to obtain multiple terms, and each term can correspond to a sequence number. The sequence number of a term can represent the order of the term in the query request. The sequence number of a term can start from 0, or start from 1, or be other numbering methods. For example, after the query request is segmented, it includes: "term A", "term B", and "term C". The sequence number of "term A" is 1, the sequence number of "term B" is 2, and the sequence number of "term C" is 3. Then the sorting of the terms is "A - B - C".
[0047] In the embodiments of the present disclosure, the sequence numbers of the first item and the last item in the continuous content can be determined through pointers and the sequence numbers of the terms, and then the continuous content can be generated. For example: As Figure 3 shown, after the query request is segmented into n terms (S301), the pointer i can be initialized first (S302), for example, i = 0. The pointer i can represent the sequence number of the first term in a certain continuous content among the multiple terms in the query request, that is, the sequence number of the first item. Determine whether the pointer i points to the last term in the query request (S303). If the pointer i points to the last term, output the continuous content list (S309). If the pointer i does not point to the last term in the query request, initialize the pointer j (S304), for example, j = i + 1. The pointer j can represent the sequence number of the last term in a certain continuous content among the multiple terms in the query request, that is, the sequence number of the last item. Determine whether the pointer j exceeds the last term (S305). If the pointer j does not exceed the last term, splice the terms between the pointer i and the pointer j and save the spliced content as a continuous content (S306). Move the sequence number pointed to by the pointer j one bit backward (S307), for example, j = j + 1. If the pointer j exceeds the last term, move the sequence number pointed to by the pointer i one bit backward (S308), for example, i = i + 1, and return to determine whether the pointer i points to the last term in the query request (S303) until i = n - 1, and output the continuous content list.
[0048] In an embodiment of the present disclosure, multiple terms in a query request, as well as the serial numbers of the first term and the last term in the consecutive content, can be used to splice each consecutive content in sequence, and then the consecutive content composed of all terms in the query request can be output in the form of a consecutive content list or a consecutive content set. For example, if the query request is segmented into: "Term A", "Term B", and "Term C", then the consecutive terms are: "Term A Term B", "Term A Term B Term C", "Term B Term C".
[0049] According to an embodiment of the present disclosure, multiple terms in a query request can be obtained by word segmentation, and consecutive content can be generated through the serial numbers of the terms, which can improve the relevance between the search results and the consecutive content, and further improve the accuracy of the search results. For example, documents with a high matching degree with the consecutive content in the query request can be sorted in the front.
[0050] Figure 4 FIG. 400 is a flowchart of a document matching method 400 according to another embodiment of the present disclosure. The method 400 can be used to implement steps S120 and S130 in the document matching method 100. In one implementation, the method 400 includes: obtaining a first matching result based on the multiple terms and the document, and further including:
[0051] S410. Obtain the first matching result based on the term frequency and inverse document frequency of the multiple terms in the document, as well as one or more of the number of terms in the document, the average document length, and a first adjustable parameter.
[0052] In an embodiment of the present disclosure, the term frequency (TF) can be the frequency of a term appearing in a document. The inverse document frequency (IDF) can be used to measure the importance or uniqueness of a term in a document collection. The inverse document frequency combined with the term frequency can be used to calculate the relevance score between a document and a term. Algorithms such as the TF-IDF algorithm and the BM25 algorithm can be used to calculate the relevance score between a document and multiple terms in a query request. The number of terms in a document can be the number of terms included in the document. The average document length can be the average length of all documents in the document collection. The average document length can be an actual statistical value, or an estimated value or an empirical value.
[0053] According to an embodiment of the present disclosure, a first matching result between a term and a document, such as a first matching score, can be obtained through multiple parameters related to the term and the document in the query request, so that the relevance between the search result and the term in the query request can be accurately obtained.
[0054] In one embodiment, the first adjustable parameter includes a term frequency saturation impact factor and / or a document length impact factor corresponding to the multiple terms in the query request.
[0055] For example, the scoring formula of a BM25 algorithm is as follows:
[0056]
[0057] Among them, score(Q, D) represents the relevance score between the multiple terms of the query request and the document, Q represents the query request, which is composed of several terms, that is, Q = {q 1 , q 2 ,..., q n}; D represents the document; f(q i , D) represents the term frequency of term q i in document D; IDF(q i ) represents the inverse document frequency of term q i ; |D| represents the length (number of words) of document D; avgdl represents the average length of all documents in the document set; k 1 and b represent the first adjustable parameter, k 1 can represent the term frequency saturation impact factor, which is used to control the saturation of the term frequency (the impact degree of the term frequency on the document), and b can represent the document length impact factor, which is used to control the impact of the document length.
[0058] In the embodiments of the present disclosure, the term frequency saturation impact factor can be used to control the impact degree of the term frequency on the document relevance score. The term frequency saturation impact factor can represent how the contribution to the document score grows when the term frequency increases. For example, when the value of the term frequency saturation impact factor is zero, the term frequency has no impact on the score. As the value of the term frequency saturation impact factor gradually increases, the impact of the term frequency on the score gradually increases. For example, the larger the term frequency saturation impact factor of a term, the faster the first matching score grows when the term frequency of a certain term in the document is larger.
[0059] In the embodiments of the present disclosure, the document length impact factor can be used to control the impact of the document length on the document relevance score. A longer document may contain more vocabulary, so there are more opportunities to appear terms related to the query request. The larger the value of the document length impact factor, the more prominent the impact of the document length on the score. For example, when the value of the document length impact factor is zero, the document length has no impact on the score; when the value of the document length impact factor is 1, the impact of the document length reaches the maximum, and the algorithm can adjust the score significantly according to the actual length of the document. For example, the larger the document length impact factor of a term, the slower the first matching score grows when the document length is longer.
[0060] According to an embodiment of the present disclosure, the influence on the relevance score between terms in a query request and a document can be flexibly changed through a first adjustable parameter, which is applicable to richer search scenarios.
[0061] In one implementation, method 400 further includes: obtaining a second matching result based on the continuous content and the document, further including:
[0062] S420. Obtain a continuous content weight based on the term length of the continuous content and the inverse document frequency of the terms in the continuous content in the document;
[0063] S430. Obtain the second matching result based on the term frequency of the terms in the continuous content in the document, the continuous content weight, and one or more of the number of words in the document, the average document length, and a second adjustable parameter.
[0064] In an embodiment of the present disclosure, the term length of the continuous content can be obtained through a continuous content list or set. For example, if a continuous content includes "term A term B term C", then the term length of this continuous content is 3. By the term length of a continuous content and the inverse document frequency of the terms in the continuous content in the document, a corresponding weight can be assigned to the continuous content, that is, the continuous content weight.
[0065] For example, a weight formula for a continuous content is as follows:
[0066]
[0067] where len(term) is the term length included in a certain continuous content. IDF(q i ) is the inverse document frequency of each term in the continuous content in the document, and q i is the i-th term of the continuous content.
[0068] When the continuous content includes "term A term B", the weight corresponding to this continuous content is:
[0069] weight AB = 1·(IDF 词项A + IDF 词项B )
[0070] When the continuous content includes "term A term B term C", the weight corresponding to this continuous content is:
[0071] weight ABC = 1.5·(IDF 词项A + IDF 词项B + IDF 词项C )
[0072] In the embodiments of the present disclosure, the weight of continuous content can be calculated based on the term length in the continuous content and the term frequency of each term in the continuous content in a certain document. Based on the weight of the continuous content, the number of words in the document, the average document length, and the second adjustable parameter, BM25 can be enhanced to calculate a second matching result, such as a second matching score.
[0073] For example, the formula of an enhanced BM25 algorithm is as follows:
[0074]
[0075] Among them, score(P, D) represents the relevance score between multiple continuous contents of the query request Q and the document; P represents a list of continuous contents, which consists of several continuous contents, that is, P = {p 1 , p 2 ,..., p n}, n represents the number of continuous contents in the query request, and the value of n, which is the number of terms in the query request, can be the same, different, or represented by other symbols, such as m; D represents the document; f(p i , D) represents the term frequency of the continuous content p i in the document D; |D| represents the length (number of words) of the document D; avgdl represents the average length of all documents in the document set; k 1 and b represent the second adjustable parameters, and the values of k 1 and b in the above first adjustable parameters can be the same or different.
[0076] According to the embodiments of the present disclosure, a second matching result, such as a second matching score, between the continuous content in the query request and the document can be obtained through multiple parameters related to the continuous content and the document, so as to improve the relevance between the search result and the continuous content in the query request, and further improve the accuracy of the search result.
[0077] In one implementation, the second adjustable parameter includes a term frequency saturation influence factor and / or a document length influence factor corresponding to the continuous content. For the meanings and functions of the term frequency saturation influence factor and the document length influence factor, reference can be made to the relevant descriptions in the first adjustable parameter. For example, the greater the term frequency saturation influence factor of the continuous content, the faster the second matching score increases when the term frequency of a certain continuous content in the document is greater. Another example is that the greater the document length influence factor of the continuous content, the slower the second matching score increases when the document length is longer.
[0078] According to the embodiments of the present disclosure, the second adjustable parameter can flexibly change the influence on the relevance score between the continuous content in the query request and the document, is applicable to richer search scenarios, and improves the accuracy and relevance between the search result and the continuous content.
[0079] In an embodiment of the present disclosure, a third matching result can be obtained by combining the first matching result and the second matching result. For example, the third matching result is obtained by means such as addition or weighting.
[0080] An example of an addition formula includes: Score = score(P, D) + score(Q, D)
[0081] An example of a weighting formula includes: Score = w 1 score(P, D) + w 2 score(Q, D), where w 1 and w 2 correspond to the weights of the first matching result and the second matching result respectively.
[0082] Figure 5 is a schematic flowchart of a document matching method 500 according to another embodiment of the present disclosure. This method may include one or more features of the above-mentioned document matching method. In one implementation, the method further includes:
[0083] S510. Obtain a diversity weight based on one or more of the number of terms in the query request, the number of terms in the plurality of terms that match the document, and a third adjustable parameter;
[0084] S520. Obtain a fourth matching result based on the diversity weight and the third matching result.
[0085] In an embodiment of the present disclosure, the relevance between a document and a query request can be represented according to the number of each term in the document that includes the query request. Generally, the document content that includes more terms in the query request often has richer semantic information. Therefore, a diversity weighting logic can be introduced.
[0086] An example of a diversity weight formula in a diversity weighting logic is as follows:
[0087]
[0088] Score final = diversity weight * Score
[0089] where diversity weight represents the diversity weight; len(match q ) represents the number of terms in the document that match the word segmentation result of the query request Q; len(Q) represents the number of all terms after word segmentation of the query request; k represents an optional parameter that determines the degree of influence of diversity, for example, the value range is 1-10;
[0090] According to an embodiment of the present disclosure, the diversity weight can more specifically reflect the relevance between each term of the query request and the document, thereby affecting the relevance score of the term, which is beneficial to determining a document that is more relevant to the term in the query request and improving the accuracy and precision of the search results.
[0091] In one implementation, the third adjustable parameter includes a diversity impact factor. The diversity impact factor can be used to change the degree of influence of the diversity weight on the relevance score. The value range of the diversity impact factor can be 1 - 10. As the value of the diversity impact factor gradually increases, the influence of the diversity weight on the score gradually increases.
[0092] According to an embodiment of the present disclosure, by flexibly changing the influence of the term hit situation of the query request on the relevance score through the third adjustable parameter, the accuracy of the search results can be improved.
[0093] Figure 6 FIG. 600 is a schematic structural diagram of a document matching device according to an embodiment of the present disclosure. The device 600 may include:
[0094] A continuous content module 610, configured to obtain continuous content based on multiple terms in the query request;
[0095] A first matching module 620, configured to obtain a first matching result based on the multiple terms and the document;
[0096] A second matching module 630, configured to obtain a second matching result based on the continuous content and the document;
[0097] A third matching module 640, configured to obtain a third matching result based on the first matching result and the second matching result.
[0098] Figure 7 FIG. 700 is a schematic structural diagram of a document matching device according to another embodiment of the present disclosure. The device 700 includes: a continuous content module 710, a first matching module 720, a second matching module 730, and a third matching module 740. The functions of the above modules can refer to the functions of the respective modules of the document matching device in the above embodiment. In one implementation, the continuous content module 710 includes:
[0099] A word segmentation sub-module 711, configured to perform word segmentation processing on the query request to obtain the multiple terms;
[0100] A determination sub-module 712, configured to determine the serial numbers of the first item and the last item in the continuous content based on the serial numbers of the multiple terms;
[0101] A generation sub-module 713, configured to generate the continuous content based on the multiple terms, and the serial numbers of the first item and the last item.
[0102] In one embodiment, the first matching module 720 includes: a first matching sub-module 721, configured to obtain the first matching result based on one or more of the term frequencies and inverse document frequencies of the multiple terms in the document, and the number of words in the document, the average document length, and a first adjustable parameter.
[0103] In one embodiment, the first adjustable parameter includes a term frequency saturation influence factor and / or a document length influence factor corresponding to the multiple terms.
[0104] In one embodiment, the second matching module 730 includes:
[0105] a weight sub-module 731, configured to obtain a continuous content weight based on the term length of the continuous content and the inverse document frequency of the terms in the continuous content in the document;
[0106] a second matching sub-module 732, configured to obtain the second matching result based on the term frequency of the terms in the continuous content in the document, the continuous content weight, and one or more of the number of words in the document, the average document length, and a second adjustable parameter.
[0107] In one embodiment, the second adjustable parameter includes a term frequency saturation influence factor and / or a document length influence factor corresponding to the continuous content.
[0108] In one embodiment, as Figure 7 shown, the document matching device further includes:
[0109] a diversity module 750, configured to obtain a diversity weight based on one or more of the number of multiple terms in the query request, the number of terms in the multiple terms that match the document, and a third adjustable parameter.
[0110] a fourth matching module 760, configured to obtain a fourth matching result based on the diversity weight and the third matching result.
[0111] In one embodiment, the third adjustable parameter includes a multi-directional influence factor.
[0112] The BM25 algorithm is used to calculate the literal matching score. The BM25 algorithm can intuitively reflect the literal features of the document and calculate the relevance score between the query request and the document simply and efficiently. Its core idea is to evaluate the matching degree between the query request and the document based on the term frequency (TF) and the inverse document frequency (IDF). Although the BM25 algorithm is widely used in the enterprise internal search system, the limitations of the BM25 algorithm will lead to the search results not conforming to the user's intention, and even obvious sorting effect problems, thus affecting the user experience.
[0113] For example, the BM25 algorithm lacks rewards for continuously matching phrase content: The reward mechanism for continuously hitting content is not considered in the BM25 algorithm, which often results in the inability to rank the results of exact matches or high hit rates at the top, seriously affecting the user experience. Another example is that the BM25 algorithm has insufficient diversity rewards: The importance of multi-term collaboration is under-considered in the algorithm, resulting in the results of some terms with excessive matches occupying the top positions, while the results containing more terms in the content are ranked behind, reducing the relevance of the search results.
[0114] The matching method of the embodiments of the present disclosure may include an enhanced BM25 algorithm based on continuous matching optimization and diversity weighting. Through this method, the overall retrieval ranking effect can be improved to a certain extent, and the actual user experience of the enterprise internal search system can be optimized.
[0115] I. Regarding the problem of the lack of a reward mechanism for continuously matching phrase content, a phrase continuous matching reward algorithm can be implemented based on the BM25 algorithm.
[0116] First, the query request (query) is segmented to obtain multiple terms. Secondly, the BM25 algorithm is used to calculate the relevance scores of the multiple terms of the query request and the document.
[0117] An example of the formula for calculating the relevance scores of the multiple terms of the query request and the document based on the BM25 algorithm is as follows:
[0118]
[0119] Among them, the meanings of the parameters are as follows:
[0120] Q: Query, which consists of several terms, that is, Q = {q 1 , q 2 ,..., q n};
[0121] D: Document;
[0122] f(q i , D): The term frequency (TF) of the term q i in the document D;
[0123] IDF(q i ): The inverse document frequency (IDF) of the term q i ;
[0124] |D|: The length (number of words) of the document D;
[0125] avgdl: The average length of all documents in the document collection;
[0126] k 1a and b: adjustable parameters, usually k 1 controls the saturation (influence degree) of the word frequency, and b controls the influence of the document length.
[0127] In the above formula, the term frequency (TF) represents the number of times a term appears in a document. The higher the term frequency, the stronger the relevance between the term in the query and the document. To prevent the score from increasing excessively due to too high a term frequency, the parameter k is introduced in the BM25 algorithm 1 to perform saturation processing on the term frequency. The inverse document frequency (IDF) can be used to measure the importance of each term. The higher the IDF value, the more important the term and the more effectively it can distinguish documents.
[0128] Next, extract the continuously matching phrase content from the query request. The flowchart for specifically extracting continuously matching phrases from the query is as Figure 8 shown, and this process may include the following steps:
[0129] S801. For example, the input original query request (query) is: Search for product feature introduction.
[0130] S802. After the query request is segmented by the word segmentation module, the segmentation result can be obtained as: [Search, product, feature, introduction].
[0131] S803. Initialize the pointer i, for example, i = 0.
[0132] S804. Determine whether the pointer i points to the last term of the query request. If so, execute S811; otherwise, execute S805.
[0133] S805. Initialize the pointer j, for example, j = i + 1.
[0134] S806. Determine whether the pointer j has passed the last term of the query request. If so, execute S810; otherwise, execute S807.
[0135] S807. Concatenate the terms between the pointer i and the pointer j in order.
[0136] S808. Save the concatenated content as a continuous content (or called continuously matching phrase content). For example, Search for product.
[0137] S809. Move the serial number pointed to by the pointer j one position backward, for example, j = j + 1, and return to execute S806.
[0138] S810. Move the serial number pointed to by the pointer i one position backward, for example, i = i + 1, and return to execute S804.
[0139] S811. Output the list of continuous content after decomposition. For example, the continuously matched phrase content is: [Search for products, search for product functions, search for product function introductions, product functions, product function introductions, function introductions].
[0140] Then, a corresponding weight value can be assigned to each continuous content in the query request. The goal is that the more terms contained in the continuous content, the higher the corresponding weight value; and as the content containing terms increases, the corresponding weight increase is greater. To achieve the above goal, the weight of a continuous content is as follows:
[0141]
[0142] where: len(term) is the length of the terms contained in the continuous content, and n = len(term) in this formula. IDF(q i ) is the IDF of the i-th term of the continuous content.
[0143] For the above formula and query statement, an example of weight calculation is as follows:
[0144] weight 搜索产品 = 1·(IDF 搜索 + IDF 产品 )
[0145] weight 搜索产品功能 = 1.5·(IDF 搜索 + IDF 产品 + IDF 功能 )
[0146] In summary, the final literal matching score can be the sum of the BM25 score and the continuous content score:
[0147]
[0148] Score = score(P, D) + score(Q, D)
[0149] where the parameter meanings of the above formula are as follows:
[0150] P: The list of continuous content, composed of several continuous contents in the query request, that is, P = {p 1 , p 2 , …, p n};
[0151] weight pi : The weight corresponding to the continuous content p i ;
[0152] f(p i , D): The continuous content p iWord frequency in document D;
[0153] Score: The total score after combining the BM25 algorithm score and the continuous content score.
[0154] In summary, after introducing the continuous content score, the overall literal matching algorithm comprehensively considers the frequency of word terms in the document after word segmentation, and also considers the frequency of continuous content in the document. And by assigning a higher weight to the continuous content, the overall literal matching algorithm will give a higher score to the document that hits the continuous content, that is, through the above reward mechanism, it further ensures that the results of exact matches or high hit rates are ranked at the top.
[0155] II. Diversity weighting mechanism: As can be seen from the above formula, the relevant BM25 algorithm focuses on the frequency of each word term in the document. However, it lacks the consideration of the overall word term diversity.
[0156] In actual usage scenarios, it is found that the document content containing more word terms in the query often has richer semantic information. Therefore, a diversity weighting logic is introduced, and its formula is as follows:
[0157]
[0158] Score final = diversity weight * Score
[0159] Where:
[0160] len(match q ): The number of word terms that match the word segmentation result (multiple word terms) of the query request Q in the retrieved document;
[0161] len(Q): The total number of word terms after query word segmentation;
[0162] k: An optional parameter that determines the degree of influence of diversity, and the value range is 1-10;
[0163] diversity weight : Diversity weight.
[0164] In summary, after introducing the diversity weighting logic, more semantically rich information can be further obtained during the literal matching process.
[0165] This disclosure can further ensure that documents with exact matches or high hit rates can be ranked at the top during the sorting process by introducing a continuous content reward mechanism and a diversity weighting logic, and reduce the ranking of results with too many single word term matches in the query. Through the above improvements, the accuracy, diversity, and user experience of the search results in the enterprise internal search system are significantly improved, meeting the retrieval requirements of actual business scenarios.
[0166] For the specific functions and examples of the modules and sub-modules of the device according to the embodiments of the present disclosure, reference may be made to the relevant descriptions of the corresponding steps in the foregoing method embodiments, which will not be elaborated herein.
[0167] In the technical solution of the present disclosure, the acquisition, storage, and application of the user's personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.
[0168] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0169] Figure 9 FIG. shows a schematic block diagram of an exemplary electronic device 900 that can be used to implement embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, for example, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, for example, a personal digital assistant, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0170] As Figure 9 shown, the device 900 includes a computing unit 901, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 902 or a computer program loaded from a storage unit 908 into a random access memory (RAM) 903. In the RAM 903, various programs and data required for the operation of the device 900 can also be stored. The computing unit 901, the ROM 902, and the RAM 903 are connected to each other via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.
[0171] A plurality of components in the device 900 are connected to the I / O interface 905, including: an input unit 906, such as a keyboard, a mouse, etc.; an output unit 907, such as various types of displays, speakers, etc.; a storage unit 908, such as a magnetic disk, an optical disk, etc.; and a communication unit 909, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 909 allows the device 900 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0172] The computing unit 901 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 901 executes the various methods and processes described above, such as the document matching method. For example, in some embodiments, the document matching method can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 908. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 900 via the ROM 902 and / or the communication unit 909. When the computer program is loaded into the RAM 903 and executed by the computing unit 901, one or more steps of the document matching method described above can be executed. Alternatively, in other embodiments, the computing unit 901 can be configured to execute the document matching method in any other suitable way (e.g., by means of firmware).
[0173] Various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuitry, integrated circuit systems, field-programmable gate arrays (FPGA), application-specific integrated circuits (ASIC), application-specific standard products (ASSP), systems-on-chip (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a dedicated or general-purpose programmable processor that receives data and instructions from a storage system, at least one input device, and at least one output device, and transmits the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0174] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program codes can be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0175] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0176] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0177] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0178] A computer system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The client-server relationship is generated by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, can also be a server of a distributed system, or a server incorporating a blockchain.
[0179] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and no limitations are imposed herein.
[0180] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub - combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the principles of this disclosure shall be included within the protection scope of this disclosure.
Claims
1. A document matching method, comprising: Obtaining continuous content based on multiple terms in a query request; Based on the multiple terms and documents, obtain a first matching result; Based on the continuous content and the document, obtain a second matching result; A third matching result is obtained based on the first matching result and the second matching result.
2. The method according to claim 1, wherein: Get continuous content based on multiple terms in the query request, including: Performing word segmentation processing on the query request to obtain the multiple terms; Determining the sequence numbers of the first and last items in the continuous content based on the sequence numbers of the plurality of terms; The continuous content is generated based on the plurality of terms and the sequence numbers of the first term and the last term.
3. The method according to claim 1, wherein: Based on the multiple terms and documents, a first matching result is obtained, including: The first matching result is obtained based on the term frequency and inverse document frequency of the multiple terms in the document, and one or more of the number of words in the document, the average document length and a first adjustable parameter.
4. The method according to claim 3, wherein: The first adjustable parameter includes a term frequency saturation influence factor and / or a document length influence factor corresponding to the multiple terms.
5. The method according to any one of claims 1 to 4, wherein: Based on the continuous content and the document, a second matching result is obtained, including: Obtaining a continuous content weight based on the term length of the continuous content and the inverse document frequency of the terms in the continuous content in the document; The second matching result is obtained based on the word frequency of the terms in the continuous content in the document and the continuous content weight, and one or more of the number of words in the document, the average document length and a second adjustable parameter.
6. The method according to claim 5, wherein: The second adjustable parameter includes a word frequency saturation influence factor and / or a document length influence factor corresponding to the continuous content.
7. The method according to any one of claims 1 to 6, further comprising: Determining a diversity weight based on the number of terms in the query request, the number of terms that match the documents, and one or more of a third adjustable parameter; A fourth matching result is obtained based on the diversity weight and the third matching result.
8. The method according to claim 7, wherein: The third adjustable parameter includes a diversity impact factor.
9. A document matching device, comprising: A continuous content module, for obtaining continuous content based on multiple terms in a query request; A first matching module, configured to obtain a first matching result based on the plurality of terms and documents; A second matching module, used for obtaining a second matching result based on the continuous content and the document; The third matching module is used to obtain a third matching result based on the first matching result and the second matching result.
10. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 8.
11. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-8.
12. A computer program product, comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 8.