Text retrieval method and device, electronic equipment and readable storage medium
By generating corresponding search models for the demand title, combining cosine similarity and DTW calculation formulas, the problem of ignoring word frequency differences in the existing technology is solved, the search accuracy and applicability are improved, and more efficient and accurate text retrieval is achieved.
Patent Information
- Application Number
- CN202510258683.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-05
- Publication Date
- 2025-06-06
AI Technical Summary
When searching the required articles from massive information, the prior art ignores the word frequency differences caused by uneven lengths and unclear vocabulary and term descriptions, resulting in insufficient retrieval accuracy and not wide application scope.
By generating a corresponding search model for each requirement title, the matching between the requirement title and the material title is evaluated by using the model adaptive weight, cosine similarity calculation formula and the dynamic time regularization (DTW) calculation formula, and the target material title is determined and displayed.
It improves the correlation between the search model and the demand title, improves the accuracy and applicability of the search model, and thus improves the search efficiency and accuracy of the demand title.
Smart Images

Figure CN120104774A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of information retrieval technology, and in particular to a text retrieval method, device, electronic device and computer-readable storage medium. Background Art
[0002] With the advent of the big data era, the amount of information on the Internet has exploded. How to quickly and accurately find the required articles from the massive information has become a challenge faced by many people. Therefore, in the process of finding the required articles from the massive information, the method of searching and retrieving related fields is very important.
[0003] The existing technology usually uses NLP (Natural Language Processing) technology and cosine similarity algorithm to calculate the similarity between the required title and the material title to match the required article.
[0004] However, the existing technology only focuses on the differences in directions between word sets, ignoring the matching gap caused by word frequency differences due to uneven length and unclear description of vocabulary terms. It is not accurate enough and has a narrow scope of application, resulting in poor retrieval results. Summary of the invention
[0005] The technical problem to be solved by the present invention is to provide a text retrieval method, device, electronic device and computer-readable storage medium in view of the above-mentioned deficiencies in the prior art. The method can realize convenient and effective retrieval of text, enhance the correlation between the retrieval model and the demand title, improve the accuracy and applicability of the retrieval model, and thereby improve the retrieval efficiency and accuracy of the demand title.
[0006] In a first aspect, the present invention provides a text retrieval method, comprising: obtaining a material title, a requirement title, and a retrieval model corresponding to the requirement title; based on the retrieval model corresponding to the requirement title, evaluating the matching degree between the requirement title and the material title; based on the matching degree between the requirement title and the material title, determining and displaying the target material title.
[0007] Preferably, the obtaining of the retrieval model corresponding to the demand title specifically includes: determining the model adaptive weight corresponding to the demand title; generating the retrieval model corresponding to the demand title based on the model adaptive weight corresponding to the demand title, the cosine similarity calculation formula and the dynamic time warping DTW calculation formula.
[0008] Preferably, the method of determining the model adaptive weight corresponding to the demand title specifically includes: converting the demand title into a proper noun to obtain a converted title corresponding to the demand title; vectorizing the material title and the converted title to obtain a demand word set and a material word set; calculating the word set length of the demand word set and the material word set as the complexity of the demand title and the material title; and determining the model adaptive weight corresponding to the demand title based on a mapping relationship between a preset complexity and the model adaptive weight.
[0009] Preferably, the conversion of the demand title into a proper noun to obtain a converted title corresponding to the demand title specifically includes: performing word segmentation processing on the demand title to obtain a word set corresponding to the demand title; performing coreference resolution on the word set corresponding to the demand title, and determining the coreference words corresponding to the demand title, wherein the coreference words are used to represent multiple words referring to the same object; judging whether the coreference words corresponding to the demand title are included in a preset coreference chain; in response to the coreference words corresponding to the demand title being included in the preset coreference chain, replacing the coreference words corresponding to the demand title with target object words to obtain a converted title corresponding to the demand title, wherein the target object words are used to represent preset words for the objects referred to by the coreference words corresponding to the demand title in the preset coreference chain; in response to the coreference words corresponding to the demand title not being included in the preset coreference chain, generating an object and its corresponding entity based on the coreference words corresponding to the demand title to update the preset coreference chain.
[0010] Preferably, the material title and the conversion title are vectorized to obtain a demand word set and a material word set, which specifically includes: based on a vectorization algorithm, the material title and the conversion title are vectorized to obtain a large demand word set and a large material word set, wherein the vectorization algorithm includes the jieba algorithm; the large demand word set and the large material word set are aggregated to obtain a vocabulary union; the word frequency of each word in the vocabulary union is calculated, and based on the word frequency of each word, each word in the large demand word set and the large material word set are sorted respectively to obtain a demand word set and a material word set.
[0011] Preferably, the model adaptive weight includes a first weight and a second weight, and the retrieval model corresponding to the demand title is based on evaluating the matching degree between the demand title and the material title, specifically including:
[0012] According to formulas (1) and (2), the matching degree between the required title and the material title is calculated:
[0013]
[0014] D(i,j)=d(ai,bj)+min(D(i-1,j),D(i,j-1),D(i-1,j-1))(2),
[0015] Where A represents the first weight, B represents the second weight, Wk represents the kth word in the demand word set a, Yk represents the kth word in the material word set b, n represents the total number of words in the demand word set a or the material word set b, D(·) represents the DTW calculation formula, ai represents the i-th word in the demand word set a, bj represents the j-th word in the material word set b, d(ai,bj) represents the distance between ai and bj, D(i,j) represents the DTW distance between the i-th word in the demand word set a and the j-th word in the demand word set b, the i-th word refers to the first word to the i-th word, the j-th word refers to the first word to the j-th word, and D(a,b) represents the DTW distance between all words in the demand word set a and all words in the material word set b.
[0016] Preferably, determining and displaying the target material title based on the matching degree between the demand title and the material title specifically includes: judging whether the matching degree between the demand title and the material title is greater than a preset threshold; in response to the matching degree between the demand title and the material title being greater than the preset threshold, determining the material title to be the target material title, and displaying the target material title.
[0017] In a second aspect, the present invention also provides a text retrieval device, including an acquisition module, an evaluation module and a determination module. The acquisition module is used to acquire a material title, a requirement title and a retrieval model corresponding to the requirement title. The evaluation module is connected to the acquisition module and is used to evaluate the matching degree between the requirement title and the material title based on the retrieval model corresponding to the requirement title. The determination module is connected to the evaluation module and is used to determine and display the target material title based on the matching degree between the requirement title and the material title.
[0018] In a third aspect, the present invention further provides an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to implement the text retrieval method provided in the first aspect.
[0019] In a fourth aspect, the present invention further provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the text retrieval method provided in the first aspect above is implemented.
[0020] The present invention provides a text retrieval method, device, electronic device and computer-readable storage medium, which generates a corresponding retrieval model for each demand title, and the correlation between the retrieval model and the demand title is strong, so that the retrieval model is suitable for demand title retrieval in various situations, improves the accuracy and applicability of the retrieval model, and thus improves the retrieval accuracy and efficiency of each demand title. Therefore, the present invention can realize convenient and effective retrieval of text, enhance the correlation between the retrieval model and the demand title, improve the accuracy and applicability of the retrieval model, and thus improve the retrieval efficiency and accuracy of the demand title. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 A flowchart of a text retrieval method according to Embodiment 1 of the present invention;
[0022] Figure 2 A flowchart of a text retrieval method according to Embodiment 2 of the present invention;
[0023] Figure 3 This is a schematic diagram of the structure of a text retrieval device according to Embodiment 3 of the present invention;
[0024] Figure 4 This is a schematic diagram of the structure of a text retrieval device according to Embodiment 4 of the present invention. DETAILED DESCRIPTION
[0025] In order to enable those skilled in the art to better understand the technical solution of the present invention, the embodiments of the present invention will be further described in detail below with reference to the accompanying drawings.
[0026] It should be understood that the specific embodiments and drawings described herein are only used to explain the present invention rather than to limit the present invention.
[0027] It can be understood that, in the absence of conflict, the various embodiments of the present invention and the various features in the embodiments can be combined with each other.
[0028] It can be understood that, for the convenience of description, the drawings of the present invention only show the parts related to the present invention, while the parts irrelevant to the present invention are not shown in the drawings.
[0029] It can be understood that each unit and module involved in the embodiments of the present invention may correspond to only one physical structure, or may be composed of multiple physical structures, or multiple units and modules may be integrated into one physical structure.
[0030] It can be understood that, without conflict, the functions and steps marked in the flowcharts and block diagrams of the present invention may occur in an order different from that marked in the drawings.
[0031] It is understood that the flowcharts and block diagrams of the present invention illustrate the possible architectures, functions, and operations of the systems, devices, equipment, and methods according to the various embodiments of the present invention. Each box in the flowchart or block diagram may represent a unit, module, program segment, or code, which contains executable instructions for implementing the specified functions. Moreover, each box or combination of boxes in the block diagram and flowchart may be implemented by a hardware-based system that implements the specified functions, or may be implemented by a combination of hardware and computer instructions.
[0032] It can be understood that the units and modules involved in the embodiments of the present invention can be implemented by software or hardware. For example, the units and modules can be located in a processor.
[0033] Embodiment 1:
[0034] like Figure 1 As shown, this embodiment provides a text retrieval method.
[0035] In this embodiment, application scenarios of the text retrieval method include, but are not limited to: academic research, internal document management, and tender document preparation. This embodiment takes the application of the text retrieval method to tender document preparation as an example.
[0036] Text retrieval methods include:
[0037] S101, obtaining a material title, a requirement title, and a search model corresponding to the requirement title.
[0038] In this embodiment, the material title refers to the title and remarks in the preset bidding proposal library, and the requirement title refers to the title and remarks that the user needs to retrieve. Before the user makes a bidding document, the user extracts the technical solution name and specific description from the bidding document to generate the requirement title.
[0039] Specifically, obtaining the search model corresponding to the required title includes steps S1011 and S1012:
[0040] S1011, determining the model adaptive weight corresponding to the demand title.
[0041] Specifically, the model adaptive weight includes a first weight and a second weight.
[0042] In this embodiment, the first weight refers to the weight of the cosine similarity calculation formula, and the second weight refers to the weight of the DTW (Dynamic Time Warping) calculation formula.
[0043] Specifically, S1011: determining the model adaptive weight corresponding to the demand title, including: converting the demand title into a proper noun to obtain a converted title corresponding to the demand title; vectorizing the material title and the converted title to obtain a demand word set and a material word set; calculating the word set length of the demand word set and the material word set as the complexity of the demand title and the material title; determining the model adaptive weight corresponding to the demand title based on the mapping relationship between the preset complexity and the model adaptive weight.
[0044] In this embodiment, after generating the demand title, the text retrieval method also optimizes the demand title to obtain a converted title corresponding to the demand title, wherein the optimization processing includes but is not limited to: proper noun conversion, English abbreviation spelling error processing, and redundant word disorder processing. This embodiment helps to reduce ambiguity and improve the accuracy and consistency of the demand title by optimizing the demand title, thereby improving the accuracy of subsequent matching.
[0045] It should be noted that the conversion of professional terms refers to replacing words with common pronouns with the same meaning, such as Multi-Service Transport Platform (MSTP) and optical transport network (OTN). The processing of English abbreviation spelling errors refers to matching and replacing common word errors based on the preset dictionary, such as 0NT-OTN and MTSP-MSTP. The processing of redundant and disordered words refers to matching and finding errors based on the vocabulary library rules, and providing prompts or automatic corrections.
[0046] Similarly, when obtaining the material title, this embodiment also performs corresponding optimization processing on the material title.
[0047] Specifically, the requirement title is converted into a proper noun to obtain a converted title corresponding to the requirement title, including: performing word segmentation processing on the requirement title to obtain a word set corresponding to the requirement title; performing coreference resolution on the word set corresponding to the requirement title, and determining the coreference words corresponding to the requirement title, wherein the coreference words are used to represent multiple words referring to the same object; judging whether the coreference words corresponding to the requirement title are included in a preset coreference chain; in response to the coreference words corresponding to the requirement title being included in the preset coreference chain, replacing the coreference words corresponding to the requirement title with target object words to obtain a converted title corresponding to the requirement title, wherein the target object words are used to represent preset words for the objects referred to by the coreference words corresponding to the requirement title in the preset coreference chain; in response to the coreference words corresponding to the requirement title not being included in the preset coreference chain, generating objects and their corresponding entities based on the coreference words corresponding to the requirement title to update the preset coreference chain.
[0048] It should be noted that the objects of the coreference words and their corresponding entities can be updated to the preset coreference chain in sequence according to the order in which the coreference words corresponding to the requirement title appear in the requirement title. For example, if the order in which the coreference words appear in the requirement title is: optical transport network, network based on optical technology, OTN, optical transport network-network based on optical technology-OTN can be updated to the preset coreference chain.
[0049] Specifically, the material title and the conversion title are vectorized to obtain the demand word set and the material word set, including: based on the vectorization algorithm, the material title and the conversion title are vectorized to obtain the demand large word set and the material large word set, wherein the vectorization algorithm includes the jieba algorithm; the demand large word set and the material large word set are aggregated to obtain a vocabulary union; the word frequency of each word in the vocabulary union is calculated, and based on the word frequency of each word, each word in the demand large word set and the material large word set are sorted respectively to obtain the demand word set and the material word set.
[0050] In this embodiment, the original dictionary provided by the jieba algorithm is used to identify the keyword vectors corresponding to the material title and the conversion title, and the keyword vectors in the material title and the conversion title are extracted with the help of the two jieba algorithm built-in functions jieba.cut and jieba.lcut, and the stop word list of the jieba algorithm is used to remove the stop word vectors in the material title and the conversion title to obtain the demand large word set and the material large word set; the union of the demand large word set and the material large word set is calculated to obtain the vocabulary union; the number of times each word in the vocabulary union appears is counted through a for loop to obtain the word frequency of each word, for example: word_count = {}; for word in words: word_count[word] = word_count.get(word,0); based on the word frequency of each word, the words in the demand large word set and the material large word set are sorted to generate the corresponding demand word set and material word set. Words with high word frequencies usually have higher relevance or importance in the text, which helps to highlight key concepts. The cosine similarity calculation formula calculates the similarity between words in the same order. This embodiment calculates and concentrates the word frequencies of each word for sorting, ensuring that all words are compared based on the same standard and are consistent. This helps to clearly identify the most important words in the overall context and facilitates deep retrieval, that is, to retrieve material texts that are similar in semantics and importance to each word in the demand title.
[0051] Since the cosine similarity calculation formula is only applicable and accurate when the word set lengths of the demand word set and the material word set are the same, and the DTW calculation formula is not limited to the demand word set and the material word set having the same word set length, after sorting the words in the demand large word set and the material large word set based on the word frequency of each word, this embodiment can generate demand word sets and material word sets with the same or different word set lengths according to the retrieval requirements, wherein the retrieval requirements include but are not limited to: breadth retrieval requirements and depth retrieval requirements, the depth retrieval requirements refer to focusing on the accuracy and depth of the material title, and the breadth retrieval requirements refer to focusing on the comprehensiveness and diversity of the material title. For example: if the retrieval requirements include depth retrieval requirements, it is necessary to focus on the accuracy of the cosine similarity calculation formula, and the demand word set a = (W1, W2, ..., Wn) and the material word set b = (Y1, Y2, ..., Yn) with the same word set length can be generated, wherein n represents the word set length of the demand word set a.
[0052] The inequality between the lengths of the demand word set and the material word set will lead to inequality in the accuracy of the cosine similarity calculation formula and the DTW calculation formula. In addition, since the complexity of the cosine similarity calculation formula is O(n), and the complexity of the DTW calculation formula is O(n 2 ), the word set lengths of the demand word set and the material word set are equal, but if the word set length is too large, the complexity of the cosine similarity calculation formula and the DTW calculation formula will be too large, and the accuracy of the cosine similarity calculation formula and the DTW calculation formula will not be high. Therefore, the word set lengths of the demand word set and the material word set are used as the complexity of the demand title and the material title, and the model adaptive weight corresponding to the demand title is determined according to the mapping relationship shown in Table 1. This embodiment effectively improves the applicability between the retrieval model and the demand title by evaluating the complexity of each demand title for subsequent dynamic determination of the retrieval model corresponding to each demand title, thereby improving the accuracy of subsequent retrieval.
[0053] Table 1 Mapping relationship between preset complexity and model adaptive weight
[0054]
[0055]
[0056] It should be noted that the vectorization algorithms include but are not limited to: Jieba algorithm, TF-IDF (Term Frequency-Inverse Document Frequency, a commonly used text feature vectorization method), and BERT (Bidirectional Encoder Representations from Transformers, a contextual word embedding algorithm based on the Transformer model).
[0057] Before identifying the keyword vectors corresponding to the material title and the conversion title through the original dictionary provided by the jieba algorithm, this embodiment also adds the keyword vectors of the corresponding industry, such as MSTP and OTN, to the original dictionary provided by the jieba algorithm.
[0058] In addition to determining the model adaptive weight corresponding to the requirement title based on the mapping relationship between the preset complexity and the model adaptive weight, the present embodiment can also obtain additional search terms corresponding to the requirement title, determine whether the material title includes the additional search terms corresponding to the requirement title, add the additional search terms to the requirement term set and the material term set in response to the material title including the additional search terms corresponding to the requirement title, and determine the first weight and the second weight to be the first preset value and the second preset value, for example: 0.3 and 0.7.
[0059] In addition, the present embodiment can also set the model adaptive weights corresponding to the demand titles. For example, if the user only wants to retrieve material titles that are strongly related to the demand titles, that is, the retrieval requirement is only a deep retrieval requirement, the first weight can be set to 1 and the second weight can be set to 0 to perform a simple deep retrieval. If the user is not very clear about the retrieval requirement or the vocabulary of the demand title is unclear, the first weight can be set to 0.5 and the second weight can be set to 0.5 to integrate deep retrieval and breadth retrieval.
[0060] S1012, based on the model adaptive weight, cosine similarity calculation formula and dynamic time warping DTW calculation formula corresponding to the demand title, generate a retrieval model corresponding to the demand title.
[0061] In this embodiment, the cosine similarity calculation formula and the DTW calculation formula are fused through the model adaptive weights corresponding to the demand title to generate a retrieval model, which combines the respective calculation advantages of the cosine similarity calculation formula and the DTW calculation formula to improve the accuracy and adaptability of the retrieval model.
[0062] S102, based on the retrieval model corresponding to the demand title, evaluating the matching degree between the demand title and the material title.
[0063] Specifically, S102: based on the retrieval model corresponding to the demand title, evaluating the matching degree between the demand title and the material title, including: calculating the matching degree between the demand title and the material title according to formulas (1) and (2):
[0064]
[0065] D(i,j)=d(ai,bj)+min(D(i-1,j),D(i,j-1),D(i-1,j-1))(2),
[0066] Where A represents the first weight, B represents the second weight, Wk represents the kth word in the demand word set a, Yk represents the kth word in the material word set b, n represents the total number of words in the demand word set a or the material word set b, D(·) represents the DTW calculation formula, ai represents the i-th word in the demand word set a, bj represents the j-th word in the material word set b, d(ai,bj) represents the distance between ai and bj, D(i,j) represents the DTW distance between the i-th word in the demand word set a and the j-th word in the demand word set b, the i-th word refers to the first word to the i-th word, the j-th word refers to the first word to the j-th word, and D(a,b) represents the DTW distance between all words in the demand word set a and all words in the material word set b.
[0067] In this embodiment, according to The cosine similarity between the demand word set a and the material word set b is calculated, and the DTW matching degree between the demand word set a and the material word set b is calculated according to D(a,b), where the DTW matching degree between the demand word set a and the material word set b is D(n,n) calculated according to formula (2).
[0068] It should be noted that when the word set lengths of the required word set and the material word set are different, n represents the maximum total number of words in the required word set a or the material word set b. For example, when the word set length of the required word set a is 3 and the word set length of the material word set is 4, n is 4. Therefore, the cosine similarity calculation formula of the retrieval model replaces W4 with a preset non-blank character when calculating W4 and Y4.
[0069] S103, based on the matching degree between the required title and the material title, determine and display the target material title.
[0070] Specifically, S103: based on the matching degree between the required title and the material title, determining and displaying the target material title, including steps S1031 and S1032:
[0071] S1031, determining whether the matching degree between the requirement title and the material title is greater than a preset threshold.
[0072] S1032: In response to the matching degree between the requirement title and the material title being greater than a preset threshold, determining the material title to be a target material title, and displaying the target material title.
[0073] In this embodiment, the material title with a matching degree greater than a preset threshold is determined as the target material title, and the material title with a matching degree less than or equal to the preset threshold is removed, and based on the matching degree between the required title and the material title, the target material title is directly displayed on the interface from high to low, wherein the preset threshold includes but is not limited to 60%.
[0074] It should be noted that after directly displaying the target material titles from high to low on the interface, this embodiment can also output and store the target material titles in a table form.
[0075] The present embodiment provides a text retrieval method, which generates a corresponding retrieval model for each demand title. The correlation between the retrieval model and the demand title is strong, so that the retrieval model is suitable for demand title retrieval in various situations, improves the precision and applicability of the retrieval model, and thus improves the retrieval accuracy and efficiency of each demand title, realizes convenient and effective retrieval of text, enhances the correlation between the retrieval model and the demand title, improves the precision and applicability of the retrieval model, and thus improves the retrieval efficiency and accuracy of the demand title.
[0076] Embodiment 2:
[0077] like Figure 2 As shown, this embodiment provides a text retrieval method. The text retrieval method includes:
[0078] S201, obtaining material title and requirement title.
[0079] S202, converting the demand title into a proper noun to obtain a converted title corresponding to the demand title.
[0080] In this embodiment, proper noun conversion is Figure 2 Optimization processing in .
[0081] S203, based on a vectorization algorithm, vectorize the material title and the conversion title to obtain a large demand word set and a large material word set, wherein the vectorization algorithm includes a Jieba algorithm.
[0082] In this embodiment, the jieba algorithm is Figure 2 Jieba technical word segmentation in the demand large word set and material large word set Figure 2 A large independent word set.
[0083] S204, summarizing the demand large word set and the material large word set to obtain a vocabulary union; calculating the word frequency of each word in the vocabulary union, and sorting each word in the demand large word set and the material large word set based on the word frequency of each word, to obtain a demand word set and a material word set.
[0084] S205, calculate the word set length of the demand word set and the material word set as the complexity of the demand title and the material title; determine the model adaptive weight corresponding to the demand title based on the mapping relationship between the preset complexity and the model adaptive weight; generate the retrieval model corresponding to the demand title based on the model adaptive weight corresponding to the demand title, the cosine similarity calculation formula and the dynamic time warping DTW calculation formula.
[0085] S206, based on the retrieval model corresponding to the demand title, evaluating the matching degree between the demand title and the material title.
[0086] S207, based on the matching degree between the required title and the material title, determine and display the target material title.
[0087] In this embodiment, the matching degree between the demand title and the material title is Figure 2 The similarity match in .
[0088] The present embodiment provides a text retrieval method, which generates a corresponding retrieval model for each demand title. The correlation between the retrieval model and the demand title is strong, so that the retrieval model is suitable for demand title retrieval in various situations, improves the precision and applicability of the retrieval model, and thus improves the retrieval accuracy and efficiency of each demand title, realizes convenient and effective retrieval of text, enhances the correlation between the retrieval model and the demand title, improves the precision and applicability of the retrieval model, and thus improves the retrieval efficiency and accuracy of the demand title.
[0089] Embodiment 3:
[0090] like Figure 3 As shown, this embodiment provides a text retrieval device, including a demand processing module, a word set construction module, a model matching module and a result display module.
[0091] Demand processing module, used to obtain material title and demand title,
[0092] The word set building module is connected to the demand processing module and is used to convert the demand title into a proper noun to obtain the converted title corresponding to the demand title.
[0093] The word set building module is also used to vectorize the material title and the conversion title to obtain the required word set and the material word set.
[0094] The model matching module is connected to the word set building module to calculate the word set length of the demand word set and the material word set as the complexity of the demand title and the material title.
[0095] The model matching module is also used to determine the model adaptive weight corresponding to the demand title based on the mapping relationship between the preset complexity and the model adaptive weight.
[0096] The model matching module is also used to generate a retrieval model corresponding to the demand title based on the model adaptive weight, cosine similarity calculation formula and dynamic time warping DTW calculation formula corresponding to the demand title.
[0097] The model matching module is also used to evaluate the matching degree between the demand title and the material title based on the retrieval model corresponding to the demand title.
[0098] The result display module is connected to the model matching module and is used to determine and display the target material title based on the matching degree between the required title and the material title.
[0099] A text retrieval device provided in this embodiment generates a corresponding retrieval model for each demand title. The correlation between the retrieval model and the demand title is strong, so that the retrieval model is suitable for demand title retrieval in various situations, improves the precision and applicability of the retrieval model, and thus improves the retrieval accuracy and efficiency of each demand title, realizes convenient and effective retrieval of text, enhances the correlation between the retrieval model and the demand title, improves the precision and applicability of the retrieval model, and thus improves the retrieval efficiency and accuracy of the demand title.
[0100] Embodiment 4:
[0101] like Figure 4 As shown, this embodiment provides a text retrieval device, including an acquisition module 41, an evaluation module 42 and a determination module 43. The acquisition module 41 is used to acquire a material title, a requirement title and a retrieval model corresponding to the requirement title. The evaluation module 42 is connected to the acquisition module 41 and is used to evaluate the matching degree between the requirement title and the material title based on the retrieval model corresponding to the requirement title. The determination module 43 is connected to the evaluation module 42 and is used to determine and display the target material title based on the matching degree between the requirement title and the material title.
[0102] Specifically, the acquisition module 41 includes: a first determination unit 411 and a generation unit 412. The first determination unit 411 is used to determine the model adaptive weight corresponding to the demand title, and the generation unit 412 is used to generate a retrieval model corresponding to the demand title based on the model adaptive weight corresponding to the demand title, the cosine similarity calculation formula and the dynamic time warping DTW calculation formula.
[0103] Specifically, the first determination unit 411 includes: a conversion subunit, a vectorization subunit, a calculation subunit and a determination subunit. The conversion subunit is used to convert the demand title into a proper noun to obtain a converted title corresponding to the demand title. The vectorization subunit is used to vectorize the material title and the converted title to obtain a demand word set and a material word set. The calculation subunit is used to calculate the word set length of the demand word set and the material word set as the complexity of the demand title and the material title. The determination subunit is used to determine the model adaptive weight corresponding to the demand title based on the mapping relationship between the preset complexity and the model adaptive weight.
[0104] Specifically, the conversion subunit includes: a word segmentation minimum unit, a coreference resolution minimum unit, a judgment minimum unit, a replacement minimum unit and an update minimum unit. The word segmentation minimum unit is used to perform word segmentation processing on the demand title to obtain a word set corresponding to the demand title. The coreference resolution minimum unit is used to perform coreference resolution on the word set corresponding to the demand title and determine the coreference words corresponding to the demand title, wherein the coreference words are used to represent multiple words referring to the same object. The judgment minimum unit is used to determine whether the preset coreference chain includes the coreference words corresponding to the demand title. The replacement minimum unit is used to replace the coreference words corresponding to the demand title with target object words in response to the preset coreference chain including the coreference words corresponding to the demand title, and obtain the conversion title corresponding to the demand title, wherein the target object word is used to represent the preset word of the object referred to by the coreference words corresponding to the demand title in the preset coreference chain. The update minimum unit is used to generate an object and its corresponding entity based on the coreference words corresponding to the demand title in response to the preset coreference chain not including the coreference words corresponding to the demand title, so as to update the preset coreference chain.
[0105] Specifically, the vectorization sub-unit includes: a vectorization minimum unit, a summary minimum unit and a sorting minimum unit. The vectorization minimum unit is used to vectorize the material title and the conversion title based on a vectorization algorithm to obtain a large demand word set and a large material word set, wherein the vectorization algorithm includes the jieba algorithm. The summary minimum unit is used to summarize the large demand word set and the large material word set to obtain a vocabulary union. The sorting minimum unit is used to calculate the word frequency of each word in the vocabulary union, and based on the word frequency of each word, sort each word in the large demand word set and the large material word set respectively to obtain a demand word set and a material word set.
[0106] Specifically, the evaluation module 42 includes an evaluation unit 421, which is used to evaluate the matching degree between the demand title and the material title based on the retrieval model corresponding to the demand title, specifically including:
[0107] According to formulas (1) and (2), the matching degree between the required title and the material title is calculated:
[0108]
[0109] D(i,j)=d(ai,bj)+min(D(i-1,j),D(i,j-1),D(i-1,j-1))(2),
[0110] Where A represents the first weight, B represents the second weight, Wk represents the kth word in the demand word set a, Yk represents the kth word in the material word set b, n represents the total number of words in the demand word set a or the material word set b, D(·) represents the DTW calculation formula, ai represents the i-th word in the demand word set a, bj represents the j-th word in the material word set b, d(ai,bj) represents the distance between ai and bj, D(i,j) represents the DTW distance between the i-th word in the demand word set a and the j-th word in the demand word set b, the i-th word refers to the first word to the i-th word, the j-th word refers to the first word to the j-th word, and d(a,b) represents the DTW distance between all words in the demand word set a and all words in the material word set b.
[0111] Specifically, the determination module 43 includes: a judgment unit 431 and a second determination unit 432, the judgment unit 431 is used to judge whether the matching degree between the required title and the material title is greater than a preset threshold, and the second determination unit 432 is used to determine that the material title is a target material title in response to the matching degree between the required title and the material title being greater than the preset threshold, and display the target material title.
[0112] It can be understood that the text retrieval device provided above executes the text retrieval method corresponding to the embodiment 1 provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects of the scheme corresponding to the text retrieval method of the embodiment 1 above, and will not be repeated here.
[0113] Embodiment 5:
[0114] This embodiment provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to implement the text retrieval method in the above-mentioned embodiment 1 or embodiment 2.
[0115] Embodiment 6:
[0116] This embodiment provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the text retrieval method in the above-mentioned embodiment 1 or embodiment 2 is implemented.
[0117] It is to be understood that the above embodiments are merely exemplary embodiments used to illustrate the principles of the present invention, but the present invention is not limited thereto. For those of ordinary skill in the art, various modifications and improvements can be made without departing from the spirit and essence of the present invention, and these modifications and improvements are also considered to be within the scope of protection of the present invention.
Claims
1. A text retrieval method, characterized in that: include: Obtain the material title, requirement title and the corresponding retrieval model of the requirement title; Based on the retrieval model corresponding to the demand title, evaluate the matching degree between the demand title and the material title; Based on the matching degree between the requirement title and the material title, the target material title is determined and displayed.
2. The text retrieval method according to claim 1, characterized in that: The retrieval model corresponding to the required title is obtained, specifically including: Determine the model adaptation weight corresponding to the requirement title; Based on the model adaptive weights, cosine similarity calculation formula and dynamic time warping DTW calculation formula corresponding to the demand title, a retrieval model corresponding to the demand title is generated.
3. The text retrieval method according to claim 2, characterized in that: The step of determining the model adaptive weight corresponding to the demand title specifically includes: Perform proper noun conversion on the demand title to obtain a converted title corresponding to the demand title; Vectorizing the material title and the conversion title to obtain a demand word set and a material word set; Calculate the word set length of the demand word set and the material word set to serve as the complexity of the demand title and the material title; Based on the mapping relationship between the preset complexity and the model adaptive weight, the model adaptive weight corresponding to the requirement title is determined.
4. The text retrieval method according to claim 3, characterized in that: The step of converting the demand title into a proper noun to obtain a converted title corresponding to the demand title specifically includes: Perform word segmentation on the demand title to obtain a word set corresponding to the demand title; Performing coreference resolution on the word set corresponding to the demand title, and determining the coreference words corresponding to the demand title, wherein the coreference words are used to represent multiple words referring to the same object; Determine whether the preset coreference chain includes the coreference words corresponding to the requirement title; In response to the preset coreference chain including the coreference words corresponding to the requirement title, the coreference words corresponding to the requirement title are replaced with the target object words to obtain a converted title corresponding to the requirement title, wherein the target object words are used to represent the preset words of the objects referred to by the coreference words corresponding to the requirement title in the preset coreference chain; In response to the fact that the preset coreference chain does not include the coreference word corresponding to the requirement title, an object and its corresponding entity are generated based on the coreference word corresponding to the requirement title to update the preset coreference chain.
5. The text retrieval method according to claim 3, characterized in that: The vectorization of the material title and the conversion title to obtain the required word set and the material word set specifically includes: Based on a vectorization algorithm, the material title and the conversion title are vectorized to obtain a large demand word set and a large material word set, wherein the vectorization algorithm includes a Jieba algorithm; Summarize the demand word set and the material word set to obtain the vocabulary union; The word frequency of each word in the vocabulary set is calculated, and based on the word frequency of each word, the words in the demand word set and the material word set are sorted respectively to obtain the demand word set and the material word set.
6. The text retrieval method according to claim 3, characterized in that: The model adaptive weight includes a first weight and a second weight, The step of evaluating the matching degree between the demand title and the material title based on the retrieval model corresponding to the demand title specifically includes: According to formulas (1) and (2), the matching degree between the required title and the material title is calculated: D(i,j)=d(ai,bj)+min(D(i-1,j),D(i,j-1),D(i-1,j-1))(2), Where A represents the first weight, B represents the second weight, Wk represents the kth word in the demand word set a, Yk represents the kth word in the material word set b, n represents the total number of words in the demand word set a or the material word set b, D(·) represents the DTW calculation formula, ai represents the i-th word in the demand word set a, bj represents the j-th word in the material word set b, d(ai,bj) represents the distance between ai and bj, D(i,j) represents the DTW distance between the i-th word in the demand word set a and the j-th word in the demand word set b, the i-th word refers to the first word to the i-th word, the j-th word refers to the first word to the j-th word, and D(a,b) represents the DTW distance between all words in the demand word set a and all words in the material word set b.
7. The text retrieval method according to claim 1, characterized in that: The determining and displaying of the target material title based on the matching degree between the required title and the material title specifically includes: Determine whether the matching degree between the requirement title and the material title is greater than a preset threshold; In response to the matching degree between the requirement title and the material title being greater than a preset threshold, the material title is determined to be a target material title, and the target material title is displayed.
8. A text retrieval device, characterized in that: It includes acquisition module, evaluation module and determination module. The acquisition module is used to obtain the material title, the requirement title and the corresponding retrieval model of the requirement title. The evaluation module is connected to the acquisition module and is used to evaluate the matching degree between the demand title and the material title based on the retrieval model corresponding to the demand title. The determination module is connected with the evaluation module and is used to determine and display the target material title based on the matching degree between the demand title and the material title.
9. An electronic device, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to implement a text retrieval method as claimed in any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, a text retrieval method as described in any one of claims 1 to 7 is implemented.