Text associated information retrieval method and device, electronic equipment and storage medium
By matching multiple features and fusing penalties in text-related information retrieval, dynamically adjusting the matching weights, the problem of low retrieval accuracy caused by relying on a single feature in the prior art is solved, and more efficient and accurate text information retrieval is achieved.
Patent Information
- Application Number
- CN202510122106.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-06-06
AI Technical Summary
Existing search methods based on text-related information usually rely only on a single feature, resulting in low retrieval accuracy.
By receiving a search request for the target content, it is determined that a variety of features match the features of the first text information set, the values of the pre-selected penalty factor are fused, and the matching weights of the target content and each text information are determined, and the second text information set is output.
By matching multiple features and fusing penalties and dynamically adjusting the matching weights, it can more comprehensively capture the similarity and correlation between text information, avoiding mismatch caused by a single feature, thereby improving retrieval accuracy.
Smart Images

Figure CN120104783A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer application technology, and in particular to a text-related information retrieval method, device, electronic device and storage medium. Background Art
[0002] With the rapid development of information technology, text information covering knowledge and content in various fields has shown an exponential growth trend, providing users with rich information resources while also bringing great challenges to text information retrieval. How to quickly and accurately find the information users need in massive text information has become an important research topic in the current information retrieval field.
[0003] Traditional information retrieval methods are mainly divided into two categories: content-based retrieval and retrieval based on text-related information. Content-based retrieval mainly relies on factors such as term information, location, and features in text information, and finds relevant information by comparing and analyzing these factors. When faced with complex and changeable text information, this method often fails to achieve ideal retrieval results. Retrieval based on text-related information is to perform conditional queries by calculating the similarity between text feature vectors. This method has a simple and effective operation process. In practical applications, users often pay more attention to the correlation and similarity between text information rather than the specific content in the text information, so retrieval based on text-related information is more in line with the actual needs of users.
[0004] However, in the related art, the retrieval method based on text-related information usually only uses a single feature as a reference for retrieval, and the retrieval accuracy is low. Summary of the invention
[0005] The purpose of this application is to provide a text-related information retrieval method, device, electronic device and storage medium to improve retrieval accuracy.
[0006] In order to solve the above technical problems, this application provides the following technical solutions:
[0007] In a first aspect, a text-related information retrieval method is provided, comprising:
[0008] receiving a retrieval request for target content;
[0009] Determining a first text information set, wherein the first text information set includes text information of at least one information category;
[0010] After performing feature matching on multiple features corresponding to the target content and multiple features corresponding to each text information in the first text information set, fusing the value of the pre-selected penalty factor to determine the matching weight between the target content and each text information in the first text information set;
[0011] Determine a second text information set according to a matching weight between the target content and each text information in the first text information set, wherein the second text information set is a complete set or a subset of the first text information set;
[0012] The second text information set is outputted to respond to the search request.
[0013] Optionally, determining the first text information set according to the target content includes:
[0014] Determining the category of information to be retrieved based on multiple features corresponding to the target content;
[0015] A set consisting of text information in the information category to be retrieved is determined as a first text information set.
[0016] Optionally, after performing feature matching on the multiple features corresponding to the target content and the multiple features corresponding to each text information in the first text information set, fusing the value of a predetermined penalty factor to determine the matching weight between the target content and each text information in the first text information set includes:
[0017] Determine the target feature among the multiple features corresponding to the target content and the multiple features corresponding to each text information in the first text information set as an independent variable of the penalty factor;
[0018] For each text information in the first text information set, determining a value of a penalty factor corresponding to the current text information according to a matching degree between the target feature corresponding to the target content and the target feature corresponding to the current text information;
[0019] Based on the value of the penalty factor corresponding to the current text information and the matching results of the multiple features corresponding to the target content and the multiple features corresponding to the current text information, the matching weight of the target content and the current text information is determined.
[0020] Optionally, the determining the matching weight between the target content and the current text information based on the value of the penalty factor corresponding to the current text information and the matching results between the multiple features corresponding to the target content and the multiple features corresponding to the current text information includes:
[0021] Calculating a similarity measure between the target content and the current text information based on a matching result between multiple features corresponding to the target content and multiple features corresponding to the current text information;
[0022] A matching weight between the target content and the current text information is determined according to the similarity measurement and the value of the penalty factor corresponding to the current text information.
[0023] Optionally, the higher the matching degree between the target feature corresponding to the target content and the target feature corresponding to the current text information, the smaller the value of the penalty factor corresponding to the current text information.
[0024] Optionally, determining the second text information set according to the matching weight between the target content and each text information in the first text information set includes:
[0025] For each text information in the first text information set, if the matching weight between the target content and the current text information is greater than or equal to a first threshold, determining the current text information as the target text information;
[0026] The determined set of target text information is determined as a second text information set.
[0027] Optionally, the outputting the second text information set includes:
[0028] The second text information set is outputted in the order of the matching weights between the target content and each text information in the second text information set.
[0029] In a second aspect, a text-related information retrieval device is provided, comprising:
[0030] A receiving module, used for receiving a search request for target content;
[0031] A first determining module, configured to determine a first text information set according to the target content;
[0032] a second determination module, configured to determine a matching weight between the target content and each text information in the first text information set by fusing a value of a predetermined penalty factor after feature matching is performed on the multiple features corresponding to the target content and the multiple features corresponding to each text information in the first text information set;
[0033] A third determination module, configured to determine a second text information set according to a matching weight between the target content and each text information in the first text information set, wherein the second text information set is a complete set or a subset of the first text information set;
[0034] An output module is used to output the second text information set in response to the search request.
[0035] In a third aspect, an electronic device is provided, including:
[0036] Memory for storing computer programs;
[0037] A processor is used to implement the steps of the text-related information retrieval method as described in the first aspect when executing the computer program.
[0038] In a fourth aspect, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of the text-related information retrieval method as described in the first aspect are implemented.
[0039] In a fifth aspect, a computer program product is provided, which includes computer instructions, which are stored in a computer-readable storage medium and are suitable for being read and executed by a processor, so that a computer device having the processor performs the steps of the text-related information retrieval method as described in the first aspect.
[0040] By applying the technical solution provided in the embodiment of the present application, after receiving a retrieval request for target content, first, a first text information set is determined. After feature matching is performed on multiple features corresponding to the target content and multiple features corresponding to each text information in the first text information set, the value of a predetermined penalty factor is integrated to determine the matching weight between the target content and each text information in the first text information set. Based on the matching weight, a second text information set is determined, and the second text information set is output to respond to the retrieval request. By matching multiple features and integrating the penalty factor to dynamically adjust the matching weight, the similarity and correlation between text information can be captured more comprehensively, avoiding mismatches caused by relying solely on a single feature, and improving retrieval accuracy.
[0041] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the related technologies, the drawings required for use in the embodiments or the related technical descriptions are briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0043] Figure 1 A schematic diagram of the composition architecture of a system to which the embodiments of the present application are applicable;
[0044] Figure 2 This is a flowchart of an implementation method of a text-related information retrieval method in an embodiment of the present application;
[0045] Figure 3 This is a schematic diagram of text information feature classification in an embodiment of the present application;
[0046] Figure 4 A schematic diagram of a text-related information retrieval architecture in an embodiment of the present application;
[0047] Figure 5 This is a first schematic diagram of the comparison results in the embodiments of the present application;
[0048] Figure 6 This is a second schematic diagram of the comparison results in the embodiments of the present application;
[0049] Figure 7 This is a third schematic diagram of the comparison results in the embodiments of the present application;
[0050] Figure 8 This is a fourth schematic diagram of the comparison results in the embodiments of the present application;
[0051] Fig. 9 This is a schematic diagram of the structure of a text-related information retrieval device in an embodiment of the present application;
[0052] Fig.10 It is a schematic diagram of the structure of an electronic device in an embodiment of the present application. DETAILED DESCRIPTION
[0053] The technical solutions in the embodiments of the present application will be described in detail below in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments in the present application belong to the scope of protection of the present application.
[0054] The terms "first", "second", etc. of the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the terms used in this way can be interchanged where appropriate, so that the embodiments of the present application can be implemented in an order other than those illustrated or described here, and the objects distinguished by "first" and "second" are generally of one type, and the number of objects is not limited, for example, the first object can be one or more. In addition, "or" in the present application represents at least one of the connected objects. For example, the protection scope of "A or B" covers at least three schemes, that is, scheme one: including A and excluding B; scheme two: including B and excluding A; scheme three: including both A and B. In addition, the terms "A and / or B", "at least one of A and B", and "at least one of A or B" also cover at least the above three schemes respectively. The character " / " generally indicates that the objects associated before and after are in an "or" relationship.
[0055] The core of this application is to provide a text-related information retrieval method, which can be applied to information retrieval related scenarios, such as information retrieval scenarios in natural language processing.
[0056] For ease of understanding, the following first introduces the composition architecture of the system to which the technical solution of this application is applicable. Figure 1 The system includes an application server and an information base. The information base can store multiple text information, each of which has multiple features. The application server can receive a user's search request for target content, determine a first text information set in the information base, and then perform feature matching on multiple features corresponding to the target content and multiple features corresponding to each text information in the first text information set, integrate penalty factors, determine the matching weight between the target content and each text information in the first text information set, and then determine a second text information set based on the matching weight between the target content and each text information in the first text information set, and finally output the second text information set to respond to the search request.
[0057] By matching multiple features and integrating penalty factors to dynamically adjust matching weights, it is possible to more comprehensively capture the similarities and correlations between text information, avoid mismatches caused by relying solely on a single feature, and improve retrieval accuracy.
[0058] It should be noted that the above description is based on an example in which the application server is an independent server, but it is understandable that in actual applications, the application server can also be replaced by an application server cluster, or a distributed cluster composed of multiple application servers. Figure 1 In the present invention, the application server can also be replaced by an application platform composed of multiple application servers.
[0059] See also Figure 2 FIG. 1 is a flowchart of an implementation of an information retrieval method provided in an embodiment of the present application. The method may include the following steps:
[0060] S210: Receive a search request for target content.
[0061] In the embodiment of the present application, the user can issue a search request for the target content according to actual needs. The search request can be understood as a query request, and the target content can be understood as the query content.
[0062] After receiving the retrieval request for the target content, the subsequent steps can be continued.
[0063] S220: Determine a first text information set.
[0064] In the embodiment of the present application, the information library may include multiple text information, and the text information often contains complex semantic relationships, such as the association between words, the logical relationship between sentences, etc. Each text information corresponds to multiple features, and each text information can be pre-classified by multiple features. By classifying the text information by multiple features, the semantic features of the text information can be more carefully portrayed, the similarities and differences between the text information can be captured, which helps to more accurately understand the meaning of the text information and provide strong support for subsequent feature matching.
[0065] In view of the multi-level and multi-attribute characteristics of text information, a dynamic classification tree can be established, and each branch stores a type of information, such as title, abstract, text, keywords, etc., to reduce the disadvantage that a single type cannot reflect global information. Based on the structure and logical relationship of the classification tree, the spatial distribution of information categories can be obtained. In the feature space, large categories and small categories of information are inclusion patterns. The directional similarity evaluation method can be used as the similarity criterion of the classification tree, which can reflect the similarity between text information even in high-dimensional space.
[0066] The basic idea of building a dynamic classification tree is to select a group of clustering points or give an initial cluster, and then make the text information cluster to the clustering points according to a certain principle, and continuously modify or iterate the clustering points until the classification is reasonable or the iteration is stable. This method allows text information to move from one category to another, and in the classification process, the total number of categories and category centers are repeatedly calculated according to certain principles, so that the classification results gradually tend to be reasonable until certain conditions are met and the classification is completed. Figure 3 As shown in FIG. 1 , a classification result of text information is shown, where the solid circle represents the feature points of the large category and the solid triangle represents the feature points of the small category.
[0067] For example, for the feature vector d of the i-th text information i , calculate its difference with the large target feature vector q i The similarity between sim(d i ,q i ), i = 1, 2, ..., N, N represents the total amount of text information, thereby realizing the multi-feature classification of text information, such as word feature, sentence feature, paragraph feature, etc.
[0068]
[0069] in, The feature vector representing the text information in high-dimensional space, Represents a large target feature vector in a high-dimensional space; l represents the dimension of the text information sample.
[0070] The text information sample dimension refers to a set of features or attributes used to represent text information, such as keywords, titles, etc. The large target feature vector can be understood as a feature vector of a large category or a small category. For example, a large category can be for text, and a small category can be for a text segment.
[0071] After receiving a search request for target content, a first text information set may be determined. Optionally, text information belonging to a corresponding information category or multiple information categories may be determined based on the target content. The first text information set may include text information of at least one information category.
[0072] S230: After performing feature matching on multiple features corresponding to the target content and multiple features corresponding to each text information in the first text information set, a predetermined penalty factor value is integrated to determine a matching weight between the target content and each text information in the first text information set.
[0073] When users search for information, the target content they provide may not accurately express their search intent, and there are multiple possible interpretations, which may result in the final retrieved text information not matching the search intent. After classifying text information using the similarity of text information features, the regionality of each category of text information in space is more obvious.
[0074] After determining the first text information set according to the target content, feature matching can be performed on the multiple features corresponding to the target content and the multiple features corresponding to each text information in the first text information set, so that not only the matching degree of a single word is taken into account, but also the association between words, contextual relationship and other factors can be comprehensively considered, so as to more comprehensively understand the user's search intention. Even if the target content targeted by the user's search request is ambiguous or vague, the text information required by the user can be more accurately retrieved by integrating information of multiple features.
[0075] After feature matching is performed on multiple features corresponding to the target content and multiple features corresponding to each text information in the first text information set, the matching weight between the target content and each text information in the first text information set can be determined by integrating the value of the predetermined penalty factor.
[0076] S240: Determine a second text information set according to a matching weight between the target content and each text information in the first text information set.
[0077] After feature matching is performed on multiple features corresponding to the target content and multiple features corresponding to each text information in the first text information set, the value of the predetermined penalty factor is integrated to determine the matching weight between the target content and each text information in the first text information set. The matching degree between the target content and each text information in the first text information set can be determined based on the size of the matching weight between the target content and each text information in the first text information set, and then the second text information set can be determined.
[0078] For each text information in the first text information set, if the matching weight between the target content and the current text information is greater, it is considered that the matching degree between the target content and the current text information is higher, and the possibility that the current text information satisfies the search request for the target content is greater. Conversely, if the matching weight between the target content and the current text information is smaller, it is considered that the matching degree between the target content and the current text information is lower, and the possibility that the current text information satisfies the search request for the target content is smaller. The current text information refers to the text information targeted by the current operation.
[0079] The determined second text information set is the entire set or a subset of the first text information set, that is, part or all of the text information in the first text information set constitutes the second text information set.
[0080] S250: Output a second text information set to respond to the search request.
[0081] After the second text information set is determined, the second text information set may be further output to respond to the user's search request.
[0082] By applying the technical solution provided in the embodiment of the present application, after receiving a retrieval request for target content, a first text information set is first determined, and feature matching is performed between multiple features corresponding to the target content and multiple features corresponding to each text information in the first text information set. A penalty factor is integrated to determine the matching weight between the target content and each text information in the first text information set. Based on the matching weight, a second text information set is determined, and the second text information set is output to respond to the retrieval request. By matching multiple features and integrating the penalty factor to dynamically adjust the matching weight, the similarity and correlation between text information can be captured more comprehensively, avoiding mismatches caused by relying solely on a single feature, and improving retrieval accuracy.
[0083] In some embodiments of the present application, determining the first text information set may include the following steps:
[0084] Determine the category of information to be retrieved based on multiple features corresponding to the target content;
[0085] A set consisting of text information in the information category to be retrieved is determined as a first text information set.
[0086] For the convenience of description, the above steps are combined for explanation.
[0087] In an embodiment of the present application, after receiving a search request for target content, feature extraction can be performed on the target content to obtain a variety of features corresponding to the target content, such as word features, sentence features, paragraph features, etc., or semantic features, word frequency and inverse document frequency value features, vocabulary features, etc. Based on the various features corresponding to the target content, the information category to be retrieved can be determined. The determined information category to be retrieved may include one information category or multiple information categories. For example, the target content corresponds to feature 1 and feature 2, the information category corresponding to feature 1 is category 1, and the information category corresponding to feature 2 is category 2, and the determined information category to be retrieved may include category 1 and category 2.
[0088] A set consisting of text information in the information category to be retrieved is determined as a first text information set.
[0089] Based on multiple features corresponding to the target content of the search request, the information category to be searched is first determined, and then the set consisting of text information in the information category to be searched is determined as the first text information set, which can make the search more targeted and help improve the search efficiency.
[0090] In some embodiments of the present application, after performing feature matching on multiple features corresponding to the target content and multiple features corresponding to each text information in the first text information set, fusing the value of a predetermined penalty factor to determine the matching weight between the target content and each text information in the first text information set may include the following steps:
[0091] Determine the target feature among the multiple features corresponding to the target content and the multiple features corresponding to each text information in the first text information set as an independent variable of the penalty factor;
[0092] For each text information in the first text information set, determining a value of a penalty factor corresponding to the current text information according to a matching degree between a target feature corresponding to the target content and a target feature corresponding to the current text information;
[0093] Based on the value of the penalty factor corresponding to the current text information and the matching results of the multiple features corresponding to the target content and the multiple features corresponding to the current text information, the matching weight of the target content and the current text information is determined.
[0094] For the convenience of description, the above steps are combined for explanation.
[0095] In an embodiment of the present application, after receiving a search request for target content and determining the first text information set, feature extraction can be performed on the target content and each text information in the first text information set to obtain multiple features corresponding to the target content and multiple features corresponding to each text information in the first text information set. The target feature among the multiple features corresponding to the target content and the multiple features corresponding to each text information in the first text information set can be determined as an independent variable of the penalty factor.
[0096] For each text information in the first text information set, the target feature corresponding to the target content is matched with the target feature corresponding to the current text information, and the value of the penalty factor corresponding to the current text information can be dynamically adjusted according to the degree of matching. The values of the penalty factors corresponding to different text information in the first text information set may be the same or different. Optionally, the higher the degree of matching between the target feature corresponding to the target content and the target feature corresponding to the current text information, the smaller the value of the penalty factor corresponding to the current text information, and conversely, the lower the degree of matching between the target feature corresponding to the target content and the target feature corresponding to the current text information, the larger the value of the penalty factor corresponding to the current text information.
[0097] Based on the value of the penalty factor corresponding to the current text information and the matching results of the multiple features corresponding to the target content and the multiple features corresponding to the current text information, the matching weight of the target content and the current text information is determined. Optionally, when the matching degree between the target feature corresponding to the target content and the target feature corresponding to the current text information is higher than the second threshold, the value of the penalty factor corresponding to the current text information is reduced to increase the matching weight between the target content and the current text information. Optionally, when the matching degree between the target feature corresponding to the target content and the target feature corresponding to the current text information is lower than the third threshold, the value of the penalty factor corresponding to the current text information is increased to reduce the matching weight between the target content and the current text information.
[0098] The second threshold and the third threshold can be set and adjusted according to actual conditions.
[0099] Optionally, the similarity measure between the target content and the current text information can be calculated based on the matching results of multiple features corresponding to the target content and multiple features corresponding to the current text information, and then the matching weight between the target content and the current text information can be determined based on the similarity measure and the value of the penalty factor corresponding to the current text information.
[0100] The current text information is the text information targeted by the current operation. The multiple features corresponding to the target content are matched with the multiple features corresponding to the current text information to obtain a matching result. Based on the matching result, the similarity measure between the target content and the current text information can be calculated. The similarity measure is adjusted using the value of the penalty factor corresponding to the current text information. Optionally, if the value of the penalty factor corresponding to the current text information is large, the similarity measure is adjusted downward to decrease it, so that the possibility of the current text information satisfying the retrieval request is reduced. If the value of the penalty factor corresponding to the current text information is small, the similarity measure is adjusted upward to increase it, so that the possibility of the current text information satisfying the retrieval request is increased.
[0101] Based on the adjusted similarity measurement, the matching weight between the target content and the current text information can be obtained. Optionally, the adjusted similarity measurement can be determined as the matching weight between the target content and the current text information.
[0102] The fusion penalty factor can dynamically adjust the similarity measure according to the features corresponding to the target content and the text information, accurately determine the matching weight between the target content and each text information in the first text information set, and help improve the overall retrieval performance.
[0103] In some embodiments of the present application, determining the second text information set according to the matching weight between the target content and each text information in the first text information set may include the following steps:
[0104] For each text information in the first text information set, if the matching weight between the target content and the current text information is greater than or equal to a first threshold, the current text information is determined as the target text information;
[0105] The determined set of target text information is determined as a second text information set.
[0106] For the convenience of description, the above steps are combined for explanation.
[0107] In the embodiment of the present application, after determining the matching weight between the target content and each text information in the first text information set, the matching weight between the target content and the current text information can be compared with the first threshold for each text information in the first text information set. If it is greater than or equal to the first threshold, it is considered that the current text information is more likely to meet the search request, and the current text information can be determined as the target text information. If it is less than the first threshold, it is considered that the current text information is less likely to meet the search request, and the current text information can be ignored. The first threshold can be set and adjusted according to actual conditions.
[0108] The determined set of target text information is determined as a second text information set, so as to output the second text information set.
[0109] Determining a set consisting of text information corresponding to a larger matching weight as the second text information set can improve the accuracy of the retrieval result.
[0110] In some embodiments of the present application, outputting the second text information set may include the following steps:
[0111] The second text information set is outputted in the order of the matching weights between the target content and each text information in the second text information set.
[0112] When outputting the second text information set, the text information is displayed in order of the matching weights, so that the user can quickly obtain the text information with a larger matching weight, thereby improving the user experience.
[0113] For ease of understanding, the process of fusing multiple feature penalty factors for information retrieval is further explained below.
[0114] Fusion of multi-feature penalty factors is a type of machine learning. The penalty factor is used as a hyperparameter to adjust the degree of penalty for text information. A key feature or some key features of the target content are used as independent variables of the penalty factor to limit or constrain the matching effect. By establishing a feature training set, you can find related matching items in the feature item set. The matching item can be understood as a feature item with similarity calculated in the feature training set, which improves the accuracy and efficiency of subsequent information retrieval.
[0115] Set the feature training set L MDSL1 and L MDSL2 , for a specific set of feature items M MDSL1 and M MDSL2 , taking one or some key features α as the independent variable of the penalty factor, and adjusting the value of the penalty factor. The larger the value of the penalty factor, the higher the strictness of the matching, or the higher the matching requirement.
[0116] The feature training set includes both text information and features corresponding to the text information. This is because when performing multi-feature matching and penalty factor calculation, it is necessary to consider both the specific content of the text information (such as vocabulary, sentences, etc.) and the corresponding features (such as word frequency, word frequency and inverse document frequency value, semantic relationship, etc.). The feature training set is used to train the model so that it can learn the association between text information or features corresponding to text information, thereby achieving accurate matching and efficient retrieval of text information. The feature item set can be understood as feature data extracted from the feature training set.
[0117] When α is 1, (1-α)×MMDSL2 , (1-α)×L MDSL2 The values are all 0, and this condition is regarded as an equivalent condition for matching. A suitable penalty factor is selected through training optimization and changed according to the current text information environment.
[0118] The conditions that the penalty factor needs to meet are:
[0119] (1)M MDSL2 Sufficient features should be retained so that the number of features can meet the minimum matching conditions for retrieval requirements and cover the core content and global characteristics of the text information;
[0120] (2)M MDSL2 The number of features in the feature library should be greater than the total number of all features in the feature library (including large category features (global features) and small category features (local features)).
[0121] Only when the above two conditions are met at the same time can the matching value of the feature library be detected in the text information set. The feature library here can also be understood as the information library. The matching length of the feature corresponding to the target content and the feature corresponding to the text information is proportional to the feature weight. The longer the matching length, the greater the feature weight. According to this feature, the following L MDSL1 The matching calculation formula is:
[0122]
[0123] Among them, ξ is the minimum matching condition that two feature matching items need to meet simultaneously, which means the similarity result or other related metrics obtained in the calculation process, which can characterize the similarity between the text information and the target content; R is the text value, which can be understood as a threshold value to determine whether the text information matching meets the requirements; w 1 、w 2 is the weight of the two matching items.
[0124] Get in M MDSL1 The feature matching results in the text information are used to achieve feature matching of the text information. After the matching is completed, key instructions are generated, and then efficient text-related information retrieval is achieved according to the key instructions.
[0125] The formula is:
[0126]
[0127] Among them, generating key instructions may include the following steps:
[0128] Extract key features from the target content, such as keywords, word frequency, topics, etc.
[0129] Generate instructions based on matching conditions, and determine whether the matching conditions are met by formula (2);
[0130] Generate instructions based on the matched keywords in a specific format.
[0131] Exemplarily, the target content includes words such as "AA", "BB", and "CC", and these keywords are extracted to generate instructions, such as MATCH ("AA", "BB", "CC").
[0132] The command can be used for feature library or information library query or feature retrieval.
[0133] The amount of text information is growing explosively. Whether it is academic papers, news reports or corporate data, they are all accumulating. This information overload makes it difficult for users to quickly find the information they really need from the massive data. Text information feature matching can capture the similarity and relevance between the query and the document in terms of content, and then obtain the matching degree of these features between the query and the document, ensuring that the retrieval results are highly relevant to the user's query in terms of content. Therefore, the feature matching results of the text information M MDSL1 As input, a vector space model is established, combined with similarity calculation to achieve information query under given conditions.
[0134] This model divides a text into multiple text segments according to its structure, and each level is divided into multiple text segments according to the feature matching result M. MDSL1 Establish and set the feature vector and weight vector. The target content of the user's search request is QS 2 The text string is obtained by the similarity of the text information with S(t 1 , t 2 ,…,t N ) has a high similarity in text information, and the Boolean model parameter θ is introduced k With the vector weight g k , obtain the similarity Sim(QS 2 , S k )for:
[0135]
[0136] Among them, S k is the total number of features of k in the data set; w ij is the feature point weight. Calculate the target content QS 2 After the similarity between each text segment is calculated, it is converted into a query password and a keyword, and the relationship between the two is obtained, thereby realizing text-related information retrieval:
[0137]
[0138] The technical solution provided in the embodiments of the present application will be described below through specific examples.
[0139] A multi-level convolutional deep learning network framework is used to build a text-related information retrieval architecture. Figure 4 The text information in the embodiment of the present application may include text information, or text information obtained by converting Chinese voice information into text, or text information obtained by extracting features from video information. After feature matching of the text information, visualization processing may be performed, and the retrieved information may be output in a visualization mode.
[0140] in, Figure 4 The Chinese image and symbol library is used to store images, symbols and graphic information associated with text information, such as charts, special characters, etc.; the vocabulary library is used to store common vocabulary, domain terms and related information, such as word frequency, synonyms, etc.; the knowledge base contains deep semantic knowledge of text, voice, and video, and supports semantic reasoning and context association; the index library: index data used for fast retrieval, including keyword index and feature vector index.
[0141] The tone and context information in Chinese speech information helps to further enrich the retrieval features.
[0142] Video information is extracted through feature extraction to generate text descriptions or image features, which participate in retrieval together with text features.
[0143] The text-related information retrieval method integrating multiple feature penalty factors proposed in the embodiment of the present application includes three stages: a preprocessing stage, a feature matching stage, and an information retrieval stage.
[0144] Preprocessing stage: Multi-feature classification of text information can more carefully characterize the semantic features of text information and capture the similarities and differences between text information. The specific steps are as follows:
[0145] (1.1) Establish a dynamic classification tree, each branch stores a type of information;
[0146] (1.2) Based on the structure and logical relationship of the classification tree, the spatial distribution of information categories is obtained;
[0147] (1.3) Using the directional similarity evaluation method as the similarity criterion of the classification tree, for the feature vector d of the i-th text information i , calculate its difference with the large target feature vector q i The similarity sim(d i ,q i ), thereby achieving feature classification of text information.
[0148] Feature matching stage: Multi-feature matching is performed on the classified text information, not only considering the matching degree of a single word, but also comprehensively considering factors such as the association between words and contextual relationships, capturing multiple key features in the target content, and thus more comprehensively understanding the user's search intent. The specific steps are as follows:
[0149] (2.1) A key feature is used as the independent variable of the penalty factor to restrict or constrain a specific set of feature items M. MDSL1 and M MDSL2 , set the feature training set L MDSL1 and L MDSL2 ;
[0150] (2.2) Training optimization selects a penalty factor that satisfies the following two conditions: M MDSL2 To retain enough features; MDSL2 The number of features in the formula should be greater than the total number;
[0151] (2.3) Calculate L MDSL1 , get MDSL1 The feature matching results in the text information are used to achieve feature matching;
[0152] (2.4) After matching is completed, key instructions are generated and M is calculated. MDSL1 , and then implement efficient text-related information retrieval according to key instructions.
[0153] Information retrieval stage: Text information feature matching can capture the similarity and relevance between the target content and the text information, and then obtain the matching degree of these features between the query and the document. The specific steps are as follows:
[0154] (3.1) The feature matching result M of the text information MDSL1 As input, a vector space model is established to divide a text into multiple text segments according to its structure. Each level is matched according to the feature matching result M MDSL1 Establish and set the feature vector and weight vector;
[0155] (3.2) The target content of the user's search request is QS 2 Text string, through the similarity of text information to obtain the S(t 1 , t 2 ,…,t N );
[0156] (3.3) Introducing Boolean model parameter θ k With the vector weight g k , obtain the similarity Sim(QS 2 , Sk );
[0157] (3.4) The query text QS 2 Each text segment is converted into a query password and a keyword, and the relationship between the two is obtained, thereby realizing text-related information retrieval.
[0158] The following is a retrieval experiment comparing the technical solution provided in the embodiment of the present application with related technology 1 and related technology 2 to compare the actual retrieval performance to illustrate the effectiveness of the technical solution provided in the embodiment of the present application.
[0159] 1. Experimental environment
[0160] To verify the effectiveness of the simulation application of text-related information retrieval with multi-feature penalty factors proposed in this paper, the widely used public datasets Flickr30K and MS COCO are selected as experimental datasets to expand the experimental dataset. The dataset includes a variety of data such as finance, dating, and entertainment. They are all described in the form of images and texts. In order to expand the experimental dataset, the Flickr30K dataset contains about 31,000 pictures, each of which has 5 English sentences. The MS COCO dataset contains more than 330,000 pictures, and each picture has an average of 5 English sentences. The document query rate is set to 0.5-1.0, the number of information deletions is 1000, and the text similarity weight is 0.6. According to the above settings, the following experiments are carried out.
[0161] 2. Related technologies
[0162] Related technology 1: Text retrieval method based on matrix weighted association rules. First, the user queries and retrieves the document set, and constructs the initial user-related document set. Then, the item set weights and frequencies are merged with the total weights of the feature words and the total number of documents in the initial user-related document set, and the frequent item sets containing the original query terms are mined. The candidate item sets are pruned by item weight sorting, and the confidence-correlation coefficient evaluation framework is used to mine association rules from the frequent item sets. Finally, the association rule antecedents whose consequents are the original query terms and the association rule consequents whose antecedents are the original query terms are used as expansion terms. The expansion terms are combined with the original query terms as new queries to retrieve the document set again to obtain the final retrieval result document and return it to the user.
[0163] Related technology 2: Text retrieval method based on weight sorting. First, the user query is used to retrieve text documents, and the initial user feedback document set is constructed. Then, the item set weights and frequencies are integrated with the total weights of the feature words and the total number of documents in the initial user feedback document set. The support-confidence-correlation coefficient evaluation framework is used to mine the feature word weighted association rules for the initial user feedback document set. The feature word weighted association rule antecedent is the original query term set, and the consequent is composed of non-query terms. The consequent of the weighted association rule is extracted as an expansion term. The expansion term is combined with the original query term into a new query to retrieve the text document again to obtain the final retrieval result and return it to the user.
[0164] 3. Classification of text-related information features
[0165] In order to directly observe and measure the performance of the text-related information retrieval method after integrating multiple feature penalty factors in practical applications, a text-related information feature classification experiment is designed to demonstrate the effect of the technical solution provided by the embodiment of this application in improving retrieval efficiency. Large categories in text usually contain more diverse information, and the boundaries between categories are more blurred, while small categories require higher discrimination because the similarities between categories are higher. The classification results are shown in Figure 2. Figure 3 shown.
[0166] Experimental results show that the technical solution provided in the embodiment of the present application has high practical value in actual applications, can effectively cope with data sets of different sizes and complexities, classify features of different categories, and can significantly improve the efficiency and accuracy of text retrieval, especially when dealing with large categories and small categories.
[0167] 4. Comparison of Normalized Discounted Cumulative Gain (NDCG) values based on penalty factors
[0168] NDCG not only considers the relevance of the search results, but also the position of these relevant results in the sorted list. In information retrieval, users tend to pay more attention to the top-ranked results. Therefore, NDCG gives higher weights to the top-ranked relevant results, thereby more accurately reflecting the degree to which the retrieval algorithm meets user needs. At the same time, by adjusting the value of the penalty factor, the performance of the retrieval method in different situations is evaluated. As the value of the penalty factor changes, the NDCG value on the vertical axis will also change accordingly, reflecting the changing trend of the retrieval performance. The experimental results are shown in Figure 2. Figure 5 shown.
[0169] from Figure 5It can be seen that the NDCG of the technical solution provided in the embodiment of the present application is up to 0.87, which more accurately reflects the degree to which the retrieval algorithm satisfies user needs. As the penalty factor value changes, the NDCG value on the ordinate will also change accordingly, and increasing the penalty factor value will increase the NDCG value. However, when the penalty factor is greater than 0.36, the NDCG value fluctuates around 0.85, which shows that an excessively high penalty factor will lead to overfitting and reduce the NDCG value.
[0170] 5. Experimental results based on the comparison of recall and precision
[0171] The precision rate refers to the proportion of correct search results in the algorithm, that is, the ratio of the retrieved relevant information to the total information retrieved, which reflects the true ability of the search algorithm and whether it can eliminate irrelevant information; the recall rate refers to the ratio of the retrieved relevant information to the total information in the data set, which reflects the ability of the search algorithm to identify relevant information. Among them, the recall rate can verify the comprehensiveness of the algorithm, and the precision rate can verify the quality and accuracy of the algorithm. Taking the SogouCA data set as the object, the results of the analysis of the technical solution provided in the embodiment of this application and related technologies 1 and 2 are as follows Figure 6 , Figure 7 shown.
[0172] It can be seen that whether it is based on the precision rate or the recall rate curve, the curve distribution of the technical solution provided in the embodiment of the present application is the highest. From the recall rate results, based on limited key information conditions, the technical solution provided in the embodiment of the present application can identify more relevant information, and there is a certain gap with related technology 1 and related technology 2; from the distribution of the precision rate curve, the technical solution provided in the embodiment of the present application is more obviously different from related technology 1 and related technology 2, and it can guarantee high accuracy regardless of which information association feature is based on. This shows that the extracted text information is highly consistent with the conditions, the retrieval algorithm has a high accuracy rate, and has certain practical value.
[0173] 6. Retrieval communication overhead
[0174] Since the amount of text information is large, the algorithm that can retrieve related information more accurately with less consumption has better comprehensive performance. 24,000 texts are randomly selected from the two experimental data sets, and the test results of the three algorithms are given as follows: Figure 8 It can be seen that no matter what kind of text data set is used, the retrieval communication overhead of the technical solution provided by the embodiment of the present application is the lowest. Under the same amount of text, the technical solution provided by the embodiment of the present application can achieve efficient text information retrieval with less communication consumption, and the comprehensive performance is excellent.
[0175] In general, the embodiment of the present application first derives the similarity between text vectors by establishing a classification tree and formulating corresponding classification tree similarity criteria. Then, in order to calculate the matching weight, a specific training set and feature set are set, a key feature in the text is used as the independent variable of the penalty factor, and the matching weight is obtained by calculating the text feature length that conforms to the change of the independent variable. Further, a feature vector space is also constructed. In this space, user query conditions are substituted into to generate passwords, and the passwords are converted into keywords to extract features. Finally, the similarity weights of matching samples are calculated by iterative comparison, and text-related information retrieval is continuously searched and efficiently realized. The embodiment of the present application can improve the recall and precision of text-related information retrieval, and can realize accurate and efficient query to cope with data sets of different information sizes and complexity.
[0176] The embodiments of the present application have the advantages of high accuracy, high efficiency, flexibility and strong adaptability, and can meet the diverse needs of users for text information retrieval.
[0177] Specifically include the following aspects:
[0178] 1. Improve the accuracy of text-related information retrieval. By integrating multiple text features and introducing penalty factors, the differences between text features are considered more comprehensively to avoid retrieval errors caused by a single feature. At the same time, by establishing a classification tree and feature vector space, effective organization and efficient management of text information can be achieved, further improving the accuracy of retrieval;
[0179] 2. Improve the efficiency of text-related information retrieval. Traditional text information retrieval methods often require one-by-one comparison in massive texts, which is time-consuming and laborious. The present invention, by establishing a classification tree and feature vector space, can achieve rapid positioning and efficient retrieval of text information;
[0180] 3. More flexibility and adaptability. According to different text features and retrieval requirements, the weight of the penalty factor can be flexibly adjusted, so that the retrieval method can better adapt to various complex text information retrieval scenarios. It has a wider range of use value and promotion prospects in practical applications.
[0181] Corresponding to the above method embodiment, the embodiment of the present application further provides a text-related information retrieval device. The text-related information retrieval device described below and the text-related information retrieval method described above can be referenced to each other.
[0182] See also Fig. 9 As shown, the text-related information retrieval device 900 may include the following modules:
[0183] A receiving module 910 is used to receive a search request for target content;
[0184] A first determining module 920, configured to determine a first text information set;
[0185] A second determination module 930 is used to determine the matching weight between the target content and each text information in the first text information set after performing feature matching on the multiple features corresponding to the target content and the multiple features corresponding to each text information in the first text information set, fusing the value of the predetermined penalty factor;
[0186] A third determination module 940 is used to determine a second text information set according to a matching weight between the target content and each text information in the first text information set, where the second text information set is a complete set or a subset of the first text information set;
[0187] The output module 950 is used to output the second text information set in response to the search request.
[0188] By using the device provided in the embodiment of the present application, after receiving a search request for target content, a first text information set is first determined. After feature matching is performed on multiple features corresponding to the target content and multiple features corresponding to each text information in the first text information set, the value of a predetermined penalty factor is integrated to determine the matching weight between the target content and each text information in the first text information set. Based on the matching weight, a second text information set is determined, and the second text information set is output to respond to the search request. By matching multiple features and integrating the penalty factor to dynamically adjust the matching weight, the similarity and correlation between text information can be captured more comprehensively, avoiding mismatches caused by relying solely on a single feature, and improving the search accuracy.
[0189] In some embodiments of the present application, the first determining module 920 is specifically configured to:
[0190] Determine the category of information to be retrieved based on multiple features corresponding to the target content;
[0191] A set consisting of text information in the information category to be retrieved is determined as a first text information set.
[0192] In some embodiments of the present application, the second determining module 930 is specifically configured to:
[0193] Determine the target feature among the multiple features corresponding to the target content and the multiple features corresponding to each text information in the first text information set as an independent variable of the penalty factor;
[0194] For each text information in the first text information set, determining a value of a penalty factor corresponding to the current text information according to a matching degree between a target feature corresponding to the target content and a target feature corresponding to the current text information;
[0195] Based on the value of the penalty factor corresponding to the current text information and the matching results of the multiple features corresponding to the target content and the multiple features corresponding to the current text information, the matching weight of the target content and the current text information is determined.
[0196] In some embodiments of the present application, the second determining module 930 is specifically configured to:
[0197] Calculate the similarity measure between the target content and the current text information according to the matching results of the multiple features corresponding to the target content and the multiple features corresponding to the current text information;
[0198] According to the similarity measurement and the value of the penalty factor corresponding to the current text information, the matching weight between the target content and the current text information is determined.
[0199] In some embodiments of the present application, the higher the matching degree between the target feature corresponding to the target content and the target feature corresponding to the current text information, the smaller the value of the penalty factor corresponding to the current text information.
[0200] In some embodiments of the present application, the third determination module 940 is specifically configured to:
[0201] For each text information in the first text information set, if the matching weight between the target content and the current text information is greater than or equal to a first threshold, the current text information is determined as the target text information;
[0202] The determined set of target text information is determined as a second text information set.
[0203] In some embodiments of the present application, the output module 950 is specifically used to:
[0204] The second text information set is outputted in the order of the matching weights between the target content and each text information in the second text information set.
[0205] Regarding the device in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.
[0206] Corresponding to the above method embodiment, the present application embodiment further provides an electronic device, including:
[0207] Memory for storing computer programs;
[0208] A processor is used to implement the steps of the above-mentioned text-related information retrieval method when executing a computer program.
[0209] like Fig.10, which is a schematic diagram of the composition structure of an electronic device, the electronic device may include: a processor 10, a memory 11, a communication interface 12 and a communication bus 13. The processor 10, the memory 11 and the communication interface 12 all communicate with each other through the communication bus 13.
[0210] In the embodiment of the present application, the processor 10 may be a central processing unit (CPU), an application specific integrated circuit, a digital signal processor, a field programmable gate array or other programmable logic devices, etc.
[0211] The processor 10 may call a program stored in the memory 11. Specifically, the processor 10 may execute operations in the embodiment of the text-related information retrieval method.
[0212] The memory 11 is used to store one or more programs, which may include program codes, and the program codes include computer operation instructions. In the embodiment of the present application, the memory 11 at least stores programs for implementing the following functions:
[0213] receiving a retrieval request for target content;
[0214] Determining a first text information set according to the target content;
[0215] Performing feature matching on multiple features corresponding to the target content and multiple features corresponding to each text information in the first text information set, integrating penalty factors, and determining a matching weight between the target content and each text information in the first text information set;
[0216] Determine a second text information set according to the matching weight between the target content and each text information in the first text information set, wherein the second text information set is a complete set or a subset of the first text information set;
[0217] A second text information set is outputted in response to the retrieval request.
[0218] In a possible implementation, the memory 11 may include a program storage area and a data storage area, wherein the program storage area may store an operating system and an application required for at least one function, etc.; the data storage area may store data created during use.
[0219] In addition, the memory 11 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device or other volatile solid-state storage device.
[0220] The communication interface 12 may be an interface of a communication module, and is used to connect to other devices or systems.
[0221] Of course, it should be noted that Fig.10 The structure shown does not constitute a limitation on the electronic device in the embodiment of the present application. In actual applications, the electronic device may include Fig.10 More or fewer components than shown, or combinations of certain components.
[0222] Corresponding to the above method embodiment, the embodiment of the present application also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the above-mentioned text-related information retrieval method are implemented.
[0223] In addition, it should be noted that: the embodiment of the present application also provides a computer program product or a computer program, which may include a computer instruction, which may be stored in a computer-readable storage medium. The processor of the computer device reads the computer instruction from the computer-readable storage medium, and the processor may execute the computer instruction so that the computer device executes the description of the text-related information retrieval method in the corresponding embodiment of the foregoing text, and therefore, it will not be repeated here. In addition, the description of the beneficial effects of the same method will not be repeated. For technical details not disclosed in the computer program product or computer program embodiment involved in the present application, please refer to the description of the method embodiment of the present application.
[0224] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.
[0225] It should be noted that, in this article, the terms "comprise", "include" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprises one..." does not exclude the presence of other identical elements in the process, method, article or device including the element. In addition, it should be pointed out that the scope of the method and device in the embodiment of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in reverse order according to the functions involved, for example, the described method may be performed in an order different from that described, and various steps may also be added, omitted or combined. In addition, the features described with reference to certain examples may be combined in other examples.
[0226] Through the description of the above implementation methods, those skilled in the art can also clearly understand that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented with electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in the above description according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0227] The steps of the method or algorithm described in conjunction with the embodiments disclosed herein can be directly implemented using hardware, a software module executed by a processor, or a combination of the two. The software module can be placed in a random access memory (RAM), a memory, a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a register, a hard disk, a removable disk, a compact disc read-only memory (CD-ROM), or any other form of storage medium known in the art, including several instructions to execute the methods described in the various embodiments of the present application.
[0228] The embodiments of the present application are described above in conjunction with the accompanying drawings, and the description of the above embodiments is only used to help understand the technical solution and its core idea of the present application. It should be pointed out that the present application is not limited to the above-mentioned specific implementation methods, which are merely illustrative and not restrictive. For ordinary technicians in this field, many forms of implementation methods can be made without departing from the scope of protection of the purpose of the present application and the claims, and the present application can also be improved and modified in a number of ways, and these implementation methods, improvements and modifications are all within the scope of protection of the present application.
Claims
1. A text-related information retrieval method, characterized in that: include: receiving a retrieval request for target content; Determining a first text information set, wherein the first text information set includes text information of at least one information category; After performing feature matching on multiple features corresponding to the target content and multiple features corresponding to each text information in the first text information set, fusing the value of a predetermined penalty factor to determine a matching weight between the target content and each text information in the first text information set; Determine a second text information set according to a matching weight between the target content and each text information in the first text information set, wherein the second text information set is a complete set or a subset of the first text information set; The second text information set is outputted to respond to the search request.
2. The method according to claim 1, characterized in that The determining of the first text information set includes: Determining the category of information to be retrieved based on multiple features corresponding to the target content; A set consisting of text information in the information category to be retrieved is determined as a first text information set.
3. The method according to claim 1, characterized in that After performing feature matching on the multiple features corresponding to the target content and the multiple features corresponding to each text information in the first text information set, fusing the value of a predetermined penalty factor to determine the matching weight between the target content and each text information in the first text information set, including: Determine the target feature among the multiple features corresponding to the target content and the multiple features corresponding to each text information in the first text information set as an independent variable of the penalty factor; For each text information in the first text information set, determining a value of a penalty factor corresponding to the current text information according to a matching degree between the target feature corresponding to the target content and the target feature corresponding to the current text information; Based on the value of the penalty factor corresponding to the current text information and the matching results of the multiple features corresponding to the target content and the multiple features corresponding to the current text information, the matching weight of the target content and the current text information is determined.
4. The method according to claim 3, characterized in that The determining the matching weight between the target content and the current text information based on the value of the penalty factor corresponding to the current text information and the matching results between the multiple features corresponding to the target content and the multiple features corresponding to the current text information includes: Calculating a similarity measure between the target content and the current text information based on a matching result between multiple features corresponding to the target content and multiple features corresponding to the current text information; A matching weight between the target content and the current text information is determined according to the similarity measurement and the value of the penalty factor corresponding to the current text information.
5. The method according to claim 3, characterized in that: The higher the matching degree between the target feature corresponding to the target content and the target feature corresponding to the current text information is, the smaller the value of the penalty factor corresponding to the current text information is.
6. The method according to claim 1, characterized in that The determining the second text information set according to the matching weight between the target content and each text information in the first text information set includes: For each text information in the first text information set, if the matching weight between the target content and the current text information is greater than or equal to a first threshold, determining the current text information as the target text information; The determined set of target text information is determined as a second text information set.
7. The method according to any one of claims 1 to 6, characterized in that The outputting the second text information set includes: The second text information set is outputted in the order of the matching weights between the target content and each text information in the second text information set.
8. A text-related information retrieval device, characterized in that: include: A receiving module, used for receiving a search request for target content; A first determining module, configured to determine a first text information set according to the target content; a second determination module, configured to determine a matching weight between the target content and each text information in the first text information set by fusing a value of a predetermined penalty factor after feature matching is performed on the multiple features corresponding to the target content and the multiple features corresponding to each text information in the first text information set; A third determination module, configured to determine a second text information set according to a matching weight between the target content and each text information in the first text information set, wherein the second text information set is a complete set or a subset of the first text information set; An output module is used to output the second text information set in response to the search request.
9. An electronic device, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the steps of the text-related information retrieval method according to any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the text-related information retrieval method according to any one of claims 1 to 7 are implemented.