Artificial Intelligence-Based Document Search Method, Device, Equipment and Storage Medium
By constructing a document vector matrix and a document matching vector, the problem of large amount of stop word filtering calculation in the prior art is solved, and the efficiency of document search matching is improved.
Patent Information
- Application Number
- CN202111276318.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-29
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2041-10-29
AI Technical Summary
Existing document search methods have a lot of calculations when filtering stop words, resulting in inefficient matching.
Using an artificial intelligence-based method, by constructing a document vector matrix and a document matching vector, the document data is filtered using a preset stop word list to reduce the calculation amount and improve matching efficiency.
It effectively reduces the calculation amount of stop word filtering during document search and improves the matching efficiency of document search.
Smart Images

Figure CN114003712B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular, to a document search method, apparatus, electronic device, and storage medium based on artificial intelligence. Background Art
[0002] The general process for searching documents in a database is as follows: input a retrieval document, the search engine matches the document resources in the database according to the input document information, and finally presents the document resources with a relatively high similarity.
[0003] For the stop words that need to be filtered out in the document resources, they are often filtered by establishing a stop word list and matching it with the searched document resources one by one. However, there are often a large number of stop words in the document resources. Therefore, filtering stop words by the one-by-one matching method requires a large amount of computation, thus reducing the matching efficiency of document search. Summary of the Invention
[0004] In view of the above, it is necessary to propose a document search method, apparatus, electronic device, and storage medium based on artificial intelligence to solve the technical problem of how to improve the matching efficiency of document search.
[0005] This application provides a document search method based on artificial intelligence, including:
[0006] Obtain a first document data set;
[0007] Construct a document vector matrix based on the first document data set;
[0008] Obtain a second document data set based on the document vector matrix, where the second document data set includes stop words;
[0009] Match the stop words in the second document data set according to a preset stop word list to construct a document matching vector and obtain a document similarity value;
[0010] Filter the second document data set based on the document similarity value to obtain a filtered document.
[0011] In this way, by constructing a word matching matrix and a document matching vector, the stop words are converted into elements of the document matching vector to filter the searched document data for stop words to obtain a filtered document. This solution only needs to modify the stop words of the retrieved characters to achieve the filtering of stop words for all document resources, reducing the computation amount in the stop word filtering stage of document search and improving the matching efficiency of document search.
[0012] In some embodiments, the constructing a document vector matrix based on the first document data set includes:
[0013] Partition the first document dataset to obtain a first word segmentation dataset;
[0014] Match the first word segmentation dataset with a preset retrieval character to obtain a document representation vector;
[0015] Construct the document vector matrix based on the document representation vector.
[0016] In this way, by more accurately partitioning the first document dataset, a more accurate document representation vector can be obtained based on the first word segmentation dataset, and then the document vector matrix is constructed as the data source for subsequent steps, which is beneficial to further improving the accuracy of the document filtering result.
[0017] In some embodiments, obtaining the second document dataset based on the document vector matrix includes:
[0018] Calculate the document vector matrix to obtain a similarity item score value;
[0019] Calculate the similarity item score value according to a preset weight to obtain a first document score value;
[0020] Filter the first document dataset based on the first document score value to obtain the second document dataset.
[0021] In this way, a more accurate score value is obtained through preset similarity items to obtain a second document dataset with a higher matching degree, further improving the accuracy of the document filtering result.
[0022] In some embodiments, the similarity item score value includes a cosine distance similarity item score value, an Euclidean distance similarity item score value, a retrieval similarity item score value, and an edit distance similarity item score value, and the search method satisfies the relational expression:
[0023] S = 0.3 * S1 + 0.3 * S2 + 0.2 * S3 + 0.2 * S4
[0024] where S represents the first document score value, S1 represents the cosine distance similarity item score value, S2 represents the Euclidean distance similarity item score value, S3 represents the retrieval similarity item score value, and S4 represents the edit distance similarity item score value.
[0025] In this way, by selecting multiple similarity items for scoring and assigning similar weights to each similarity item, the first document score value can be comprehensively calculated from multiple dimensions, improving the accuracy of the first document score value. At the same time, the error of the first document score value will not be too large due to the inaccuracy of a single similarity item score value, enhancing the stability and reliability of the first document score value.
[0026] In some embodiments, matching stop words in the second document dataset according to a preset stop word list to construct a document matching vector and obtain a document similarity value includes:
[0027] Partition the second document dataset to obtain a second word segmentation result;
[0028] Match the second word segmentation result and a preset retrieval character to obtain a word matching matrix;
[0029] Update the word matching matrix according to a preset stop word list to obtain a word updated matrix;
[0030] Construct a document matching vector based on the word updated matrix;
[0031] Obtain the document similarity value based on the document matching vector.
[0032] In this way, partitioning the second document dataset can obtain a more accurate word segmentation result, effectively distinguish stop words, which is convenient for more accurately filtering stop words in subsequent processes to obtain a more accurate word matching matrix. By constructing a document matching vector, the rapid calculation of the document similarity value can be realized, without word-by-word matching of document resources, effectively improving the document matching efficiency.
[0033] In some embodiments, updating the word matching matrix according to a preset stop word list to obtain a word updated matrix includes:
[0034] Match a preset retrieval character and a preset stop word list to obtain retrieval character stop words;
[0035] Update the dimension value corresponding to the retrieval character stop words in the word matching matrix to obtain the word updated matrix.
[0036] In this way, only by modifying the stop words of the retrieval character can the stop words of all document resources be filtered, effectively reducing the calculation amount in the stop word filtering stage of the document search process and improving the matching efficiency of the document search.
[0037] In some embodiments, filtering the second document dataset based on the document similarity value to obtain a filtered document includes:
[0038] If the document similarity value is 0, filter the corresponding document;
[0039] If the document similarity value is not 0, retain the corresponding document, and use all the retained documents as the filtered document.
[0040] In this way, by using the numerical value 0 as an index for filtering the second document dataset, documents that only match the stop words in the preset retrieval characters in the retrieved documents can be quickly filtered out, and the retrieved documents can be sorted according to their relevance to the preset retrieval characters based on the scores, thereby improving the efficiency of user retrieval.
[0041] The embodiment of the present application further provides an artificial intelligence-based document search device, including:
[0042] An acquisition unit, configured to acquire a first document dataset;
[0043] A construction unit, configured to construct a document vector matrix based on the first document dataset;
[0044] A scoring unit, configured to obtain a second document dataset based on the document vector matrix, where the second document dataset includes stop words;
[0045] A filtering unit, configured to match the stop words in the second document dataset according to a preset stop word list to construct a document matching vector and obtain a document similarity value;
[0046] The filtering unit is further configured to filter the second document dataset based on the document similarity value to obtain a filtered document.
[0047] The embodiment of the present application further provides an electronic device, including:
[0048] A memory, storing at least one instruction;
[0049] A processor, configured to execute the instruction stored in the memory to implement the artificial intelligence-based document search method.
[0050] The embodiment of the present application further provides a computer-readable storage medium, where at least one instruction is stored in the computer-readable storage medium, and the at least one instruction is executed by a processor in an electronic device to implement the artificial intelligence-based document search method. Description of the Drawings
[0051] Figure 1 is a flowchart of a preferred embodiment of the artificial intelligence-based document search method involved in the present application.
[0052] Figure 2 is a flowchart of a preferred embodiment of constructing a document matching vector based on the second document dataset and a preset stop word list to obtain a document similarity value involved in the present application.
[0053] Figure 3 is a functional module diagram of a preferred embodiment of the artificial intelligence-based document search device involved in the present application.
[0054] Figure 4 It is a schematic structural diagram of an electronic device which is a preferred embodiment of the artificial intelligence-based document search method involved in the present application.
[0055] Figure 5 It is the TF-IDF score table of the document data involved in the present application.
[0056] Figure 6 It is the BM25 score table of the document data involved in the present application.
[0057] Figure 7 It is the 2-gram value table of the document data involved in the present application. Detailed implementation manners
[0058] In order to more clearly understand the purpose, features and advantages of the present application, the present application will be described in detail below with reference to the accompanying drawings and specific embodiments. It should be noted that, without conflict, the embodiments of the present application and the features in the embodiments can be combined with each other. Many specific details are set forth in the following description in order to fully understand the present application. The described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments.
[0059] In addition, the terms "first" and "second" are only used for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of the described features. In the description of the present application, "a plurality of" means two or more, unless otherwise specifically defined.
[0060] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present application belongs. The terms used in the specification of the present application herein are only for the purpose of describing specific embodiments, and are not intended to limit the present application. The term "and / or" used herein includes any and all combinations of one or more of the related listed items.
[0061] The embodiment of the present application provides an artificial intelligence-based document search method, which can be applied to one or more electronic devices. An electronic device is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes but is not limited to a microprocessor, an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a digital signal processor (DSP), an embedded device, etc.
[0062] An electronic device can be any kind of electronic product that can perform human-computer interaction with a user. For example, a personal computer, a tablet computer, a smart phone, a personal digital assistant (PDA), a game console, an Internet Protocol Television (IPTV), a smart wearable device, etc.
[0063] The electronic device may further include a network device and / or a user device. Among them, the network device includes, but is not limited to, a single network server, a server group composed of multiple network servers, or a cloud composed of a large number of hosts or network servers based on cloud computing.
[0064] The network where the electronic device is located includes, but is not limited to, the Internet, a wide area network, a metropolitan area network, a local area network, a virtual private network (VPN), etc.
[0065] As Figure 1 shown, it is a flowchart of a preferred embodiment of the document search method based on artificial intelligence in this application. According to different requirements, the order of steps in this flowchart can be changed, and some steps can be omitted.
[0066] S10. Obtain a first document dataset.
[0067] In an optional embodiment, obtaining the first document dataset includes:
[0068] S101. Establish a document database, and the database is an Elasticsearch database.
[0069] In this optional embodiment, since the sources of the obtained document data information are diverse and the structures are different, unified and standardized tags can be assigned according to the business system, type, structure, length, etc. of the data to identify the attributes such as the source and type of the data.
[0070] In this optional embodiment, Elasticsearch is a distributed, highly scalable, and highly real-time search and data analysis engine. The implementation principle of Elasticsearch is mainly divided into the following steps. First, the user submits data to the Elasticsearch database, and then the corresponding statements are tokenized through a tokenization controller, and the weights and tokenization results are stored in the data together. When the user searches for data, the results are sorted and scored according to the weights, and then the returned results are presented to the user.
[0071] S102. Obtain the first document dataset based on the document database and a preset search character.
[0072] In this optional embodiment, the preset retrieval character is the document information input by the user. The Elasticsearch search relevance ranking (ES) algorithm can be used to obtain the first document dataset. The ES algorithm adopts the word-based term frequency-inverse document frequency (TF-IDF) method. Elasticsearch first analyzes the document and then uses the results to build an inverted index to find the resources related to the retrieval character to participate in the recall calculation. The calculation process is as follows:
[0073] Construct an inverted index table: that is, count which documents contain certain words. Exemplarily, for three sentence sequences 1) "ABC", 2) "BCD", 3) "CDW", the documents for the dictionary ("A", "B", "C", "D", "W") can be constructed. Then there are A: [1], B: [1, 2], C: [1, 2, 3], D: [2, 3], W: [3].
[0074] Initial resource recall: that is, find the hit words among the resources constructed in the above steps. For example, for a user search of "AW", the user's retrieval character is first split by word, and then the document where A appears is found to be 1) the document, and the document where W appears is found to be 3) the document. Then the documents participating in the calculation will be documents 1) and 3). In actual recall, punctuation marks will be removed to avoid recalling all documents;
[0075] Recall data calculation: Through the above document calculation, the basic resource pool can be found. Two metrics are calculated based on the resource pool, 1) TF value, 2) IDF value. Among them, TF is the term frequency value, and IDF is the inverse document frequency value. The calculation formulas are as follows:
[0076] TF = the number of times a certain word appears in the document / the total number of words in the document;
[0077] IDF = log(total number of documents in the corpus / (number of documents containing the word + 1)).
[0078] The above formula shows that if a word appears more frequently, then the denominator is larger, and the inverse document frequency is closer to 0; the reason for adding 1 to the denominator is to avoid the denominator being 0 (that is, the word is not contained in all documents); log means taking the logarithm of the obtained value.
[0079] Then the TF-IDF calculation expression is: TF-IDF = TF * IDF.
[0080] According to the above expression, the TF-IDF scores of a single document in each word dimension can be calculated. All words are used as a dictionary, and the scores of a single document in each word dimension can be obtained.
[0081] Exemplarily, in this example as Example 1, assume that the documents in the document database are "I love Shenzhen", "I am in Qianhai Bay", and "He loves Shenzhen more", and the search character input by the user is "I am in Shenzhen". Then, according to the TF-IDF calculation method, a score table can be obtained, as shown in Figure 5 the TF-IDF score table of the document data shown.
[0082] In the above manner, the matching scores between the search character and each document in the document database can be obtained, and the documents with a matching score of 0 are filtered out, and the remaining documents are retained as the first document dataset.
[0083] In this way, by establishing a document database, the first document dataset can be obtained quickly and accurately. The reason is that the Elasticsearch database can automatically segment the preset search character based on the Elasticsearch analysis engine, assign different weights to each segment according to the relevance degree to sort the matching data, and finally present the best matching result to the user according to the sorting.
[0084] S11. Construct a document vector matrix based on the first document dataset.
[0085] In an optional embodiment, constructing a document vector matrix based on the first document dataset includes:
[0086] S111. Divide the first document dataset to obtain a first segmented dataset.
[0087] In this optional embodiment, a word segmentation tool can be used to divide the first document dataset. The word segmentation tool can use jieba. Since the first document dataset is obtained word by word in the form of character segmentation and cannot effectively match the semantic information of the document dataset, jieba is used to segment the first document dataset and then match it with the input document of the user to obtain a more accurate document matching result. Jieba has three word segmentation modes, namely the exact mode, the full mode, and the search engine mode. Among them, the exact mode attempts to cut the sentence most precisely and is suitable for document analysis. Therefore, the exact mode is selected in this solution to segment the first document dataset to obtain the first segmented dataset.
[0088] S112. Match the first segmented dataset and the preset search character to obtain a document representation vector.
[0089] In this alternative embodiment, the pluggable similarity algorithm (BM25) can be used to match the first tokenized dataset to obtain the document representation vector. The BM25 algorithm incorporates the adjustment of the resource sentence length and can further calculate the first tokenized dataset. BM25 adds several adjustable parameters on the basis of the traditional TF-IDF, making it more flexible and powerful in application and highly practical. The expression of this algorithm is:
[0090] BM25 = (idf * (k + 1) * tf) / (k * (1.0 - b + b * L) + tf)
[0091] Among them, the parameter b represents the degree of restricting the influence of the document length on the calculation result; L is the ratio of the current document length to the average document length; the constant k is used to limit the growth limit of the TF value. When TF increases, the TF score increases accordingly, but the TF score of BM25 will be restricted between 0 and k + 1. It can approach k + 1 infinitely, but can never reach it. This can be understood in business as the influence intensity of a certain factor cannot be infinite but has a maximum value, which also conforms to our understanding of the document relevance logic. In this solution, k = 1.2 is selected for use.
[0092] In this alternative embodiment, the document representation vector of each document in the first tokenized dataset is calculated through the BM25 algorithm. The dimension of the document representation vector is the total number of word occurrences, and the data is stored in a nested list manner List<List <double>, nested lists in a list, and the length of each sub - list is fixed, which is the length of the vocabulary.
[0093] Exemplarily, in Example 1, after all the documents are counted, there are 8 different words, so the length of the dictionary is 8, the dimension of the vector is 8, and the length of the sub - list is the fixed value 8. Through the above BM25 calculation formula, the BM25 value of each word in each document can be calculated, and then a data table is constructed, such as Figure 6 the BM25 score table of the document data shown.
[0094] Figure 6 Each value in it is the BM25 value corresponding to each word in each document, and each row corresponds to the vector representation of the document. For example, the document representation vector of the retrieved character is [0.3, 0, 0.4, 0.33, 0, 0, 0, 0]. Among them, since the BM25 value is obtained by matching the relevance between each word in the retrieved character and each document, different retrieved characters will result in different calculated BM25 values, and thus the document representation vectors of each document composed of BM25 values are also different.
[0095] S113, construct a document vector matrix based on the document representation vectors.
[0096] In this optional embodiment, a document vector matrix is constructed based on the document representation vectors corresponding to all the documents in the first document dataset.
[0097] In this way, by performing more accurate word segmentation on the first document dataset to obtain the first word - segmented dataset, and more accurately obtaining the document representation vectors by matching the first word - segmented dataset with the preset retrieved word segments to construct the document vector matrix as the data source for subsequent steps, it is beneficial to further improve the accuracy of the document filtering results.
[0098] S12, obtain a second document dataset based on the document vector matrix, and the second document dataset contains stop words.
[0099] In an optional embodiment, obtaining the second document dataset based on the document vector matrix includes:
[0100] S121, calculate the document vector matrix to obtain the similarity item score value.
[0101] In this optional embodiment, the similarity items are respectively the cosine distance similarity item score value, the Euclidean distance similarity item score value, the retrieval similarity item score value, and the edit distance similarity item score value between each document vector and a preset retrieval character vector. Among them, the cosine distance similarity item score value is the cosine similarity score value S1 between each document vector and the retrieval character vector; the Euclidean distance similarity item score value is the Euclidean distance score value S2 between each document vector and the retrieval character vector; the retrieval similarity item score value is the similarity item score value S3 obtained by querying and calculating based on the retrieval character in the elastic search database; the edit distance similarity item score value is the edit distance score value S4 between each document vector and the retrieval character vector.
[0102] Among them, S1, S2, S3, and S4 are respectively calculated between the preset retrieval character vector and each document vector. Exemplarily, let the retrieval character vector corresponding to the preset retrieval character be A T , and the two document vectors be B T and C T . Then, calculate the similarity item score values between the retrieval character vector A T and the document vector B T to obtain S1 = 0.6, S2 = 0.3, S3 = 0.8, S4 = 0.2. Similarly, the similarity item score values between the retrieval character vector A T and the document vector C T can be calculated to be S1 = 0.5, S2 = 0.2, S3 = 0.6, S4 = 0.3.
[0103] In this optional embodiment, the edit distance refers to the number of single-character edit operations required to change one string to another. Single-character operations include: insertion, deletion, and replacement. If the number of operations is c, then the edit distance similarity score value is:
[0104] S4 = (Max(lenth(doc1), length(doc2)) - c) / Max(lenth(doc1), length(doc2))
[0105] Since the number of operations may be greater than the length of the document, the final S4 may be negative, and the larger value between this value and 0 needs to be taken, that is, S = Max(0, S4).
[0106] Exemplarily, if the preset retrieval character is "I am in Shenzhen", then the edit distance between the retrieval character and document 1 is:
[0107] Changing "I am in Shenzhen" to "I love Shenzhen" requires one step of changing "am" to "love".
[0108] Therefore, S4 = (Max(4, 4) - 1) / Max(4, 4) = 3 / 4 = 0.75.
[0109] The edit distance between the retrieval character and Document 2 is:
[0110] Changing "I am in Shenzhen" to "I am in Qianhaiwan" requires three steps: changing "Shen" to "Qian", "Zhen" to "Hai", and then inserting "Wan".
[0111] Therefore, S4 = (Max(4, 5) - 3) / Max(4, 5) = 2 / 5 = 0.4.
[0112] S122. Calculate the similarity item score value according to the preset weight to obtain the first document score value.
[0113] In an optional embodiment, the preset weight of S1 is 0.3, the weight of S2 is 0.3, the weight of S3 is 0.2, and the weight of S4 is 0.2. The calculation formula for the first document score value S is:
[0114] S = 0.3 * S1 + 0.3 * S2 + 0.2 * S3 + 0.2 * S4
[0115] Through the above embodiments, the first document score value S of each document in the first document dataset can be obtained.
[0116] S123. Screen the first document dataset based on the first document score value to obtain the second document dataset.
[0117] In an optional embodiment, according to the first document score value, the S values of each document are sorted from high to low to sort the documents in the first document dataset. A screening threshold of 0.2 is set, and the documents with values less than the threshold in the sorting result are filtered out, and the remaining documents are retained to obtain the second document dataset.
[0118] Among them, since the above process has not filtered the stop words contained in the first document dataset, the second document dataset obtained still contains the stop words in each document.
[0119] In this way, by constructing multiple similarity items, a more accurate score value is obtained, and document data with a higher matching degree with the retrieval character is obtained from the first document dataset as the second document dataset, further improving the accuracy of the document filtering result.
[0120] S13. Match the stop words in the second document dataset according to the preset stop word list to construct a document matching vector and obtain a document similarity value.
[0121] Please refer to Figure 2 , in an optional embodiment, constructing a document matching vector based on the second document dataset and a preset stop word list to obtain a document similarity value includes:
[0122] S131, partitioning the second document dataset to obtain a second word segmentation result.
[0123] In an optional embodiment, the binary grammar (2-gram) algorithm can be used to split the preset retrieval characters and the documents in the second document dataset, and the result of splitting the documents in the second document dataset by the 2-gram algorithm is used as the second word segmentation result. Among them, 2-gram word segmentation is a maximum probability word segmentation. When calculating the probability of a word, it not only considers itself but also its predecessor.
[0124] Exemplarily, for example, if a sentence is segmented into [I, am, Aries, of, man] based on the Jieba word segmentation tool, based on the 2-gram method, words with more than 2 characters can be further split into 2-character words, and we will get: [I, am, Aries, Ari, ries, of, man, man, an].
[0125] S132, matching the second word segmentation result and the preset retrieval characters to obtain a word matching matrix.
[0126] In an optional embodiment, the Elasticsearch search and data analysis engine are used to match each word after 2-gram word segmentation with each document in the second document dataset one by one. If the word exists in the word segmentation set of the document, the corresponding value of the word in the word matching matrix is 1, otherwise it is 0.
[0127] Exemplarily, let the retrieval characters be: "Am I a man of Aries?", and after segmentation, it is [I, am, Aries, of, man,?]. Based on the 2-gram method, the words are split into 2-character words, and we will get: [I, am, Aries, Ari, ries, of, man, man, an,?], which is a set obtained by splitting words with more than 2 characters. For document A: [I, of, constellation, am, Ari], for document B: [Ari, no, black, good-looking,?]. Among them, the dimension of the matrix is a set of all word dimensions, and the size of the dimension is the total number of all words in the set. The dictionary obtained by splitting the retrieval characters according to 2-gram is as Figure 7 shown in the 2-gram numerical table of the document data, and the corresponding dimension values are also filled into Figure 7 and finally obtain the word matching matrix as shown in Figure 7 .
[0128] S133, updating the word matching matrix according to the preset stop word list to obtain a word updated matrix.
[0129] In an optional embodiment, the preset stop word list is constructed by adopting a basic word list plus manual collation, and includes: punctuation marks, interrogative words, and connecting words. Among them, the punctuation marks are common symbols in documents such as commas, periods, and pauses; the interrogative words are words of the same kind such as "why", "how", "ne", "ma", etc.; the connecting words indicate the semantic connection. For example, words such as "then", "furthermore", "immediately afterwards" are of this kind, and words indicating a turn, such as "but", "again", "suddenly", "however" and other words of the same kind. The collation of this kind of words adopts the method of manual collation. The collation method is as follows:
[0130] Perform word segmentation on all existing data, mark the part-of-speech after word segmentation, remove some basic words with part-of-speech, such as nouns, verbs, adjectives, etc., and then collect the remaining words together, and manually mark whether they need to be sorted together as stop words.
[0131] The stop words are finally sorted into a document format (txt) document and placed together with other basic data. For example, they are placed on a disk or on a Network Attached Storage (NAS) disk. When the system starts, the data is loaded from the disk into the memory, and then stays resident in the memory for processing in the memory.
[0132] In an optional embodiment, the process of updating the word matching matrix according to the preset stop word list to obtain the word updated matrix is: set the stop word dimension in the word matching matrix to 0 according to the preset stop word list. Exemplarily, Figure 7 The search character "ma" in the middle matches the "ma" in Document B, but in the filtering stage here, the dimension value of the stop word "ma" will be changed to 0. Here, changing it to 0 only needs to modify the search character. Since there is only the computational amount of one source document, the computational amount is very small.
[0133] S134, construct a document matching vector based on the word updated matrix.
[0134] In an optional embodiment, the Jaccard metric algorithm can be used to calculate the similarity between the word updated matrix and the preset search character. The Jaccard metric algorithm is jac = (A ∩ B) / (A ∪ B), where A is the set of words of the preset search character, and B is the set of words of a certain document. If the result of A intersection B is 0, then jac is 0.
[0135] Exemplarily, for example, if the preset search character is "Am I a man of Aries?", after splitting the words based on the 2-gram method, we will get: [I, am, Aries, white ram, ram constellation, of, man, male, son of sweat]. Then for the document [I, of, constellation, am, white ram], when calculating A intersection B, the word "white ram" will be hit, and the jac result is not 0.
[0136] In an alternative embodiment, after updating the values corresponding to the stop words, the document matching vectors corresponding to each document can be obtained based on the word update matrix. Since the values of the document matching vectors are only 1 and 0, the Jaccard similarity between two documents can be obtained by directly multiplying the two document matching vectors.
[0137] Exemplarily, Figure 7 the document vector of the retrieved characters in [1, 1, 1, 1, 1, 1, 1, 1, 1, 0, 0, 0, 0, 0, 0], and the retrieved characters of document A are [1, 1, 0, 1, 0, 1, 0, 0, 0, 0, 1, 0, 0, 0, 0].
[0138] S135, obtaining a document similarity value based on the document matching vector.
[0139] Exemplarily, the result after multiplying the document matching vector of the retrieved characters and the document matching vector of document A is:
[0140] [1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 0, 0, 0, 0, 0, 0] * [1, 1, 0, 1, 0, 1, 0, 0, 0, 0, 1, 0, 0, 0, 0] = 4
[0141] The result after multiplying the document matching vector of the retrieved characters and the document matching vector of document B is:
[0142] [1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 0, 0, 0, 0, 0, 0] * [0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 0, 1, 1, 1, 1] = 0
[0143] The result obtained after multiplying the above vectors is the document similarity value.
[0144] S14, filtering the second document dataset based on the document similarity value to obtain a filtered document.
[0145] In an alternative embodiment, the Jaccard similarity results between the retrieved character vector and each document vector in the second document dataset are calculated respectively. The documents with a result of 0 are filtered out, and the remaining documents are retained to obtain a document filtering result. The documents in the document filtering result are sorted from high to low according to the Jaccard similarity results and then displayed to the user.
[0146] In an alternative embodiment, considering that the user may directly search for stop words. For example, when the user searches for "lalalala", to avoid the user's search results returning no data, it is first determined whether the number of stop words in the retrieved characters is equal to the total number of words in the retrieved characters before filtering, so as to determine whether all the words in the retrieved characters are stop words. If all are stop words, no filtering calculation is required.
[0147] In this way, by constructing a word matching matrix and a document matching vector to filter the retrieved document data for stop words to obtain filtered documents, it is possible to achieve the purpose of filtering the stop words of all document resources only by modifying the stop words of the retrieved characters, effectively reducing the computational complexity in the stop word filtering stage of document search and improving the matching efficiency of document search.
[0148] It can be seen from the above technical solutions that this application can construct a word matching matrix and a document matching vector, convert the stop words into elements of the document matching vector to filter the retrieved document data for stop words to obtain filtered documents, and only by modifying the stop words of the retrieved characters can the stop words of all document resources be filtered, effectively reducing the computational complexity in the stop word filtering stage of document search and improving the matching efficiency of document search.
[0149] Please refer to Figure 3 , Figure 3 FIG. is a functional module diagram of a preferred embodiment of a document search device based on artificial intelligence provided by an embodiment of this application. The document search device 11 based on artificial intelligence includes an acquisition unit 110, a construction unit 111, a scoring unit 112, and a filtering unit 113. The modules / units referred to in this application refer to a series of computer program segments that can be executed by a processor 13 and can complete fixed functions, and are stored in a memory 12. In this embodiment, the functions of each module / unit will be described in detail in subsequent embodiments.
[0150] In an alternative embodiment, the acquisition unit 110 is configured to acquire a first document data set.
[0151] In this alternative embodiment, a document database is first established. The database is an Elasticsearch database. Elasticsearch is a distributed, highly scalable, and highly real-time search and data analysis engine. The implementation principle of Elasticsearch mainly includes the following steps. First, the user submits data to the Elasticsearch database, and then the corresponding statements are tokenized by a tokenization controller, and the weights and tokenization results are stored in the data together. When the user searches for data, the results are sorted and scored according to the weights, and then the returned results are presented to the user.
[0152] In this optional embodiment, the obtaining unit 110 obtains the first document data set based on the document database and a preset retrieval character, where the preset retrieval character is the document information input by the user. The Elasticsearch search relevance ranking (ES) algorithm can be used to obtain the first document data set. The ES algorithm adopts the word-based term frequency-inverse document frequency (TF-IDF) method. Elasticsearch first analyzes the document, then uses the results to establish an inverted index, and finds the resources related to the retrieval character to participate in the recall calculation.
[0153] In an optional embodiment, the constructing unit 111 is configured to construct a document vector matrix based on the first document data set.
[0154] Divide the first document data set to obtain the first word segmentation data set;
[0155] Match the first word segmentation data set with the preset retrieval character to obtain a document representation vector;
[0156] Construct a document vector matrix based on the document representation vector.
[0157] In this optional embodiment, a word segmentation tool can be used to divide the first document data set. The word segmentation tool can use Jieba. Since the first document data set is obtained word by word in the form of character segmentation and cannot effectively match the semantic information of the document data set, Jieba is used to perform word segmentation on the first document data set and then match it with the user's input document, so as to obtain a more accurate document matching result. Jieba has three word segmentation modes, namely the accurate mode, the full mode, and the search engine mode. Among them, the accurate mode attempts to cut the sentence most accurately and is suitable for document analysis. Therefore, the accurate mode is selected in this solution to perform word segmentation on the first document data set to obtain the first word segmentation data set.
[0158] In this optional embodiment, the pluggable similarity algorithm (BM25) can be used to match the first word segmentation data set to obtain a document representation vector. The BM25 algorithm adds the harmonic of the resource sentence length and can further calculate the first word segmentation data set. BM25 adds several adjustable parameters on the basis of the traditional TF-IDF, making it more flexible and powerful in application and having high practicability. The expression of this algorithm is:
[0159] BM25 = (idf * (k + 1) * tf) / (k * (1.0 - b + b * L) + tf)
[0160] Among them, the parameter b represents the degree of influence of the restricted document length on the calculation result; L is the ratio of the current document length to the average document length; the constant k is used to limit the growth limit of the TF value. When TF increases, the TF score increases accordingly, but the TF score of BM25 will be restricted between 0 and k + 1. It can approach k + 1 infinitely, but it can never reach it. This can be understood in business as the influence intensity of a certain factor cannot be infinite, but has a maximum value, which also conforms to our understanding of the document relevance logic. In this solution, k = 1.2 is selected for use.
[0161] In this optional embodiment, the document representation vector of each document in the first tokenized dataset is calculated through the BM25 algorithm. The dimension of the document representation vector is the total number of word occurrences, and the data is stored in a nested list manner List<List <double>, a list nested within a list, where the length of each sub - list is fixed and equal to the length of the vocabulary.
[0162] In this optional embodiment, a document vector matrix is constructed based on the document representation vectors corresponding to all the documents in the first document dataset.
[0163] In an optional embodiment, the scoring unit 112 is used to obtain the second document dataset based on the document vector matrix, specifically including:
[0164] Calculate the document vector matrix to obtain the similarity item scoring values;
[0165] Calculate the first document scoring value by calculating the similarity item scoring values according to the preset weights;
[0166] Filter the first document dataset based on the first document scoring value to obtain the second document dataset.
[0167] In this optional embodiment, the similarity items are respectively the cosine - distance similarity item scoring values, Euclidean - distance similarity item scoring values, retrieval similarity item scoring values, and edit - distance similarity item scoring values between each document vector and a preset retrieval character vector. Among them, the cosine - distance similarity item scoring value is the cosine similarity scoring value S1 between each document vector and the retrieval character vector; the Euclidean - distance similarity item scoring value is the Euclidean - distance scoring value S2 between each document vector and the retrieval character vector; the retrieval similarity item scoring value is obtained by querying and calculating in the elastic search database based on the retrieval character, getting the similarity item scoring value S3; the edit - distance similarity item scoring value is the edit - distance scoring value S4 between each document vector and the retrieval character.
[0168] Among them, S1, S2, S3, and S4 between the preset retrieval character vector and each document vector are calculated respectively. Exemplarily, let the retrieval character vector corresponding to the preset retrieval character be A T , and the two document vectors be B T and C T , then calculate the similarity item scoring values between the retrieval character vector A T and the document vector B T to get S1 = 0.6, S2 = 0.3, S3 = 0.8, S4 = 0.2. Similarly, the similarity item scoring values between the retrieval character vector A T and the document vector C T can be calculated to be S1 = 0.5, S2 = 0.2, S3 = 0.6, S4 = 0.3.
[0169] In this optional embodiment, the edit distance refers to the number of single-character edit operations required to change one string to another. Single-character operations include: insertion, deletion, and replacement. If the number of operations is c, the edit distance similarity score is:
[0170] S4 = (Max(length(doc1), length(doc2)) - c) / Max(length(doc1), length(doc2))
[0171] Since the number of operations may be greater than the length of the document, the final S4 may be negative, and the larger value between this value and 0 needs to be taken, that is, S = Max(0, S4).
[0172] Exemplarily, if the preset search string is "I am in Shenzhen", the edit distance between the search string and Document 1 is:
[0173] Changing "I am in Shenzhen" to "I love Shenzhen" requires one step of changing "am" to "love",
[0174] So S4 = (Max(4, 4) - 1) / Max(4, 4) = 3 / 4 = 0.75.
[0175] The edit distance between the search string and Document 2 is:
[0176] Changing "I am in Shenzhen" to "I am in Qianhaiwan" requires 3 steps: changing "Shen" to "Qian", "zhen" to "hai", and then inserting "wan",
[0177] So S4 = (Max(4, 5) - 3) / Max(4, 5) = 2 / 5 = 0.4.
[0178] Among them, the order relationship between the search string and each document during calculation does not affect the corresponding similarity scores, because S1, S2, S3, and S4 are all calculated by taking the whole of the search string and each document.
[0179] Exemplarily, if the preset search string is "I am in Shenzhen", the result of calculating the edit distance between the search string and Document 1 and the result of calculating the edit distance between Document 1 and the search string are the same, because there is only one character difference between "I am in Shenzhen" and "I love Shenzhen", so the calculation formula is the same: S4 = (Max(4, 4) - 1) / Max(4, 4) = 3 / 4 = 0.75.
[0180] In an optional embodiment, the preset weight of S1 is 0.3, the weight of S2 is 0.3, the weight of S3 is 0.2, and the weight of S4 is 0.2. The calculation formula for the weighted sum result S is:
[0181] S = 0.3 * S1 + 0.3 * S2 + 0.2 * S3 + 0.2 * S4
[0182] In an optional embodiment, according to the first document score value, the S values of each document are sorted in descending order for the documents in the first document dataset, a screening threshold is set to 0.2, the documents with values less than the threshold in the sorting result are filtered out, and the remaining documents are retained to obtain the second document dataset.
[0183] In an optional embodiment, the filtering unit 113 is used to match the stop words in the second document dataset according to a preset stop word list to construct a document matching vector and obtain a document similarity value, which specifically includes:
[0184] Partition the second document dataset to obtain a second word segmentation result;
[0185] Match the second word segmentation result with a preset retrieval character to obtain a word matching matrix;
[0186] Update the word matching matrix according to the preset stop word list to obtain a word update matrix;
[0187] Construct a document matching vector based on the word update matrix;
[0188] Obtain a document similarity value based on the document matching vector.
[0189] In an optional embodiment, the binary grammar (2-gram) algorithm can be used to split a preset retrieval character and the documents in the second document dataset, and the result of splitting the documents in the second document dataset by the 2-gram algorithm is used as the second word segmentation result. Among them, 2-gram word segmentation is a maximum probability word segmentation. When calculating the probability of a word, it not only considers itself but also its predecessor.
[0190] Exemplarily, for example, if a sentence is segmented by the Jieba word segmentation tool as [I, am, an, Aries, man], based on the 2-gram method, words with more than 2 characters can be further split into 2-character words, and the result will be: [I, am, an, Aries, Ari, ries, man, man, an].
[0191] In an optional embodiment, the Elasticsearch search and data analysis engine is used to match each word / character after 2-gram word segmentation with each document in the second document dataset one by one. If the word exists in the word segmentation set of the document, the corresponding value of the word in the word matching matrix is 1, otherwise it is 0.
[0192] Exemplarily, let the retrieval character be: "Am I a man of Aries?", after word segmentation, it becomes [I, am, Aries, of, man, am], and based on the 2-gram method, that is, splitting the words into 2-character words, we will get: [I, am, Aries, white sheep, sheep constellation, of, man, man, son sweat, am], here is a set obtained after splitting words with more than 2 characters. For document A: [I, of, constellation, am, white sheep], for document B: [white, not, black, good-looking, am], the dimension of the matrix is a set of the dimensions of all words, and the size of the dimension is the total number of all words in the set. The dictionary obtained after splitting the retrieval character according to 2-gram is as Figure 7 the 2-gram numerical table of the document data shown, and the values of the corresponding dimensions are also filled into Figure 7 it.
[0193] In an optional embodiment, the preset stop word list is constructed by adopting the basic word list plus manual collation, and includes: punctuation marks, interrogative words, and connecting words. Among them, the punctuation marks are common symbols in documents such as commas, periods, and semicolons; interrogative words such as why, how, ne, am, etc.; connecting words indicate semantic connection, such as words like "then", "furthermore", "immediately following", and words indicating turning, such as words like "but", "again", "suddenly", "however", etc. This type of words is sorted manually. The sorting method is as follows:
[0194] Perform word segmentation on all existing data, mark the part-of-speech after word segmentation, remove some basic words with part-of-speech, such as nouns, verbs, adjectives, etc., and then collect the remaining words together and manually mark whether they need to be sorted together as stop words.
[0195] The stop words are finally sorted into a document format (txt) document and placed together with other basic data, such as on a disk or on a Network Attached Storage (NAS) disk. When the system starts, the data is loaded from the disk into the memory and then stays in the memory for processing.
[0196] In an optional embodiment, the process of updating the word matching matrix according to the preset stop word list to obtain the word update matrix is: set the dimension of the stop words in the word matching matrix to 0 according to the preset stop word list. Exemplarily, Figure 7 in the retrieval character "am" and the "am" in document B are matched, but in the filtering stage here, the dimension value of the stop word "am" will be changed to 0. Here, changing it to 0 only needs to modify the retrieval character. Since there is only the computational amount of one source document, the computational amount is very small.
[0197] In an alternative embodiment, the Jaccard metric algorithm can be used to calculate the similarity between the word update matrix and the preset retrieval characters. The calculation expression of the Jaccard metric algorithm is jac = (A ∩ B) / (A ∪ B), where A is the set of words in the retrieved document, and B is the set of words in a single retrieved resource. If the result of A intersection B is 0, then jac is 0.
[0198] In an alternative embodiment, after updating the values corresponding to the stop words, based on the word update matrix, the document matching vectors corresponding to each document can be obtained. Since the values of the document matching vectors are only 1 and 0, the Jaccard similarity between two documents can be obtained by directly multiplying the two document matching vectors.
[0199] The filtering unit 113 is further configured to filter the second document dataset based on the document similarity value to obtain filtered documents.
[0200] In an alternative embodiment, the Jaccard similarity results between the document matching vector of the retrieval characters and the document matching vectors of each document in the second document dataset are calculated respectively. The documents with a result of 0 are filtered out, and the remaining documents are retained to obtain the document filtering result. The documents in the document filtering result are sorted from high to low according to the Jaccard similarity results and then displayed to the user.
[0201] In an alternative embodiment, considering that the user may directly search for stop words. For example, when the user searches for "lalalala", to avoid the user's search results returning no data, before filtering, it is first determined whether the number of stop words in the retrieval characters is equal to the total number of words in the retrieval characters, so as to determine whether all the words in the retrieval characters are stop words. If all are stop words, there is no need to perform filtering calculations.
[0202] From the above technical solutions, it can be seen that this application can filter the stop words in the searched document data by constructing a word matching matrix and a document matching vector, converting the stop words into elements of the document matching vector to obtain filtered documents. Only by modifying the stop words of the retrieval characters can the stop words of all document resources be filtered, reducing the computational complexity in the stop word filtering stage of the document search process and improving the matching efficiency of the document search.
[0203] Please refer to Figure 4 , Figure 4 is a schematic structural diagram of an electronic device provided by an embodiment of this application. The electronic device 1 includes a memory 12 and a processor 13. The memory 12 is used to store computer-readable instructions, and the processor 13 is used to execute the computer-readable instructions stored in the memory to implement the artificial intelligence-based document search method of any of the above embodiments.
[0204] In an alternative embodiment, the electronic device 1 further includes a bus and a computer program stored in the memory 12 and executable on the processor 13, such as an artificial intelligence-based document search program.
[0205] Figure 4 Only the electronic device 1 with the memory 12 and the processor 13 is shown. Those skilled in the art can understand that Figure 4 the shown structure does not constitute a limitation on the electronic device 1, and it may include fewer or more components than shown, or combine certain components, or have different component arrangements.
[0206] In combination with Figure 1 , the memory 12 in the electronic device 1 stores a plurality of computer-readable instructions to implement an artificial intelligence-based document search method, and the processor 13 can execute the plurality of instructions to implement:
[0207] Obtain a first document data set;
[0208] Construct a document vector matrix based on the first document data set;
[0209] Obtain a second document data set based on the document vector matrix, and the second document data set contains stop words;
[0210] Match the stop words in the second document data set according to a preset stop word list to construct a document matching vector and obtain a document similarity value;
[0211] Filter the second document data set based on the document similarity value to obtain a filtered document.
[0212] Specifically, for the specific implementation method of the processor 13 for the above instructions, reference may be made to Figure 1 the description of the relevant steps in the corresponding embodiment, which will not be elaborated here.
[0213] Those skilled in the art can understand that the schematic diagram is only an example of the electronic device 1 and does not constitute a limitation on the electronic device 1. The electronic device 1 can be either a bus structure or a star structure. The electronic device 1 may also include more or fewer other hardware or software than shown, or different component arrangements. For example, the electronic device 1 may also include input / output devices, network access devices, etc.
[0214] It should be noted that the electronic device 1 is only an example. Other existing or future electronic products that can be adapted to this application should also be included in the protection scope of this application and are hereby incorporated by reference.
[0215] Among them, the memory 12 includes at least one type of readable storage medium, which can be non-volatile or volatile. The readable storage medium includes flash memory, mobile hard disks, multimedia cards, card-type memories (such as SD or DX memories, etc.), magnetic memories, magnetic disks, optical disks, etc. The memory 12 can be an internal storage unit of the electronic device 1 in some embodiments, such as the mobile hard disk of the electronic device 1. The memory 12 can also be an external storage device of the electronic device 1 in other embodiments, such as a plug-in mobile hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the electronic device 1. Further, the memory 12 can also include both the internal storage unit and the external storage device of the electronic device 1. The memory 12 can be used not only to store application software installed in the electronic device 1 and various types of data, such as the code of an artificial intelligence-based document search program, etc., but also to temporarily store data that has been output or will be output.
[0216] In some embodiments, the processor 13 can be composed of integrated circuits. For example, it can be composed of a single packaged integrated circuit, or can be composed of multiple integrated circuits with the same or different functions, including the combination of one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips, etc. The processor 13 is the control core (Control Unit) of the electronic device 1, connecting various components of the entire electronic device 1 through various interfaces and lines, and by running or executing programs or modules stored in the memory 12 (such as executing an artificial intelligence-based document search program, etc.), and calling data stored in the memory 12, to perform various functions of the electronic device 1 and process data.
[0217] The processor 13 executes the operating system of the electronic device 1 and various installed application programs. The processor 13 executes the application programs to implement the steps in the above-mentioned various embodiments of the artificial intelligence-based document search method, such as Figure 1 the steps shown.
[0218] Exemplarily, the computer program may be divided into one or more modules / units, and the one or more modules / units are stored in the memory 12 and executed by the processor 13 to complete the present application. The one or more modules / units may be a series of computer-readable instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer program in the electronic device 1. For example, the computer program may be divided into an acquisition unit 110, a construction unit 111, a scoring unit 112, and a filtering unit 113.
[0219] The integrated units implemented in the form of software function modules may be stored in a computer-readable storage medium. The above-mentioned software function modules stored in a storage medium include several instructions for causing a computer device (which may be a personal computer, a computer device, or a network device, etc.) or a processor to execute parts of the method for artificial intelligence-based document search described in various embodiments of the present application.
[0220] If the modules / units integrated in the electronic device 1 are implemented in the form of software function units and sold or used as independent products, they may be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above-mentioned embodiment methods of the present application, it may also be completed by a computer program instructing relevant hardware devices. The computer program may be stored in a computer-readable storage medium, and when the computer program is executed by a processor, the steps of the above-mentioned various method embodiments may be implemented.
[0221] Among them, the computer program includes computer program code, and the computer program code may be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a mobile hard disk, a magnetic disk, an optical disc, a computer memory, a read-only memory (ROM, Read-Only Memory), a random access memory, and other memories, etc.
[0222] Furthermore, the computer-readable storage medium mainly includes a program storage area and a data storage area. Among them, the program storage area may store an operating system, application programs required for at least one function, etc.; the data storage area may store data created according to the use of the blockchain node, etc.
[0223] The blockchain referred to in this application is a new application mode of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, and encryption algorithms. Blockchain, in essence, is a decentralized database, a series of data blocks generated by using cryptographic methods. Each data block contains information about a batch of network transactions, which is used to verify the validity of the information and generate the next block. The blockchain can include the blockchain underlying platform, the platform product service layer, and the application service layer, etc.
[0224] The bus can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, in Figure 4 it is only represented by one arrow, but it does not mean that there is only one bus or one type of bus. The bus is arranged to realize the connection and communication between the memory 12 and at least one processor 13, etc.
[0225] Although not shown, the electronic device 1 may further include a power source (such as a battery) for powering each component. Preferably, the power source can be logically connected to the at least one processor 13 through a power management device, so as to realize functions such as charge management, discharge management, and power consumption management through the power management device. The power source may further include any components such as one or more DC or AC power sources, a recharge device, a power failure detection circuit, a power converter or inverter, a power status indicator, etc. The electronic device 1 may further include various sensors, a Bluetooth module, a Wi-Fi module, etc., which will not be elaborated here.
[0226] Furthermore, the electronic device 1 may further include a network interface. Optionally, the network interface may include a wired interface and / or a wireless interface (such as a WI-FI interface, a Bluetooth interface, etc.), which is usually used to establish a communication connection between the electronic device 1 and other electronic devices.
[0227] Optionally, the electronic device 1 may further include a user interface, which may be a display, an input unit (such as a keyboard), and optionally, the user interface may also be a standard wired interface or a wireless interface. Optionally, in some embodiments, the display may be an LED display, a liquid crystal display, a touch liquid crystal display, and an OLED (Organic Light-Emitting Diode) toucher, etc. Among them, the display may also be appropriately referred to as a display screen or a display unit, which is used to display the information processed in the electronic device 1 and to display a visual user interface.
[0228] The embodiment of the present application also provides a computer-readable storage medium (not shown in the figure), in which computer-readable instructions are stored. The computer-readable instructions are executed by a processor in the electronic device to implement the document search method based on artificial intelligence described in any of the above embodiments.
[0229] It should be understood that the above embodiments are only for illustration purposes and are not limited by this structure in the scope of the application.
[0230] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of the modules is only a logical function division, and there may be other division methods in actual implementation.
[0231] The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0232] In addition, in each embodiment of the present application, the functional modules can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a hardware plus a software functional module.
[0233] In addition, obviously, the term "including" does not exclude other units or steps, and the singular does not exclude the plural. The multiple units or devices described in the specification can also be implemented by one unit or device through software or hardware. The terms such as "first" and "second" are used to represent names and do not represent any specific order.
[0234] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application and not to limit them. Although the present application has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present application can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present application.< / double> < / double>
Claims
1. A document search method based on artificial intelligence, characterized in that Including: Obtain a first document data set; Construct a document vector matrix based on the first document data set; Obtain a second document data set based on the document vector matrix, where the second document data set contains stop words; Match the stop words in the second document data set according to a preset stop word list to construct a document matching vector and obtain a document similarity value, including: dividing the second document data set based on a bigram algorithm to obtain a second word segmentation result; matching the second word segmentation result and a preset retrieval character to obtain a word matching matrix; updating the word matching matrix according to the preset stop word list to obtain a word updated matrix; constructing a document matching vector based on the word updated matrix; obtaining the document similarity value based on the document matching vector; the step of updating the word matching matrix according to the preset stop word list to obtain a word updated matrix includes: matching the preset retrieval character and the preset stop word list to obtain a retrieval character stop word; updating the dimension value corresponding to the retrieval character stop word in the word matching matrix to obtain the word updated matrix; setting the stop word dimension in the word matching matrix to 0 according to the preset stop word list; Filter the second document data set based on the document similarity value to obtain a filtered document, including: if the document similarity value is 0, filter the corresponding document; if the document similarity value is not 0, retain the corresponding document, and use all the retained documents as the filtered document.
2. The method for document search based on artificial intelligence according to claim 1, wherein The step of constructing a document vector matrix based on the first document data set includes: Divide the first document data set to obtain a first word segmentation data set; Match the first word segmentation data set and a preset retrieval character to obtain a document representation vector; Construct the document vector matrix based on the document representation vector.
3. The method for document search based on artificial intelligence according to claim 1, wherein The step of obtaining a second document data set based on the document vector matrix includes: Calculate the document vector matrix to obtain a similarity item score value; Calculate the similarity item score value according to a preset weight to obtain a first document score value; Screen the first document data set based on the first document score value to obtain the second document data set.
4. The method for document search based on artificial intelligence according to claim 3, wherein The similarity item score value includes a cosine distance similarity item score value, an Euclidean distance similarity item score value, a retrieval similarity item score value, and an edit distance similarity item score value, and the search method satisfies the relationship: S = 0.3 * S1 + 0.3 * S2 + 0.2 * S3 + 0.2 * S4 where S represents the first document score value, S1 represents the cosine distance similarity item score value, S2 represents the Euclidean distance similarity item score value, S3 represents the retrieval similarity item score value, and S4 represents the edit distance similarity item score value.
5. An artificial intelligence-based document search device, characterized in that, Including: An obtaining unit, configured to obtain a first document data set; A constructing unit, configured to construct a document vector matrix based on the first document data set; A scoring unit, configured to obtain a second document data set based on the document vector matrix, where the second document data set contains stop words; A filtering unit, configured to match stop words in the second document dataset according to a preset stop word list to construct a document matching vector and obtain a document similarity value, including: dividing the second document dataset based on a bigram algorithm to obtain a second word segmentation result; matching the second word segmentation result and a preset retrieval character to obtain a word matching matrix; updating the word matching matrix according to the preset stop word list to obtain a word updated matrix; constructing a document matching vector based on the word updated matrix; obtaining the document similarity value based on the document matching vector; the updating the word matching matrix according to the preset stop word list to obtain a word updated matrix includes: matching the preset retrieval character and the preset stop word list to obtain a retrieval character stop word; updating the dimension value corresponding to the retrieval character stop word in the word matching matrix to obtain the word updated matrix; setting the stop word dimension in the word matching matrix to 0 according to the preset stop word list; The filtering unit is further configured to filter the second document dataset based on the document similarity value to obtain a filtered document, including: if the document similarity value is 0, filtering the corresponding document; if the document similarity value is not 0, retaining the corresponding document, and taking all the retained documents as the filtered document.
6. An electronic device, characterized in that, Including: A memory for storing computer-readable instructions; And A processor, configured to execute the computer-readable instructions stored in the memory to implement the artificial intelligence-based document search method according to any one of claims 1 to 4.
7. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-readable instructions, and the computer-readable instructions are executed by a processor in an electronic device to implement the artificial intelligence-based document search method according to any one of claims 1 to 4.
Citation Information
Patent Citations
Speech-based retrieval method, server, and computer-readable storage medium
CN109522392A
A question and answer mode intelligent retrieval system and method for online education
CN109885672A