An article retrieval method, program product, device, and storage medium

By obtaining the word frequency vector after dimensionality reduction from the article library and calculating the similarity, the problems of high memory usage and inability to sort in traditional full-text search technology are solved, and efficient article retrieval and sorting are achieved.

CN119621940BActive Publication Date: 2025-06-27BEIJING STARSHINE DIGITAL SYST CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510131542.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-06
Publication Date
2025-06-27
Estimated Expiration
2045-02-06

AI Technical Summary

Technical Problem

When the article has a large vocabulary, the index volume is huge, resulting in high memory usage and cannot meet the user needs such as sorting based on word frequency.

Method used

By obtaining the first target word frequency vector of each article from the preset article library, the vector is obtained from the initial word frequency vector after dimensionality reduction, and the second target word frequency vector is determined based on the number of occurrences of each word segment in the search content in the word segmentation dictionary. Then, the matching target article is determined from the article library based on the similarity between the two.

Benefits of technology

It reduces memory usage, improves user retrieval experience, and realizes full-text retrieval and sorting simultaneously, without the need to introduce additional sorting algorithms, and has low adaptability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119621940B_ABST
    Figure CN119621940B_ABST
Patent Text Reader

Abstract

An embodiment of the present application provides an article retrieval method, a program product, a device, and a storage medium. The method includes: in response to the input retrieval content, obtaining a first target word frequency vector of each article from an article library; the first target word frequency vector is obtained by reducing the dimension of a first initial word frequency vector; the first initial word frequency vector of each article is determined based on the number of occurrences of each word segment in the article; determining a second initial word frequency vector of the retrieval content based on the number of occurrences of each word segment in the retrieval content, and reducing the dimension of the second initial word frequency vector to obtain a second target word frequency vector; the second target word frequency vector has the same dimension as the first target word frequency vector; determining a target article matching the retrieval content from the article library based on the similarity between each first target word frequency vector and the second target word frequency vector. The first target word frequency vector greatly reduces the memory occupancy and carries the word segment word frequency information, so that retrieval and sorting can be performed simultaneously.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of retrieval, and more particularly, to an article retrieval method, a program product, a device, and a storage medium. Background Art

[0002] In the related art, the traditional full-text retrieval technology adopts the inverted index method. Although the principle of the inverted index method is simple, when the number of words involved in the article is large, there will be a huge index volume, resulting in a high memory occupancy. At the same time, the inverted index can only determine that the retrieved articles include the segmented words in the retrieval content, but such retrieval results cannot meet other user requirements such as sorting based on word frequency. Summary of the Invention

[0003] The purpose of the embodiments of the present application is to provide an article retrieval method, a program product, a device, and a storage medium to achieve the technical effect of reducing memory occupancy and improving the user's retrieval experience.

[0004] The first aspect of the embodiments of the present application provides an article retrieval method, and the method includes:

[0005] In response to the input retrieval content, obtain the first target word frequency vector corresponding to each article from a preset article library; wherein, the first target word frequency vector is obtained by dimensionality reduction of the first initial word frequency vector; the first initial word frequency vector corresponding to each article is determined based on the number of occurrences of each segmented word in the preset segmentation dictionary in the article;

[0006] Based on the number of occurrences of each segmented word in the retrieval content in the segmentation dictionary, determine the second initial word frequency vector of the retrieval content, and perform dimensionality reduction on the second initial word frequency vector to obtain a second target word frequency vector; the second target word frequency vector has the same dimension as the first target word frequency vector;

[0007] Based on the similarity between each of the first target word frequency vectors and the second target word frequency vector, determine the target article that matches the retrieval content from the article library.

[0008] In the above implementation process, the first target word frequency vector carries the word segmentation information in the article and the information on the number of occurrences of the word segmentation in the article. At the same time, the memory occupancy of the first target word frequency vector after dimensionality reduction is greatly reduced. During the retrieval process, the matching target article is found through the similarity calculation between the first target word frequency vector and the second target word frequency vector, thereby converting part of the storage pressure into calculation pressure. Moreover, to implement full-text retrieval, there is no need to newly build a retrieval library. It only requires performing word frequency statistics on the original text of the articles in the article library, and the adaptation requirements for the database are relatively low. At the same time, since the first target word frequency vector can represent the number of occurrences of the word segmentation in the article, retrieval and sorting can be performed simultaneously without the need to introduce other sorting algorithms additionally, improving the user's retrieval experience.

[0009] Further, the word segmentation dictionary includes multiple word segmentation sets; the positions of the word segmentations in each word segmentation set in the word segmentation dictionary are related to the semantics of the word segmentations, and the positional distance of the multiple word segmentations in the word segmentation set in the word segmentation dictionary is negatively correlated with the semantic similarity between the multiple word segmentations.

[0010] In the above implementation process, by dividing the word segmentation sets, multiple word segmentations in the same domain or multiple word segmentations with similar semantics are divided into the same word segmentation set. In each word segmentation set, the positions of the multiple word segmentations are determined based on the semantic relevance, making the word segmentations with similar semantics in the word segmentation set closer to each other, which is beneficial to retaining text features in subsequent dimensionality reduction operations.

[0011] Further, the word segmentation dictionary is constructed through the following steps:

[0012] Obtain the word segmentations included in all articles in the article library, and cluster the word segmentations to obtain multiple cluster centers and the word segmentation sets corresponding to each cluster center;

[0013] Determine the interval positions of each word segmentation set in the word segmentation dictionary;

[0014] For each word segmentation set, evaluate the semantic similarity between the multiple word segmentations in the word segmentation set;

[0015] Based on the semantic similarity, determine the positions of each word segmentation in the word segmentation set in the interval positions to obtain the word segmentation dictionary.

[0016] In the above implementation process, by performing semantic clustering on the word segmentations, the word segmentations involved in the article are divided into multiple word segmentation sets, and then the word segmentations with similar semantics in the word segmentation set are set at close positions, which is beneficial to retaining text feature information in subsequent dimensionality reduction operations.

[0017] Further, each element in the first target word frequency vector is obtained through the following steps:

[0018] Obtain a first initial word frequency vector; the first initial word frequency vector includes a plurality of first sub-vectors, and the first sub-vectors correspond to the word segmentation set; each first sub-vector includes a plurality of elements, and the number of elements in the first sub-vector is the same as the number of word segmentations included in the word segmentation set;

[0019] For each target dimension in each second sub-vector, obtain the weight of each word segmentation under the target dimension, and perform weighted summation on the corresponding elements in the first sub-vector based on the weight corresponding to each word segmentation to obtain the element of the target dimension; wherein, the second sub-vector is obtained by dimensionality reduction of the first sub-vector; the weight of each word segmentation under the first target dimension is positively correlated with the importance degree of the word segmentation in the word segmentation set.

[0020] In the above implementation process, by distributing the weights of the word segmentations in each dimension of the second sub-vector, and at the same time mainly distributing the weights of the word segmentations important to the word segmentation set to the first target dimension, the elements under the first target dimension can retain more semantic information of the word segmentations with high importance degrees, avoiding the loss of semantic information of important word segmentations during dimensionality reduction. In addition, by arranging the word segmentations with similar semantics in adjacent positions in the word segmentation dictionary, the first target word frequency vector after dimensionality reduction can retain the semantic information of the word segmentations to a greater extent, reducing the loss of semantic information caused by dimensionality reduction.

[0021] Further, the first initial word frequency vector is normalized.

[0022] In the above implementation process, normalizing the first initial word frequency vector can improve the distribution characteristics of the first initial word frequency vector and enhance the dimensionality reduction effect.

[0023] Further, the determining the target article matching the retrieval content from the article library based on the similarity between the first target word frequency vector and the second target word frequency vector includes:

[0024] Determine the target sub-vector corresponding to the target field to which the retrieval content belongs in each of the first target word frequency vectors;

[0025] For each first target word frequency vector, increase the target sub-vector according to a preset domain weight factor to obtain a weighted first target word frequency vector;

[0026] Based on the similarity between the weighted first target word frequency vector and the second target word frequency vector, determine the target article matching the retrieval content from the article library.

[0027] In the above implementation process, since the target sub-vector is increased in the weighted first target word frequency vector, the role of word segmentation in the target field in similarity calculation is strengthened, and the role of word segmentation in other fields in similarity calculation is weakened. The obtained similarity result will be more affected by the target field, effectively reducing the interference of the high-frequency occurrence of word segmentation in other fields, thereby improving the retrieval accuracy.

[0028] Further, determining a target article matching the retrieval content from the article library based on the similarity between the first target word frequency vector and the second target word frequency vector includes:

[0029] Combining each of the first target word frequency vectors into a word frequency matrix;

[0030] Parallelly calculating the dot product of the first target word frequency matrix in each row of the word frequency matrix and the second target word frequency vector to determine the calculation result as the similarity;

[0031] Determining a target article matching the retrieval content from the article library based on the similarity.

[0032] In the above implementation process, the first target word frequency vectors of each article are combined into a word frequency matrix. Matrix operations do not require serial operations and can use parallel computing resources, such as the parallel logic of a GPU, to implement matrix operations, thereby comprehensively improving the operation speed of similarity, improving the retrieval efficiency, and making full use of computing power resources.

[0033] A second aspect of the embodiments of the present application provides a computer program product, where the computer program product includes a computer program, and when the computer program is executed by a processor, it implements any method of the first aspect.

[0034] A third aspect of the embodiments of the present application provides an electronic device, where the electronic device includes:

[0035] A processor;

[0036] A memory for storing executable instructions of the processor;

[0037] Wherein, when the processor calls the executable instructions, it implements the operations of any method of the first aspect.

[0038] A fourth aspect of the embodiments of the present application provides a computer-readable storage medium, on which computer instructions are stored, and when the computer instructions are executed by a processor, they implement the steps of any method of the first aspect. Description of the Drawings

[0039] To more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the accompanying drawings required for use in the embodiments of the present application. It should be understood that the following drawings only show certain embodiments of the present application and should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other relevant drawings can also be obtained based on these drawings.

[0040] Figure 1 It is a schematic flow chart of an article retrieval method provided by an embodiment of the present application;

[0041] Figures 2 - 3 It is a schematic flow chart of another article retrieval method provided by an embodiment of the present application;

[0042] Figure 4 It is a schematic diagram of the dimensionality reduction processing of the first initial word frequency vector provided by an embodiment of the present application;

[0043] Figures 5 - 6 It is a schematic flow chart of another article retrieval method provided by an embodiment of the present application;

[0044] Figure 7 It is a hardware structure diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0045] The following will describe the technical solutions in the embodiments of the present application in conjunction with the accompanying drawings in the embodiments of the present application.

[0046] It should be noted that: similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. At the same time, in the description of the present application, terms such as "first", "second", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.

[0047] The reverse indexing method means that after the article is segmented, each word in the article is used as an index to point to the article, and an index library is established accordingly. For example, if the words "segment 1" and "segment 3" appear in article A, and the words "segment 2" and "segment 3" appear in article B, then "segment 1" is indexed to article A, "segment 2" is indexed to article B, and "segment 3" is indexed to both article A and article B. When performing a full-text search later, after segmenting the input search content, the articles involved in each segment can be found through the segment index of the search content to achieve the full-text search function. Although the principle of the reverse indexing method is simple, when the vocabulary of the article is large, the indexing method based on segments will have a huge problem of the amount of index, which occupies a large amount of memory. At the same time, the reverse index can only determine that the retrieved articles include the segments in the search content, and cannot count the number of times the segments appear in the article. If the user needs to sort the search results according to the word frequency, it is necessary to re-count the word frequency in the articles in the search results for sorting, or introduce other sorting algorithms. Therefore, the reverse indexing method cannot meet other user needs such as sorting based on word frequency.

[0048] To solve at least one of the above technical problems, the present application provides an article retrieval method for performing a full-text search on all articles in a preset article library. Specifically, it includes steps 110-step 130 as Figure 1 shown.

[0049] Step 110: In response to the input search content, obtain the first target word frequency vector corresponding to each article from the preset article library.

[0050] Among them, the first target word frequency vector is obtained by dimensionality reduction of the first initial word frequency vector; the first initial word frequency vector corresponding to each article is determined based on the number of occurrences of each word in the preset word segmentation dictionary in the article.

[0051] Exemplarily, the article library includes multiple articles and the first target word frequency vector corresponding to each article. There is a one-to-one correspondence between the article and the first target word frequency vector. The first target word frequency vector is obtained by dimensionality reduction of the first initial word frequency vector. There is also a one-to-one correspondence between the first initial word frequency vector and the article.

[0052] In addition, a word segmentation dictionary is also established in advance, and the word segmentation dictionary includes the words involved in all articles in the article library. The first initial word frequency vector corresponding to each article is determined based on the number of occurrences of each word in the word segmentation dictionary in the article. The dimension of the first initial word frequency vector is the same as the number of words included in the word segmentation dictionary. It can be seen that the first word frequency vector corresponding to the article can not only represent the words included in the article, but also represent the number of times the words appear in the article.

[0053] For example, if the word segmentation dictionary includes m word segments, namely a, b, …, m, and article 1 includes word segments a, b, and m, and word segment a appears 1 time, word segment b appears 5 times, and word segment m appears 3 times, then the first initial word frequency vector is an m-dimensional vector and can be expressed as (1, 5, 0, 0, …, 3).

[0054] It can be known that before executing the retrieval method, the first initial word frequency vector corresponding to each article is generated in advance, and then the first initial word frequency vector corresponding to each article is dimensionally reduced to obtain the first target word frequency vector. The dimensionality reduction process of the first target word frequency vector will be elaborated below. The first target word frequency vector is stored in the article library and associated with the corresponding article. In subsequent use, it is possible to periodically monitor whether there are new articles in the article library and automatically update the word segmentation dictionary and the corresponding first target word frequency vector for the new articles to maintain the timeliness and accuracy of the retrieval algorithm.

[0055] In this way, when executing step 110, when the input retrieval content is received, for example, the retrieval content input by the user, the first target word frequency vector corresponding to each article is obtained from the article library.

[0056] Step 120: Determine the second initial word frequency vector of the retrieval content based on the number of occurrences of each word segment in the retrieval content in the word segmentation dictionary, and dimensionally reduce the second initial word frequency vector to obtain the second target word frequency vector; the second target word frequency vector has the same dimension as the first target word frequency vector.

[0057] Exemplarily, the input retrieval content can be first segmented to determine the word segments included in the retrieval content. Among them, the word segmentation algorithms used can include but are not limited to the JIEBA algorithm, the THULAC algorithm, and the SnowNLP, etc. Subsequently, based on the number of occurrences of each word segment in the retrieval content in the word segmentation dictionary, the second initial word frequency vector of the retrieval content is determined. Subsequently, the second initial word frequency vector is dimensionally reduced to obtain the second target word frequency vector of the retrieval content, and the second target word frequency vector has the same dimension as the first target word frequency vector. For example, both the first target word frequency vector and the second target word frequency vector are n-dimensional vectors. The dimensionality reduction process of the second target word frequency vector will be elaborated below.

[0058] Step 130: Determine the target article matching the retrieval content from the article library based on the similarity between each of the first target word frequency vectors and the second target word frequency vector.

[0059] Exemplarily, the similarity between the first target word frequency vector and the second target word frequency vector corresponding to each article can be calculated separately. For example, cosine similarity, Euclidean distance, Manhattan distance, etc. can be used, but are not limited to these, to characterize the similarity between the first target word frequency vector and the second target word frequency vector. Subsequently, based on the similarity, the target article matching the retrieved content is determined from the article library. For example, the article corresponding to the first target word frequency vector with the highest similarity can be determined as the target article. Another example is that the article corresponding to the first target word frequency vector with a similarity higher than a preset similarity threshold can be determined as the target article.

[0060] It can be understood that the traditional inverted index needs to establish the index relationship between as many word segments and articles as there are word segments involved in the article, and cannot perform dimensionality reduction, resulting in more memory occupation. In this application, the first initial word frequency vector of the article is determined based on the number of occurrences of each word segment in the article in the word segment dictionary, and then the first initial word frequency vector is dimensionally reduced to obtain the first target word frequency vector, so that the first target word frequency vector not only carries the word segment information in the article, but also carries the information of the number of occurrences of the word segment in the article. Each element in each dimension of the first target word frequency vector can be associated with a word segment, and at the same time, the memory occupation of the dimensionally reduced first target word frequency vector is greatly reduced. During the retrieval process, the matching target article is found through the similarity calculation between the first target word frequency vector and the second target word frequency vector, thus converting part of the storage pressure into calculation pressure. Moreover, to implement full-text retrieval, there is no need to newly build an additional retrieval library, and only the word frequency statistics of the original articles in the article library need to be performed, with a relatively low adaptation requirement for the database. At the same time, since the first target word frequency vector can represent the number of occurrences of a word segment in an article, retrieval and sorting can be performed simultaneously without the need to additionally introduce other sorting algorithms, improving the user's retrieval experience.

[0061] The following provides a detailed introduction to steps 110 - 130.

[0062] Regarding the word segment dictionary, in some embodiments, the word segment dictionary includes multiple word segment sets. Each word segment set corresponds to a different interval position in the word segment dictionary. Among them, different word segment sets can correspond to vocabularies in different fields. The fields include but are not limited to the medical field, the catering field, the construction field, and so on.

[0063] The position of each word segment in the word segment set in the word segment dictionary, or rather, the specific position in the interval position, is related to the semantics of the word segment. Specifically, for each word segment set, the position distance of multiple word segments in the word segment set in the word segment dictionary, or rather, the position distance in the interval position, is negatively correlated with the semantic relevance between the multiple word segments. That is, if the semantic relevance between two word segments is higher, then the position distance between these two word segments in the word segment dictionary is smaller, and the two word segments are closer.

[0064] Optionally, the word segmentation set may include a sensitive word segmentation set, and the sensitive word segmentation set may include sensitive word segmentations. The sensitive word segmentation set is placed at a specific position in the word segmentation dictionary to facilitate blocking the search for sensitive word segmentations when needed.

[0065] It can be seen that in this embodiment, by dividing the word segmentation set, multiple word segmentations in the same field or multiple word segmentations with similar semantics are divided into the same word segmentation set, and the positions of multiple word segmentations are determined based on semantic relevance in each word segmentation set, so that word segmentations with similar semantics in the word segmentation set are closer to each other, which is beneficial to retaining text features in subsequent dimensionality reduction operations.

[0066] According to some embodiments of the present application, the word segmentation dictionary can be constructed through steps 210 - 240 as shown in Figure 2 shown.

[0067] Step 210: Obtain the word segmentations included in all articles in the article library, and cluster the word segmentations to obtain multiple cluster centers and the word segmentation sets corresponding to each cluster center.

[0068] Exemplarily, each article in the article library can be first segmented to obtain all the word segmentations included in all articles. Then, all the word segmentations are clustered. The clustering of word segmentations can be based on semantics. The specific clustering process can refer to related technologies and will not be elaborated in this embodiment. After clustering, multiple cluster centers and the word segmentation sets corresponding to each cluster center can be obtained. Exemplarily, one or more word segmentation sets correspond to the vocabulary of one field.

[0069] Step 220: Determine the interval position of each word segmentation set in the word segmentation dictionary.

[0070] Exemplarily, after obtaining multiple word segmentation sets, the interval position corresponding to each word segmentation set in the word segmentation dictionary can be determined, so that each word segmentation set in the word segmentation dictionary is independent of each other.

[0071] Step 230: For each word segmentation set, evaluate the semantic similarity between multiple word segmentations in the word segmentation set.

[0072] Exemplarily, for each word segmentation set, the semantic similarity between each pair of word segmentations in the word segmentation set can be evaluated respectively. The specific evaluation method of semantic similarity can refer to related technologies and will not be elaborated here.

[0073] Step 240: Based on the semantic similarity, determine the position of each word segmentation in the interval position in the word segmentation set to obtain the word segmentation dictionary.

[0074] Finally, based on the semantic similarity between multiple word segmentations in the word segmentation set, determine the position of each word segmentation in the interval position, so as to arrange and obtain the word segmentation dictionary.

[0075] It can be seen that in this embodiment, by performing semantic clustering on the segmented words, the segmented words involved in the article are divided into multiple segmented word sets, and then the segmented words with similar semantics in the segmented word sets are set at adjacent positions, which is beneficial to retaining text feature information in subsequent dimensionality reduction operations.

[0076] Based on any of the above embodiments, each element in the first target word frequency vector is obtained through steps 310 - 320 as shown below. Figure 3 as shown.

[0077] Step 310: Obtain a first initial word frequency vector. Among them, the first initial word frequency vector includes multiple first sub - vectors, and the first sub - vectors correspond to the segmented word sets; each first sub - vector includes multiple elements, and the number of elements in the first sub - vector is the same as the number of segmented words included in the corresponding segmented word set.

[0078] As described above, since the segmented word dictionary includes multiple segmented word sets, the first initial word frequency vector can also be divided into multiple first sub - vectors. The first sub - vectors and the segmented word sets are in a one - to - one correspondence. Each first sub - vector includes one or more elements, and the number of elements in each first sub - vector is the same as the number of segmented words included in the segmented word set corresponding to the first sub - vector.

[0079] For example, as shown Figure 4 if the segmented word dictionary includes segmented word set A and segmented word set B, where segmented word set A includes segmented words a1 to a i a total of i segmented words, and segmented word set B includes segmented words b1 to b j a total of j segmented words, then the first initial word frequency vector of each article is an (i + j) - dimensional vector, as shown Figure 4 The first initial word frequency vector includes a first sub - vector and a first sub - vector The first sub - vector corresponds to segmented word set A and includes i elements, and the first sub - vector corresponds to segmented word set B and includes j elements.

[0080] Step 320: For each target dimension in each second sub - vector, obtain the weight of each segmented word in the target dimension, and perform a weighted sum on the corresponding elements in the first sub - vector based on the weight corresponding to each segmented word to obtain the element of the target dimension. Among them, the second sub - vector is obtained by reducing the dimension of the first sub - vector; the weight of each segmented word in the first dimension is positively correlated with the importance of the segmented word in the segmented word set.

[0081] It can be understood that each first sub-vector in the first initial word frequency vector is dimensionally reduced separately without interference. In this way, each first sub-vector is dimensionally reduced to obtain a second sub-vector after dimensional reduction. The first target word frequency vector includes multiple second sub-vectors. Then, when calculating the elements of each target dimension in each second sub-vector, first determine the word segmentation set corresponding to the corresponding first sub-vector, and then obtain the weights of each word segmentation in the target dimension. Then, based on the weights corresponding to each word segmentation, perform weighted summation on the corresponding elements in the first sub-vector to obtain the elements of the target dimension.

[0082] Continue to refer to Figure 4 , the first target word frequency vector includes multiple second sub-vectors, and each second sub-vector is obtained by dimensional reduction of the corresponding first sub-vector in the first initial word frequency vector. Among them, the second sub-vector is obtained by dimensional reduction of the first sub-vector , where p < i; the second sub-vector is obtained by dimensional reduction of the first sub-vector , where q < j.

[0083] Taking the second sub-vector as an example, the second sub-vector is a p-dimensional vector and includes a total of p elements. For the first target dimension in the second sub-vector , that is, the dimension where the element T 11 is located, first determine the word segmentation set A corresponding to the first sub-vector , and then obtain the weights of each word segmentation in the word segmentation set A in the first target dimension, that is, the weight w 11 of the word segmentation a1 in the first target dimension, …… the weight w i of the word segmentation a 1i in the first target dimension. Subsequently, based on the weights, perform weighted summation on the corresponding elements in the first sub-vector to obtain the elements of the first target dimension, that is, . Similarly, for the elements of the pth target dimension , and so on.

[0084] It can be seen from this that if the rth first sub-vector in the first initial word frequency vector includes v elements, then the element T rs of the sth target dimension in the rth second sub-vector in the first target word frequency vector is calculated as follows: .

[0085] In addition, the weight of each word segment under the first target dimension is positively correlated with the importance of the word segment in the word segment set. The importance refers to the degree of correlation between the word segment and the field indicated by the word segment set. For example, if a word segment set indicates the medical field, then the word segments "vaccine" and "pneumonia" have a higher importance in the word segment set, and correspondingly, their weights under the first target dimension are also higher. In addition, the sum of the weights of the word segments under all target dimensions is 1. It can be seen that by distributing the weights of the word segments in each dimension of the second sub-vector, and mainly distributing the weights of the word segments important to the word segment set to the first target dimension, the elements under the first target dimension can retain more semantic information of the word segments with high importance, avoiding the loss of semantic information of important word segments during dimensionality reduction. In addition, by arranging word segments with similar semantics in adjacent positions in the word segment dictionary, the first target word frequency vector after dimensionality reduction can retain the semantic information of the word segments to a greater extent, reducing the loss of semantic information caused by dimensionality reduction.

[0086] In addition, according to some embodiments of the present application, before reducing the dimension of the first initial word frequency vector, the first initial word frequency vector can be first normalized to improve the distribution characteristics of the first initial word frequency vector and improve the dimensionality reduction effect.

[0087] In addition, for the generation process of the second initial word frequency vector and the dimensionality reduction process of the second initial word frequency vector, reference can be made to the dimensionality reduction process of the first initial word frequency vector, which will not be elaborated herein.

[0088] Based on any of the above embodiments, step 130 may specifically include steps 510-step 530 as Figure 5 shown.

[0089] Step 510: Determine the target sub-vector corresponding to the target field to which the retrieved content belongs in each of the first target word frequency vectors.

[0090] Exemplarily, after obtaining the input retrieved content, the field to which the retrieved content belongs can be determined. The method for determining the field to which the retrieved content belongs can refer to the related art and will not be elaborated in this embodiment. As described above, one or more word segment sets in the word segment dictionary can refer to a field. Therefore, after determining the target field to which the retrieved content belongs, the target sub-vector corresponding to the target field can be determined in each first target word frequency vector. In addition, since the first target word frequency vector includes multiple second sub-vectors, the target sub-vector is actually one or more of the second sub-vectors.

[0091] Step 520: For each first target word frequency vector, increase the target sub-vector according to a preset domain weight factor to obtain a weighted first target word frequency vector.

[0092] Exemplarily, after determining the target sub-vector, each element in the target sub-vector can be multiplied by a preset domain weight factor to increase the target sub-vector and obtain a weighted first target word frequency vector. The domain weight factor is greater than 1.

[0093] Step 530: Based on the similarity between the weighted first target word frequency vector and the second target word frequency vector, determine a target article matching the retrieved content from the article library.

[0094] It can be seen that since the target sub-vector is increased in the weighted first target word frequency vector, the role of word segmentation in the target domain in similarity calculation is strengthened, and the role of word segmentation in other domains in similarity calculation is weakened. The obtained similarity result will be more affected by the target domain, effectively reducing the interference of the high-frequency occurrence of word segmentation in other domains, thereby improving the retrieval accuracy.

[0095] In addition, based on any of the above embodiments, step 130 may specifically include steps 610-step 630 as Figure 6 shown.

[0096] Step 610: Combine each of the first target word frequency vectors into a word frequency matrix.

[0097] Exemplarily, each first target word frequency vector can be used as a row vector in the word frequency matrix, thereby obtaining a word frequency matrix composed of multiple first target word frequency vectors.

[0098] Step 620: Calculate the dot product product of the first target word frequency matrix in each row of the word frequency matrix and the second target word frequency vector in parallel, and determine the calculation result as the similarity.

[0099] Exemplarily, the second target word frequency vector can be a column vector. Multiplying the word frequency matrix by the second target word frequency vector is actually performing a dot product operation between each row first target word frequency vector in the word frequency matrix and the second target word frequency vector. Thus, the dot product product of the first target word frequency matrix in each row of the word frequency matrix and the second target word frequency vector can be calculated in parallel to obtain a similarity vector. The similarity vector is a column vector, and the elements it includes are similarities.

[0100] Optionally, the parallel logic of the GPU (Graphics Processing Unit) can be used to execute step 620, so as to make full use of computing resources.

[0101] Step 630: Based on the similarity, determine a target article matching the retrieved content from the article library.

[0102] Finally, the article corresponding to the first target word frequency vector with the highest similarity can be determined as the target article. For another example, the article corresponding to the first target word frequency vector with a similarity higher than a preset similarity threshold can be determined as the target article.

[0103] It can be seen that in this embodiment, the first target word frequency vectors of each article are combined into a word frequency matrix. Matrix operations do not require serial operations and can use parallel computing resources, such as the parallel logic of a GPU, to implement matrix operations. Thus, the operation speed of similarity can be comprehensively improved, the retrieval efficiency can be enhanced, and computing power resources can be fully utilized.

[0104] To better understand this solution, the following is illustrated by an example.

[0105] Before performing the retrieval, an article library and a word segmentation dictionary are pre-constructed. Specifically, all word segments involved in all articles in the article library can be counted to construct the word segmentation dictionary. For example, the article library includes three articles. The content of article A is "test case", the content of article B is "This article is a test case", and the content of article C is "article article". By performing word segmentation on the three articles, the word segmentation dictionary can be constructed as (1. test 2. case 3. this 4. article 5. is).

[0106] Subsequently, based on the number of occurrences of each word segment in the dictionary in the article, the first initial word frequency vector of each article is determined. For example, the first initial word frequency vector of article A is (1, 1, 0, 0, 0), the first initial word frequency vector of article B is (1, 1, 1, 1, 1), and the first initial word frequency vector of article C is (0, 0, 0, 2, 0).

[0107] Subsequently, normalization processing is performed on the first initial word frequency vector of each article so that the modulus lengths of all first initial word frequency vectors are 1. The normalized first initial word frequency vector of article A is (0.707, 0.707, 0, 0, 0), the normalized first initial word frequency vector of article B is (0.447, 0.447, 0.447, 0.447, 0.447), and the normalized first initial word frequency vector of article C is (0, 0, 0, 1, 0).

[0108] In this example, since there are only 5 word segments in the word segmentation dictionary and the storage and operation pressure is not high, dimension reduction is not required. However, in actual use, the word segmentation dictionary usually has hundreds of thousands of word segments, and the first initial word frequency vector has hundreds of thousands of dimensions. Therefore, a dimension reduction algorithm is needed to reduce the dimension of the first initial word frequency vector and relieve the storage and calculation pressure.

[0109] When starting the retrieval, all the term frequency vectors in the article library (the first initial term frequency vector in this example, or the first target term frequency vector if dimensionality reduction has been performed) can be merged into a term frequency matrix, as shown below:

[0110]

[0111] Obtain the input retrieval content "article", and determine the second initial term frequency vector of the retrieval content as (0, 0, 0, 1, 0) according to the number of occurrences of each token in the retrieval content in the tokenization dictionary. Similarly, in actual use, the second initial term frequency vector needs to be normalized and dimensionality-reduced. However, in this example, since the dimensionality reduction process is omitted, the term frequency matrix and the second initial term frequency vector are then subjected to a dot product operation, and the operation process is as shown below:

[0112]

[0113] Thus, a similarity vector (0, 0.447, 1) can be obtained T , and each element in the similarity vector is the similarity between the corresponding article and the retrieval content. That is, the similarity between the retrieval content "article" and article A is 0, the similarity with article B is 0.447, and the similarity with article C is 1. Finally, according to the similarity results, the target articles matching the retrieval content are obtained, and the multiple target articles are sorted in descending order of similarity, so as to achieve the full-text retrieval function integrating retrieval and sorting.

[0114] Based on an article retrieval method described in any of the above embodiments, the present application further provides a computer program product, and the computer program product includes one or more computer programs or instructions. The computer programs or instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. When the computer program is executed by a processor, it implements an article retrieval method described in any of the above embodiments.

[0115] Based on an article retrieval method described in any of the above embodiments, the present application further provides a schematic structural diagram of an electronic device as Figure 7 shown. As Figure 7 , at the hardware level, the electronic device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory. Of course, it may also include other hardware required for other services. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to implement an article retrieval method described in any of the above embodiments.

[0116] The present application also provides a computer storage medium storing a computer program, which when executed by a processor can be used to execute an article retrieval method described in any of the above embodiments.

[0117] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions, and operations of devices, methods, and computer program products according to multiple embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.

[0118] In addition, in each embodiment of the present application, the various functional modules can be integrated together to form an independent part, or each module can exist alone, or two or more modules can be integrated to form an independent part.

[0119] If the above functions are implemented in the form of software function modules and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0120] The above are only embodiments of the present application and are not intended to limit the protection scope of the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application. It should be noted that similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.

[0121] As described above, this is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or replacements within the technical scope disclosed by the present application, and all should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

[0122] It should be noted that in this text, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the existence of additional identical elements in the process, method, article or device comprising the said element.

Claims

1. An article retrieval method, characterized in that: The method comprises: In response to the input search content, a first target word frequency vector corresponding to each article is obtained from a preset article library; wherein the first target word frequency vector is obtained by reducing the dimension of the first initial word frequency vector; the article and the first initial word frequency vector are in a one-to-one correspondence, and the first initial word frequency vector corresponding to each article is determined based on the number of occurrences of each word in a preset word segmentation dictionary in the article; the word segmentation dictionary includes the words included in all articles in the article library; the dimension of the first initial word frequency vector is consistent with the number of words included in the word segmentation dictionary; the word segmentation dictionary includes multiple word segmentation sets; Based on the number of occurrences of each word segment in the word segmentation dictionary in the search content, determine a second initial word frequency vector of the search content, and perform dimension reduction on the second initial word frequency vector to obtain a second target word frequency vector; the dimension of the second target word frequency vector is consistent with that of the first target word frequency vector; Calculate the similarity between the first target word frequency vector and the second target word frequency vector corresponding to each of the articles respectively, and determine that the article corresponding to the first target word frequency vector with the highest similarity is the target article matching the search content, or determine that the article corresponding to the first target word frequency vector with a similarity higher than a preset similarity threshold is the target article; The first initial word frequency vector includes multiple first sub-vectors; the first sub-vectors correspond to the word segmentation sets one by one, and the number of elements in each first sub-vector is consistent with the number of word segmentations included in the word segmentation set corresponding to the first sub-vector; the first target word frequency vector includes multiple second sub-vectors obtained by reducing the dimension of the first sub-vector; the element of the sth target dimension in the rth second sub-vector is , w su is the weight of the word corresponding to the u-th dimension in the r-th first sub-vector under the s-th target dimension, t ru is the element of the uth dimension in the rth first subvector, and v is the number of elements in the rth first subvector.

2. The method according to claim 1, characterized in that The word segmentation dictionary includes multiple word segmentation sets; the position of each word segmentation in the word segmentation dictionary is related to the semantics of the word segmentation, and the position distance of multiple word segments in the word segmentation set in the word segmentation dictionary is negatively correlated with the semantic similarity between the multiple word segments.

3. The method according to claim 2, characterized in that The word segmentation dictionary is constructed by the following steps: Obtaining the segmented words included in all articles in the article library, and clustering the segmented words to obtain multiple cluster centers and a segmented word set corresponding to each cluster center; Determine the interval position of each word segmentation set in the word segmentation dictionary; For each word segmentation set, evaluating the semantic similarity between multiple word segments in the word segmentation set; The position of each word in the word set in the interval position is determined based on the semantic similarity to obtain a word segmentation dictionary.

4. The method according to claim 2 or 3, characterized in that The weight of each word segment under the first target dimension is positively correlated with the importance of the word segment in the word segment set.

5. The method according to claim 4, characterized in that The first initial word frequency vector is normalized.

6. The method according to claim 2, characterized in that The respectively calculating the similarity between the first target word frequency vector and the second target word frequency vector corresponding to each of the articles comprises: Determine a target subvector corresponding to the target field to which the search content belongs in each of the first target word frequency vectors; For each first target word frequency vector, increase the target sub-vector according to a preset domain weight factor to obtain a weighted first target word frequency vector; The similarity between the weighted first target word frequency vector and the second target word frequency vector corresponding to each of the articles is calculated respectively.

7. The method according to claim 1 or 6, characterized in that: The respectively calculating the similarity between the first target word frequency vector and the second target word frequency vector corresponding to each of the articles comprises: Merge each of the first target word frequency vectors into a word frequency matrix; The dot product of the first target word frequency vector and the second target word frequency vector in each row of the word frequency matrix is ​​calculated in parallel, and the calculation result is determined as the similarity.

8. A computer program product, characterized in that The computer program product comprises a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.

9. An electronic device, characterized in that: The electronic device comprises: processor; a memory for storing processor-executable instructions; Wherein, when the processor calls the executable instructions, the operation of any method described in claims 1-7 is implemented.

10. A computer-readable storage medium, characterized in that: Computer instructions are stored thereon, and when the computer instructions are executed by a processor, the steps of any method described in claims 1-7 are implemented.

Citation Information

Patent Citations

  • Article search method, device and electronic device

    CN109241238A

  • Associated lexicon generation method and device, text retrieval method and device, equipment and medium

    CN114519350A