Content search method and device, equipment and medium
By performing attention analysis and vocabulary expansion on the search condition content entered by the user, and feature extraction combined with the contextual context, the problem that keyword context in the prior art is not considered is solved, and the matching accuracy of search results is improved.
Patent Information
- Application Number
- CN202411185543.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-27
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2044-08-27
AI Technical Summary
The prior art fails to effectively consider the context of keywords in text search, resulting in low accuracy of matching results and the inability to effectively distinguish ambiguity and puns in natural language.
By performing attention analysis on the search conditions input by the user, the keyword collection is determined, and vocabulary expansion is performed, feature extraction is performed in combination with the contextual context to improve the matching degree.
Improve the matching accuracy of search results, can better understand the semantics of search criteria content, and reduce ambiguity and mismatch of puns.
Smart Images

Figure CN119988595A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of text processing technology, and in particular to a content search method, device, equipment, and medium. Background Art
[0002] The search system provides users with search results based on the search criteria they input, helping them quickly retrieve desired content, such as writing materials, information articles, etc.
[0003] In the related art, the user inputs text content as search conditions, and the search system analyzes the search conditions, determines the keywords in the search conditions, matches the keywords with various contents in the database, and feeds back the matching results that meet the preset requirements to the user.
[0004] However, the matching method based on surface keywords does not take into account the context of the keywords, cannot distinguish ambiguity, puns, etc. in natural language, and has insufficient analysis of the context, resulting in low accuracy of matching results and poor content feedback to users. Summary of the invention
[0005] The embodiments of the present application provide a content search method, apparatus, device, and medium, which can improve the matching degree of search results. The technical solution is as follows:
[0006] In one aspect, a content search method is provided, the method comprising:
[0007] Acquire search condition content, where the search condition content is content input by a user for searching;
[0008] When the length of the search condition content indicates that the search condition content belongs to sentence content, performing attention analysis on the search condition content to determine a first keyword set of the search condition content, wherein the first keyword set includes at least one keyword in the search condition content, and the at least one keyword refers to a word in the search condition content that meets a preset attention condition;
[0009] Performing vocabulary expansion on at least one keyword in the first keyword set to obtain an expanded vocabulary and a second keyword set including the first keyword set and the expanded vocabulary, wherein the expanded vocabulary meets a preset correlation requirement with the keywords in the first keyword set;
[0010] Performing feature extraction on the second keyword set to obtain keyword feature representation;
[0011] Based on the matching degree between the keyword feature representation and the candidate content in the candidate content library, at least one first content is obtained from the candidate content library as a search result of the search condition content.
[0012] In another aspect, a content search device is provided, the device comprising:
[0013] An acquisition module, used to acquire search condition content, where the search condition content is content input by a user for searching;
[0014] an attention analysis module, configured to, when the length of the search condition content indicates that the search condition content belongs to sentence content, perform attention analysis on the search condition content to determine a first keyword set of the search condition content, wherein the first keyword set includes at least one keyword in the search condition content, and the at least one keyword refers to a word in the search condition content that meets a preset attention condition;
[0015] a vocabulary expansion module, configured to perform vocabulary expansion on at least one keyword in the first keyword set to obtain an expanded vocabulary, and a second keyword set including the first keyword set and the expanded vocabulary, wherein the expanded vocabulary meets a preset relevance requirement with the keywords in the first keyword set;
[0016] A feature extraction module, used to extract features from the second keyword set to obtain keyword feature representation;
[0017] The search result determination module is used to obtain at least one first content from the candidate content library as the search result of the search condition content based on the matching degree between the keyword feature representation and the candidate content in the candidate content library.
[0018] In an optional embodiment, the attention analysis module is further used to perform word segmentation on the search condition content to obtain a word sequence arranged in order after the segmentation, and each word in the word sequence corresponds to a serial number; the word sequence is processed by multiple transformer layers in a deep learning model, and a self-attention weight matrix is output, wherein the i-th row element in the self-attention weight matrix is used to indicate the attention distribution of the word with serial number i, and the j-th column element in the self-attention weight matrix is used to indicate the attention weight assigned to each word in the word sequence to the word with serial number j, and the i-th row and j-th column element in the self-attention weight matrix is used to indicate the attention weight assigned to each word in the word sequence to the word with serial number j. The element represents the attention weight assigned by the word with sequence number i to the word with sequence number j in the word sequence; i and j are positive integers; the self-attention weight matrix is adjusted to obtain a weighted self-attention weight matrix; the attention score of each word in the word sequence is determined based on the weighted self-attention weight matrix, and the attention score is used to indicate the importance of each word in the word sequence in the search condition content; the first keyword set of the search condition content is determined based on the attention score of each word in the word sequence, wherein the attention score of the words in the first keyword set is greater than a preset score threshold.
[0019] In an optional embodiment, the attention analysis module is further used to perform feature extraction on the search condition content to obtain a first sentence feature representation; and determine a weighted self-attention weight matrix based on the product of the first sentence feature representation and the self-attention weight matrix.
[0020] In an optional embodiment, the vocabulary expansion module is also used to perform feature extraction on the at least one keyword in the first keyword set to obtain at least one corresponding word feature representation; obtain a vocabulary and a feature representation library corresponding to the vocabulary, and there is a mapping relationship between multiple word feature representations in the feature representation library and multiple words in the vocabulary, wherein the kth word feature representation in the feature representation library is obtained after feature extraction of the kth word in the vocabulary, and k is a positive integer; determine a synonym feature representation corresponding to the at least one word feature representation in the feature representation library; determine a corresponding synonym in the vocabulary as the extended vocabulary based on the synonym feature representation, wherein the semantics of the synonym and the semantics of the at least one keyword meet a preset similarity requirement.
[0021] In an optional embodiment, the vocabulary expansion module is also used to calculate the similarity between the mth word feature representation in the at least one word feature representation and the multiple word feature representations in the feature representation library, and determine the word feature representation whose similarity reaches a preset threshold as the synonym feature representation corresponding to the mth word feature representation; m is a positive integer.
[0022] In an optional embodiment, the feature extraction module is further used to extract features from the search condition content to obtain a first sentence feature representation;
[0023] The device also includes:
[0024] A vocabulary screening module is used to process the first sentence feature representation and the keyword feature representation through a text sorting algorithm to obtain a semantic similarity image, wherein the semantic similarity image includes a main node representing the search condition content and multiple sub-nodes representing the vocabulary in the second keyword set, and multiple lines between the main node and the multiple sub-nodes respectively correspond to node weights, and the node weights are used to indicate the semantic similarity between the vocabulary in the second keyword set and the search condition content; iteratively calculate and update the node weights in the semantic similarity image through the text sorting algorithm; when the number of iterations reaches a preset iteration number threshold, the second keyword set is screened based on the node weights to obtain a third keyword set;
[0025] The feature extraction module is further used to extract features from the third keyword set to obtain the keyword feature representation.
[0026] In an optional embodiment, the search result determination module is further used to perform feature extraction on the candidate content in the candidate content library to obtain a feature representation pool corresponding to the candidate content library, the candidate content belongs to sentence content, and there is a mapping relationship between the nth sentence feature representation in the feature representation pool and the nth sentence content in the candidate content, where n is a positive integer; based on the similarity between the keyword feature representation and the sentence feature representation in the feature representation pool, a candidate content list is determined from the candidate content library, and the candidate content list includes multiple first contents; based on the matching degree between the first sentence feature representation and the feature representations of the multiple first contents, the at least one first content is determined from the candidate content list as the search result of the search condition content.
[0027] On the other hand, a computer device is provided, comprising a processor and a memory, wherein the memory stores at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement a content search method as described in any of the above-mentioned embodiments of the present application.
[0028] On the other hand, a computer-readable storage medium is provided, wherein at least one instruction, at least one program, a code set or an instruction set is stored in the storage medium, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by a processor to implement a content search method as described in any of the above-mentioned embodiments of the present application.
[0029] On the other hand, a computer program product or a computer program is provided, the computer program product or the computer program includes computer instructions, the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes any of the content search methods described in the above embodiments.
[0030] The beneficial effects brought by the technical solution provided by the embodiment of the present application include at least:
[0031] By analyzing the search condition content input by the user, for the search condition content belonging to the sentence type, the attention each word receives from the other words is determined to determine the importance of each word in the sentence, and keywords are selected as the basis for expansion. Combining the context to expand the keywords, the semantics of the search condition content can be more accurately understood based on the expanded vocabulary, and the ambiguity and puns that may exist in the text content can be fully considered. Compared with the related technology that only relies on surface keywords for matching and feedback of search results, it can improve the accuracy of matching results and feedback results that are more in line with the search requirements. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0033] Figure 1 is a schematic diagram of a content search system provided by an exemplary embodiment of the present application;
[0034] Figure 2is a flow chart of a content search method provided by an exemplary embodiment of the present application;
[0035] Figure 3 is a flow chart of a content search method provided by another exemplary embodiment of the present application;
[0036] Figure 4 is a structural block diagram of a content search device provided by an exemplary embodiment of the present application;
[0037] Figure 5 is a structural block diagram of a content search device provided by another exemplary embodiment of the present application;
[0038] Figure 6 It is a structural block diagram of a computer device provided by an exemplary embodiment of the present application. DETAILED DESCRIPTION
[0039] In order to make the objectives, technical solutions and advantages of the present application clearer, the implementation methods of the present application will be further described in detail below with reference to the accompanying drawings.
[0040] Exemplary embodiments will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The implementations described in the following exemplary embodiments do not represent all implementations consistent with the present application. Instead, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.
[0041] The terms used in this application are for the purpose of describing specific embodiments only and are not intended to limit this application. The singular forms of "a", "said" and "the" used in this application and the appended claims are also intended to include plural forms unless the context clearly indicates other meanings. It should also be understood that the term "and / or" used herein refers to and includes any or all possible combinations of one or more associated listed items.
[0042] It should be noted that the information and data involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant laws, regulations and standards of relevant countries and regions.
[0043] It should be understood that, although the terms first, second, etc. may be used in the present application to describe various information, these information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of the present application, the first parameter may also be referred to as the second parameter, and similarly, the second parameter may also be referred to as the first parameter. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".
[0044] First, a brief introduction is given to the terms involved in the embodiments of this application:
[0045] Word2Vec model (Word to Vector): The main purpose is to generate word vectors, that is, to map words into vector space. Each word is represented as a vector of fixed dimension, and these vectors can capture the semantic relationship between words.
[0046] In this application, feature representation refers to quantity. The process of extracting features from text to obtain feature representation can be regarded as the process of converting text into a vector. Feature extraction is performed on words to obtain word vectors (i.e., word feature representation), and feature extraction is performed on sentences to obtain sentence vectors (sentence feature representation).
[0047] BERT model (Bidirectional Encoder Representations from Transformers): A pre-trained language model based on the Transformer architecture that can also convert text into vector form. Compared with the Word2Vector model, the BERT model generates context-related word vectors. It can simultaneously consider the contextual information on the left and right sides of a word through a multi-layer self-attention mechanism to generate deep bidirectional representations.
[0048] Among them, the deep learning model for processing word sequences in this application refers to the BERT model.
[0049] TextRank algorithm (text ranking algorithm): It is a graph-based ranking algorithm that divides the text into sentences and performs word segmentation, using some statistical indicators to evaluate the importance of words or phrases. The co-occurrence relationship between words or phrases is used to construct a graph, where the nodes represent words or sentences, and the weights of the edges represent the similarity or co-occurrence frequency between them. The importance scores of the nodes in the graph are calculated iteratively until convergence, and finally the ranking of keywords or sentences is obtained.
[0050] In the present application, a text sorting algorithm is used to sort the words in the search content condition and the second keyword set (including the words in the first keyword set and the extended words) to determine several keywords that are most similar in semantics to the search content condition.
[0051] The search system is an efficient tool that can help users quickly retrieve the information they want based on the search conditions entered by the users.
[0052] In the related technology, the system analyzes the text content input by the user, identifies and extracts the key words therein, and then accurately matches them with the content in the database, and presents the content that meets the preset standards to the user as search results, thereby improving the efficiency and accuracy of user retrieval.
[0053] However, relying solely on keyword matching has certain limitations. The system matches content based on the literal meaning of keywords, ignoring the specific context of the keywords, and cannot effectively identify and distinguish common phenomena such as ambiguity and puns in natural language. Lacking context analysis, the system may not be able to accurately grasp the user's true intentions, resulting in low accuracy of matching results, and the search results fed back to users do not meet their expectations.
[0054] The present application provides a content search method that can reasonably expand the search conditions, analyze the proportion of attention each word in the search conditions receives from other words, determine the important keywords in the search conditions in combination with the context, and correctly understand the semantics of the search conditions when the search conditions belong to the sentence content (that is, the search conditions have a high dimension). Based on the keywords, the vocabulary is expanded, and the content that meets the search conditions is matched to the expanded vocabulary as the search results for feedback. It can be appropriately expanded on the basis of correctly understanding the text semantics of the search conditions, and the content with a high matching degree is determined to be presented to the user as the search results.
[0055] Figure 1 It is a schematic diagram of a content search system provided by an exemplary embodiment of the present application, wherein the content search system can meet the needs of users to search based on high-dimensional search conditions and low-dimensional search conditions.
[0056] The content search system 100 involves a user end 110 and a server end 120 , and the user end 110 and the server end 120 are connected via a communication network 130 .
[0057] The user terminal 110 inputs the search condition content and sends it to the server terminal 120 . The server terminal 120 analyzes the search condition content, determines the search result and feeds the search result back to the user terminal 110 .
[0058] Optionally, when the server 120 receives the search condition content sent by the user 110, it first performs dimension judgment on the search condition content. Wherein, the search condition content is text content, and this embodiment performs dimension judgment based on Chinese characters.
[0059] If the number of Chinese characters contained in the search condition content does not reach 5, it is determined that the search condition content belongs to low-dimensional content, that is, the search condition is determined to be vocabulary content; if the number of Chinese characters contained in the search condition content reaches 5, it is determined that the search condition content belongs to high-dimensional content, that is, the search condition is determined to be sentence content.
[0060] (1) When the search condition is a sentence content, the search condition content is split into multiple words, and the attention mechanism of the BERT model is used to analyze the search condition content. The attention amount allocated to each word in the search condition content from other words is determined based on the self-attention weight matrix of each transformer layer in the BERT model, and the attention of each word is normalized. The attention weight ratio obtained after normalization is used as a benchmark to rank the importance of each word, and some words with higher rankings are selected as keywords to obtain a first keyword set. The first keyword set is used for vocabulary expansion to obtain expanded words with similar or identical semantics to the first keyword set, and the expanded words and the first keyword set are combined into a second keyword set.
[0061] Feature extraction is performed on the second keyword set and the search condition content respectively to obtain word feature representations corresponding to each word in the second keyword set and the first sentence feature representation corresponding to the search condition content.
[0062] The server 120 obtains a preset feature representation pool in the preparation phase (ie, the phase before receiving the search condition content), wherein the feature representation pool contains multiple sentence feature representations, each of which is obtained by mapping a section of text content (sentence text content).
[0063] The similarity between the sentence feature representations in the feature representation pool and the word feature representations corresponding to the second keyword set is calculated one by one, and the sentence feature representations with the highest similarity are used as candidate feature representations. The candidate feature representations are matched with the first sentence feature representation, and the text content corresponding to the candidate feature representation with the highest matching degree is fed back to the user terminal 110 as the search result.
[0064] (2) When the search condition is a vocabulary content, firstly, homophone correction is performed on the search condition to correct the homophone typos in the search condition to obtain a corrected vocabulary, keyword association is performed on the corrected vocabulary to obtain an expanded vocabulary, and the expanded vocabulary is combined with the corrected vocabulary to obtain an expanded vocabulary set.
[0065] Perform feature extraction on the extended vocabulary set to obtain a corresponding word feature representation set, that is, the word feature representation set is the result of mapping the extended vocabulary set in the vector space, and determine several word feature representations that are closest to the word feature representation set from the vector space based on the distance between the feature representations, obtain synonyms of the extended vocabulary set, and add the synonyms to the extended vocabulary set.
[0066] After obtaining the extended vocabulary set, the sentence feature representations in the feature representation pool are calculated for similarity with the word feature representations corresponding to the extended vocabulary set one by one, and the sentence feature representations with the highest similarity are used as candidate feature representations. Matching is performed based on the candidate feature representations and the feature representations corresponding to the search condition content, and the text content corresponding to the candidate feature representation with the highest matching degree is fed back to the user terminal 110 as the search result.
[0067] To sum up, the content search system and content search method in this application can simultaneously deal with search condition content of different dimensions, reasonably expand the search condition content, fully understand the semantics of the search condition content in combination with the context, and then feedback the search results, thereby improving the accuracy of the search results.
[0068] In combination with the above-mentioned noun introduction and application scenarios, the content search method provided by the present application is described. The method can be executed by a server or a terminal, or can be executed by a server and a terminal together. In the embodiment of the present application, the method is described by taking the execution of the terminal as an example. Figure 2 As shown, Figure 2 1 is a flow chart of a content search method provided by an exemplary embodiment of the present application. The method includes the following steps.
[0069] Step 210, obtaining search condition content.
[0070] The search condition content is the content input by the user for searching. The search condition content is text content with a certain length, including but not limited to characters, punctuation marks, numbers, etc.
[0071] For example, the search condition content is as follows: "human health report", which contains 6 characters, indicating that the user expects to retrieve information related to human health report.
[0072] Step 220: When the length of the search condition content indicates that the search condition content belongs to sentence content, an attention analysis is performed on the search condition content to determine a first keyword set for the search condition content.
[0073] After obtaining the search condition content, the dimension of the search condition content is first determined.
[0074] Users can enter words, sentence fragments, sentences, etc. of any dimension as search condition content, and judge based on the length of the search condition. If the length is less than 5, it is judged as a word; if the length is greater than or equal to 5, the input is judged as a sentence or sentence fragment. Hereinafter, sentences or sentence fragments are referred to as sentence content.
[0075] In some embodiments, a vocabulary is prepared in advance, which contains a variety of words, including search conditions whose user input frequency reaches a preset frequency threshold (such as 5 searches out of every 100 searches contain the same word), associated search conditions (words with a high degree of association determined by combining the search conditions entered by the user within a historical time period), etc. When judging the dimension of the search condition content, the dimension of the search condition can also be determined based on whether the search condition content is included in the vocabulary. For example, if the length of the search condition content is greater than or equal to 5 and is not in the constructed vocabulary, the search condition content is judged to be sentence content.
[0076] For example, if the search condition content is as follows: "health, diet, recipes", the length of the search condition content is based on the number of Chinese characters, and if it contains 6 Chinese characters, then the search condition content is a sentence content.
[0077] For example, if the search condition content is as follows: "health, recipe", the length of the search condition content is based on the number of Chinese characters, and if it contains 4 Chinese characters, then the search condition content is vocabulary content.
[0078] Optionally, the first keyword set includes at least one keyword in the search condition content, and the at least one keyword refers to a word in the search condition content that meets a preset attention condition.
[0079] The search condition content is segmented to obtain a sequence of words arranged in order after the segmentation, and each word in the word sequence corresponds to a sequence number.
[0080] For example, the search condition content is: "Cultivation methods of desert plants", and after splitting, the word sequence obtained is as follows: "desert", "plant", "of", "cultivation", "method", and the sequence numbers are 1 to 5.
[0081] The word sequence is processed by multiple transformer layers in the deep learning model (BERT model), and the self-attention weight matrix is output, where the i-th row element in the self-attention weight matrix is used to indicate the attention distribution of the word with sequence number i, and the j-th column element in the self-attention weight matrix is used to indicate the attention weight assigned to each word in the word sequence to the word with sequence number j, and the i-th row and j-th column element in the self-attention weight matrix represents the attention weight assigned to the word with sequence number i to the word with sequence number j in the word sequence. i and j are positive integers.
[0082] The self-attention weight matrix is adjusted to obtain the weighted self-attention weight matrix.
[0083] The search condition content is feature extracted to obtain the first sentence feature representation, and the weighted self-attention weight matrix is determined based on the product between the first sentence feature representation and the self-attention weight matrix.
[0084] The attention score of each word in the word sequence is determined based on the weighted self-attention weight matrix, and the attention score is used to indicate the importance of each word in the word sequence in the search condition content.
[0085] A first keyword set of the search condition content is determined based on the attention score of each word in the word sequence, wherein the attention score of the words in the first keyword set is greater than a preset score threshold.
[0086] Exemplarily, the search condition content is split into the following sequence [t1, t2...ta] and input into the BERT model, and each tb (vocabulary) is tokenized and converted into a token sequence that the BERT model can understand, where a is a positive integer and b is a positive integer not exceeding a. A special tag [CLS] is added at the beginning of the sequence to aggregate sentence-level semantic information. A [SEP] tag is added at the end of the sequence to indicate the end of the sentence.
[0087] The tokenization result of the input sequence is [101, t1, t2...ta, 102], where 101 and 102 are the tag IDs (identifiers) of [CLS] and [SEP] respectively.
[0088] The output of the BERT model is a vector matrix that contains the feature representation of each tag ID in the sentence. These feature representations are used to represent the hidden state of each tag in the sentence in the last layer of the model.
[0089] Through the BERT model, each tag ID is mapped to the corresponding word feature representation, and an embedding matrix E is obtained with a size of (a+2)×d, where d is the dimension of the word feature representation.
[0090] There are several ways to obtain the first sentence feature representation that can represent the entire search condition content from the output of the BERT model.
[0091] (1) Obtain the feature representation corresponding to the [CLS] tag as the first sentence feature representation: The first sentence feature representation is S, S = E[1,:]. “E[1,:]” is used to access all columns of the second row of the matrix E, “1” means selecting the second row (the index usually starts from 0, and the index of the [CLS] tag in the matrix is 0), and “:” means selecting all columns of this row.
[0092] (2) Aggregate all word feature representations (e.g., by average pooling) to obtain the feature representation of the first sentence: S = mean(E[1:a+1,:], 1), ignoring the [SEP] tag. "[1:a+1,:]" is used to select all columns from the 1st row to the a+1th row from the matrix E. Here a is a specific integer representing the number of tags in the sequence (excluding the end-of-sequence tag). "[1:a+1]" includes rows from the first tag to the a+1th tag, and ":" means selecting all columns of these rows. The mean function is used to calculate the average value of the matrix E in the specified dimension. The parameter "1" in mean(..., 1) means to operate along the second dimension (column) of the matrix E, that is, to average the elements of each row.
[0093] Exemplarily, the output weights of the self-attention layer are extracted from each Transformer layer of the BERT model, and for each tb (vocabulary), the weighted sum of the self-attention weights of all layers is calculated as the attention score of each vocabulary in the word sequence.
[0094] Assume that the self-attention weight matrix of the lth (lowercase L) layer Transformer is A^l, and its size is (a+2)×(a+2). After multiplying the first sentence vector by the self-attention weight matrix, we get the weighted self-attention weight matrix.
[0095] On this basis, for each word tb, the weighted sum of its self-attention weights in all layers (that is, the attention score) can be expressed as: importance(tb)=sum(l=1toL,sum(A^l[:,i])), where L is the number of Transformer layers.
[0096] Apply the Softmax normalization method, convert the attention score of each word into a percentage and sort them, and take the words with the highest scores as the first keyword set.
[0097] For example, the search condition content is divided into 10 words, and the three words with the highest scores are determined as the first keyword set.
[0098] Step 230: perform vocabulary expansion on at least one keyword in the first keyword set to obtain an expanded vocabulary and a second keyword set including the first keyword set and the expanded vocabulary.
[0099] The extended vocabulary and the keywords in the first keyword set meet the preset correlation requirement.
[0100] Optionally, feature extraction is performed on at least one keyword in the first keyword set to obtain at least one corresponding word feature representation.
[0101] A vocabulary and a feature representation library corresponding to the vocabulary are obtained, and a mapping relationship exists between multiple word feature representations in the feature representation library and multiple words in the vocabulary, wherein the k-th word feature representation in the feature representation library is obtained after feature extraction of the k-th word in the vocabulary, and k is a positive integer.
[0102] A synonym feature representation corresponding to at least one word feature representation is determined in the feature representation library.
[0103] Among them, for the mth word feature representation in at least one word feature representation, the similarity between the mth word feature representation and multiple word feature representations in the feature representation library is calculated, and the word feature representation whose similarity reaches a preset threshold is determined as the synonym feature representation corresponding to the mth word feature representation. m is a positive integer.
[0104] Based on the synonym feature representation, a corresponding synonym is determined in the vocabulary as an extended vocabulary, wherein the semantics of the synonym and the semantics of at least one keyword meet a preset similarity requirement. The preset similarity requirement means that for each keyword in the at least one keyword, a vocabulary with the closest vector distance to the keyword is used as a synonym of the keyword.
[0105] Exemplarily, the first keyword set includes two keywords: "fitness" and "recipe". After feature extraction, the corresponding word feature representations are C1 (1, 0, 4, 5) and C2 (5, 2, 1, 3). The feature representation library includes word feature representations K1, K2...K909.
[0106] The vector distances between C1 and K1, K2...K909 are calculated respectively, and it is determined that the vector distance between K501 (1, 0, 4, 4) and C1 is the smallest. Then K501 has the highest similarity with C1, which is the synonym feature representation T1; the vector distances between C2 and K1, K2...K909 are calculated respectively, and the vector distance between K128 (5, 2, 1, 4) and C2 is the smallest. Then K128 has the highest similarity with C2, which is the synonym feature representation T2.
[0107] Among them, the synonym feature indicates that the vocabulary corresponding to T1 is "exercise", and the synonym feature indicates that the vocabulary corresponding to T2 is "recipe". Taking "exercise" and "recipe" as extended vocabulary, the second keyword set obtained includes the following vocabulary: "fitness", "recipe", "exercise", and "recipe".
[0108] Step 240: extract features from the second keyword set to obtain keyword feature representation.
[0109] In some embodiments, the second keyword set may be further screened to improve the accuracy of the matching result, and features may be extracted from the search condition content to obtain the first sentence feature representation.
[0110] The first sentence feature representation and keyword feature representation are processed by a text sorting algorithm to obtain a semantic similarity image.
[0111] The semantic similarity image includes a main node representing the search condition content and multiple sub-nodes representing the vocabulary in the second keyword set. The multiple lines between the main node and the multiple sub-nodes correspond to node weights respectively, and the node weights are used to indicate the semantic similarity between the vocabulary in the second keyword set and the search condition content.
[0112] The node weights in the semantic similarity image are iteratively calculated and updated through the text sorting algorithm.
[0113] When the number of iterations reaches a preset iteration number threshold, the second keyword set is screened based on the node weight to obtain a third keyword set; and feature extraction is performed on the third keyword set to obtain a keyword feature representation.
[0114] Optionally, the semantic similarity (node weight) between each keyword in the second keyword set and the search condition content can be obtained through a text sorting algorithm, and the words in the second keyword set are sorted in order of similarity from high to low based on the semantic similarity, and the words in the top 50% of the sequence are screened out as the words in the third keyword set.
[0115] For example, the second keyword set includes 8 words, namely word 1, word 2...word 8, and the semantic similarities between them and the search condition content are 0.81, 0.43, 0.55, 0.90, 0.91, 0.72, 0.77, and 0.98, respectively. The 8 words are ranked according to the similarity as follows: word 8 (0.98), word 5 (0.91), word 4 (0.90), word 1 (0.91), word 7 (0.77), word 6 (0.72), word 3 (0.55), and word 2 (0.43). The first 50% of the words (including word 8, word 5, word 4, and word 1) are determined as the third keyword set.
[0116] Exemplarily, the feature representation set corresponding to the second keyword set is V = {v1, v2, ..., vn}, where each vi is the d-dimensional feature representation of keyword i, the feature representation of the first sentence is S, and the merged feature representation set is C = {S, v1, v2, ..., vn}.
[0117] The above set C is used as the input of the text sorting algorithm to construct a semantic similarity image, in which nodes represent keywords or search condition content. According to the semantic similarity or co-occurrence relationship between nodes, the edges between nodes are constructed and the initial weights of the edges are assigned. The weight of each node is iteratively calculated and updated through the text sorting algorithm, and the iterative calculation is repeated until the node weight converges or reaches the preset number of iterations. Finally, several keywords with the highest weight are selected to obtain the third keyword set.
[0118] For example, for two keyword feature representations vi and vj, the cosine similarity sim(vi,vj) between them is calculated. In the text sorting algorithm, the node weight Score(vi) is updated as follows:
[0119]
[0120] Where d is the damping coefficient, which is usually set to 0.85, In(vi) is the set of nodes pointing to the vi node, outDegree(vj) is the out-degree of the vj node (the set of all nodes that can be reached from the node vj), and the above formula is iterated until the preset convergence condition is met or the preset number of iterations is reached.
[0121] Step 250: Based on the matching degree between the keyword feature representation and the candidate content in the candidate content library, at least one first content is obtained from the candidate content library as a search result of the search condition content.
[0122] Optionally, feature extraction is performed on candidate contents in the candidate content library to obtain a feature representation pool corresponding to the candidate content library, the candidate contents belong to sentence contents, and there is a mapping relationship between the nth sentence feature representation in the feature representation pool and the nth sentence content in the candidate contents, where n is a positive integer.
[0123] Based on the similarity between the keyword feature representation and the sentence feature representation in the feature representation pool, a candidate content list is determined from the candidate content library, and the candidate content list includes a plurality of first contents.
[0124] Exemplarily, a pre-trained HNSW model (Hierarchical Navigable Small World, high-speed channel model) is used to calculate the keyword feature representation to determine the candidate content list.
[0125] The training process is as follows: obtain a corpus containing multiple text contents, including but not limited to articles, news, poems, etc., perform text preprocessing on the corpus, including removing special contents in the text content (such as punctuation, numbers, special characters, etc.), converting uppercase characters to lowercase, deleting stop words, etc., split the text content into single sentence text [u1, u2, ..., um], use the BERT model to vectorize each sentence, and obtain the sentence vector representation [s1, s2, ..., sm]. Input the vector representation [s1, s2, ..., sm] into the HNSW model for training.
[0126] At the end of the training process, the corpus is processed to obtain a candidate content library, which contains multiple sentence contents. These sentence contents are mapped during the HNSW model training process to form a feature representation pool.
[0127] Optionally, the keyword feature representation is input into a pre-trained HNSW model to calculate a candidate content list.
[0128] Exemplarily, for each feature representation si in the keyword feature representation, the HNSW model is used to search for the sentence feature representation set R(si) most similar to si from the feature representation pool. The cosine similarity between each sentence feature representation vj in R(si) and the keyword feature representation si is calculated, the search results are sorted according to the cosine similarity score, and the K most similar sentence feature representations with the highest ranking are returned, and the candidate content list is determined from the candidate content library based on the sentence content corresponding to the K sentence feature representations.
[0129] Optionally, based on the matching degree between the feature representation of the first sentence and the feature representations of the plurality of first contents, at least one first content is determined from the candidate content list as the search result of the search condition content.
[0130] Exemplarily, the degree of matching between the feature representation of the first sentence and the feature representations of the plurality of first contents is determined based on the cosine distance, the Pearson distance, and the BM25 correlation value.
[0131] Among them, cosine distance is a method to measure the difference between two vectors. The distance is calculated based on the cosine similarity of the vectors. Cosine similarity measures the degree of similarity between two vectors in direction. Pearson distance refers to the complement of the Pearson correlation coefficient, which is used to measure the degree of correlation between two variables. The BM25 (Best Matching 25) algorithm is a relevance scoring algorithm based on the TF-IDF (Term Frequency-Inverse Document Frequency) model, and improves it to consider factors such as document length. The BM25 algorithm evaluates the relevance between a document and a query by calculating the frequency (TF) and inverse document frequency (IDF) of the query terms in the document, and adjusting it based on the document length.
[0132] The sentence feature representation of each first content is obtained, and the first content feature representation set R1 is represented, wherein the i-th feature in R1 is represented as R_i. The search condition content and the multiple first contents are respectively encoded using the BERT model, and the matching score match_score is calculated.
[0133] Calculate the cosine distance cosine, Pearson distance pearson and BM25 correlation value bm25 between S and each feature representation R_i in the R1 set.
[0134] The matching score match_score, cosine distance cosine, Pearson distance pearson and BM25 correlation value bm25 are used as features for the next step of fine ranking and input into the pre-trained fine ranking model LGBMRanker to obtain the sorting results after fine ranking, and the highest several first contents in the sorting results are returned to the user as search results.
[0135] In summary, the method provided by the present application analyzes the search condition content input by the user, and for the search condition content belonging to the sentence type, determines the amount of attention each word obtains from the other words, so as to determine the importance of each word in the sentence, and selects keywords as the basis for expansion. By expanding the vocabulary of keywords in combination with the context, it is possible to more accurately understand the semantics of the search condition content based on the expanded vocabulary, and fully consider the ambiguity, puns and other problems that may exist in the text content. Compared with the related art that only relies on surface keywords for matching and then feeds back search results, it can improve the accuracy of the matching results and feed back results that are more in line with the search requirements.
[0136] In some embodiments, the solution provided by the present application can also expand the vocabulary search conditions input by the user, and return search results to the user based on the expanded vocabulary set. Figure 3It is a flowchart of a content search method provided by another exemplary embodiment of the present application, which includes the following steps.
[0137] Step 310, obtaining search condition content and determining the dimension of the search condition content.
[0138] Users can enter words, sentence fragments, sentences, etc. of any dimension as search condition content, and judge based on the length of the search condition. If the length is less than 5, it is judged as a word; if the length is greater than or equal to 5, the input is judged as a sentence or sentence fragment. Hereinafter, sentences or sentence fragments are referred to as sentence content.
[0139] For example, if the search condition content is as follows: "health, diet, recipes", the length of the search condition content is based on the number of Chinese characters, and if it contains 6 Chinese characters, then the search condition content is a sentence content.
[0140] For example, if the search condition content is as follows: "health, recipe", the length of the search condition content is based on the number of Chinese characters, and if it contains 4 Chinese characters, then the search condition content is vocabulary content.
[0141] Step 321 , when the length of the search condition content indicates that the search condition content belongs to sentence content, attention analysis is performed on the search condition content to determine a first keyword set for the search condition content.
[0142] The search condition content is segmented to obtain a sequence of words arranged in order after the segmentation, and each word in the word sequence corresponds to a sequence number.
[0143] The word sequence is processed through multiple transformer layers in the deep learning model (BERT model) and the self-attention weight matrix is output.
[0144] The BERT model is used to extract features from the search condition content to obtain the feature representation of the first sentence. The self-attention weight matrix is adjusted based on the feature representation of the first sentence. The weighted attention result obtained by each word from other words is calculated as the attention score and the scores are sorted. A preset number of words with higher scores are determined as the first keyword set.
[0145] Step 331: perform vocabulary expansion on at least one keyword in the first keyword set to obtain an expanded vocabulary and a second keyword set including the first keyword set and the expanded vocabulary.
[0146] The first keyword set is mapped into a vector space, and an expanded vocabulary is determined based on distances between a plurality of vectors in the vector space and feature representations of the vocabulary in the first keyword set.
[0147] Step 341: extract features from the second keyword set to obtain keyword feature representation.
[0148] In some embodiments, the words in the second keyword set may be further screened based on a text sorting algorithm to obtain a third keyword set, and feature extraction may be performed on the third keyword set to obtain a keyword feature representation.
[0149] Optionally, the semantic similarity between each keyword in the second keyword set and the search condition content can be obtained through a text sorting algorithm, and the words in the second keyword set are sorted in order of similarity from high to low based on the semantic similarity, and the top 50% of the words in the sequence are screened out as the words in the third keyword set.
[0150] For example, the second keyword set includes 8 words, namely word 1, word 2...word 8, and the semantic similarities between them and the search condition content are 0.81, 0.43, 0.55, 0.90, 0.91, 0.72, 0.77, and 0.98, respectively. The 8 words are ranked according to the similarity as follows: word 8 (0.98), word 5 (0.91), word 4 (0.90), word 1 (0.91), word 7 (0.77), word 6 (0.72), word 3 (0.55), and word 2 (0.43). The first 50% of the words (including word 8, word 5, word 4, and word 1) are determined as the third keyword set.
[0151] Step 351: based on the matching degree between the keyword feature representation and the candidate content in the candidate content library, at least one first content is obtained from the candidate content library as a search result of the search condition content.
[0152] Based on the keyword feature representation, a candidate content list is screened out from the candidate content library, and the candidate content list contains multiple first contents. The multiple first contents in the candidate content list are finely ranked based on the cosine distance, Pearson distance and BM25 correlation value. According to the fine ranking result, the first content with the highest ranking is used as the search result that needs to be finally returned to the user end.
[0153] Step 322: When the length of the search condition content indicates that the search condition content belongs to vocabulary content, the search condition content is modified to obtain a modified vocabulary.
[0154] Firstly, homophone correction is performed on the search conditions to correct the homophone typos in the search conditions and obtain the corrected vocabulary.
[0155] For example, the search condition content is the word "nutritious and good food", obtain a pre-prepared word list, which contains a variety of common words, and find homophones in the word list based on the pronunciation of "nutritious diet" to determine whether there are typos.
[0156] After searching, it was found that "nutrition" existed in the vocabulary, but "good food" did not exist. Based on the pronunciation of "good food", the homonym "diet" was found, and the search conditions were corrected, and the corrected vocabulary was "nutrition diet".
[0157] Step 332, expand the revised vocabulary to obtain an expanded vocabulary and an expanded vocabulary set including the expanded vocabulary and search condition content.
[0158] Keyword association is performed on the revised vocabulary to obtain an expanded vocabulary, and the expanded vocabulary is combined with the revised vocabulary to obtain an expanded vocabulary set.
[0159] The semantic similarity or identity between the extended vocabulary and the revised vocabulary can be extended in terms of feature representation as well as text semantics.
[0160] For example, the revised vocabulary is split into "nutrition" and "diet", and articles containing the keywords "nutrition" and "diet" are searched in the preset corpus to further determine whether the revised vocabulary appears in the same context with a frequency reaching a first preset threshold, and other vocabulary with a higher frequency (the frequency of appearance reaches a second preset threshold) is searched for in the corresponding articles according to the context as the extended vocabulary.
[0161] For example, "nutrition" and "diet" frequently appear in the context of articles on health topics. Other common high-frequency words in this type of topic include "balance" and "vegetables". In this case, "balance" and "vegetables" can be used as extended words.
[0162] In some embodiments, the expanded vocabulary can also be determined by the distance between the feature representations.
[0163] Optionally, feature extraction is performed on the corrected vocabulary to obtain a corrected word feature representation. A vocabulary and a feature representation library corresponding to the vocabulary are obtained, and a mapping relationship exists between multiple word feature representations in the feature representation library and multiple words in the vocabulary, wherein the k-th word feature representation in the feature representation library is obtained after feature extraction of the k-th word in the vocabulary, and k is a positive integer. A synonym feature representation corresponding to the corrected word feature representation is determined in the feature representation library.
[0164] Among them, for the mth word feature representation in the modified word feature representation, the similarity between the mth word feature representation and multiple word feature representations in the feature representation library is calculated, and the word feature representation whose similarity reaches a preset threshold is determined as the synonym feature representation corresponding to the mth word feature representation. m is a positive integer.
[0165] Based on the synonym feature representation, corresponding synonyms are determined in the vocabulary as extended vocabulary, wherein the semantics of the synonyms and the semantics of the revised vocabulary meet a preset similarity requirement.
[0166] Exemplarily, the pre-trained Word2Vec model is used to extract features of the modified vocabulary to obtain a modified word feature representation. Since the number of modified word feature representations is at least one, the modified word feature representation set is represented as [w1, w2, ..., wm]. For any word in the set, the Word2Vec model maps it to a d-dimensional feature representation.
[0167] Based on the similarity calculation function of the Word2Vec model, the feature representation with the highest similarity to each feature representation in the modified word feature representation set is determined from the feature representation library as the synonym feature representation.
[0168] Based on the vocabulary corresponding to the synonym feature representation, the synonyms of the revised vocabulary are determined, and the synonyms and the revised vocabulary are used as an extended vocabulary set.
[0169] It is worth noting that since the semantics of each revised vocabulary may be similar, it is necessary to filter the synonyms obtained based on the synonym feature representation to avoid synonym duplication.
[0170] Step 342: extract features from the extended vocabulary set to obtain extended word feature representation.
[0171] In some embodiments, the expanded vocabulary set may be further screened to improve the accuracy of the matching results.
[0172] Feature extraction is performed on all words in the extended vocabulary set to obtain extended word feature representations corresponding to the extended vocabulary and search word feature representations corresponding to the search condition content.
[0173] The search term feature representation and the extended term feature representation are processed by a text sorting algorithm to obtain a semantic similarity image.
[0174] The semantic similarity image includes a main node representing the search condition content and multiple sub-nodes representing extended vocabulary. Multiple links between the main node and the multiple sub-nodes correspond to node weights respectively, and the node weights are used to indicate the semantic similarity between the extended vocabulary and the search condition content.
[0175] The node weights in the semantic similarity image are iteratively calculated and updated through the text sorting algorithm.
[0176] When the number of iterations reaches a preset iteration number threshold, the expanded vocabulary set is screened based on the node weights to obtain a matching vocabulary set.
[0177] Feature extraction is performed on the matching vocabulary set to obtain matching word feature representations, which can be used to match candidate content in the candidate content library, thereby determining search results to be fed back to the user.
[0178] Step 352: Based on the matching degree between the extended word feature representation and the candidate content in the candidate content library, at least one second content is obtained from the candidate content library as a search result of the search condition content.
[0179] A candidate content list is screened out from the candidate content library based on the extended word feature representation (or the matching word feature representation), the candidate content list includes a plurality of second contents, and the second contents may be expressed in the form of words, sentences or sentence fragments.
[0180] The multiple second-person contents in the candidate content list are finely ranked based on the cosine distance, Pearson distance and BM25 correlation value, and the second content with the highest ranking is used as the search result that finally needs to be returned to the user end according to the fine ranking result.
[0181] It is worth noting that the contents executed from step 321 to step 351 are in parallel with the contents executed from step 322 to step 352, that is, the contents from step 321 to step 351 may be executed first, and then the contents from step 322 to step 352; or the contents from step 322 to step 352 may be executed first, and then the contents from step 321 to step 351 may be executed; or the contents from step 321 to step 351 and the contents from step 322 to step 352 may be executed at the same time.
[0182] To sum up, the content search method provided by the present application can simultaneously meet the needs of result feedback for multi-dimensional search conditions, analyze the vocabulary in the search condition content in combination with the context, fully understand the possible puns, ambiguities and other phenomena in the text content, and expand the search condition content to obtain search results with higher matching degrees, thereby improving the content search feedback effect.
[0183] Figure 4 is a structural block diagram of a content search device provided by an exemplary embodiment of the present application, such as Figure 4 As shown, the device includes the following parts.
[0184] An acquisition module 410 is used to acquire search condition content, where the search condition content is content input by a user for searching;
[0185] The attention analysis module 420 is used to perform attention analysis on the search condition content when the length of the search condition content indicates that the search condition content belongs to sentence content, and determine a first keyword set of the search condition content, wherein the first keyword set includes at least one keyword in the search condition content, and the at least one keyword refers to a word in the search condition content that meets a preset attention condition;
[0186] A vocabulary expansion module 430, configured to perform vocabulary expansion on at least one keyword in the first keyword set to obtain an expanded vocabulary and a second keyword set including the first keyword set and the expanded vocabulary, wherein the expanded vocabulary meets a preset relevance requirement with the keywords in the first keyword set;
[0187] A feature extraction module 440 is used to extract features from the second keyword set to obtain a keyword feature representation;
[0188] The search result determination module 450 is configured to obtain at least one first content from the candidate content library as a search result of the search condition content based on the matching degree between the keyword feature representation and the candidate content in the candidate content library.
[0189] In an optional embodiment, the attention analysis module 420 is further used to perform word segmentation on the search condition content to obtain a word sequence arranged in order after the segmentation, and each word in the word sequence corresponds to a serial number; the word sequence is processed by multiple transformer layers in a deep learning model, and a self-attention weight matrix is output, wherein the i-th row element in the self-attention weight matrix is used to indicate the attention distribution of the word with serial number i, and the j-th column element in the self-attention weight matrix is used to indicate the attention weight assigned to each word in the word sequence to the word with serial number j, and the i-th row and j-th column in the self-attention weight matrix are The element represents the attention weight assigned by the word with sequence number i to the word with sequence number j in the word sequence; i and j are positive integers; the self-attention weight matrix is adjusted to obtain a weighted self-attention weight matrix; the attention score of each word in the word sequence is determined based on the weighted self-attention weight matrix, and the attention score is used to indicate the importance of each word in the word sequence in the search condition content; the first keyword set of the search condition content is determined based on the attention score of each word in the word sequence, wherein the attention score of the words in the first keyword set is greater than a preset score threshold.
[0190] In an optional embodiment, the attention analysis module 420 is further used to perform feature extraction on the search condition content to obtain a first sentence feature representation; and determine a weighted self-attention weight matrix based on the product of the first sentence feature representation and the self-attention weight matrix.
[0191] In an optional embodiment, the vocabulary expansion module 430 is further used to perform feature extraction on the at least one keyword in the first keyword set to obtain at least one corresponding word feature representation; obtain a vocabulary and a feature representation library corresponding to the vocabulary, wherein a mapping relationship exists between multiple word feature representations in the feature representation library and multiple words in the vocabulary, wherein the kth word feature representation in the feature representation library is obtained after feature extraction of the kth word in the vocabulary, and k is a positive integer; determine a synonym feature representation corresponding to the at least one word feature representation in the feature representation library; determine a corresponding synonym in the vocabulary as the extended vocabulary based on the synonym feature representation, wherein the semantics of the synonym and the semantics of the at least one keyword meet a preset similarity requirement.
[0192] In an optional embodiment, the vocabulary expansion module 430 is further used to calculate the similarity between the mth word feature representation in the at least one word feature representation and the multiple word feature representations in the feature representation library, and determine the word feature representation whose similarity reaches a preset threshold as the synonym feature representation corresponding to the mth word feature representation; m is a positive integer.
[0193] In an optional embodiment, the feature extraction module 440 is further used to extract features from the search condition content to obtain a first sentence feature representation;
[0194] like Figure 5 As shown, the device also includes:
[0195] The vocabulary screening module 460 is used to process the first sentence feature representation and the keyword feature representation through a text sorting algorithm to obtain a semantic similarity image, wherein the semantic similarity image includes a main node representing the search condition content and multiple sub-nodes representing the vocabulary in the second keyword set, and multiple links between the main node and the multiple sub-nodes correspond to node weights respectively, and the node weights are used to indicate the semantic similarity between the vocabulary in the second keyword set and the search condition content respectively; iteratively calculate and update the node weights in the semantic similarity image through the text sorting algorithm; when the number of iterations reaches a preset iteration number threshold, the second keyword set is screened based on the node weights to obtain a third keyword set;
[0196] The feature extraction module 440 is further configured to extract features from the third keyword set to obtain the keyword feature representation.
[0197] In an optional embodiment, the search result determination module 450 is further used to perform feature extraction on the candidate content in the candidate content library to obtain a feature representation pool corresponding to the candidate content library, the candidate content belongs to sentence content, and there is a mapping relationship between the nth sentence feature representation in the feature representation pool and the nth sentence content in the candidate content, where n is a positive integer; based on the similarity between the keyword feature representation and the sentence feature representation in the feature representation pool, a candidate content list is determined from the candidate content library, and the candidate content list includes multiple first contents; based on the matching degree between the first sentence feature representation and the feature representations of the multiple first contents, the at least one first content is determined from the candidate content list as the search result of the search condition content.
[0198] In summary, the content search device provided by the present application can analyze the search condition content input by the user, and for the search condition content belonging to the sentence type, determine the amount of attention each word has obtained from the other words, so as to determine the importance of each word in the sentence, and select keywords as the basis for expansion. Combining the context to expand the keywords, it is possible to more accurately understand the semantics of the search condition content based on the expanded vocabulary, and fully consider the possible ambiguity and pun problems in the text content. Compared with the related art that only relies on surface keywords for matching and feedback of search results, it can improve the accuracy of the matching results and feedback results that are more in line with the search requirements.
[0199] It should be noted that the content search device provided in the above embodiment is only illustrated by the division of the above functional modules. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the content search device provided in the above embodiment and the content search method embodiment belong to the same concept. The specific implementation process is detailed in the method embodiment and will not be repeated here.
[0200] Figure 6The block diagram of a computer device 600 provided by an exemplary embodiment of the present application is shown. The computer device 600 may be: a smart phone, a tablet computer, an MP3 player (Moving Picture Experts Group Audio Layer III, Moving Picture Experts Group Audio Layer 3), an MP4 player (Moving Picture Experts Group Audio Layer IV, Moving Picture Experts Group Audio Layer 4), a laptop computer or a desktop computer. The computer device 600 may also be called a user device, a portable terminal, a laptop terminal, a desktop terminal or other names.
[0201] Typically, the computer device 600 includes a processor 601 and a memory 602 .
[0202] The processor 601 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 601 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). The processor 601 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the awake state, also known as a CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 601 may be integrated with a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 601 may also include an AI (Artificial Intelligence) processor, which is used to process computing operations related to machine learning.
[0203] The memory 602 may include one or more computer-readable storage media, which may be non-transitory. The memory 602 may also include a high-speed random access memory, and a non-volatile memory, such as one or more disk storage devices, flash memory storage devices. In some embodiments, the non-transitory computer-readable storage medium in the memory 602 is used to store at least one instruction, which is used to be executed by the processor 601 to implement the content search method provided in the method embodiment of the present application.
[0204] In some embodiments, the computer device 600 further includes some other components 603, and the type and quantity of the other components 603 can be selected based on the functional requirements of the computer device 600. Those skilled in the art will appreciate that Figure 6 The structure shown in the figure does not constitute a limitation on the computer device 600, and the computer device 600 may include more or less components than those shown in the figure, or combine some components, or adopt a different arrangement of components.
[0205] Optionally, the computer readable storage medium may include: a read-only memory (ROM), a random access memory (RAM), a solid state drive (SSD), or an optical disk. Among them, the random access memory may include a resistance random access memory (ReRAM) and a dynamic random access memory (DRAM). The serial numbers of the above embodiments of the present application are only for description and do not represent the advantages and disadvantages of the embodiments.
[0206] An embodiment of the present application also provides a computer device, which includes a processor and a memory, wherein the memory stores at least one instruction, at least one program, a code set, or an instruction set, and the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by the processor to implement a content search method as described in any of the above embodiments of the present application.
[0207] An embodiment of the present application also provides a computer-readable storage medium, in which at least one instruction, at least one program, a code set or an instruction set is stored. The at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by a processor to implement a content search method as described in any of the above embodiments of the present application.
[0208] The embodiment of the present application also provides a computer program product or a computer program, which includes a computer instruction stored in a computer-readable storage medium. The processor of the computer device reads the computer instruction from the computer-readable storage medium, and the processor executes the computer instruction, so that the computer device executes the content search method described in any of the above embodiments.
[0209] A person skilled in the art will understand that all or part of the steps to implement the above embodiments may be accomplished by hardware or by instructing related hardware through a program, and the program may be stored in a computer-readable storage medium, and the above-mentioned storage medium may be a read-only memory, a disk or an optical disk, etc.
[0210] The above description is only an optional embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.
Claims
1. A content search method, characterized in that: The method comprises: Acquire search condition content, where the search condition content is content input by a user for searching; When the length of the search condition content indicates that the search condition content belongs to sentence content, performing attention analysis on the search condition content to determine a first keyword set of the search condition content, wherein the first keyword set includes at least one keyword in the search condition content, and the at least one keyword refers to a word in the search condition content that meets a preset attention condition; Performing vocabulary expansion on at least one keyword in the first keyword set to obtain an expanded vocabulary and a second keyword set including the first keyword set and the expanded vocabulary, wherein the expanded vocabulary meets a preset correlation requirement with the keywords in the first keyword set; Performing feature extraction on the second keyword set to obtain keyword feature representation; Based on the matching degree between the keyword feature representation and the candidate content in the candidate content library, at least one first content is obtained from the candidate content library as a search result of the search condition content.
2. The method according to claim 1, characterized in that The performing attention analysis on the search condition content to determine a first keyword set of the search condition content includes: Perform word segmentation on the search condition content to obtain a word sequence arranged in order after segmentation, wherein each word in the word sequence corresponds to a sequence number; The word sequence is processed by multiple transformer layers in a deep learning model, and a self-attention weight matrix is output, wherein the i-th row element in the self-attention weight matrix is used to indicate the attention distribution of the word with sequence number i, the j-th column element in the self-attention weight matrix is used to indicate the attention weight assigned by each word in the word sequence to the word with sequence number j, and the i-th row and j-th column element in the self-attention weight matrix represents the attention weight assigned by the word with sequence number i to the word with sequence number j in the word sequence; i and j are positive integers; Adjusting the self-attention weight matrix to obtain a weighted self-attention weight matrix; Determining an attention score for each word in the word sequence based on the weighted self-attention weight matrix, wherein the attention score is used to indicate the importance of each word in the word sequence in the search condition content; The first keyword set of the search condition content is determined based on the attention score of each word in the word sequence, wherein the attention score of the words in the first keyword set is greater than a preset score threshold.
3. The method according to claim 2, characterized in that The adjusting the self-attention weight matrix to obtain a weighted self-attention weight matrix includes: Extracting features from the search condition content to obtain a first sentence feature representation; A weighted self-attention weight matrix is determined based on the product between the first sentence feature representation and the self-attention weight matrix.
4. The method according to claim 1, characterized in that The step of performing vocabulary expansion on at least one keyword in the first keyword set to obtain an expanded vocabulary includes: Performing feature extraction on the at least one keyword in the first keyword set to obtain at least one corresponding word feature representation; Obtain a vocabulary and a feature representation library corresponding to the vocabulary, wherein a mapping relationship exists between multiple word feature representations in the feature representation library and multiple words in the vocabulary, wherein the k-th word feature representation in the feature representation library is obtained after feature extraction of the k-th word in the vocabulary, and k is a positive integer; Determining, in the feature representation library, a synonym feature representation corresponding to the at least one word feature representation; Based on the synonym feature representation, a corresponding synonym is determined in the vocabulary as the extended vocabulary, wherein the semantics of the synonym and the semantics of the at least one keyword meet a preset similarity requirement.
5. The method according to claim 4, characterized in that The determining, in the feature representation library, a synonym feature representation corresponding to the at least one word feature representation comprises: For the mth word feature representation in the at least one word feature representation, calculate the similarity between the mth word feature representation and the multiple word feature representations in the feature representation library, and determine the word feature representation whose similarity reaches a preset threshold as the synonym feature representation corresponding to the mth word feature representation; m is a positive integer.
6. The method according to any one of claims 1 to 5, characterized in that: After extracting features from the second keyword set to obtain keyword feature representation, the method further includes: Extracting features from the search condition content to obtain a first sentence feature representation; Processing the first sentence feature representation and the keyword feature representation by a text sorting algorithm to obtain a semantic similarity image, wherein the semantic similarity image includes a main node representing the search condition content and a plurality of sub-nodes representing words in the second keyword set, and a plurality of links between the main node and the plurality of sub-nodes respectively correspond to node weights, and the node weights are used to indicate the semantic similarities between the words in the second keyword set and the search condition content; Iteratively calculating and updating the node weights in the semantic similarity image by using the text sorting algorithm; When the number of iterations reaches a preset iteration number threshold, the second keyword set is screened based on the node weight to obtain a third keyword set; The step of extracting features from the second keyword set to obtain keyword feature representation includes: Feature extraction is performed on the third keyword set to obtain the keyword feature representation.
7. The method according to any one of claims 1 to 5, characterized in that: The acquiring at least one first content from the candidate content library as a search result of the search condition content based on the matching degree between the keyword feature representation and the candidate content in the candidate content library comprises: Extracting features from the candidate content in the candidate content library to obtain a feature representation pool corresponding to the candidate content library, wherein the candidate content belongs to sentence content, and an n-th sentence feature representation in the feature representation pool has a mapping relationship with an n-th sentence content in the candidate content, where n is a positive integer; Determining a candidate content list from the candidate content library based on the similarity between the keyword feature representation and the sentence feature representation in the feature representation pool, wherein the candidate content list includes a plurality of first contents; Based on the matching degree between the first sentence feature representation and the feature representations of the plurality of first contents, the at least one first content is determined from the candidate content list as the search result of the search condition content.
8. A content search device, characterized in that: The device comprises: An acquisition module, used to acquire search condition content, where the search condition content is content input by a user for searching; an attention analysis module, configured to, when the length of the search condition content indicates that the search condition content belongs to sentence content, perform attention analysis on the search condition content to determine a first keyword set of the search condition content, wherein the first keyword set includes at least one keyword in the search condition content, and the at least one keyword refers to a word in the search condition content that meets a preset attention condition; a vocabulary expansion module, configured to perform vocabulary expansion on at least one keyword in the first keyword set to obtain an expanded vocabulary, and a second keyword set including the first keyword set and the expanded vocabulary, wherein the expanded vocabulary meets a preset relevance requirement with the keywords in the first keyword set; A feature extraction module, used to extract features from the second keyword set to obtain keyword feature representation; The search result determination module is used to obtain at least one first content from the candidate content library as the search result of the search condition content based on the matching degree between the keyword feature representation and the candidate content in the candidate content library.
9. A computer device, characterized in that: The computer device includes a processor and a memory, wherein the memory stores at least one program, and the at least one program is loaded and executed by the processor to implement the content search method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The storage medium stores at least one program, and the at least one program is loaded and executed by the processor to implement the content search method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Text classification method based on multi-source features, terminal equipment and storage medium
CN114444497A
Video search result-free processing method, system and equipment and medium
CN117453950A
Data processing method and device, computer equipment and storage medium
CN118210874A
Image retrieval method and apparatus
KR102201390B1
System and methods for semiautomatic generation and tuning of natural language interaction applications
US20130268260A1