Page association method, device and storage medium
By determining the associated information through semantic analysis and implicit semantic encoding models of the displayed pages, the problem of insufficient accuracy in page association was solved, resulting in higher accuracy in page association and improved website ranking.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- ALIBABA CLOUD COMPUTING CO LTD
- Filing Date
- 2022-06-06
- Publication Date
- 2026-07-24
AI Technical Summary
The accuracy of page association in existing technologies is insufficient, which affects the website's ranking and exposure in search engines.
By performing semantic analysis on the page information to be displayed, semantic vectors are obtained, implicit semantic coding models are used to determine related information, and page interlinking is performed to generate related content aggregation pages.
It improved the accuracy of page associations, thereby increasing the website's ranking and traffic in search engines.
Smart Images

Figure CN115168685B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of Internet technology, and in particular to a page association method, device and storage medium. Background Technology
[0002] Search Engine Optimization (SEO) is a technique that analyzes the ranking patterns of search engines to understand how various search engines search, crawl web pages, and determine the ranking of search results for specific keywords. Page relevance is a crucial aspect of SEO. Proper relevance between website pages improves page quality and helps improve a website's ranking in search engines. Therefore, improving the accuracy of page relevance is a technical issue that website developers continuously research. Summary of the Invention
[0003] This application provides a page association method, device, and storage medium to improve the accuracy of page association.
[0004] This application provides a page association method, including:
[0005] Get the page information of the page to be displayed;
[0006] Semantic extraction is performed on the page information to determine the semantic vector of the page information;
[0007] Based on the semantic vector of the page information, determine the associated information of the page information;
[0008] Based on the association information of the page information, determine the associated pages of the page to be displayed;
[0009] Link the page to be displayed and the associated page to obtain a related content aggregation page.
[0010] This application embodiment also provides a computing device, including: a memory and a processor; wherein, the memory is used to store computer programs;
[0011] The processor is coupled to the memory and is used to execute the computer program for performing the steps in the page association method described above.
[0012] This application also provides a computer-readable storage medium storing computer instructions, which, when executed by one or more processors, cause the one or more processors to perform the steps in the page association method described above.
[0013] In this embodiment, semantic analysis is performed on the page information of the page to be displayed to obtain a semantic vector of the page information; then, the association information of the page information is determined based on the semantic vector of the page information. Therefore, the determined association information of the page information incorporates the semantics of the page information. The association information of the page information is related to the semantics of the page information, making the relevance between the association information and the content of the page information more accurate. Furthermore, the accuracy of the associated pages of the page to be displayed determined by the relationship information of the page information is relatively high. Thus, the content relevance of the related content aggregation page obtained by linking the page to be displayed and the associated pages is more accurate. Therefore, the page association method provided in this embodiment helps to improve the accuracy of page association. Attached Figure Description
[0014] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:
[0015] Figure 1 A flowchart illustrating the page association method provided in an embodiment of this application;
[0016] Figure 2 This is a schematic diagram of an aggregation page provided in an embodiment of this application;
[0017] Figure 3 This is a schematic diagram of the architecture of the page association system provided in the embodiments of this application;
[0018] Figure 4a and Figure 4b This is a schematic diagram of the vector recall process provided in an embodiment of this application;
[0019] Figure 5 A schematic diagram of compound words provided for embodiments of this application;
[0020] Figure 6 A schematic diagram of the structure of a computing device provided in an embodiment of this application. Detailed Implementation
[0021] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0022] To improve the accuracy of page association, in some embodiments of this application, semantic analysis is performed on the page information of the page to be displayed to obtain a semantic vector of the page information; then, the association information of the page information is determined based on the semantic vector of the page information. Therefore, the determined association information of the page information incorporates the semantics of the page information. The association information of the page information is related to the semantics of the page information, making the relevance between the association information and the content of the page information more accurate. Furthermore, the accuracy of the associated pages of the page to be displayed determined through the relationship information of the page information is relatively high. Thus, the content relevance of the related content aggregation page obtained by linking the page to be displayed and the associated pages is more accurate. Therefore, the page association method provided in the embodiments of this application helps to improve the accuracy of page association.
[0023] The technical solutions provided by the various embodiments of this application are described in detail below with reference to the accompanying drawings.
[0024] It should be noted that the same reference numerals denote the same object in the following figures and embodiments. Therefore, once an object is defined in one figure or embodiment, it does not need to be discussed further in subsequent figures and embodiments.
[0025] Figure 1 This is a flowchart illustrating the page association method provided in an embodiment of this application. Figure 1 As shown, the main methods for linking pages include:
[0026] 101. Obtain the page information of the page to be displayed.
[0027] 102. Perform semantic extraction on the page information to determine the semantic vector of the page information.
[0028] 103. Determine the associated information of the page based on the semantic vector of the page information.
[0029] 104. Based on the association information of the page information, determine the associated pages of the page to be displayed.
[0030] 105. Link the pages to be displayed and related pages to obtain a related content aggregation page.
[0031] Page linking refers to embedding other pages within a single page, facilitating information retrieval and automatic crawling by web crawlers. Proper page linking and the logical categorization of related content help improve webpage quality and website ranking in search engines, thereby increasing website exposure and traffic. Therefore, the accuracy and appropriateness of page linking are crucial aspects of SEO.
[0032] In this embodiment, to improve the accuracy of page association, page information of the page to be displayed is obtained in step 101. In this embodiment, the page to be displayed can be any page of the website. Pages may include: the website's homepage, detail pages, and aggregation pages, etc. An aggregation page is a new list page where the website rearranges and combines existing articles or product pages into a new list based on a certain theme or keywords; it can group web pages explaining the same topic onto the same page. For example... Figure 2 The web pages explaining SEO techniques are grouped into a new list page to facilitate user information retrieval and automatic crawling by web crawlers. Figure 2 This is a schematic diagram of an aggregated page. Figure 2 In the diagram, area 1 is the display area corresponding to the page to be displayed; area 2 is the identification information of the associated page of the page to be displayed.
[0033] The page information of the page to be displayed refers to the page content information (i.e., the page text "Page") of the page to be displayed. This page content information may include: the page's category, title, URL, and summary. In some embodiments, the page information may further include: context information. This context information may include: the inbound anchors of the page, user behavior information related to the page, corresponding keywords, and time information of the page content. User behavior-related information may include: click query information. For example, the context information includes the inbound anchors of the webpage and user behavior information related to the webpage.
[0034] For a page to be displayed, its Uniform Resource Locator (URL) can be used to retrieve the page's content information from the page's server or search results. Alternatively, the URL can be used to collect the page's inbound anchor information from the search engine, incorporating this information into the page's contextual information. Correspondingly, the query string and corresponding click information for the page to be displayed can also be collected from the search engine, serving as user behavior information related to the page. For example, for page 1, if a user enters the string "xxxx" into a search engine and page 1 appears in the search results, then "xxxx" is the query string for page 1. If the user clicks a link to page 1 in the search results, the click count for the query string "xxxx" is incremented by 1.
[0035] In this embodiment of the application, the page information of the page to be displayed can be carried in the access request for the page to be displayed. Accordingly, the page information of the page to be displayed can be obtained from the access request.
[0036] In step 102, semantic extraction can be performed on the page information of the page to be displayed to determine the semantic vector of the page information. In this embodiment, the context information may include semantically extractable information, such as continuation words and inbound anchor information; of course, it may also include information that cannot be semantically extracted, such as time information. Based on this, in step 102, semantic extraction can be performed on the page content information of the page to be displayed and the vectorizable context information to determine the semantic vector of the page information.
[0037] In this application, the specific implementation method for semantic extraction of page information is not limited. In some embodiments, a semantic extraction model can be used to extract semantics from page information. The semantic extraction model can be a latent semantic encoding model. Latent semantic encoding models can uncover potential features between page information. Network machine learning techniques can be used to discover relationships between topics from a text collection.
[0038] In this application, the specific implementation of the latent semantic encoding model is not limited. In some embodiments, the latent semantic encoding model can be implemented as a representation-focused architecture model or an interaction-focused architecture model. The basic assumption of the representation-focused architecture model is that relevance depends on the combined meaning of the input text. The basic assumption of the interaction-focused architecture model is that relevance is essentially a relationship between the input texts, therefore learning directly from interactions rather than from individual representations is more effective. Furthermore, the representation-focused architecture model can be a Deep Structured Semantic Model (DSSM), an Architecture-I (Arc-I) model, a Convolutional Neural Tensor Network (CNTN), or a Convolutional Latent Semantic Model (CLSM), but is not limited to these. The interactive central architecture model can be a deep relevance matching model (DRMN), a deep text matching model (K-NRM) based neural model for document ranking, Architecture-II (Arc-II), or Match-SRNN, but is not limited to these.
[0039] Based on the implicit semantic coding model, step 102 can be implemented as follows: using the implicit semantic coding model to extract the semantic information of the page to be displayed, so as to obtain the semantic vector of the page information.
[0040] In some embodiments, the latent semantic encoding model adopts the DSSM dual-tower model architecture. In the DSSM dual-tower model architecture, the page information is mapped to a low-dimensional vector space through the first input layer, and then transformed into a low-dimensional vector for the deep learning network. In the representation layer, a vector transformation model (such as the bag-of-words model) is used to map the low-dimensional vector to a high-dimensional space. In addition, two hidden layers are added to enhance the model's expressive power, transforming the high-dimensional vector into a semantic vector (such as a 128-dimensional semantic vector).
[0041] In this embodiment of the application, before using the Latent Semantic Coding (LSC) model to extract semantic information from the page information of the page to be displayed, it is necessary to train the LSC model. The training process of the LSC model can be an offline process. The training process of the LSC model is illustrated below.
[0042] In this embodiment, known semantically related positive sample pairs and known semantically unrelated negative sample pairs can be obtained. A positive sample pair includes two texts that are known to be semantically related; a negative sample pair includes two texts that are known to be semantically unrelated. These positive and negative sample pairs can be sample pairs obtained from any corpus. For example, they can be positive and negative sample pairs obtained from any sentence, any article, or any website content, etc. Preferably, the positive and negative sample pairs are positive and negative sample pairs obtained from corpus related to the application scenario of the website where the webpage to be displayed is located, which can improve the accuracy of subsequent model training.
[0043] In some embodiments, for a corpus text, central words can be extracted to obtain the central words of the corpus text; then, duplicate deletion and merging of the central words can be performed to obtain a text corpus. Further, positive and negative sample pairs are determined based on the semantics of each text in the text corpus. These positive and negative sample pairs can be determined through manual annotation, etc.
[0044] After obtaining positive and negative sample pairs, the training objective is to minimize the loss function. The initial latent semantic encoding model is then trained using both positive and negative sample pairs to obtain the final latent semantic encoding model. The loss function is determined based on the differences between the relevance of positive sample pairs and their ground truth values, as well as the differences between the relevance of negative sample pairs and their ground truth values. For positive sample pairs, the ground truth value for the relevance between the text within the pair is 1; for negative sample pairs, the ground truth value for the relevance between the text within the pair can be 0.
[0045] Optionally, the loss function can be expressed as the correlation between the positive and negative sample pairs output by the model training, the mean square error of the correlation truth values of the corresponding positive and negative sample pairs, etc.
[0046] For the DSSM dual-tower model, during model training, for any pair of positive or negative samples, the text in the sample pair is mapped to a low-dimensional vector space through the first input layer, transforming it into a low-dimensional vector for the deep learning network. In the representation layer, a vector transformation model (such as the bag-of-words model) is used to map the low-dimensional vector to a high-dimensional space. Two hidden layers are then added to enhance the model's expressive power, transforming the high-dimensional vector into semantic vectors of the text (such as 128-dimensional semantic vectors). Finally, a matching layer calculates the similarity between the two corresponding semantic vectors of the sample pair. Optionally, the similarity between semantic vectors can be represented by calculating the distance between the two semantic vectors. The smaller the distance between semantic vectors, the higher the similarity. The distance between semantic vectors can be cosine distance, Euclidean distance, Mahalanobis distance, Mahalanobis distance, or Hamming distance, but is not limited to these.
[0047] In this embodiment, to measure the correlation between semantic vectors, the distance between two semantic vectors can be normalized and converted into a posterior probability. For example, an activation function can be used to normalize the distance between two semantic vectors, resulting in a posterior probability that characterizes the correlation between the semantic vectors. The activation function can be a SoftMax function, a Sigmoid function, or a ReLU function, etc.
[0048] After training the aforementioned model to obtain the latent semantic encoding model, the latent semantic encoding model can be used to extract semantics from the text in the aforementioned text corpus to obtain a text vector library. In this embodiment, the text in the aforementioned text corpus may contain known aggregate word text. To improve the richness of the text vector library and the accuracy of subsequently determining the association information of page information, the text corpus for obtaining positive and negative sample pairs may include: the content of each page on the website where the page to be displayed is located, etc. Furthermore, an index for the text vector library can be established through a semantic vector retrieval system. In this way, semantic vectors can be retrieved subsequently through the semantic vector retrieval system.
[0049] The above step 102, in its specific implementation, can be derived by... Figure 3 The recommendation system is implemented in the above. The recommendation system may include an offline training module and an online recommendation module. The offline training module can be used to train the implicit semantic encoding model described above; the online recommendation module can be used to perform real-time semantic extraction of page information from the page to be displayed in step 102 above.
[0050] In this embodiment of the application, after obtaining the latent semantic encoding model through the above-mentioned model training, in the specific implementation of step 102, the page information of the page to be displayed can be input into the latent semantic encoding model, and the latent semantic encoding model can be used to extract the semantics of the page information to obtain the semantic vector of the page information. In the DSSM dual-tower model, the semantic vector of the page information is a multi-dimensional vector output by the expression layer of the DSSM dual-tower model, such as a 128-dimensional vector.
[0051] Since the page information to be displayed may include redundant data, it can be preprocessed before inputting it into the latent semantic encoding model. For example, the page information can be used to extract central words to obtain the central words contained in the page information; further, the central words contained in the page information can be cleaned, and then the cleaned central words can be input into the latent semantic encoding model; the latent semantic encoding model can then be used to extract the semantics of the cleaned central words to obtain the semantic vector corresponding to the page information.
[0052] Furthermore, in step 103, the associated information of the page information can be determined based on the semantic vector of the page information. In this embodiment, the specific implementation of determining the associated information of the page information based on the semantic vector is not limited.
[0053] Optionally, the semantic vectors of the page information can be used to perform vector retrieval in a text vector library to select candidate text vectors corresponding to the page information from the text vector library. In this embodiment, the text vector library can be obtained by semantic extraction from a known aggregated page thesaurus using a latent semantic coding model. To improve the richness of the text vector library and the accuracy of determining the association information of the page information, the text corpus for obtaining positive and negative sample pairs may include: the content of each page of the website where the page to be displayed is located, etc. The generation process of the text vector library can be an offline process, which can be performed by... Figure 3 The offline training module in the recommendation system completes the task. Step 103, real-time vector recall, can be performed by the online recommendation module in the recommendation system.
[0054] Specifically, in some embodiments, the similarity between the semantic vector of the page information and each text vector in the text vector library can be calculated; and candidate text vectors can be selected from the text vector library based on the similarity between the semantic vector of the page information and each text vector in the text vector library. Preferably, text vectors from the text vector library whose similarity to the semantic vector of the page information is greater than or equal to a preset similarity threshold can be selected as candidate text vectors. Alternatively, based on a specified number M of candidate text vectors, M text vectors can be selected sequentially as candidate text vectors according to the order of their similarity between the semantic vector of the page information and each text vector in the text vector library from high to low. Wherein, M is a positive integer. Preferably, M≥2. Regarding the calculation method of the similarity between the semantic vector of the page information and each text vector in the text vector library, please refer to the relevant content on calculating the similarity between the corresponding sample vectors of positive and negative sample pairs, which will not be repeated here.
[0055] In other embodiments, considering that matching the semantic vectors of page information one by one in the text vector library to obtain candidate text vectors is complex and inefficient, this embodiment utilizes an approximate nearest neighbor (ANN) algorithm to spatially partition the text vector library, resulting in a multi-layered text vector space. The ANN algorithm can be the Annoy algorithm, Locality Sensitive Hash algorithm, Vector Quantization algorithm, or K-Nearest Neighbor (KNN) algorithm, but is not limited to these.
[0056] Accordingly, based on the aforementioned multi-layered text vector space, when retrieving the semantic vectors corresponding to page information from the text vector library, the target text subspace to which the semantic vectors corresponding to the page information belong can be determined from the multi-layered text vector space according to the specified number of candidate text vectors; further, candidate text vectors are selected from the text vectors contained in the target text subspace. This retrieval method can quickly find the target text subspace to which the semantic vectors corresponding to the page information belong, which helps improve vector retrieval efficiency, thereby improving the efficiency of subsequently determining the associated pages of the page to be displayed, and helping to improve the real-time nature of obtaining the associated information of the page information.
[0057] Furthermore, the similarity between each text vector contained in the target text subspace and the semantic vector corresponding to the page information can be calculated. Based on these similarities, candidate text vectors are selected from the text feature vectors in the target text subspace. In this way, the semantic vector corresponding to the page information only requires similarity calculations between the text vectors in the target text subspace, reducing computational complexity and improving vector recall efficiency. This, in turn, further improves the efficiency of subsequently determining the associated pages of the page to be displayed. For specific implementation methods regarding the calculation of the similarity between each text vector contained in the target text subspace and the semantic vector corresponding to the page information, please refer to the relevant content in the above embodiments, which will not be repeated here.
[0058] In some embodiments, text vectors whose similarity to the semantic vectors of the page information is greater than or equal to a preset similarity threshold can be selected from the text vectors contained in the target text subspace as candidate text vectors. Alternatively, M text vectors can be selected sequentially from the target text subspace as candidate text vectors according to the specified number M of candidate text vectors, in descending order of similarity between the semantic vectors of the page information and the text vectors contained in the target text subspace. Here, M is a positive integer.
[0059] The following example illustrates the specific implementation of spatial partitioning of the text vector library and vector retrieval of semantic vectors of page information from the text vector library using the Annoy algorithm.
[0060] The Annoy algorithm can construct a text vector library as a binary tree, meaning the multi-level text vector space has a binary tree structure. This reduces the time complexity of vector retrieval to O(logQ) when semantic vectors of page information are retrieved from the text vector library, where Q is the number of vectors in the text vector library. The following section combines... Figure 4a and Figure 4b This paper uses the Annoy algorithm as an example to illustrate the process of spatial partitioning of a text vector library. Figure 4a and Figure 4b In the text, for ease of illustration, dots are used to identify the text vector library.
[0061] Step 1: Randomly select two text vectors from the text vector library, and use these two text vectors as initial center nodes to perform a K-means clustering process with a cluster size of 2. This will ultimately produce two converged cluster center vectors, as shown below. Figure 4a As shown in the middle pentagram.
[0062] Step 2: The two cluster center vectors form a hyperplane ( Figure 4a (As shown by the dashed line in the middle), establish an equidistant perpendicular hyperplane for this hyperplane, as follows: Figure 4a As shown by the solid line, this equidistant vertical hyperplane divides the entire space composed of the text vector library into two parts, A and B, that is, into two text subspaces, A and B. In other words, the equidistant vertical hyperplane is hyperplane 1 corresponding to text subspaces A and B.
[0063] Step 3: Within the divided text subspaces, continue recursively dividing them using the method described in Step 2 until the number of remaining text vectors in each subspace is less than or equal to K. Here, K is the maximum number of text vectors that can be contained in each text subspace.
[0064] At this point, the text vector library is divided into multiple text spaces, forming a binary tree. The root node of the binary tree represents the entire space comprised of the text vector library. In each level of the binary tree, the left and right subtrees corresponding to the same parent node represent the two subspaces of the text space corresponding to that parent node. For example, as... Figure 4b As shown, hyperplane 2 further divides text subspace A into text subspaces A1 and A2, and hyperplane 3 further divides text subspace B into text subspaces B1 and B2. Therefore, the left and right subtrees corresponding to parent node A represent text subspaces A1 and A2, respectively; and the left and right subtrees corresponding to parent node B represent text subspaces B1 and B2, respectively.
[0065] In this embodiment, the hyperplane that divides the text space corresponding to the parent node into two text subspaces is defined as the hyperplane corresponding to the two text subspaces. That is, the hyperplane corresponding to text subspaces A1 and A2 is hyperplane 2, and the hyperplane corresponding to text subspaces B1 and B2 is hyperplane 3.
[0066] Based on the aforementioned binary tree, when determining the target text subspace to which the semantic vectors of page information belong from a multi-layered text space, the traversal direction of the binary tree can be determined according to the geometric relationship between the semantic vectors of the page information and the hyperplanes corresponding to the left and right subtrees at each level of the binary tree. The binary tree is then traversed along this direction until a target subtree is found whose represented text subspace contains a number of text vectors less than the specified number of candidate texts M. The text subspace represented by the parent node of the target subtree is then taken as the target text subspace to which the semantic vectors of the page information belong. Here, the target subtree whose represented text subspace contains a number of text vectors less than the specified number of candidate texts M refers to the first subtree in the traversal process where the represented text subspace contains a number of text vectors less than the specified number of candidate texts M.
[0067] For example, such as Figure 4a and Figure 4bAs shown, text subspaces A and B represent the left and right subtrees of the root node, respectively; text subspaces A1 and A2 represent the left and right subtrees of parent node A, respectively; and text subspaces B1 and B2 represent the left and right subtrees of parent node B, respectively. When determining the traversal direction of the semantic vector of page information through the binary tree, we can first determine the geometric relationship between the semantic vector of page information and hyperplane 1 corresponding to text subspaces A and B. Further, based on the geometric relationship between the semantic vector of page information and hyperplane 1 corresponding to text subspaces A and B, we can determine whether the semantic vector of page information belongs to text subspace A or text subspace B. Further, assuming the semantic vector of page information belongs to text subspace A, we can further determine the geometric relationship between the semantic vector of page information and hyperplane 2 corresponding to text subspaces A1 and A2, and based on this geometric relationship, determine whether the semantic vector of page information belongs to text subspace A1 or text subspace A2. Furthermore, assuming the semantic vectors of the page information belong to text subspace A2, the traversal direction of the binary tree is: the entire space composed of the text vector library -> text subspace A -> text subspace A2. This process is repeated until a text subspace containing fewer than the specified number of candidate texts M is found. The upper-level subspace of this text subspace is then designated as the target text subspace to which the problem text belongs. Further, candidate text vectors are selected from the text vectors contained in the target text subspace.
[0068] Furthermore, a target text vector can be determined based on the candidate text vectors. The target text vector can be a subset or all of the candidate text vectors. In some embodiments, candidate text vectors can be determined as the target text vectors. In other embodiments, the similarity between the semantic vector of the page information and the candidate text vectors can be calculated; and the target text vector can be selected from the candidate text vectors based on the similarity between the semantic vector of the page information and the candidate text vectors. For example, candidate text vectors whose similarity to the semantic vector of the page information is greater than or equal to a set similarity threshold can be selected as the target text vectors. Another example is that a set number N candidate text vectors can be selected as target text vectors according to the order of similarity between the candidate text vectors and the semantic vector of the page information from high to low. N is a positive integer. For example, 2 ≤ N ≤ M. M is the number of candidate text vectors. The number of target text vectors and the similarity threshold can be determined by... Figure 3 The “Protal page” control.
[0069] In some other embodiments, the target text vector may be selected from the candidate text vectors according to the context information of the page to be displayed and the similarity between the semantic vector of the page information and the candidate text vectors. Optionally, the similarity between the semantic vector of the page information and the candidate text vectors may be weighted according to the time interval between the time of the text information corresponding to the candidate text vector and the time information included in the context information, so as to obtain the weighted similarity between the semantic vector of the page information and the candidate text vectors. Among them, the shorter the time interval between the time of the text information corresponding to the candidate text vector and the time information included in the context information, the greater the weight corresponding to the candidate text vector. After that, the target text vector may be selected from the candidate text vectors according to the weighted similarity between the semantic vector of the page information and the candidate text vectors. For the specific implementation manner of selecting the target text vector from the candidate text vectors according to the weighted similarity between the semantic vector of the page information and the candidate text vectors, reference may be made to the relevant content of selecting the target text vector from the candidate text vectors according to the similarity between the semantic vector of the page information and the candidate text vectors as described above, which will not be elaborated here.
[0070] After determining the target text vector, the text information corresponding to the target text vector may be determined as the associated information of the page information of the page to be displayed.
[0071] Further, in step 104, the associated page of the page to be displayed may be determined according to the associated information of the page information of the page to be displayed. Optionally, the associated information of the page information may be used to query in the pre-generated correspondence between the combined words and the pages, so as to obtain the page corresponding to the associated information of the page information as the associated page of the page to be displayed. The relevant content of step 104 may be implemented by the Figure 3 ""s "Related Content Collection Page Production" module.
[0072] The above correspondence between the combined words and the pages may be pre-generated offline. In some embodiments, the词性 analysis may be performed on the words in the known word library to determine the词性 of the words in the known word library. Among them, the known word library may be the classification standard of the词性 of the above words, which may be flexibly set according to the actual application scenario. For example, in some application scenarios, the词性 of the words may be divided into: core words, non-core words, stop words, etc., for the optimization of word segmentation. The core words are flexibly set according to the specific application scenario. For example, for a technology website, the core words may be technical terms, technical product names, etc., such as "Java parameters". The non-core words include professional words, auxiliary words, etc. The stop words include meaningless adverbs, prepositions, etc. such as "的", "地", "得".
[0073] Furthermore, words in the known lexicon can be combined based on their parts of speech to obtain candidate words. For example, in some embodiments, words in the known lexicon can be combined based on their parts of speech and conventional grammatical organization. Conventional grammatical organization can include subject-verb-object grammar, etc. Figure 5 In Chinese, the words "Java", "settings", and "parameters" can be combined to form the compound word "Java setting parameters".
[0074] Considering that some combined words do not fit the website's application scenario, to reduce the efficiency of subsequent combined word queries, validity identification can be performed on candidate combined words to determine valid combined words from the candidate combined words. Optionally, candidate combined words can be used to search within the website content; the candidate combined words found in the website content are determined to be valid combined words. And / or, words used by the user on the website can be obtained; candidate combined words corresponding to the words used by the user on the website can be selected from the candidate combined words as valid combined words. Here, the website is the website where the page to be displayed is located. Words used by the user on the website can be words associated with the user's behavior on the website. For example, search terms used by the user in the website search, words clicked by the user on the website, and words in the content of pages visited by the user on the website, etc.
[0075] In this application embodiment, the specific implementation of selecting candidate words corresponding to words used by the user on the website is not limited. In some embodiments, candidate words that are the same as words used by the user on the website can be selected as valid words. In other embodiments, the aforementioned latent semantic network model can be used to extract semantics from the candidate words and the words used by the user on the website to obtain semantic vectors corresponding to the candidate words and the words used by the user on the website respectively; further, the similarity between the semantic vector of the candidate words and the words used by the user on the website can be calculated; and based on the similarity between the semantic vector of the candidate words and the words used by the user on the website, candidate words with a similarity greater than or equal to a set similarity threshold can be selected as valid words, etc.
[0076] The above embodiments illustrating the selection of valid combination words from candidate combination words are merely illustrative and do not constitute a limitation. Furthermore, after determining the valid combination words, the page corresponding to the valid combination words can be determined. In this application embodiment, the specific implementation method for determining the page corresponding to the valid combination words is not limited. Several optional implementation methods are described below as examples.
[0077] Implementation Method 1: Search the website's page content using effective keyword combinations; identify the pages corresponding to the search results for effective keyword combinations.
[0078] Implementation Method 2: Obtain pages accessed by the user using valid keyword combinations from the website's pages; determine the pages corresponding to the valid keyword combinations based on the pages accessed by the user using valid keyword combinations. In some embodiments, pages accessed by the user using valid keyword combinations can be determined as pages corresponding to valid keyword combinations. In other embodiments, the frequency of user access to pages using valid keyword combinations can be obtained; based on the frequency of user access to pages using valid keyword combinations, pages whose access frequency meets a set requirement can be obtained from the pages accessed by the user using valid keyword combinations as pages corresponding to valid keyword combinations. For example, a set number of pages can be obtained sequentially from the pages accessed by the user using valid keyword combinations in descending order of frequency, as pages corresponding to valid keyword combinations. Another example is that pages whose access frequency is greater than or equal to a set access frequency threshold can be obtained from the pages accessed by the user using valid keyword combinations as pages corresponding to valid keyword combinations, and so on.
[0079] Implementation Method 3: Semantic extraction is performed on effective combined words to determine their semantic vectors; based on the semantic vectors of the effective combined words, related words are determined; based on the related words, the page corresponding to the effective combined words is determined. For a detailed implementation method of semantic extraction of effective combined words, please refer to the relevant content on semantic extraction of page information to be displayed above, which will not be repeated here. For a detailed implementation method of determining related words of effective combined words based on their semantic vectors, please refer to the relevant content on step 103 above, which will not be repeated here. For a detailed implementation method of determining the page corresponding to the effective combined words based on their related words, please refer to implementation methods 1 and 2 above, and implementation method 4 above.
[0080] Implementation method 4: Determine the application scenario and / or category information of the effective combination words; further, determine the page under the application scenario and / or category information as the page corresponding to the effective combination words.
[0081] After identifying the pages corresponding to valid compound words, a correspondence between the compound words and their corresponding pages can be generated. In this embodiment, the compound words in the correspondence between compound words and pages can also be input into a latent semantic encoding model; semantic extraction is performed on the compound words in the correspondence between compound words and pages in the latent semantic encoding model to determine the text vectors of the compound words in the correspondence between compound words and pages; and the text vectors of the compound words in the correspondence between compound words and pages are added to a text vector library.
[0082] Based on the correspondence between compound words and pages, after determining the associated information of usable page information in step 103, step 104 can be specifically implemented as follows: querying the pre-generated correspondence between compound words and pages to obtain the pages corresponding to the associated information of the page information, which are then used as the associated pages of the page to be displayed. Specifically, from the correspondence between compound words and pages, the pages corresponding to the compound words that match the associated information of the page information are obtained and used as the associated pages of the page to be displayed.
[0083] Further, in step 105, the related pages of the page to be displayed can be interconnected to obtain the related content aggregation page corresponding to the page to be displayed. Optionally, the display element information and URL information of the related pages of the page to be displayed are obtained. Specifically, the display element information of the related pages of the page to be displayed refers to the display element information of the related pages within the page to be displayed, including but not limited to: the display content and display attribute information of the related pages within the page to be displayed. The display attribute information includes: the display attributes of the display content of the related pages within the page to be displayed. The display attributes of the display content of the related pages within the page to be displayed may include: display position on the page to be displayed, font (e.g., KaiTi, SongTi), color, size, and format. Figure 3 The Portal page allows you to control the number of associated pages to be displayed and their display attributes.
[0084] Optionally, the associated information of the page information can be used as the display content of the associated page; and the set display attribute information can be determined as the display attribute information of the associated information.
[0085] Furthermore, the display element information and URL information of the related pages can be embedded into the HTML code of the page to be displayed, so as to realize the mutual linking between the page to be displayed and the related pages, and obtain the HTML code of the related content aggregation page of the page to be displayed.
[0086] In this embodiment, semantic analysis is performed on the page information of the page to be displayed to obtain a semantic vector of the page information. Then, based on the semantic vector, the association information of the page information is determined. Since the association information of the page information is determined through the semantic vector obtained from the semantic analysis of the page information, the determined association information incorporates the semantics of the page information. Therefore, the association information of the page information is related to the semantics of the page information, making the relevance between the association information and the content of the page information more accurate. Furthermore, the accuracy of the associated pages of the page to be displayed determined through the relationship information of the page information is high. Thus, the content relevance of the related content aggregation pages obtained by linking the page to be displayed and the associated pages is more accurate. Therefore, the page association method provided in this embodiment helps to improve the accuracy of page association.
[0087] Because the page association accuracy is high, the page association method provided in this application embodiment can improve the quality of related content aggregation pages, help improve the weight of website pages in search engines, and achieve search engine optimization (SEO).
[0088] After obtaining the HTML code of the related content aggregation page of the page to be displayed, in this embodiment of the application, the HTML code of the related content aggregation page can also be executed; the display content of the associated page is displayed on the page to be displayed according to the display attribute information, so as to display the related content aggregation page, that is... Figure 3 The page in the document. The display process of the related content aggregation page can be as described above. Figure 3 The data dashboard is used for execution. Because the relationship information of the page information determines the related pages of the page to be displayed with high accuracy, users are more likely to obtain the expected information through the related content aggregation page, which helps to shorten the user's information acquisition path and thus optimize the user information acquisition path.
[0089] It should be noted that the execution subject of each step of the method provided in the above embodiments can be the same device, or the method can be executed by different devices. For example, the execution subject of steps 101 and 102 can be device A; or the execution subject of step 101 can be device A, and the execution subject of step 102 can be device B; and so on.
[0090] Furthermore, some processes described in the above embodiments and accompanying drawings include multiple operations that appear in a specific order. However, it should be clearly understood that these operations may not be executed in the order they appear herein, or they may be executed in parallel. The operation numbers, such as 101, 102, etc., are merely used to distinguish different operations and do not represent any execution order. In addition, these processes may include more or fewer operations, and these operations may be executed sequentially or in parallel.
[0091] Accordingly, embodiments of this application also provide a computer-readable storage medium storing computer instructions, which, when executed by one or more processors, cause one or more processors to perform the steps in the page association method described above.
[0092] Figure 6 A schematic diagram of the structure of a computing device provided in an embodiment of this application. Figure 6 As shown, the computing device includes a memory 60a and a processor 60b. The memory 60a is used to store computer programs;
[0093] The processor 60b is coupled to the memory 60a and is used to execute a computer program for: obtaining page information of a page to be displayed; performing semantic extraction on the page information to determine the semantic vector of the page information; determining the associated information of the page information based on the semantic vector of the page information; determining the associated pages of the page to be displayed based on the associated information of the page information; and linking the page to be displayed and the associated pages to obtain a related content aggregation page.
[0094] Optionally, when extracting semantics from page information, the processor 60b specifically performs the following: extracts semantics from page information using a latent semantic coding model to obtain a semantic vector of the page information.
[0095] Optionally, the processor 60b is further configured to: extract the central words from the page information before inputting the page information into the implicit semantic coding model, so as to obtain the central words contained in the page information; and input the central words contained in the page information into the implicit semantic coding model.
[0096] Optionally, the processor 60b is further configured to: obtain known semantically related positive sample pairs and known semantically unrelated negative sample pairs before using the latent semantic coding model to extract semantic information from the page; the positive sample pairs include: semantically related text; the negative sample pairs include: semantically unrelated text; and train the initial latent semantic coding model using the positive sample pairs and negative sample pairs with the loss function minimization as the training objective, so as to obtain the latent semantic coding model.
[0097] The loss function is determined based on the difference between the correlation of positive sample pairs output by the model training and the true correlation value of positive sample pairs, and the difference between the correlation of negative sample pairs output by the model training and the true correlation value of negative sample pairs.
[0098] In some embodiments, when the processor 60b determines the associated information of the page information based on the semantic vector of the page information, it is specifically used to: perform vector retrieval in a text vector library using the semantic vector of the page information to select candidate text vectors corresponding to the page information from the text vector library; determine the target text vector based on the candidate text vectors; and determine the text information corresponding to the target text vector as the associated information of the page information.
[0099] Optionally, when the processor 60b performs vector retrieval in the text vector library using the semantic vector of the page information, it specifically performs the following: using an approximate nearest neighbor algorithm to spatially partition the text vector library to obtain a multi-level text vector space; determining the target text vector subspace to which the semantic vector of the page information belongs from the multi-level text vector space based on the specified number of candidate text vectors; and selecting candidate text vectors from the text vectors contained in the target text vector subspace.
[0100] Optionally, when determining the target text vector based on the candidate text vectors, the processor 60b specifically performs the following: calculates the similarity between the semantic vector of the page information and the candidate text vectors; and selects a predetermined number of candidate text vectors as the target text vectors based on the order of the similarity between the semantic vector of the page information and the candidate text vectors from high to low.
[0101] Alternatively, when determining the target text vector based on the candidate text vectors, the processor 60b specifically performs the following: calculates the similarity between the semantic vector of the page information and the candidate text vectors; and selects the target text vector from the candidate text vectors based on the context information of the page to be displayed contained in the page information and the similarity between the semantic vector of the page information and the candidate text vectors.
[0102] In other embodiments, when the processor 60b determines the associated page of the page to be displayed based on the association information of the page information, it specifically performs the following: using the association information of the page information, it queries the pre-generated correspondence between the combined words and the page to obtain the page corresponding to the association information of the page information, which is then used as the associated page of the page to be displayed.
[0103] Optionally, the processor 60b is further configured to: perform part-of-speech analysis on words in a known dictionary to determine the part of speech of the words in the known dictionary; combine words in the known dictionary according to the part of speech of the words in the known dictionary to obtain candidate combined words; identify the validity of the candidate combined words to determine the valid combined words from the candidate combined words; determine the page corresponding to the valid combined words; and generate a correspondence between the combined words and the page corresponding to the valid combined words.
[0104] Furthermore, when identifying the validity of candidate combination words, the processor 60b is specifically used to: use candidate combination words to search in the website content; determine that the candidate combination words found in the website content are valid combination words; and / or, obtain the words used by the user on the website; and select candidate combination words from the candidate combination words that correspond to the words used by the user on the website as valid combination words.
[0105] Optionally, when determining the page corresponding to a valid combination of words, the processor 60b specifically performs the following: searches within the website's page content using the valid combination of words; determines the page corresponding to the searched page content containing the valid combination of words; and / or, retrieves the page accessed by the user using the valid combination of words from the website's pages; determines the page corresponding to the valid combination of words based on the page accessed by the user using the valid combination of words; and / or, performs semantic analysis on the valid combination of words to determine the semantic vector of the valid combination of words; determines the related words of the valid combination of words based on the semantic vector of the valid combination of words; determines the page corresponding to the valid combination of words based on the related words of the valid combination of words; and / or, determines the application scenario and / or category information of the valid combination of words; and determines the page under the application scenario and / or category information as the page corresponding to the valid combination of words.
[0106] In some embodiments, the processor 60b is further configured to: input the combined words in the correspondence between combined words and pages into a latent semantic coding model; perform semantic extraction on the combined words in the correspondence between combined words and pages in the latent semantic coding model to determine the text vectors of the combined words in the correspondence between combined words and pages; and add the text vectors of the combined words in the correspondence between combined words and pages to a text vector library.
[0107] In some other embodiments, when the processor 60b performs page linking between the page to be displayed and the associated page, it is specifically used to: obtain the display element information and the Uniform Resource Locator (URL) information of the associated page; embed the display element information and the URL information of the associated page into the HTML code of the page to be displayed, so as to obtain the HTML code of the related content aggregation page by performing page linking between the page to be displayed and the associated page.
[0108] Optionally, when the processor 60b obtains the display element information of the associated page, it specifically performs the following: using the association information of the page information as the display content of the associated page; and determining the set display attribute information as the display attribute information of the associated page.
[0109] Accordingly, processor 60b is also used to: execute the HTML code of the related content aggregation page; and display the display content of the associated page on the page to be displayed according to the display attribute information through display component 60c, so as to display the related content aggregation page through display component 60c.
[0110] In some alternative implementations, such as Figure 6 As shown, the computing device may also include optional components such as a communication component 60d, a power supply component 60e, and an audio component 60f. Figure 6 The diagram only shows some components and does not mean that the computing device must contain them. Figure 6 The inclusion of all components does not imply that a computing device can only include... Figure 6 The components shown.
[0111] The computing device provided in this application embodiment performs semantic analysis on the page information of the page to be displayed to obtain a semantic vector of the page information; then, based on the semantic vector of the page information, it determines the association information of the page information. Since the association information of the page information is determined through the semantic vector obtained from the semantic analysis of the page information, the determined association information incorporates the semantics of the page information. Therefore, the association information of the page information is related to the semantics of the page information, making the relevance between the association information and the content of the page information more accurate. Furthermore, the accuracy of the associated pages of the page to be displayed determined through the relationship information of the page information is high. Thus, the content relevance of the related content aggregation pages obtained by linking the page to be displayed and the associated pages is more accurate. Therefore, the page association method provided in this application embodiment helps to improve the accuracy of page association.
[0112] In this embodiment, the memory is used to store computer programs and can be configured to store various other data to support operation on its host device. The processor can execute the computer programs stored in the memory to implement corresponding control logic. The memory can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
[0113] In the embodiments of this application, the processor can be any hardware processing device capable of executing the above-described method logic. Optionally, the processor can be a central processing unit (CPU), a graphics processing unit (GPU), or a microcontroller unit (MCU); it can also be a programmable device such as a field-programmable gate array (FPGA), a programmable array logic (PAL), a general array logic (GAL), or a complex programmable logic device (CPLD); or it can be an advanced reduced instruction set (RISC) processor (ARM) or a system on chip (SOC), etc., but is not limited thereto.
[0114] In this embodiment, the communication component is configured to facilitate wired or wireless communication between its host device and other devices. The device housing the communication component can access wireless networks based on communication standards, such as WiFi, 2G or 3G, 4G, 5G, or combinations thereof. In one exemplary embodiment, the communication component receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In another exemplary embodiment, the communication component may also be implemented based on Near Field Communication (NFC), Radio Frequency Identification (RFID), Infrared Data Association (IrDA), Ultra Wideband (UWB), Bluetooth (BT), or other technologies.
[0115] In embodiments of this application, the display component may include a liquid crystal display (LCD) and a touch panel (TP). If the display component includes a touch panel, the display component may be implemented as a touchscreen to receive input signals from a user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors may not only sense the boundaries of touch or swipe actions but also detect the duration and pressure associated with the touch or swipe operation.
[0116] In this embodiment, a power supply component is configured to provide power to various components of the device in which it resides. The power supply component may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the device in which the power supply component resides.
[0117] In embodiments of this application, the audio component can be configured to output and / or input audio signals. For example, the audio component includes a microphone (MIC), which is configured to receive external audio signals when the device containing the audio component is in an operating mode, such as call mode, recording mode, and voice recognition mode. The received audio signals can be further stored in memory or transmitted via a communication component. In some embodiments, the audio component also includes a speaker for outputting audio signals. For example, in devices with voice interaction capabilities, voice interaction with the user can be achieved through the audio component.
[0118] It should be noted that the terms "first" and "second" in this article are used to distinguish different messages, devices, modules, etc., and do not represent a chronological order, nor do they limit "first" and "second" to different types.
[0119] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0120] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0121] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0122] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0123] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0124] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0125] Computer storage media are readable storage media, also known as removable media. Removable and non-removable media can be used to store information by any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media do not include transient media, such as modulated data signals and carrier waves.
[0126] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0127] The above description is merely an embodiment of this application and is not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.
Claims
1. A page association method, characterized in that, include: Get the page information of the page to be displayed; Semantic extraction is performed on the page information to determine the semantic vector of the page information; Based on the semantic vector of the page information, determine the associated information of the page information; Based on the association information of the page information, determine the associated pages of the page to be displayed; Link the page to be displayed and the associated page to obtain a related content aggregation page; The step of determining the associated page of the page to be displayed based on the association information of the page information includes: Using the association information of the page information, a query is performed in the pre-generated correspondence between compound words and pages to obtain the page corresponding to the association information of the page information, which is then used as the associated page of the page to be displayed. The correspondence between the combined words and the page is generated based on the valid combined words and the page corresponding to the valid combined words; the valid combined words are determined from the candidate combined words by identifying the validity of the candidate combined words; the candidate combined words are obtained by combining words in the known word library according to the part of speech of the words in the known word library.
2. The method according to claim 1, characterized in that, The step of semantically extracting the page information to determine the feature vector of the page information includes: The semantics of the page information are extracted using a latent semantic coding model to obtain the semantic vector of the page information.
3. The method according to claim 2, characterized in that, Before using the implicit semantic coding model to perform semantic extraction on the page information, the following steps are also included: Obtain known semantically related positive sample pairs and known semantically unrelated negative sample pairs; the positive sample pairs include semantically related text; the negative sample pairs include semantically unrelated text. With minimizing the loss function as the training objective, the initial latent semantic encoding model is trained using the positive sample pairs and the negative sample pairs to obtain the latent semantic encoding model. The loss function is determined based on the difference between the correlation of the positive sample pair output by the model training and the true correlation value of the positive sample pair, and the difference between the correlation of the negative sample pair output by the model training and the true correlation value of the negative sample pair.
4. The method according to claim 1, characterized in that, Determining the association information of the page information based on the semantic vector of the page information includes: The semantic vector of the page information is used to perform vector retrieval in the text vector library to select candidate text vectors corresponding to the page information from the text vector library; Based on the candidate text vectors, determine the target text vector; The text information corresponding to the target text vector is determined to be the associated information of the page information.
5. The method according to claim 4, characterized in that, The step of using the semantic vector of the page information to perform vector retrieval in a text vector library, in order to select candidate text vectors corresponding to the page information from the text vector library, includes: The text vector library is spatially partitioned using the approximate nearest neighbor algorithm to obtain a multi-layer text vector space; Based on the specified number of candidate text vectors, determine the target text vector subspace to which the semantic vector of the page information belongs from the multi-layer text vector space; The candidate text vectors are selected from the text vectors contained in the target text vector subspace.
6. The method according to claim 4, characterized in that, Determining the target text vector based on the candidate text vector includes: Calculate the similarity between the semantic vector of the page information and the candidate text vector; Based on the context information of the page to be displayed contained in the page information and the similarity between the semantic vector of the page information and the candidate text vector, the target text vector is selected from the candidate text vector.
7. The method according to any one of claims 1-6, characterized in that, Before using the association information of the page information to perform a query in the pre-generated correspondence between combined words and pages, the method further includes: Part-of-speech analysis is performed on words in a known vocabulary to determine their part of speech. Based on the part of speech of the words in the known vocabulary, the words in the known vocabulary are combined to obtain candidate combined words; The candidate words are evaluated for validity in order to identify valid words from among them. Determine the page corresponding to the valid combination of words; Based on the valid combined words and the pages corresponding to the valid combined words, a correspondence between the combined words and the pages is generated.
8. The method according to claim 7, characterized in that, The step of identifying the validity of the candidate word combinations to determine the valid word combinations from the candidate word combinations includes: The candidate words are used to search the website content; the candidate words found in the website content are determined to be the valid words. And / or, Obtain the words used by the user on the website; select the candidate words that correspond to the words used by the user on the website from the candidate words, and use them as the effective words.
9. The method according to claim 7, characterized in that, The step of determining the page corresponding to the valid combination of words includes: The effective combination of words is used to search the page content of the website; the page corresponding to the page content where the effective combination of words is found is determined as the page corresponding to the effective combination of words; And / or, Obtain the pages accessed by the user using effective keyword combinations from the website's pages; determine the pages corresponding to the effective keyword combinations based on the pages accessed by the user using effective keyword combinations; And / or, Semantic analysis is performed on the effective combined words to determine their semantic vectors; based on the semantic vectors of the effective combined words, their associated words are determined; based on the associated words of the effective combined words, the page corresponding to the effective combined words is determined. And / or, Determine the application scenario and / or category information of the effective combination words; determine the page under the application scenario and / or category information as the page corresponding to the effective combination words.
10. The method according to claim 1, characterized in that, The step of linking the page to be displayed and the associated page includes: Obtain the display element information and Uniform Resource Locator (URL) information of the associated page; The display element information and Uniform Resource Locator (URL) information of the associated page are embedded in the HTML code of the page to be displayed, so as to perform page linking between the page to be displayed and the associated page to obtain the HTML code of the related content aggregation page.
11. The method according to claim 10, characterized in that, Also includes: Execute the HTML code of the related content aggregation page to display the related content aggregation page.
12. A computing device, characterized in that, include: A memory and a processor; wherein the memory is used to store computer programs; The processor is coupled to the memory for executing the computer program to perform the steps of the method according to any one of claims 1-11.
13. A computer-readable storage medium storing computer instructions, characterized in that, When the computer instructions are executed by one or more processors, the one or more processors are caused to perform the steps of the method according to any one of claims 1-11.