Search Engine Optimization Method and Device, Electronic Device, and Readable Storage Medium
By obtaining visual information of mobile terminal App pages, using preset scripts and model optimization technology, the problem of insufficient SEO on the App page is solved, and the accurate acquisition and inclusion of page titles and core texts is achieved, which improves the performance effect of search engines.
Patent Information
- Application Number
- CN202210176355.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-02-24
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2042-02-24
AI Technical Summary
In the prior art, the search function in the application app of the mobile terminal lacks the search ability, resulting in insufficient SEO means for most pages in the website, and the inability to realize effective title and core text display in traditional search engines, affecting the search effect.
By obtaining the visual information of the target page, using preset scripts to traverse the page document, determine the page title and core text, and optimize it based on the trained model to achieve inclusion of the target page.
It improves the efficiency of search engine optimization for mobile terminal App pages, ensures accurate acquisition of page titles and core text, and improves the quality and presentation effect of search results.
Smart Images

Figure CN114528513B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of search engine optimization. Specifically, it relates to a search engine optimization method, device, electronic device, and readable storage medium. Background Art
[0002] In order to improve the accuracy of search terms, the HTML pages within a website and the native pages on the client are traversed to obtain key information therefrom, and through keyword matching, appropriate pages are selected, and information such as titles and core texts are displayed in the search list and distributed to users.
[0003] Currently, the commonly used page collection technology in the industry only collects the structural information and content information of pages. If a page only provides what-you-see-is-what-you-get information, then in most cases, this information is insufficient to determine whether the page is suitable for distribution under a search keyword, and is even less sufficient to extract appropriate information such as titles and core texts to display to users. The search effect of traditional search engines highly depends on the implementation of SEO (Search Engine Optimization) of the website itself being collected. That is, only when a website has established good SEO can it be better presented in traditional search engines, and most of the time, what a search engine has to do is just directly read the SEO information provided by the website to determine whether to include the page, and directly display the suggested title and core text in the SEO information to users.
[0004] During the implementation of the present invention by the applicant, it was found that there are at least the following technical problems in the related technologies.
[0005] As can be seen from the above-mentioned prior art solutions, whether a website can achieve good presentation of titles and core texts in search results, or even whether it can be distributed as a search result to users, largely depends on whether the website itself has good SEO. If the SEO of the page itself is not done well enough, it is often easily misfiltered by the search engine. Even if it is distributed, due to the low quality of the extracted titles and core texts, its ranking and display in search results will be relatively poor.
[0006] Since there has never been such a retrieval ability in the search of application programs (Apps) on mobile terminals, and most pages within a website do not have the need to be included by external search engines, the SEO means for most pages within a website are relatively weak; in addition, the SEO rules are all for HTML documents, and there are even no SEO specifications for native pages on the client, and naturally, page maintainers will not do SEO either. This has led to the fact that if only the commonly used technology in the industry is used to perform SEO only for the web side and not for the APP side, it is absolutely impossible to achieve the search effect of traditional search engines.
[0007] It can be seen that in the related art, no effective solution has been proposed for the above problems. Summary of the Invention
[0008] Embodiments of the present invention provide a search engine optimization method, an apparatus, an electronic device, and a readable storage medium, so as to at least solve the technical problem that the content information of the target page cannot be accurately obtained because the target page in the related art is not optimized for the search engine.
[0009] According to one aspect of the embodiments of the present invention, a search engine optimization method is provided, including: obtaining visual information of a target page; determining a page title and core text in the target page according to the visual information; and indexing the target page according to the page title and the core text.
[0010] Further, obtaining visual information of the target page includes: injecting a preset script into the target information interface; obtaining a page document of the target page according to the preset script, where the page document includes the visual information; and obtaining the visual information according to the page document.
[0011] Further, injecting a preset script into the target information interface includes: when the target page is an HTML page, injecting the preset script into a preset interface of the HTML page; or, when the target page is an application page, adding a preset script to the source code of the application corresponding to the application page.
[0012] Further, obtaining the page document of the target page according to the preset script includes: traversing a view tree in the page document; obtaining node attributes of each node in the view tree; and generating the page document according to the node attributes.
[0013] Further, obtaining the visual information according to the page document includes: determining a keyword position of the keyword in the page document in the target page; and obtaining the visual information according to the keyword position.
[0014] Further, determining the page title and core text in the target page according to the visual information includes: inputting visual information features corresponding to the visual information, page features of the target page, and text language features of the target page into a pre-trained page title prediction model to obtain the page title of the target page; and inputting the visual information features and content features of the core area in the target page into a pre-trained core text prediction model to obtain the core text of the target page.
[0015] Further, the page features include at least one of the following: the HTML page features and DOM features of the target page; the content features include at least one of the following: the text features, picture features, and link features of the core area.
[0016] According to another aspect of the embodiments of the present invention, there is also provided a search engine optimization device, including: an acquisition unit, configured to acquire visual information of a target page; a determination unit, configured to determine a page title and core text in the target page according to the visual information; and an optimization unit, configured to include the target page according to the page title and the core text.
[0017] According to another aspect of the embodiments of the present invention, there is also provided an electronic device, including a processor, a memory, and a program or instruction stored on the memory and executable on the processor. When the program or instruction is executed by the processor, the steps of the search engine optimization method described above are implemented.
[0018] According to another aspect of the embodiments of the present invention, there is also provided a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the steps of the search engine optimization method described above are implemented.
[0019] In the embodiments of the present invention, by acquiring the visual information of the target page; determining the page title and core text in the target page according to the visual information; and including the target page according to the page title and the core text, the page title and core text in the target page are determined through the target visual information, and then the target page is included through the page title and the core text, so as to optimize the search engine, and further solve the technical problem that the content information of the target page cannot be accurately obtained because the target page in the related technology is not optimized for the search engine. Description of the Drawings
[0020] The drawings described herein are used to provide a further understanding of the present invention, and constitute a part of this application. The illustrative embodiments of the present invention and their descriptions are used to explain the present invention, and do not constitute an improper limitation to the present invention. In the drawings:
[0021] Figure 1 is a schematic flowchart of an optional search engine optimization method according to an embodiment of the present invention;
[0022] Figure 2a is a schematic diagram of an optional application page according to an embodiment of the present invention;
[0023] Figure 2b is a schematic diagram of an optional page code according to an embodiment of the present invention;
[0024] Figure 2c It is a schematic diagram of an optional page visual information according to an embodiment of the present invention;
[0025] Figure 3a It is a schematic diagram of an optional application page according to an embodiment of the present invention;
[0026] Figure 3b It is a schematic diagram of another optional application page according to an embodiment of the present invention;
[0027] Figure 3c It is a schematic diagram of another optional application page according to an embodiment of the present invention;
[0028] Figure 4 It is a schematic diagram of the structure of an optional search engine optimization device according to an embodiment of the present invention. Detailed implementation manners
[0029] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0030] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0031] Embodiment 1
[0032] Before introducing the technical solution of the present invention, the following terms are first explained:
[0033] Traversal: The collection of the page content of the target web page. Here, it should be noted that during the implementation process of the technical solution of the present invention, the collection of data such as page content is carried out under the authorization and permission of the data owner or all data users.
[0034] SEO: Search Engine Optimization, which is a technology that analyzes the ranking rules of search engines, understands how various search engines conduct searches, crawl Internet pages, and determine the search result rankings of specific keywords. Search engines adopt means that are easy to be searched and referenced to optimize websites targeted.
[0035] The following introduces the process of traversing web page content. In a specific website, the process of traversing the page content of the website may include the following steps:
[0036] S1, Define an initial page set. For example, define some high-quality pages that need to be collected initially, such as portal websites, etc.;
[0037] S2, Collect these high-quality pages in sequence, as well as the pages that these pages can jump to (also called the next-level pages), and then continue to collect the next-level pages of the jump pages;
[0038] S3, Filter out low-quality pages according to rules or models, and the remaining pages are used as included pages;
[0039] S4, Mine the candidate set of titles and the candidate set of core texts (abstracts) from the included pages according to rules or models.
[0040] The following introduces the index construction process of SEO, which specifically may include the following steps:
[0041] S1, For each included page, segment its title, abstract, and content, so that it can be known which keywords each page contains, also called the forward index;
[0042] S2, Count which pages contain these keywords, so that the results of the corresponding pages can be displayed according to the keywords queried by users, also called the inverted index. At the same time, this step can also calculate the TF (term frequency) and IDF (Inverse Document Frequency) of each word.
[0043] To solve the problem in the prior art that the target page is not optimized for the search engine, resulting in the inability to accurately obtain the content information of the target page, according to the embodiments of the present invention, a search engine optimization method is provided, as Figure 1 shown. This method specifically includes the following steps:
[0044] S102, Obtain the visual information of the target page;
[0045] S104, Determine the page title and core text in the target page according to the visual information;
[0046] S106. Index the target page according to the page title and the core text.
[0047] Specifically, in this embodiment, the target page includes but is not limited to website pages and application pages of applications, etc. For example, the website page is an HTML page, and the application page is a native page of a mobile APP client.
[0048] It should be noted that the target page in this embodiment includes but is not limited to the current page content corresponding to the target page, the web page or application page corresponding to the jump link in the target page, and the page content rendered after the target page receives relevant operations, etc.
[0049] In a specific application scenario, the visual information includes but is not limited to the screen coordinates and sizes of each element in the target page; information such as the font, font size, thickness, color, and visibility of the text. In this embodiment, the visual information is used to indicate the display parameters and display effects of each element displayed in the target page.
[0050] In this embodiment, determine the page title and the core text in the target page according to the visual information of the target page. For example, determine the page title of the target page according to the thickness, color, and coordinates of the text in the web page element. Generally, the feature of the page title is located at the middle top position of the web page view, with the format of being centered, bold, and having a larger font size. Determine the page title of the page through the feature of the page title. In addition, the body text corresponding to the page title in the web page can be further determined.
[0051] Then, after determining the page title and the core text in the target page through the visual information, search engine optimization can be performed according to the page title and the core text, associate the page title and the core text with the target page, then build an index for the target page, and further include the page content of the target page in the database of the search engine.
[0052] It should be noted that through this embodiment, obtain the visual information of the target page; determine the page title and the core text in the target page according to the visual information; index the target page according to the page title and the core text. Determine the page title and the core text in the target page through the target visual information, and then realize indexing the target page through the page title and the core text, realize the optimization of the search engine, and then solve the technical problem that the content information of the target page cannot be accurately obtained due to the lack of search engine optimization in the target page in the related technology.
[0053] Optionally, in this embodiment, obtaining the visual information of the target page includes, but is not limited to: traversing the page content of the target page to obtain the page document of the target page; injecting a preset script into the page document to obtain the view tree corresponding to the target page; and obtaining the visual information according to the view tree.
[0054] In this embodiment, first, traverse the page content of the target page. The target page in this embodiment includes, but is not limited to, HTML pages and client native pages.
[0055] On the one hand, for HTML pages, use the selenium framework + WebDriver middleware + Chrome browser to collect HTML documents to obtain HTML documents, that is, page documents.
[0056] On the other hand, for the native page of the specified application APP client, by obtaining the APP source code and adding the traversal and dump logic of the layout tree of the instruction application in the source code, a page document with a structure similar to the HTML document is obtained.
[0057] After traversing the page document of the target page in the above manner, inject a preset script into the page document to obtain the view tree HTML DOM tree of the target page. The HTML DOM tree includes multiple DOM nodes, and each node has corresponding node attributes. The visual information can be calculated from the DOM nodes in the view tree.
[0058] Optionally, in this embodiment, injecting a preset script into the page document to obtain the view tree corresponding to the target page includes, but is not limited to: in the case where the target page is an HTML page, injecting the preset script into a preset interface of the HTML page; or, in the case where the target page is an application page, adding a preset script to the application source code corresponding to the application page.
[0059] Specifically, in an example, assume that the preset script is a piece of JavaScript code. For HTML pages, the selenium framework provides an interface for injecting the preset script JavaScript code. Inject a piece of JavaScript code through this interface to trigger the CSS cascade style calculation, then traverse the view tree HTML DOM tree, and read the visual information calculated on each DOM node. After obtaining the visual information, inject it into the page document in the form of node attributes, and then extract the visual document HTML document. When injecting attributes, by avoiding the W3C standard, attribute conflicts are avoided.
[0060] In another example, for the native page of the client for a specified application, visual information is obtained by acquiring the App source code and directly adding in the App source code a traversal of the view tree and reading the attributes of each view. Similarly, the visual information is injected into the XML document in the form of node attributes to obtain a visual document.
[0061] It should be noted that there are many frameworks and solutions that can inject code to obtain visual information. The solution for obtaining visual information by injecting code in this embodiment, including but not limited to the specific implementation methods listed above, will not impose any limitation on the technical solution of this embodiment.
[0062] Optionally, in this embodiment, obtaining the page document of the target page according to a preset script includes but is not limited to: traversing the view tree in the page document; obtaining the node attributes of each node in the view tree; and the visual document corresponding to the target page, where the visual document includes visual information.
[0063] Specifically, in this embodiment, the visual document is obtained by rewriting the node attributes of the view tree listed above into the page document obtained by traversing the target page, so as to obtain a visual document with visual information. Specifically, by traversing the view tree of the HTML page or the native page of the client and adding information such as the screen coordinates, size, font, font size, thickness, color, and visibility of the text of the view to each node of the view tree and placing it as an XML attribute in the corresponding node, the visual document output at this time contains visual information.
[0064] It should be noted that the visual information not only includes the attributes listed above, but also includes but is not limited to information that cannot be obtained but can be seen or perceived by page visitors; the content structure of the document is not limited to being placed as XML attributes in the corresponding nodes, but also includes any other means that can store visual information as a document and associate it with the nodes in the original document. This embodiment does not make any limitation on this.
[0065] In one example, taking Figure 2a the target page shown as an example, by acquiring the source code of the target page, as Figure 2b shown, which is the source code of the target page, and the information that can be collected by a search engine, the content and structure information directly reflected in the HTML source code. The information collected is as Figure 2c shown, Figure 2cThe content in the middle frame shows information such as the coordinates, font size, text color, thickness, and visibility of each element. Based on this information, the content processing engine can identify which content is in the upper position of the page, which content is bold and highlighted on the page, and which content is actually invisible. It should be noted that invisible content is not suitable as direct features such as keyword hits, titles, and abstracts, but is suitable as indirect features such as relevance or closeness. Therefore, it is also inappropriate to directly remove it during collection.
[0066] Through the above embodiments, traverse the page content of the target page, inject a preset script into the page document to obtain the view tree corresponding to the target page; obtain visual information according to the view tree, which can realize the collection of visual information of different types of pages and improve the collection efficiency of visual information.
[0067] Optionally, in this embodiment, after generating the visual document corresponding to the target page according to the node attributes, it further includes: obtaining the visibility information of each node in the visual document; filtering the page content corresponding to each node according to the visibility information.
[0068] In specific application scenarios, some web pages cannot display valid content due to network connection problems or content quality problems. For example, "404 Not Found", "The web page is lost", "Loading", and the content in the web page has no valid content (the displayed content is recommended content), etc.
[0069] In order to avoid the influence of low-quality web page content, in this embodiment, according to the visibility information of each DOM node obtained from the visual document, the interference of DOM nodes with invisible attributes can be eliminated, and then the core text content can be extracted. Then, content filtering is performed according to the pre-set matching content filtering rules. The visibility information includes, but is not limited to, visibility keywords such as "Loading", "404 Not Found", "The web page is lost", and the position information of the visibility keywords.
[0070] For example, if the visibility keyword "Loading" is located at the top of the application page, it means that the page content in the current application page has not been loaded. If the visibility keyword "Loading" is located at the bottom of the application page, it means that part of the page content in the application page may not have been loaded.
[0071] In specific application scenarios, such as Figure 3a , Figure 3b and Figure 3c shown, are schematic diagrams of the application pages of the application program respectively.
[0072] In one example, what is collected is Figure 3aIf a blank page is shown, the page is a low-quality page (slow loading or unable to load) and is not suitable for distribution to users. The page quality judgment system will filter out the application page based on the fact that it contains the visibility keyword "loading" and is located in the middle of the application page.
[0073] In another example, the collected Figure 3b The application page shown also has the visibility keyword "Loading", but the word is located at the bottom of the page and outside the screen (the screenshot is the scene after the page is pulled to the bottom). This is a normal application page and should not be filtered because it contains the word "Loading".
[0074] In another example, the collected Figure 3c The application page shown also has the visibility keyword "Loading", but the word is located at the top of the page. In fact, the main content of the application page is not loaded. What is loaded is personalized recommended content based on user information, which should also be filtered.
[0075] From the above content, we can know that we cannot judge whether to filter based on whether there is content on the page; nor can we judge whether to filter based on the position of the "loading" element in the HTML code, because the content of the page is at the top, which does not mean that the content corresponding to its HTML code is at the front. Therefore, in such scenarios, only the visibility keywords in the visibility information and the location information of the visibility keywords can easily and accurately filter the non-distributable application pages.
[0076] Through the above embodiment, the visibility information of each node in the visual document is obtained, and the page content corresponding to each node is screened according to the visibility information, thereby avoiding the workload of optimization work of the application page caused by invalid content in the application page and improving the optimization efficiency of the application page.
[0077] Optionally, in this embodiment, the page title and core text of the target page are determined based on the visual information, including but not limited to: inputting visual information features corresponding to the visual information, page features of the target page, and text language features of the target page into a pre-trained page title prediction model to obtain the page title of the target page; inputting visual information features and content features of the core area of the target page into a pre-trained core text prediction model to obtain the core text of the target page.
[0078] In this embodiment, on the one hand, according to the visual information features, page features, text language features, and page titles in the page content of the application page, a first training sample is constructed. Then, based on the first training sample, a title training data set is constructed. Subsequently, a page title prediction model is trained according to the title training data set until the model converges.
[0079] After the above model training is completed, the visual information features corresponding to the visual information, the page features of the target page, and the text language features of the target page are input into the pre-trained page title prediction model to obtain the page title of the target page.
[0080] On the other hand, according to the visual information features in the page content, the content features of the core area, and the core text, a second training sample is constructed. Then, based on the second training sample, a core text training data set is constructed. Subsequently, a core text prediction model is trained according to the core text training data set until the model converges.
[0081] After the above model training is completed, the visual information features and the content features of the core area in the target page are input into the pre-trained core text prediction model to obtain the core text of the target page.
[0082] Through the above embodiment, by inputting the visual information features corresponding to the visual information into the pre-trained page title prediction model and core text prediction model respectively, the page title and core text of the target page are obtained, which simplifies the operation of engine optimization and improves the optimization efficiency of the target page.
[0083] Optionally, in this embodiment, the page features include at least one of the following: the HTML page features and DOM features of the target page; the content features include at least one of the following: the text features, picture features, and link features of the core area.
[0084] On the one hand, the visual information of most page titles in the application page is usually stronger than other texts on the page. For example, the font size is larger and the color is more prominent. Therefore, the visual information of the page title can assist the page title extraction strategy.
[0085] To extract page titles more efficiently and automatically, text visual information features are introduced, and four major types of features, namely HTML page features, DOM features, and text natural language features, are added to construct a title extraction model. The steps for constructing the page title prediction model are as follows:
[0086] 1) Sample cleaning and construction: Extract the titles of pages with SEO information to form the positive samples for model training, and the other texts on the page form the negative samples for model training.
[0087] 2) Feature construction: Based on text visual information features, H5 page features, DOM features, and text natural language features, a total of 44-dimensional features are constructed. Among them, text visual features include text thickness, text page position, text font size, and text color.
[0088] 3) Model training and classification: Train the XGBOOST model on the training set, and get the weights corresponding to the 44-dimensional features. Predict all the text on the test page to get the optimal title text for the page.
[0089] On the other hand, in the search scenario, the relevance of text is very important. It is necessary to extract text information based on web page elements and then match the relevance with the search terms. However, there are many meaningless texts on web pages, such as the text in the navigation bar, header area, and bottom area. These texts are irrelevant to the main content of the page. However, if all the text on the page is used, the text in these meaningless areas will affect the calculation effect of the relevance. Therefore, it is necessary to mine the core text in the core area of the page.
[0090] Based on the visual information in the document, such as element position, combined with the position information of text elements, locate the lower structural area in the first screen, and use the text in this area as the core text.
[0091] The core text prediction model construction steps are as follows:
[0092] 1) Feature construction: text in the core area, images in the core area, jump links in the core area, etc., and construct various feature variants, such as TF-IDF, image quality score, number of external links, etc.
[0093] 2) Model training and classification: Based on the features mined from the core area, combined with models such as XGBOOST and DNN, the model is trained with the goal of optimizing the click-through rate.
[0094] Through the embodiments of the present application, visual information of a target page is obtained; a page title and a core text in the target page are determined based on the visual information; the target page is included based on the page title and the core text, and the page title and the core text in the target page are determined through the target visual information, thereby achieving inclusion of the target page through the page title and the core text, achieving search engine optimization, and thereby solving the technical problem of not being able to accurately obtain the content information of the target page due to the lack of search engine optimization in the target page in the related art.
[0095] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present invention is not limited by the described action sequence, because according to the present invention, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present invention.
[0096] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions for causing a terminal device (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods described in various embodiments of the present invention.
[0097] Embodiment 2
[0098] According to an embodiment of the present invention, there is also provided a search engine optimization device for implementing the above search engine optimization method, as Figure 4 shown. The device includes:
[0099] 1) An acquisition unit 40, configured to acquire visual information of a target page;
[0100] 2) A determination unit 42, configured to determine a page title and core text in the target page according to the visual information;
[0101] 3) An optimization unit 44, configured to perform indexing on the target page according to the page title and the core text.
[0102] Optionally, the specific examples in this embodiment can refer to the examples described in the above Embodiment 1, and this embodiment will not be elaborated here.
[0103] Embodiment 3
[0104] According to an embodiment of the present invention, there is also provided an electronic device, including a processor, a memory, and a program or instruction stored on the memory and executable on the processor. When the program or instruction is executed by the processor, the steps of the search engine optimization method described in Embodiment 1 are implemented.
[0105] Optionally, in this embodiment, the memory is configured to store program codes for executing the following steps:
[0106] S1, Obtain the visual information of the target page;
[0107] S2, Determine the page title and the core text in the target page according to the visual information;
[0108] S3, Index the target page according to the page title and the core text.
[0109] Optionally, the specific examples in this embodiment may refer to the examples described in the above Embodiment 1, and will not be elaborated herein.
[0110] Embodiment 4
[0111] An embodiment of the present invention also provides a readable storage medium. Programs or instructions are stored on the readable storage medium, and when the programs or instructions are executed by a processor, the steps of the search engine optimization method described in Embodiment 1 are implemented.
[0112] Optionally, in this embodiment, the readable storage medium is set to store program codes for executing the following steps:
[0113] S1, Obtain the visual information of the target page;
[0114] S2, Determine the page title and the core text in the target page according to the visual information;
[0115] S3, Index the target page according to the page title and the core text.
[0116] Optionally, the storage medium is also set to store program codes for executing the steps included in the method in the above Embodiment 1, which will not be elaborated herein.
[0117] Optionally, in this embodiment, the above storage medium may include, but is not limited to: various media such as USB flash drives, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), mobile hard disks, magnetic disks or optical discs that can store program codes.
[0118] Optionally, the specific examples in this embodiment may refer to the examples described in the above Embodiment 1, and will not be elaborated herein.
[0119] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.
[0120] If the integrated units in the above embodiments are implemented in the form of software functional units and sold or used as independent products, they can be stored in the above computer-readable storage media. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing one or more computer devices (which can be personal computers, servers, or network devices, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention.
[0121] In the above embodiments of the present invention, the descriptions of the various embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0122] In the several embodiments provided by the present application, it should be understood that the disclosed client can be implemented in other ways. Among them, the device embodiments described above are only illustrative. For example, the division of the units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the units or modules can be in electrical or other forms.
[0123] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0124] In addition, the functional units in the various embodiments of the present invention can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.
[0125] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. A search engine optimization method, characterized in that, including: Obtaining visual information of a target page; Determining a page title and core text in the target page according to the visual information; Including the target page according to the page title and the core text; The obtaining visual information of the target page includes: Traversing the page content of the target page to obtain a page document of the target page; Injecting a preset script into the page document to obtain a view tree corresponding to the target page; Obtaining the visual information according to the view tree; The injecting a preset script into the page document to obtain a view tree corresponding to the target page includes: when the target page is an HTML page, injecting the preset script into a preset interface of the HTML page, triggering CSS calculation and traversing the HTML DOM tree to generate a view tree; Or, when the target page is an application page, adding a preset script to the application source code corresponding to the application page and traversing the native view tree to generate a view tree.
2. The method according to claim 1, wherein Obtaining the visual information according to the view tree includes: Traversing the view tree; Obtaining node attributes of each node in the view tree; Generating a visual document corresponding to the target page according to the node attributes, where the visual document includes the visual information.
3. The method according to claim 2, wherein After generating the visual document corresponding to the target page according to the node attributes, it further includes: Obtaining visibility information of each node in the visual document; Filtering the page content corresponding to each node according to the visibility information.
4. The method according to claim 1, characterized in that, Determining the page title and core text in the target page according to the visual information includes: Inputting visual information features corresponding to the visual information, page features of the target page, and text language features of the target page into a pre-trained page title prediction model to obtain the page title of the target page; Inputting the visual information features and content features of the core area in the target page into a pre-trained core text prediction model to obtain the core text of the target page.
5. The method according to claim 4, characterized in that The page features include at least one of the following: HTML page features and DOM features of the target page; The content features include at least one of the following: Text features, picture features, and link features of the core area.
6. A search engine optimization device, characterized in that, including: An obtaining unit for obtaining visual information of a target page; A determining unit for determining a page title and core text in the target page according to the visual information; An optimizing unit for including the target page according to the page title and the core text; The obtaining visual information of the target page includes: Traversing the page content of the target page to obtain a page document of the target page; Injecting a preset script into the page document to obtain a view tree corresponding to the target page; Obtaining the visual information according to the view tree; Injecting a preset script into the page document to obtain a view tree corresponding to the target page includes: when the target page is an HTML page, injecting the preset script into a preset interface of the HTML page, triggering CSS calculation, and traversing the HTML DOM tree to generate a view tree; Alternatively, when the target page is an application page, adding a preset script to the application source code corresponding to the application page, and traversing the native view tree to generate a view tree.
7. An electronic device, characterized in that, It includes a processor, a memory, and a program or instruction stored on the memory and executable on the processor. When the program or instruction is executed by the processor, the steps of the search engine optimization method described in claims 1-5 are implemented.
8. A readable storage medium, characterized in that, A program or instruction is stored on the readable storage medium. When the program or instruction is executed by the processor, the steps of the search engine optimization method described in claims 1-5 are implemented.
Citation Information
Patent Citations
Page configuration method and device for search engine optimization
CN113779359A