Website information processing method and device

By obtaining and storing the meta-information and content of multiple web pages of the target website, the web page query function is provided, which solves the problem of low efficiency for users to obtain information through the website, and realizes the rapid search of web pages that meet their needs.

CN120104901APending Publication Date: 2025-06-06BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510182550.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-18
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

Users are less efficient in obtaining information through the website and need to browse each web page in turn to find the required information.

Method used

Provide a website information processing method, by obtaining meta information of multiple web pages of the target website, accessing and extracting the content of each web page, storing web page identifiers and contents to provide web page query function.

Benefits of technology

This improves the efficiency of users to obtain information through the website, and users can enter query requests and quickly find web pages that meet their needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120104901A_ABST
    Figure CN120104901A_ABST
Patent Text Reader

Abstract

The invention discloses a website information processing method which comprises the steps that meta-information of each web page in a plurality of web pages included in a target website is obtained, and the meta-information comprises a web page identifier and a web page link; and aiming at each webpage in the plurality of webpages, accessing each webpage based on the webpage link of each webpage, and extracting webpage contents of each webpage. And for each webpage, correspondingly storing the webpage identifier of the webpage and the webpage content of the webpage. By utilizing the scheme, aiming at each webpage in the plurality of webpages in the target website, the webpage identifier of the webpage and the webpage content of the webpage can be correspondingly stored, so that a webpage query function can be conveniently provided for a user subsequently. Specifically, when a subsequent user wants to obtain a certain piece of information through the target website, the query request can be input, and the webpage meeting the user requirement can be queried for the user by utilizing the correspondingly stored webpage identifier and the webpage content of the webpage according to the query request input by the user, so that the efficiency of obtaining the information through the target website by the user is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a website information processing method and device. Background Art

[0002] A website is a collection of web pages, that is, a website can include multiple web pages. Users can browse the web pages included in the website to obtain the information they want to obtain.

[0003] At present, the efficiency of users obtaining information through websites is low. Specifically, when users want to obtain certain information through a website, they need to browse various web pages of the website in sequence to find a web page including the information so as to obtain the information through the web page.

[0004] Therefore, a solution is urgently needed to solve the above problems. Summary of the invention

[0005] In order to solve or at least partially solve the above technical problems, the present application provides a website information processing method and device.

[0006] In the first aspect, the application provides a website information processing method, the method comprising:

[0007] Obtaining meta information of each web page among a plurality of web pages included in the target website, wherein the meta information includes: a web page identifier and a web page link;

[0008] Based on the webpage link of each webpage, access each webpage and extract the webpage content of each webpage;

[0009] For each web page among the multiple web pages, the web page identifier and the web page content of the web page are correspondingly stored, and the correspondingly stored web page identifier and the web page content of the web page are used to provide a web page query function.

[0010] Optionally, obtaining meta information of each web page among the multiple web pages included in the target website includes:

[0011] Obtaining the meta information of each web page through the meta tag information file of the target website, wherein the meta tag information file includes the meta information of each web page; or

[0012] The meta information of each web page is obtained through the navigation list interface of the target website, wherein the navigation list interface is used to return the meta information of each web page.

[0013] Optionally, accessing each web page based on the web page link of each web page includes:

[0014] Each web page is accessed based on a web page link of each web page using a headless browser.

[0015] Optionally, extracting the webpage content of each webpage includes:

[0016] Extracting non-image text from the web page; and / or,

[0017] The text included in the image in the webpage is recognized using optical character recognition technology.

[0018] Optionally, the method further includes:

[0019] For each web page, keywords of the web page content of the web page are extracted using a large language model, and the web page identifier of the web page and the keywords of the web page content are correspondingly stored.

[0020] Optionally, the meta information further includes web page description information, and the method further includes:

[0021] For each web page, the web page identifier and the web page description information of the web page are correspondingly stored.

[0022] Optionally, the method further includes:

[0023] Receiving a query request input by a user in a first webpage, wherein the target website includes the first webpage;

[0024] Extracting keywords of the query request, where the keywords of the query request include: words used to describe functions and / or words used to describe indicators;

[0025] Using the keywords of the query request, querying the page information of the website to obtain a first query score of each web page, wherein the page information of the website includes at least one of the web page content of each web page, the keywords of the web page content of each web page, and the web page description information of each web page, wherein the first query score of the web page is used to indicate the matching degree of the web page with the keywords of the query request;

[0026] Based on the first query score of each web page, a query result is determined.

[0027] Optionally, determining the query result based on the first query score of each web page includes:

[0028] A query result is determined based on the first query score of each web page and the access popularity of each web page.

[0029] Optionally, determining the query result based on the first query score of each web page and the access popularity of each web page includes:

[0030] For each web page, determining a second query score of the web page based on the first query score of the web page and the access popularity of the web page;

[0031] The second query scores are sorted in descending order, and the web pages corresponding to the top N second query scores are taken as the query results, where N is an integer greater than or equal to 1.

[0032] Optionally, the multiple web pages include a second web page, and the first query score of the second web page is determined in the following manner:

[0033] Using the keywords of the query request, respectively query the webpage content of the second webpage, the keywords of the webpage content of the second webpage, and the webpage description information of the second webpage to obtain a first query score of the webpage content of the second webpage, a first query score of the keywords of the webpage content of the second webpage, and a first query score of the webpage description information of the second webpage;

[0034] The maximum value among the first query score of the webpage content of the second webpage, the first query score of the keyword of the webpage content of the second webpage, and the first query score of the webpage description information of the second webpage is determined as the first query score of the second webpage.

[0035] Optionally, the multiple web pages are web pages used by a third party to access the target website.

[0036] In a second aspect, the present application provides a website information processing device, the device comprising:

[0037] An acquisition unit, used to acquire meta information of each web page among a plurality of web pages included in the target website, wherein the meta information includes: a web page identifier and a web page link;

[0038] An access unit, configured to access each web page based on a web page link of each web page;

[0039] A content extraction unit, used to extract the webpage content of each webpage;

[0040] The first storage unit is used to store the web page identifier and the web page content of each web page in the multiple web pages, and the correspondingly stored web page identifier and the web page content of the web page are used to provide a web page query function.

[0041] Optionally, the acquiring unit is used to:

[0042] Obtaining the meta information of each web page through the meta tag information file of the target website, wherein the meta tag information file includes the meta information of each web page; or

[0043] The meta information of each web page is obtained through the navigation list interface of the target website, wherein the navigation list interface is used to return the meta information of each web page.

[0044] Optionally, the access unit is used to:

[0045] Each web page is accessed based on a web page link of each web page using a headless browser.

[0046] Optionally, the content extraction unit is used to:

[0047] Extracting non-image text from the web page; and / or,

[0048] The text included in the image in the webpage is recognized using optical character recognition technology.

[0049] Optionally, the device further comprises:

[0050] A first keyword extraction unit, configured to extract keywords of the webpage content of each webpage by using a large language model;

[0051] The second storage unit is used to store the web page identifier of the web page and the keywords of the web page content correspondingly.

[0052] Optionally, the meta information further includes web page description information, and the device further includes:

[0053] The third storage unit is used to store, for each web page, the web page identifier and the web page description information of the web page correspondingly.

[0054] Optionally, the device further comprises:

[0055] A receiving unit, configured to receive a query request input by a user in a first webpage, wherein the target website includes the first webpage;

[0056] A second keyword extraction unit, configured to extract keywords of the query request, wherein the keywords of the query request include: words used to describe functions and / or words used to describe indicators;

[0057] A query unit, configured to query the page information of the website using the keyword of the query request to obtain a first query score of each web page, wherein the page information of the website includes at least one of the web page content of each web page, the keyword of the web page content of each web page, and the web page description information of each web page, wherein the first query score of the web page is used to indicate the matching degree of the web page with the keyword of the query request;

[0058] The determining unit is used to determine the query result based on the first query score of each web page.

[0059] Optionally, the determining unit is used to:

[0060] A query result is determined based on the first query score of each web page and the access popularity of each web page.

[0061] Optionally, determining the query result based on the first query score of each web page and the access popularity of each web page includes:

[0062] For each web page, determining a second query score of the web page based on the first query score of the web page and the access popularity of the web page;

[0063] The second query scores are sorted in descending order, and the web pages corresponding to the top N second query scores are taken as the query results, where N is an integer greater than or equal to 1.

[0064] Optionally, the multiple web pages include a second web page, and the first query score of the second web page is determined in the following manner:

[0065] Using the keywords of the query request, respectively query the webpage content of the second webpage, the keywords of the webpage content of the second webpage, and the webpage description information of the second webpage to obtain a first query score of the webpage content of the second webpage, a first query score of the keywords of the webpage content of the second webpage, and a first query score of the webpage description information of the second webpage;

[0066] The maximum value among the first query score of the webpage content of the second webpage, the first query score of the keyword of the webpage content of the second webpage, and the first query score of the webpage description information of the second webpage is determined as the first query score of the second webpage.

[0067] Optionally, the multiple web pages are web pages used by a third party to access the target website.

[0068] In a third aspect, an embodiment of the present application provides an electronic device, the electronic device comprising a processor and a memory;

[0069] The processor is used to execute instructions stored in the memory so that the electronic device performs the method as described in any one of the first aspects above.

[0070] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, comprising instructions, wherein the instructions instruct a device to execute a method as described in any one of the above first aspects.

[0071] In a fifth aspect, an embodiment of the present application provides a computer program product, which, when executed on a computer, enables the computer to execute any of the methods described in the first aspect above.

[0072] Compared with the prior art, the embodiments of the present application have the following advantages:

[0073] The present application provides a website information processing method, the method comprising: obtaining meta information of each web page in a plurality of web pages included in a target website, the meta information comprising: a web page identifier and a web page link. After obtaining the meta information of each web page in the plurality of web pages, for each web page in the plurality of web pages, the web page can be accessed based on the web page link of each web page, and the web page content of each web page can be extracted. For each web page, the web page identifier of the web page and the web page content of the web page are stored correspondingly. Using this solution, for each web page in the plurality of web pages in the target website, the web page identifier of the web page and the web page content of the web page can be stored correspondingly, so as to provide a web page query function for the user later. As a specific example, when the user subsequently wants to obtain certain information through the target website, the user can enter a query request. For the query request entered by the user, the aforementioned correspondingly stored web page identifier and the web page content of the web page can be used to query the user for a web page that meets the user's needs, thereby improving the efficiency of the user obtaining information through the target website. BRIEF DESCRIPTION OF THE DRAWINGS

[0074] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0075] Figure 1 A flowchart of a website information processing method provided in an embodiment of the present application;

[0076] Figure 2 A flowchart of a query method provided in an embodiment of the present application;

[0077] Figure 3 A flowchart of a method for determining a first query score provided in an embodiment of the present application;

[0078] Figure 4 A schematic diagram of the structure of a website information processing device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0079] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0080] The inventor of the present application has found through research that currently, the efficiency of users obtaining information through websites is low. Specifically, when users want to obtain certain information through a website, they need to browse each web page of the website in sequence to find the web page containing the information so as to obtain the information through the web page.

[0081] If a webpage query function can be provided in a website, the efficiency of users obtaining information through the website can be improved. Specifically, a user can input a query request, and the webpage query function is used to search for webpages that meet the query request from the webpages included in the website based on the query request input by the user, and return the information of the webpage obtained by the query (such as webpage links and summary information of the webpage) to the user, so that the user can further view the corresponding webpage.

[0082] However, a website may include multiple types of web pages, one of which is a web page that is fully controllable by the developer, and the other is a web page that is accessed by a third party to the website. The web page that is fully controllable by the developer refers to a web page whose content is determined by the developer. On the contrary, the web page accessed by a third party is not controllable by the developer, and the web page links of the web page are not controllable by the developer. For the convenience of description, "web pages accessed by a third party to the website" are referred to as "third-party web pages".

[0083] For the development of fully controllable web pages, it is possible to implement information query through online transaction processing (OLTP) engines such as MySQL, thereby providing users with web page query functions. Among them, MySQL is a popular relational database management system. In terms of website (WEB) applications, MySQL is one of the better relational database management system (RDBMS) application software.

[0084] However, for third-party web pages, because their information is uncontrollable, current technology cannot support retrieving web pages that meet user needs from the web page content of third-party web pages. Accordingly, when the content that the user wants to view is the content in the third-party web page, the information acquisition efficiency is low.

[0085] In order to solve the above problems, the embodiments of the present application provide a website information processing method and device.

[0086] Various non-limiting implementations of the present application are described in detail below in conjunction with the accompanying drawings.

[0087] Exemplary Methods

[0088] See also Figure 1 , which is a flow chart of a website information processing method provided in an embodiment of the present application.

[0089] The method provided in the embodiment of the present application may be executed by a client or a server, and the embodiment of the present application does not make any specific limitation.

[0090] In this embodiment, the method may include, for example, the following steps: S101 - S103 .

[0091] S101: Obtaining meta information of each web page among a plurality of web pages included in a target website, wherein the meta information includes: a web page identifier and a web page link.

[0092] In one example, the multiple web pages mentioned here may include web pages that are fully controllable by the R&D personnel, as well as third-party web pages. For example, the multiple web pages may be all web pages of the target website. Of course, the multiple web pages may also be some web pages of the target website.

[0093] In another example, considering the web pages that are currently fully controllable by the R&D personnel, an OLTP engine such as MySQL can be used to implement the web page query function. Therefore, the multiple web pages mentioned here may include third-party web pages, excluding web pages that are fully controllable by the R&D personnel.

[0094] In the present application, for any one of the aforementioned multiple web pages, the meta information of the web page may include a web page identifier and a web page link of the web page. The web page identifier may be, for example, a web page name. In addition, in one example, the meta information of the web page may include web page description information of the web page in addition to the web page identifier and the web page link, wherein the web page description information refers to information describing the content included in the web page, and the web page description information may also be understood as a summary of the content included in the web page (i.e., summary information of the web page content). The meta information of the web page may be understood in conjunction with the following Table 1.

[0095] Table 1

[0096] Page Name Web Links Web page description information A domain1.com Description of Web Page A B domain2.com Description of Web Page B

[0097] In one example, it is considered that a website generally has a meta tag information file for recording the meta information of each web page included in the website. Therefore, as an example, when S101 is implemented specifically, the meta information of each of the aforementioned web pages can be obtained through the meta tag information file of the target website. Specifically, the meta tag information file of the target website includes the meta information of each of the aforementioned web pages. Therefore, after obtaining the meta tag information file of the target website, the meta information of each of the aforementioned web pages can be extracted from the meta tag information file of the target website. In one example, the meta tag information file can be a "sitemap", and the meta tag information file can be a file in the Extensible Markup Language (XML) format.

[0098] In another example, it is considered that a website generally provides a navigation list interface for returning the meta information of each web page included in the website. Therefore, as an example, when S101 is implemented, the meta information of each web page can be obtained through the navigation list interface of the target website. Specifically, the navigation list interface of the target website can be called to obtain the information returned by the navigation list interface, and further, the meta information of each web page can be extracted from the information returned by the navigation list interface. In one example, the navigation list interface can be a menu interface.

[0099] S102: Based on the web page link of each web page, access each web page and extract the web page content of each web page.

[0100] After obtaining the meta information of each web page, for each web page in the plurality of web pages, the web page can be accessed based on the web page link of the web page, and the web page content of each web page can be further extracted.

[0101] In one example, in order to improve the efficiency of extracting the webpage content of each webpage, a headless browser can be used to access the webpage based on the webpage link of the webpage. Wherein, a headless browser refers to a browser without an interface, which can increase the webpage access rate and reduce other interferences, which is conducive to improving the efficiency of extracting the webpage content of each webpage.

[0102] In another example, a browser with an interface may also be used to access the web page based on the web link of the web page, which is not specifically limited in the embodiment of the present application.

[0103] The embodiment of the present application does not specifically limit the web page content, and the web page content may include text, wherein the text may include: text in non-pictures and / or text in pictures. In one example, the text and pictures in non-pictures in the web page content may be extracted separately. For pictures included in the web page, optical character recognition (OCR) technology may be used to recognize the text in the picture.

[0104] Regarding the specific implementation of extracting text from a web page other than in an image, the conventional method for extracting text from a web page other than in an image may be used, which will not be described in detail here.

[0105] Regarding the specific implementation of extracting images from web pages, the traditional method for extracting images from web pages can also be used, which will not be described in detail here.

[0106] S103: For each web page, a web page identifier and web page content of the web page are correspondingly stored, and the web page identifier and web page content of the web page are correspondingly stored to provide a web page query function.

[0107] After the webpage content of each webpage is acquired, the webpage identifier of the webpage and the webpage content of the webpage may be correspondingly stored, so that a webpage query function can be provided based on the stored content later.

[0108] In one example, considering that the content of a web page is generally large and may be relatively messy and disordered, the retrieved content may not be suitable for direct output to the user based on the web page content. Therefore, in one example, for each web page, the keywords of the web page content of the web page can be extracted. As a specific example, considering that the large language model (LLM) has excellent semantic recognition capabilities and can accurately extract keywords from text, the large language model can be used to extract keywords from the web page content. After extracting the keywords of the web page content, the web page identifier and the keywords of the web page content can also be stored correspondingly, so as to provide a web page query function based on the corresponding stored web page identifier and the keywords of the web page content in the future, so as to improve the efficiency of users obtaining information from the target website. In this scenario, for each web page, the corresponding stored content can be understood in conjunction with Table 2 below.

[0109] Table 2

[0110]

[0111] In an example, the content saved for each web page (eg, the content shown in the aforementioned Table 2) may be persistently stored in a database (eg, a hive table).

[0112] In the present application, after executing the above S101-S103, a web page query function can be provided based on the stored content.

[0113] In one example, S101-S103 may be executed periodically, for example, once a day (eg, at midnight), to ensure that the stored information matches the latest web page included in the target website, thereby ensuring the real-time nature of the web page query function.

[0114] Next, the specific implementation of the webpage query function is introduced in conjunction with the accompanying drawings.

[0115] See also Figure 2 , which is a flow chart of a query method provided in an embodiment of the present application. Figure 2 The method shown includes the following S201-S204.

[0116] S201: receiving a query request input by a user in a first webpage, wherein the target website includes the first webpage.

[0117] In the present application, the first webpage is one of the multiple webpages included in the target website. As an example, the first webpage may be the webpage currently being browsed by the user. As another example, the first webpage may be the homepage of the target website.

[0118] The first web page may include a query request input area, in which the user may input the aforementioned query request and trigger a query operation. Accordingly, the client (or server) may receive the query request input by the user in response to the query operation.

[0119] S202: extracting keywords of the query request, where the keywords of the query request include: words used to describe functions and / or words used to describe indicators.

[0120] In one example, in order to improve the accuracy of the query, the keywords of the query request can be extracted. Specifically, considering that there are some special words in the query request, these special words are particularly important for determining the search results. Therefore, in the present application, these special words can be extracted as keywords of the query request, and these special words can be used for retrieval. In one example, considering the words used to describe functions and the words used to describe indicators, they can often indicate what the user wants to query. Among them, the words describing the function indicate that the user wants to know a certain function, and the words describing the indicator indicate that the user wants to know a certain indicator. Therefore, the keywords of the aforementioned query request may be the words used to describe the function and / or the words used to describe the indicator in the query request.

[0121] The embodiments of the present application do not specifically limit the words used to describe functions. The functions mentioned here may be functions of an object, for example, functions of a software tool or a physical tool. The indicators mentioned here include but are not limited to business indicators (such as trading volume).

[0122] In an example, the keywords of the query request may be extracted through the LLM. For example, the keywords of the query request may be extracted through the LLM using a function call.

[0123] S203: Utilizing the keywords of the query request, query the page information of the website to obtain a first query score for each web page, wherein the page information of the website includes: the web page content of each web page, keywords of the web page content of each web page, and at least one item of web page description information of each web page, wherein the first query score of the web page is used to indicate the degree of match between the web page and the keywords of the query request.

[0124] In one example, after extracting the keywords of the query request, the keywords of the query request may be used to query page information of the website to obtain a first query score for each web page.

[0125] In one example, an index may be pre-established for the page information of the website, and when querying the page information of the website, the query may be performed based on the pre-established index to improve the query efficiency. In the specific implementation of establishing the index, for example, for any web page, keywords of the web page description information and keywords of the web page content may be identified, and added to the word segmentation dictionary respectively, and then, based on the word segmentation dictionary, the index may be constructed using a word segmenter.

[0126] In the present application, for each of the aforementioned multiple web pages, the method for determining the first query score is the same. Next, the first query score of the second web page is determined as an example. The second web page is any one of the aforementioned multiple web pages. The second web page and the first web page can be the same or different, and the present application embodiment does not make specific limitations.

[0127] See also Figure 3 , which is a flow chart of a method for determining a first query score provided in an embodiment of the present application.

[0128] Figure 3 The method shown includes the following S301-S302.

[0129] S301: Using the keywords of the query request, query the web page content of the second web page, the keywords of the web page content of the second web page, and the web page description information of the second web page respectively, to obtain a first query score for the web page content of the second web page, a first query score for the keywords of the web page content of the second web page, and a first query score for the web page description information of the second web page.

[0130] In one example, the keywords of the query request can be used to query the webpage content of the second webpage, the keywords of the webpage content of the second webpage, and the webpage description information of the second webpage by using the Elasticsearch (ES) query method. Specifically, the aforementioned pre-built index can be used to query the webpage content of the second webpage, the keywords of the webpage content of the second webpage, and the webpage description information of the second webpage by using the ES query method. Among them:

[0131] By using the keyword of the query request to query the webpage content of the second webpage, a first query score of the webpage content of the second webpage can be obtained. The first query score of the webpage content of the second webpage can represent the matching degree between the keyword of the query request and the webpage content of the second webpage.

[0132] By using the keywords of the query request to query the keywords of the webpage content of the second webpage, a first query score of the keywords of the webpage content of the second webpage can be obtained. The first query score of the keywords of the webpage content of the second webpage can represent the matching degree between the keywords of the query request and the keywords of the webpage content of the second webpage.

[0133] By using the keyword of the query request to query the webpage description information of the second webpage, a first query score of the webpage description information of the second webpage can be obtained. The first query score of the webpage description information of the second webpage can represent the matching degree between the keyword of the query request and the webpage description information of the second webpage.

[0134] When the ES query method is used, the aforementioned first query score may be a BM25 score.

[0135] S302: Determine the maximum value of the first query score of the webpage content of the second webpage, the first query score of the keyword of the webpage content of the second webpage, and the first query score of the webpage description information of the second webpage as the first query score of the second webpage.

[0136] After obtaining the first query score of the webpage content of the second webpage, the first query score of the keywords of the webpage content of the second webpage, and the first query score of the webpage description information of the second webpage, the maximum value of the three first query scores can be determined as the first query score of the second webpage. Using the maximum value of the three first query scores as the first query score of the second webpage can effectively avoid missing valid query results.

[0137] In another example, the average of the three first query scores may be used as the first query score of the second webpage. In yet another example, the three first query scores may be weighted and summed, and the result of the weighted sum may be used as the first query score of the second webpage. When the three first query scores are weighted and summed, the sum of the weights of the three first query scores is equal to 1, and the weight of any one of the three first query scores may be greater than or equal to 0 and less than or equal to 1.

[0138] S204: Determine a query result based on the first query score of each web page.

[0139] After obtaining the first query score of each web page, the query result can be determined based on the first query score of each web page. For example, the web pages can be sorted from high to low according to the first query score, and the web pages corresponding to the top N first query scores are taken as the query result, where N is an integer greater than or equal to 1.

[0140] In another example, considering that for a certain web page, the higher its access popularity, the possibility that the user needs to query the web page is relatively high. In view of this, in one implementation of the present application, when determining the query result, the access popularity of each web page can also be combined. In other words, when S204 is specifically implemented, the query result can be determined based on the first query score of each web page and the access popularity of each web page. The access popularity here can be, for example, the number of visits within a certain period of time (such as a week).

[0141] As a specific example, for each web page, the second query score of the web page can be determined based on the first query score of the web page and the access popularity of the web page. For example, a first weight is determined based on the access popularity of the web page, and the second query score is obtained by using the first weight and the first query score. For example:

[0142] Assuming that the access popularity of a web page is the number of weekly visitors (weekly access user, wau), then for any web page, the second query score = the first query score * log (1 + wau). The log (1 + wau) mentioned here corresponds to the aforementioned first weight.

[0143] After determining the second query score of each web page, the query result can be determined based on the second query score of each web page. For example, the web pages can be sorted from high to low according to the second query score, and the web pages corresponding to the top N second query scores are taken as the query result, where N is an integer greater than or equal to 1.

[0144] As another example, the access popularity of a web page may be used as a screening condition for the query result. For the N web pages obtained by sorting according to the first query score, the web page whose access popularity is higher than a certain threshold among the N web pages may be used as the query result.

[0145] Correspondingly, after the query result is determined, the query result may be displayed, for example, the web page links and web page description information of the aforementioned N web pages may be displayed, so that the user can further access other web pages based on the displayed query result.

[0146] From the above description, it can be seen that, by using this solution, for each of the multiple web pages in the target website, the web page identifier and the web page content of the web page can be stored accordingly, so as to provide the user with a web page query function later. Specifically, when the user subsequently hopes to obtain certain information through the website, he can enter a query request. For the query request entered by the user, the corresponding stored content can be used to query the user for the web page that meets the user's needs, thereby improving the efficiency of the user in obtaining information through the target website.

[0147] As described above, in one example, the multiple web pages may all be third-party web pages. In this case, for web pages included in the target website and fully controlled by the developer, an OLTP engine such as MySQL may be used to implement information query, thereby providing users with a query function for the web pages fully controlled by the developer.

[0148] Exemplary Devices

[0149] Based on the method provided in the above embodiment, the embodiment of the present application further provides a device, which is described below in conjunction with the accompanying drawings.

[0150] See also Figure 4 , Figure 4 A schematic diagram of the structure of a website information processing device provided in an embodiment of the present application. Figure 4The device 400 shown is used to execute the website information processing method provided by the above method embodiment.

[0151] In an example, the apparatus 400 may specifically include: an acquisition unit 401 , an access unit 402 , a content extraction unit 403 , and a first storage unit 404 .

[0152] The acquisition unit 401 is used to acquire meta information of each web page among a plurality of web pages included in the target website, wherein the meta information includes: a web page identifier and a web page link.

[0153] The access unit 402 is used to access each web page based on the web page link of each web page.

[0154] The content extraction unit 403 is used to extract the webpage content of each webpage.

[0155] The first storage unit 404 is used to store the web page identifier and the web page content of each web page in the multiple web pages, and the correspondingly stored web page identifier and the web page content of the web page are used to provide a web page query function.

[0156] Optionally, the acquiring unit 401 is used to:

[0157] Obtaining the meta information of each web page through the meta tag information file of the target website, wherein the meta tag information file includes the meta information of each web page; or

[0158] The meta information of each web page is obtained through the navigation list interface of the target website, wherein the navigation list interface is used to return the meta information of each web page.

[0159] Optionally, the access unit 402 is configured to:

[0160] Each web page is accessed based on a web page link of each web page using a headless browser.

[0161] Optionally, the content extraction unit 403 is used to:

[0162] Extracting non-image text from the web page; and / or,

[0163] The text included in the image in the webpage is recognized using optical character recognition technology.

[0164] Optionally, the device further comprises:

[0165] A first keyword extraction unit, configured to extract keywords of the webpage content of each webpage by using a large language model;

[0166] The second storage unit is used to store the web page identifier of the web page and the keywords of the web page content correspondingly.

[0167] Optionally, the meta information further includes web page description information, and the device further includes:

[0168] The third storage unit is used to store, for each web page, the web page identifier and the web page description information of the web page correspondingly.

[0169] Optionally, the device further comprises:

[0170] A receiving unit, configured to receive a query request input by a user in a first webpage, wherein the target website includes the first webpage;

[0171] A second keyword extraction unit, configured to extract keywords of the query request, wherein the keywords of the query request include: words used to describe functions and / or words used to describe indicators;

[0172] A query unit, configured to query the page information of the website using the keyword of the query request to obtain a first query score of each web page, wherein the page information of the website includes at least one of the web page content of each web page, the keyword of the web page content of each web page, and the web page description information of each web page, wherein the first query score of the web page is used to indicate the matching degree of the web page with the keyword of the query request;

[0173] The determining unit is used to determine the query result based on the first query score of each web page.

[0174] Optionally, the determining unit is used to:

[0175] A query result is determined based on the first query score of each web page and the access popularity of each web page.

[0176] Optionally, determining the query result based on the first query score of each web page and the access popularity of each web page includes:

[0177] For each web page, determining a second query score of the web page based on the first query score of the web page and the access popularity of the web page;

[0178] The second query scores are sorted in descending order, and the web pages corresponding to the top N second query scores are taken as the query results, where N is an integer greater than or equal to 1.

[0179] Optionally, the multiple web pages include a second web page, and the first query score of the second web page is determined in the following manner:

[0180] Using the keywords of the query request, respectively query the webpage content of the second webpage, the keywords of the webpage content of the second webpage, and the webpage description information of the second webpage to obtain a first query score of the webpage content of the second webpage, a first query score of the keywords of the webpage content of the second webpage, and a first query score of the webpage description information of the second webpage;

[0181] The maximum value among the first query score of the webpage content of the second webpage, the first query score of the keyword of the webpage content of the second webpage, and the first query score of the webpage description information of the second webpage is determined as the first query score of the second webpage.

[0182] Optionally, the multiple web pages are web pages used by a third party to access the target website.

[0183] Since the device 400 is a device corresponding to the website information processing method provided by the above method embodiment, the specific implementation of each unit of the device 400 is the same concept as the above method embodiment. Therefore, regarding the specific implementation of each unit of the device 400, reference can be made to the relevant description part of the above method embodiment, which will not be repeated here.

[0184] The embodiment of the present application also provides an electronic device, the electronic device comprising a processor and a memory;

[0185] The processor is used to execute the instructions stored in the memory, so that the electronic device executes the website information processing method provided by the above method embodiment.

[0186] An embodiment of the present application provides a computer-readable storage medium, including instructions, wherein the instructions instruct a device to execute the website information processing method provided by the above method embodiment.

[0187] The embodiment of the present application also provides a computer program product. When the computer program product is run on a computer, the computer executes the website information processing method provided by the above method embodiment.

[0188] Those skilled in the art will readily appreciate other embodiments of the present application after considering the specification and practicing the invention disclosed herein. The present application is intended to cover any modification, use or adaptation of the present application, which follows the general principles of the present application and includes common knowledge or customary techniques in the art that are not disclosed in the present disclosure. The specification and examples are intended to be exemplary only, and the true scope and spirit of the present application are indicated by the following claims.

[0189] It should be understood that the present application is not limited to the precise structures that have been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present application is limited only by the appended claims.

[0190] The above description is only a preferred embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application should be included in the protection scope of the present application.

Claims

1. A website information processing method, characterized in that: The method comprises: Obtaining meta information of each web page among a plurality of web pages included in the target website, wherein the meta information includes: a web page identifier and a web page link; Based on the webpage link of each webpage, access each webpage and extract the webpage content of each webpage; For each web page among the multiple web pages, the web page identifier and the web page content of the web page are correspondingly stored, and the correspondingly stored web page identifier and the web page content of the web page are used to provide a web page query function.

2. The method according to claim 1, characterized in that: The step of obtaining the meta information of each web page in the plurality of web pages included in the target website includes: Obtaining the meta information of each web page through the meta tag information file of the target website, wherein the meta tag information file includes the meta information of each web page; or, The meta information of each web page is obtained through the navigation list interface of the target website, wherein the navigation list interface is used to return the meta information of each web page.

3. The method according to claim 1, characterized in that The step of accessing each web page based on the web page link of each web page comprises: Each web page is accessed based on a web page link of each web page using a headless browser.

4. The method according to claim 3, characterized in that The extracting the webpage content of each webpage includes: Extracting non-image text from the web page; and / or, The text included in the image in the webpage is recognized using optical character recognition technology.

5. The method according to claim 1, characterized in that The method further comprises: For each web page, keywords of the web page content of the web page are extracted using a large language model, and the web page identifier of the web page and the keywords of the web page content are correspondingly stored.

6. The method according to claim 5, characterized in that The meta information also includes web page description information, and the method further includes: For each web page, the web page identifier and the web page description information of the web page are correspondingly stored.

7. The method according to any one of claims 1 to 6, characterized in that: The method further comprises: Receiving a query request input by a user in a first webpage, wherein the target website includes the first webpage; Extracting keywords of the query request, where the keywords of the query request include: words used to describe functions and / or words used to describe indicators; Using the keywords of the query request, querying the page information of the website to obtain a first query score of each web page, wherein the page information of the website includes at least one of the web page content of each web page, the keywords of the web page content of each web page, and the web page description information of each web page, wherein the first query score of the web page is used to indicate the matching degree of the web page with the keywords of the query request; Based on the first query score of each web page, a query result is determined.

8. The method according to claim 7, characterized in that The determining of the query result based on the first query score of each web page includes: A query result is determined based on the first query score of each web page and the access popularity of each web page.

9. The method according to claim 8, characterized in that The step of determining the query result based on the first query score of each web page and the access popularity of each web page includes: For each web page, determining a second query score of the web page based on the first query score of the web page and the access popularity of the web page; The second query scores are sorted in descending order, and the web pages corresponding to the top N second query scores are taken as the query results, where N is an integer greater than or equal to 1.

10. The method according to any one of claims 7 to 9, characterized in that: The plurality of web pages include a second web page, and the first query score of the second web page is determined in the following manner: Using the keywords of the query request, respectively query the webpage content of the second webpage, the keywords of the webpage content of the second webpage, and the webpage description information of the second webpage to obtain a first query score of the webpage content of the second webpage, a first query score of the keywords of the webpage content of the second webpage, and a first query score of the webpage description information of the second webpage; The maximum value among the first query score of the webpage content of the second webpage, the first query score of the keyword of the webpage content of the second webpage, and the first query score of the webpage description information of the second webpage is determined as the first query score of the second webpage.

11. The method according to any one of claims 1 to 10, characterized in that: The multiple web pages are web pages used by a third party to access the target website.

12. A website information processing device, characterized in that: The device comprises: An acquisition unit, used to acquire meta information of each web page among a plurality of web pages included in the target website, wherein the meta information includes: a web page identifier and a web page link; An access unit, configured to access each web page based on a web page link of each web page; A content extraction unit, used to extract the webpage content of each webpage; The first storage unit is used to store the web page identifier and the web page content of each web page in the multiple web pages, and the correspondingly stored web page identifier and the web page content of the web page are used to provide a web page query function.

13. An electronic device, characterized in that: The electronic device comprises a processor and a memory; The processor is used to execute instructions stored in the memory so that the electronic device executes the method according to any one of claims 1 to 11.

14. A computer-readable storage medium, characterized in that: The method comprises instructions, wherein the instructions instruct a device to execute the method according to any one of claims 1 to 11.