Information screening method, device and computer equipment

By using an automated filtering method and pre-stored website identifiers and keyword forms, the system automatically expands and updates website URLs and automatically identifies information. This solves the problem of low efficiency in manual website URL filtering and ensures the comprehensiveness of data collection and compliance with user needs.

CN116561456BActive Publication Date: 2026-01-13IND BANK CO +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310478584.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-28
Publication Date
2026-01-13
Estimated Expiration
2043-04-28

AI Technical Summary

Technical Problem

Manually screening website URLs is slow, lacks standardized criteria, and wastes manpower and resources, preventing companies from collecting relevant industry data in a timely and comprehensive manner.

Method used

By obtaining the information to be filtered from the websites corresponding to the target URLs, and using pre-stored website identifiers and keyword forms, the system automatically extracts and identifies URLs with monitoring value, thereby achieving automatic URL expansion and updates and automatic information identification.

Benefits of technology

It enables automatic expansion and updating of the number of monitored website URLs and automatic identification of information, avoiding waste of manpower and resources, and ensuring that the obtained website information better meets user needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116561456B_ABST
    Figure CN116561456B_ABST
Patent Text Reader

Abstract

The application relates to an information screening method and device, computer equipment, a storage medium and a computer program product. The information screening method comprises the following steps: obtaining to-be-screened information contained in a website corresponding to a target website; extracting to-be-identified websites contained in the website corresponding to the target website from the to-be-screened information according to a pre-stored website identifier; determining an identification result of preset type information in to-be-identified information contained in the website corresponding to the to-be-identified website according to a preset keyword list; and screening the to-be-identified website according to the identification result. Through the arrangement, waste of manpower and material resources caused by manual screening of irrelevant websites is avoided, the monitored websites are more in line with user requirements, the processor can automatically update and expand the websites to be monitored, and the finally obtained website information is more comprehensive.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing technology, and in particular to an information filtering method, apparatus, computer equipment, storage medium, and computer program product. Background Technology

[0002] With the development and application of the internet, various types of information in people's daily lives are more closely integrated with the internet. More and more internet companies are beginning to pay attention to the application of data, mining and analyzing user behavior and preferences from behind massive amounts of data, and adjusting and optimizing themselves in a targeted manner based on user needs.

[0003] However, monitoring and analyzing data on websites requires knowing the website addresses. Typically, internet companies need employees to manually screen websites with monitoring value and collect their URLs. But due to the rapid development of the modern internet, the sheer number of websites makes the manual screening process extremely labor-intensive. Manually collecting URLs with business opportunities is not only slow, but also makes it difficult to standardize the criteria for website selection. This not only wastes a lot of human and material resources, but also prevents companies from collecting relevant industry data in a timely and comprehensive manner. Summary of the Invention

[0004] Therefore, it is necessary to provide an information filtering method, apparatus, computer equipment, storage medium, and computer program product that can automatically filter websites with business opportunities to address the above-mentioned technical problems.

[0005] Firstly, this application provides an information filtering method, including:

[0006] Obtain the filterable information contained in the website corresponding to the target URL;

[0007] Based on the pre-stored website identifiers, extract the target URLs contained in the websites corresponding to the target URLs from the information to be filtered;

[0008] Based on the preset keyword form, determine the recognition result of the preset type information in the information to be recognized contained in the website corresponding to the URL to be recognized;

[0009] Based on the identification results, the URLs to be identified are filtered.

[0010] In one embodiment, the step of extracting the target URL contained in the website corresponding to the target URL from the filter information based on the pre-stored website identifier includes:

[0011] The information to be filtered is matched with the website identifier;

[0012] Based on the successfully matched website identifiers and the pre-stored one-to-one correspondence between the website identifiers and the URLs to be filtered, the URLs to be filtered included in the information to be filtered are determined;

[0013] URLs that meet the preset rules are selected as the URLs to be identified.

[0014] In one embodiment, the information to be identified includes text to be identified;

[0015] The step of determining the recognition result of preset type information in the information to be recognized contained in the website corresponding to the URL to be recognized, based on a preset keyword form, includes:

[0016] Match the text to be identified with the keyword form;

[0017] Based on the matched keywords, the recognition result of the preset type information in the information to be recognized is determined.

[0018] In one embodiment, the information to be identified includes an image to be identified;

[0019] The step of determining the recognition result of preset type information in the information to be recognized contained in the website corresponding to the URL to be recognized, based on a preset keyword form, includes:

[0020] Extract the text information from the image to be identified;

[0021] The extracted text information is matched with the keyword form;

[0022] Based on the matched keywords, the recognition result of the preset type information in the information to be recognized is determined.

[0023] In one embodiment, the information to be identified includes a table to be identified;

[0024] The step of determining the recognition result of preset type information in the information to be recognized contained in the website corresponding to the URL to be recognized, based on a preset keyword form, includes:

[0025] Extract the text information from the table to be identified;

[0026] The extracted text information is matched with the keyword form;

[0027] Based on the matched keywords, the recognition result of the preset type information in the information to be recognized is determined.

[0028] In one embodiment, the information to be identified includes sequentially arranged sub-tables to be identified; and the sub-tables to be identified contain at least one of the sub-tables to be identified;

[0029] Before extracting the text information from the table to be identified, the following steps are included:

[0030] The sub-tables to be identified are merged to form the table to be identified;

[0031] The step of merging the sub-tables to be identified to form the table to be identified includes:

[0032] Obtain the header recognition results of each of the sub-tables to be identified;

[0033] According to the arrangement order of the sub-tables to be identified, the sub-tables to be identified whose header identification result is that they do not contain a header are merged with the sub-tables to be identified whose header identification result is that they contain a header, to form the table to be identified.

[0034] In one embodiment, determining the recognition result of the preset type information in the information to be identified based on the matched keywords includes:

[0035] The number of matched keywords is used as the recognition result of the preset type information in the information to be identified;

[0036] The step of filtering the URLs to be identified based on the identification results includes:

[0037] Remove the URLs to be identified that do not reach the preset threshold, thus completing the filtering of the URLs to be identified.

[0038] Secondly, this application also provides an information filtering device, comprising:

[0039] The acquisition module is used to obtain the information to be filtered contained in the website corresponding to the target URL;

[0040] The first extraction module is used to extract the target URL contained in the website corresponding to the target URL from the information to be filtered based on the pre-stored website identifier;

[0041] The determination module is used to determine the recognition result of the preset type information in the information to be identified contained in the website corresponding to the URL to be identified, based on the preset keyword form;

[0042] The filtering module is used to filter the URLs to be identified based on the recognition results.

[0043] Thirdly, this application also provides a computer device. The computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the information filtering method described in any of the above embodiments.

[0044] Fourthly, this application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program thereon, which, when executed by a processor, implements the information filtering method described in any of the above embodiments.

[0045] Fifthly, this application also provides a computer program product. The computer program product includes a computer program that, when executed by a processor, implements the information filtering method described in any of the above embodiments.

[0046] The aforementioned information filtering method, apparatus, computer equipment, storage medium, and computer program product automatically extract the URLs to be identified from the websites corresponding to the pre-stored target URLs. These websites are then used as candidate monitoring websites, thus automatically expanding and updating the number of monitored websites. Subsequently, the identification information contained in the websites corresponding to the URLs to be identified is extracted, and combined with a keyword form, the identification results for each website corresponding to the URL to be identified are obtained. This achieves automatic identification of information from monitored websites. Based on the identification results, the URLs to be identified are filtered to obtain URLs with monitoring value. This avoids the waste of human and material resources caused by manual filtering of irrelevant websites, making the monitored websites more in line with user needs. The processor can automatically update and expand the websites to be monitored, resulting in more comprehensive website information. Attached Figure Description

[0047] Figure 1 This is a diagram illustrating the application environment of an information filtering method in one embodiment;

[0048] Figure 2 This is a flowchart illustrating an information filtering method in one embodiment;

[0049] Figure 3 This is a flowchart illustrating an information filtering method in one embodiment;

[0050] Figure 4 This is a flowchart illustrating an information filtering method in one embodiment;

[0051] Figure 5 This is a flowchart illustrating an information filtering method in one embodiment;

[0052] Figure 6 This is a flowchart illustrating an information filtering method in one embodiment;

[0053] Figure 7 This is a flowchart illustrating an information filtering method in one embodiment;

[0054] Figure 8 This is a flowchart illustrating an information filtering method in one embodiment;

[0055] Figure 9 This is a structural block diagram of an information filtering device in one embodiment;

[0056] Figure 10 This is a structural block diagram of the extraction module in an information filtering device in one embodiment;

[0057] Figure 11 This is a structural block diagram of the determination module in the information filtering device in one embodiment;

[0058] Figure 12 This is a structural block diagram of the determination module in the information filtering device in one embodiment;

[0059] Figure 13 This is a structural block diagram of the determination module in the information filtering device in one embodiment;

[0060] Figure 14 This is a structural block diagram of the determination module in the information filtering device in one embodiment;

[0061] Figure 15 This is a structural block diagram of the merging unit in an information filtering device in one embodiment;

[0062] Figure 16 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation

[0063] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0064] The information filtering method provided in this application embodiment can be applied to, for example, Figure 1 In the application environment shown, terminal 102 communicates with server 104 via a network.

[0065] For example, the information filtering method is applied to terminal 102. Terminal 102 first obtains the information to be filtered contained in the website corresponding to the target URL; then, based on the pre-stored website identifiers, it extracts the URLs to be identified contained in the website corresponding to the target URL from the information to be filtered; it extracts the information to be identified contained in the website corresponding to the URL to be identified; finally, terminal 102 determines the recognition result of the preset type of information in the information to be identified based on a preset keyword form, and sends the recognition result to server 104. Server 104 saves the recognition result to the data storage system. Terminal 102 can be, but is not limited to, various personal computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. IoT devices can be smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, etc. Portable wearable devices can be smartwatches, smart bracelets, head-mounted devices, etc. Server 104 can be implemented using a standalone server or a server cluster composed of multiple servers. Terminal 102 and server 104 can be directly or indirectly connected via wired or wireless communication, such as through a network connection.

[0066] For example, the information filtering method is applied to server 104. Server 104 can obtain the target URL from terminal 102, and then obtain the information to be filtered contained in the website corresponding to the target URL; according to the pre-stored website identifier, it extracts the URL to be identified contained in the website corresponding to the target URL from the information to be filtered; it extracts the information to be identified contained in the website corresponding to the URL to be identified; finally, server 104 determines the recognition result of the preset type information in the information to be identified according to the preset keyword form, and saves the recognition result to the data storage system. It can be understood that the data storage system can be an independent storage device, or the data storage system can be located on the server, or the data storage system can be located on another terminal.

[0067] In one embodiment, an information filtering method is provided. This embodiment illustrates this method by applying it to a processor. It is understood that the processor may be located on a terminal or a server. Figure 2 As shown, the information filtering method includes:

[0068] Step 202: Obtain the information to be filtered contained in the website corresponding to the target URL.

[0069] The target URL can refer to a URL pre-stored by the processor. Users can pre-store URLs that have monitoring value into the processor.

[0070] The information to be filtered can refer to information such as text, images, tables, and Uniform Resource Locators (URLs) contained in the website corresponding to the target URL.

[0071] The processor can extract the pre-stored target URL from the data storage system according to a pre-set processing frequency, and further extract the text, images, tables and other information content contained in the website corresponding to the target URL.

[0072] Step 204: Based on the pre-stored website identifiers, extract the URLs to be identified from the information to be filtered, which are contained in the website corresponding to the target URL.

[0073] A website identifier can consist of at least one of letters, characters, or numbers. The website identifier is used to uniquely identify the corresponding URL. In this embodiment, the processor pre-stores a one-to-one mapping relationship between website identifiers and multiple URLs.

[0074] As an example, after obtaining the text, images, tables, and other information content contained in the website corresponding to the target URL in step 202, the text content contained in the images and tables can be extracted. All information in various forms in the information to be filtered can be converted into text type. Then, the text content corresponding to the text, images, tables, and other information content contained in the information to be filtered is matched with the website identifier. The URL corresponding to the successfully matched website identifier is extracted from the pre-stored one-to-one mapping relationship between website identifiers and multiple URLs, and used as the URL to be identified. Alternatively, all URLs contained in the website corresponding to the target URL can be obtained, the URLs are matched with the website identifiers, and the URL corresponding to the successfully matched website identifier is extracted from the pre-stored one-to-one mapping relationship between website identifiers and multiple URLs, and used as the URL to be identified. In this embodiment, the website identifier can be in the form of a URL.

[0075] The URL to be identified refers to the URL information contained in the website corresponding to the target URL that can redirect to other websites, or the URL information involved in the website corresponding to the target URL.

[0076] As an example, regular expression matching can also be performed on the information to be filtered. The website identifier can be in the form of a regular expression. The regular expression contained in the information to be filtered is matched with the website identifier to obtain the URL to be identified.

[0077] Step 206: Based on the preset keyword form, determine the recognition result of the preset type of information in the information to be recognized contained in the website corresponding to the URL to be recognized.

[0078] The information to be identified can refer to information such as text, images, tables, and Uniform Resource Locators (URLs) contained in the website corresponding to the URL to be identified.

[0079] A pre-defined keyword form refers to a form containing various types of words. The keyword form can contain multiple pre-defined word types, each corresponding one-to-one with pre-defined information types. Pre-defined information types refer to information that has been pre-categorized by the user into different types. For example, it could be financial information types, online sales types, etc. As an example, keywords for financial information types could include "account," "income," "expenses," "deductions," and "balance," while keywords for online sales types could include "activities," "recharge," "hot-selling," "benefits," "free," and "great value," etc.

[0080] The recognition result can refer to the matching result between the information to be recognized and the keyword form.

[0081] In this embodiment, after the processor extracts the URL to be identified, it first obtains the information to be identified contained in the website corresponding to the URL, and then compares and matches the information to be identified with a preset keyword form to obtain the matching results of the information to be identified with the keywords in each preset type, and obtains the identification result.

[0082] Step 208: Based on the recognition results, filter the URLs to be recognized.

[0083] The processor filters the URLs to be identified based on the matching results of the information to be identified contained in the website corresponding to the URL to be identified and the keywords in each preset type, so as to obtain URL information with monitoring value.

[0084] In one embodiment, the filtered URLs to be identified can be used as target URLs, thereby enabling automatic updates and iterative expansion of URL information with monitoring value.

[0085] In the aforementioned information filtering method, the processor can automatically extract the URLs to be identified from the websites corresponding to the pre-stored target URLs, and use these websites as candidate monitoring websites. It then extracts the information to be identified from these websites, combines it with a keyword form, obtains the identification results for each website corresponding to the target URL, and filters the URLs based on these results to obtain those with monitoring value. This setup enables automatic expansion and updating of the number of monitored websites, as well as automatic identification of website information. It avoids the waste of manpower and resources caused by manual filtering of irrelevant websites, making the monitored websites more in line with user needs. The processor can automatically update and expand the websites to be monitored, resulting in more comprehensive website information.

[0086] like Figure 3 As shown, in some optional embodiments, step 204 includes:

[0087] Step 2042: Match the information to be filtered with the website identifier;

[0088] Step 2044: Based on the successfully matched website identifiers and the one-to-one correspondence between the pre-stored website identifiers and the URLs to be filtered, determine the URLs to be filtered that are included in the information to be filtered;

[0089] Step 2046: Select the URLs that meet the preset rules as the URLs to be identified.

[0090] As an example, the information to be identified contains all the URLs contained in the website corresponding to the URL to be identified. The website identifier can be at least one URL. The website identifier is matched with the information to be identified, and then the URL to be filtered corresponding to the successfully matched URL is obtained from the one-to-one correspondence between the website identifier and the URL to be filtered.

[0091] As an example, the preset rule can refer to the regular expression that the processor has saved in advance. When the regular expression corresponding to the URL to be filtered matches the regular expression in the preset rule, the URL to be filtered is considered to conform to the preset rule. Otherwise, the URL to be filtered is considered not to conform to the preset rule, and the URL to be filtered that conforms to the preset rule is taken as the URL to be identified.

[0092] like Figure 4 As shown, in some optional embodiments, the information to be identified includes text to be identified;

[0093] Step 206 includes:

[0094] Step 2062: Match the text to be recognized with the keyword form;

[0095] Step 2064: Based on the matched keywords, determine the recognition result of the preset type information in the information to be recognized.

[0096] The information to be identified includes all text information contained in the website corresponding to the URL to be identified. The processor matches all text information contained in the website corresponding to the URL to be identified with the keyword form, and further extracts the keywords that successfully match the keywords in the website corresponding to the URL to be identified with the keyword form, thereby determining the identification result.

[0097] like Figure 5 As shown, in some optional embodiments, the information to be identified includes an image to be identified;

[0098] Step 206 includes:

[0099] Step 2066: Extract text information from the image to be recognized;

[0100] Step 2068: Match the extracted text information with the keyword form;

[0101] Step 20610: Based on the matched keywords, determine the recognition result of the preset type information in the information to be recognized.

[0102] The information to be identified includes all the images contained in the website corresponding to the URL to be identified. The processor first extracts the text information from all the images in the website corresponding to the URL to be identified, and then matches the text information with the keyword form. Subsequently, it extracts the keywords that successfully match all the text information in the website corresponding to the URL to be identified with the keyword form, thereby determining the recognition result.

[0103] like Figure 6 As shown, in some optional embodiments, the information to be identified includes a table to be identified;

[0104] Step 206 includes:

[0105] Step 20612: Extract text information from the table to be recognized;

[0106] Step 20614: Match the extracted text information with the keyword form;

[0107] Step 20616: Based on the matched keywords, determine the recognition result of the preset type information in the information to be recognized.

[0108] The information to be identified includes all the tables contained in the website corresponding to the URL to be identified. The processor first extracts the text information from all the tables in the website corresponding to the URL to be identified, and then matches the text information with the keyword form. Subsequently, it extracts the keywords that successfully match all the text information in the website corresponding to the URL to be identified with the keyword form, thereby determining the identification result.

[0109] like Figure 7-8 As shown, in some optional embodiments, the information to be identified includes sequentially arranged sub-tables to be identified; and the table to be identified contains at least one sub-table to be identified.

[0110] Before step 20612, the following are included:

[0111] Step 20611: Merge the sub-tables to be identified to form a table to be identified;

[0112] Step 20611 includes:

[0113] Step 206112: Obtain the header recognition results of each sub-table to be recognized;

[0114] Step 206114: According to the order of the sub-tables to be identified, merge the sub-tables whose header identification result is that they do not contain a header with the sub-tables whose header identification result is that they contain a header, to form the table to be identified.

[0115] Since the tables on the websites corresponding to the URLs to be identified may be too large, the processor may divide a table into multiple sub-tables for display when extracting the information to be identified, so it is necessary to merge the table contents.

[0116] As an example, the processor can use information such as the font and pixel values ​​of the text in all the sub-tables to be identified to filter out the text whose font and pixel values ​​are different from those of all the text in the preceding and following sub-tables to be identified, and use it as the header of the current sub-table to be identified.

[0117] The header of the sub-table to be identified can be obtained using any header identification method, as long as it can identify and extract the header. The above-mentioned methods for identifying headers are merely examples and not limitations of this application. Any implementation method that determines the headers contained in the sub-table to be identified based on the information of the sub-table should be included within the protection scope of this application.

[0118] As an example, the information to be identified includes sub-tables A, B, C, D, and E arranged in sequence. The header identification result indicates that sub-tables A and E contain headers. Therefore, for sub-tables B, C, and D, the corresponding previous header identification result is that the sub-table containing the header is A. Thus, sub-tables B, C, and D are merged with sub-table A to obtain the table to be identified.

[0119] In some optional embodiments, the step of determining the recognition result of preset type information in the information to be recognized based on the matched keywords includes:

[0120] The number of matched keywords is used as the recognition result of the preset type of information in the information to be recognized;

[0121] Step 208 includes: removing URLs that do not meet the preset threshold number to be identified, thus completing the filtering of URLs to be identified.

[0122] In this embodiment, if the number of keywords successfully matched by the information to be identified in the website corresponding to the URL to be identified reaches a preset threshold, the URL to be identified can be considered to have monitoring value; if the number of keywords successfully matched by the information to be identified in the website corresponding to the URL to be identified does not reach the preset threshold, the URL to be identified can be considered to have no monitoring value.

[0123] The aforementioned information filtering method can automatically extract the URLs to be identified from the websites corresponding to the pre-stored target URLs, and use these websites as candidate monitoring websites. It then extracts the information to be identified from these websites, combines it with a keyword form, obtains the identification results for each website corresponding to the target URL, and filters the URLs based on these results to obtain those with monitoring value. This setup enables automatic expansion and updating of the number of monitored websites, as well as automatic identification of website information. It avoids the waste of manpower and resources caused by manual filtering of irrelevant websites, making the monitored websites more in line with user needs. The processor can automatically update and expand the websites to be monitored, resulting in more comprehensive website information.

[0124] It should be understood that although the steps in the flowcharts of the above embodiments are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the above embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.

[0125] Based on the same inventive concept, this application also provides a memory data access apparatus for implementing the memory data access method described above. The solution provided by this apparatus is similar to the implementation described in the above method; therefore, the specific limitations in one or more memory data access apparatus embodiments provided below can be found in the limitations of the memory data access method described above, and will not be repeated here.

[0126] In one embodiment, such as Figure 9 As shown, an information filtering device 900 is provided, including: an acquisition module 902, an extraction module 904, a determination module 906, and a filtering module 908, wherein:

[0127] The acquisition module 902 is used to acquire the filter information contained in the website corresponding to the target URL;

[0128] The extraction module 904 is used to extract the target URL contained in the website corresponding to the target URL from the information to be filtered based on the pre-stored website identifier;

[0129] The determination module 906 is used to determine the recognition result of the preset type of information in the information to be recognized contained in the website corresponding to the URL to be recognized, based on the preset keyword form;

[0130] The filtering module 908 is used to filter the URLs to be identified based on the recognition results.

[0131] like Figure 10 As shown, in some optional embodiments, the extraction module 904 includes:

[0132] The first matching unit 9042 is used to match the information to be filtered with the website identifier;

[0133] The first determining unit 9044 is used to determine the URLs to be filtered included in the information to be filtered based on the successfully matched website identifiers and the one-to-one correspondence between the pre-stored website identifiers and the URLs to be filtered.

[0134] The filtering unit 9046 is used to select URLs that meet preset rules as URLs to be identified.

[0135] like Figure 11 As shown, in some optional embodiments, the information to be identified includes text to be identified;

[0136] Module 906 includes:

[0137] The second matching unit 9062 is used to match the text to be identified with the keyword form;

[0138] The second determining unit 9064 is used to determine the recognition result of the preset type information in the information to be recognized based on the matched keywords.

[0139] like Figure 12 As shown, in some optional embodiments, the information to be identified includes an image to be identified;

[0140] Module 906 includes:

[0141] The first extraction unit 9066 is used to extract text information from the image to be recognized;

[0142] The third matching unit 9068 matches the extracted text information with the keyword form;

[0143] The third determining unit 90610 determines the recognition result of the preset type information in the information to be recognized based on the matched keywords.

[0144] like Figure 13 As shown, in some optional embodiments, the information to be identified includes a table to be identified;

[0145] Module 906 includes:

[0146] The second extraction unit 90612 extracts text information from the table to be recognized;

[0147] The fourth matching unit 90614 matches the extracted text information with the keyword form;

[0148] The fourth determining unit 90616 determines the recognition result of the preset type information in the information to be recognized based on the matched keywords.

[0149] like Figure 14-15 As shown, in some optional embodiments, the information to be identified includes sequentially arranged sub-tables to be identified; and the table to be identified contains at least one sub-table to be identified.

[0150] Module 906 also includes:

[0151] Merging unit 90611 is used to merge the sub-tables to be identified to form a table to be identified;

[0152] Merging unit 90611 includes:

[0153] Component 906112 is used to obtain the header recognition results of each sub-table to be recognized;

[0154] The merging component 906114 is used to merge the sub-tables to be identified whose header identification result is "not containing a header" with the previous sub-tables whose header identification result is "containing a header", according to the arrangement order of the sub-tables to be identified, to form a table to be identified.

[0155] Each module in the aforementioned information filtering device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in the processor of a computer device in hardware form or independent of it, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.

[0156] In one embodiment, a computer device is provided, which may be a terminal, and its internal structure diagram may be as follows: Figure 16The computer device includes a processor, memory, input / output interfaces, a communication interface, a display unit, and input devices. The processor, memory, and input / output interfaces are connected via a system bus, and the communication interface, display unit, and input devices are also connected to the system bus via the input / output interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The input / output interfaces are used for exchanging information between the processor and external devices. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it implements an information filtering method. The display unit is used to form a visually visible image and can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be an LCD screen or an e-ink screen. The input device of the computer device can be a touch layer covering the display screen, or buttons, trackballs, or touchpads set on the casing of the computer device, or external keyboards, touchpads, or mice, etc.

[0157] Those skilled in the art will understand that Figure 16 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0158] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the various steps of the information filtering method described above.

[0159] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the various steps of the information filtering method described above.

[0160] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.

[0161] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0162] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. An information filtering method, characterized in that, include: Obtain the filterable information contained in the website corresponding to the target URL; Based on the pre-stored website identifiers, extract the target URLs contained in the websites corresponding to the target URLs from the information to be filtered; Based on the preset keyword form, determine the recognition result of the preset type information in the information to be recognized contained in the website corresponding to the URL to be recognized; Based on the identification results, the URLs to be identified are filtered. The step of extracting the target URL from the filtering information based on the pre-stored website identifiers includes: The information to be filtered is matched with the website identifier; Based on the successfully matched website identifiers and the pre-stored one-to-one correspondence between the website identifiers and the URLs to be filtered, the URLs to be filtered included in the information to be filtered are determined; URLs that meet the preset rules are selected as the URLs to be identified.

2. The method according to claim 1, characterized in that, The information to be identified includes the text to be identified; The step of determining the recognition result of preset type information in the information to be recognized contained in the website corresponding to the URL to be recognized, based on a preset keyword form, includes: Match the text to be identified with the keyword form; Based on the matched keywords, the recognition result of the preset type information in the information to be recognized is determined.

3. The method according to claim 1, characterized in that, The information to be identified includes the image to be identified; The step of determining the recognition result of preset type information in the information to be recognized contained in the website corresponding to the URL to be recognized, based on a preset keyword form, includes: Extract the text information from the image to be identified; The extracted text information is matched with the keyword form; Based on the matched keywords, the recognition result of the preset type information in the information to be recognized is determined.

4. The method according to claim 1, characterized in that, The information to be identified includes a table to be identified; The step of determining the recognition result of preset type information in the information to be recognized contained in the website corresponding to the URL to be recognized, based on a preset keyword form, includes: Extract the text information from the table to be identified; The extracted text information is matched with the keyword form; Based on the matched keywords, the recognition result of the preset type information in the information to be recognized is determined.

5. The method according to claim 4, characterized in that, The information to be identified includes sequentially arranged sub-tables to be identified; and each sub-table to be identified contains at least one of the sub-tables to be identified. Before extracting the text information from the table to be identified, the following steps are included: The sub-tables to be identified are merged to form the table to be identified; The step of merging the sub-tables to be identified to form the table to be identified includes: Obtain the header recognition results of each of the sub-tables to be identified; According to the arrangement order of the sub-tables to be identified, the sub-tables to be identified whose header identification result is that they do not contain a header are merged with the sub-tables to be identified whose header identification result is that they contain a header, to form the table to be identified.

6. The method according to any one of claims 2-5, characterized in that, The step of determining the recognition result of the preset type information in the information to be recognized based on the matched keywords includes: The number of matched keywords is used as the recognition result of the preset type information in the information to be identified; The step of filtering the URLs to be identified based on the identification results includes: Remove the URLs to be identified that do not reach the preset threshold, thus completing the filtering of the URLs to be identified.

7. An information filtering device, characterized in that, include: The acquisition module is used to obtain the information to be filtered contained in the website corresponding to the target URL; The first extraction module is used to extract the target URL contained in the website corresponding to the target URL from the information to be filtered based on the pre-stored website identifier; The determination module is used to determine the recognition result of the preset type information in the information to be identified contained in the website corresponding to the URL to be identified, based on the preset keyword form; The filtering module is used to filter the URLs to be identified based on the recognition results; The first extraction module includes: The first matching unit is used to match the information to be filtered with the website identifier; The first determining unit is used to determine the URL to be filtered contained in the information to be filtered based on the successfully matched website identifier and the pre-stored one-to-one correspondence between the website identifier and the URL to be filtered. The filtering unit is used to select URLs that meet preset rules as the URLs to be identified.

8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the information filtering method according to any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the information filtering method according to any one of claims 1 to 6.

10. A computer program product, comprising a computer program, characterized in that, When executed by a processor, the computer program implements the steps of the information filtering method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Website operation state monitoring and abnormal detection based on MapReduce

    CN102724059A

  • Webpage information acquisition method and device and computer readable storage medium

    CN109902220A