Internet-based information screening matching method and system, and storage medium

By performing multi-level filtering and dynamic weight matching of internet information, the problems of information overload and inaccurate matching in internet information filtering are solved, achieving efficient and accurate information filtering and matching.

CN121350329BActive Publication Date: 2026-05-15国投人力资源服务有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
国投人力资源服务有限公司
Filing Date
2025-09-05
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

In existing technologies, the process of filtering information on the Internet suffers from problems such as information overload, insufficient dimensions of filtering results, and inaccurate matching, making it difficult for users to efficiently obtain high-value and highly relevant information.

Method used

By self-filtering data sources and types, optimizing the database structure, and performing multi-level searches and dynamic weight matching, information filtering is carried out in combination with multi-dimensional filtering conditions, including filtering condition deconstruction, data crawling, preliminary filtering, feature point comparison, and data sorting, thereby improving the efficiency and accuracy of the filtering process.

Benefits of technology

This improved the accuracy and efficiency of information filtering on the internet, ensuring high data value and high matching accuracy, and reducing filtering errors and mismatches.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121350329B_ABST
    Figure CN121350329B_ABST
Patent Text Reader

Abstract

The present application relates to the field of information screening, and aims to solve the problem of single keyword matching method in the process of screening and matching of Internet information, which may cause screening errors or matching errors, and specifically relates to an information screening and matching method and system based on the Internet and a storage medium; in the process of screening Internet information, the screening conditions are deconstructed in multiple dimensions, and the keyword and feature point extraction methods are applied according to the information categories, so as to improve the deconstruction accuracy of the screening conditions; in the process of information screening, the data sources and categories are self-screened, the database structure is optimized, and then the screening conditions are introduced, so as to screen the data in the database that meet the keywords; after the screening is completed, the dynamic weight matching degree sorting method is performed according to the corresponding relevance between the screening results and the multi-dimensional screening conditions, so as to improve the accuracy of the efficiency and relevance sorting in the data screening process from multiple angles.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of information filtering, specifically to an information filtering and matching method, system, and storage medium based on the Internet. Background Technology

[0002] As a crucial means of information acquisition, the internet offers a vast and diverse amount of information. However, it also suffers from drawbacks such as information overload and low value density. The increasing volume of messages and services pushed to users easily leads to severe information overload. Without effective methods, users struggle to find valuable information amidst this deluge. Therefore, acquiring high-value, highly relevant, and reliable data is paramount. To obtain the most effective data, significant manual data filtering is often required.

[0003] One current solution is a search system. Users have specific needs, which are transformed into search terms. These terms are then submitted to the corresponding search engine. The search engine retrieves relevant information from a vast amount of data and displays it to the user. However, because different search engines have different built-in search algorithms, the search results of a single search engine often have certain limitations. At the same time, most current search methods use a single keyword search comparison approach, resulting in insufficient dimensions of information filtering results and making it easy for filtering errors or inaccurate matching to occur.

[0004] To address the aforementioned technical problems, this application proposes a solution. Summary of the Invention

[0005] This invention optimizes the database structure by self-filtering data sources and types during the information filtering process. Then, filtering conditions are introduced, and multi-level searches are performed based on these conditions to filter data that matches the keywords in the database. After filtering, a dynamic weighted ranking method is implemented based on the correlation between the filtering results and the multi-dimensional filtering conditions. This improves the efficiency and accuracy of correlation ranking in the data filtering process from multiple perspectives, solving the problem that the keyword matching method is too simplistic in internet information filtering and matching, which easily leads to filtering omissions or inaccurate matching. Therefore, this invention proposes an internet-based information filtering and matching method, system, and storage medium.

[0006] The objective of this invention can be achieved through the following technical solutions:

[0007] Internet-based information filtering and matching methods include the following steps:

[0008] Step 1: Obtain the filtering conditions and deconstruct them to obtain the filtering feature points;

[0009] Step 2: Crawl information from the Internet and record the crawled information as the initial database;

[0010] Step 3: Perform preliminary screening of the data sources and data types in the initial database to obtain the initial screening database;

[0011] Step 4: Based on the feature points selected in Step 1, compare the feature points in the initial screening database to obtain multiple single-point databases;

[0012] Step 5: Select duplicate data from multiple single-point databases and retain the duplicate data as the final screening database;

[0013] Step Six: Compare the data in the final screening database with the screening feature points from Step One to obtain a data correlation ranking, and use the data correlation ranking as the screening matching result.

[0014] The Internet-based information filtering and matching system includes a filtering condition deconstruction module, a data information crawling module, a preliminary filtering module, a feature point comparison module, a data set comparison module, and a data sorting and output module.

[0015] The filtering condition deconstruction module can obtain filtering conditions through manual input. The information types of the filtering conditions include any one or more of text information, image information, or voice information. When deconstructing the filtering conditions, the filtering condition deconstruction module first determines the information types contained in the filtering conditions, separates the different information types, and deconstructs each information type separately to obtain feature points of a single type. Then, it merges the feature points of each single type to obtain the filtering feature points.

[0016] The data crawling module is used to crawl information via the Internet to obtain an initial database;

[0017] The preliminary screening module accesses the initial database obtained by the data crawling module and judges the data source and data type in the initial database to obtain the data source and data type corresponding to each initial data in the initial database. The preliminary screening module filters the data type and data source to obtain the preliminary screening database.

[0018] After obtaining the filtered feature points, the feature point comparison module selects one feature point from the filtered feature points one by one, matches the selected feature point with the feature points of the data in the initial screening database, and records all the matched feature points as a single-point database.

[0019] The data set comparison module selects data that appears in at least two single-point databases as the final screening database;

[0020] The data sorting and output module counts the number of times a data point appears in different single-point databases, records this as data correlation, and outputs the data based on the data correlation.

[0021] In a preferred embodiment of the present invention, when the filtering condition deconstruction module deconstructs text information, it first obtains a pre-built custom knowledge base, then performs keyword recognition through an algorithm, and uses the recognized keywords as feature points obtained from the deconstruction of text information.

[0022] When deconstructing image information, the filtering condition deconstruction module first obtains feature points through spatial transformation and extracts typical feature types in the image information. The typical feature types include: edges, corners, regions and ridges. After the typical feature types are extracted, the first k words are selected as feature points using the chi-square test.

[0023] When deconstructing speech information, the filtering condition deconstruction module first filters the speech information for noise, then converts the speech waveform into a parameter representation, then extracts multi-dimensional feature vectors for each speech signal and performs recognition to obtain the feature points corresponding to the speech information.

[0024] In a preferred embodiment of the present invention, after the preliminary screening module accesses the initial database, it judges the data types in the database and divides the same data types into a data set, wherein the data types also include text data, image data and sound source data, and the data set is a text data set, an image data set and a sound source data set;

[0025] The preliminary screening module then judges the data source in each dataset, marks the same data source, sorts the marked data sources according to the frequency of occurrence, obtains the data source sort, obtains the preset confidence level M, and selects the first m data in the data source sort for retention according to the preset confidence level.

[0026] After the preliminary screening module has retained the data for each dataset, it records all datasets as the preliminary screening database.

[0027] In a preferred embodiment of the present invention, after the feature point comparison module obtains the selected feature points, it numbers each individual feature point in the selected feature points and records it as i. The feature point i is compared with the data in the initial screening database one by one. If the data in the initial screening database matches the feature point, it is recorded as single point database i.

[0028] After obtaining multiple single-point databases, the data set comparison module sequentially selects data from one single-point database and compares it with data from other single-point databases. If the data in the selected single-point database appears at least once in other single-point databases, it is retained; if the data in the selected single-point database does not appear in other single-point databases, it is deleted.

[0029] The data set comparison module iterates through each data point in each single-point database to obtain the final screening database.

[0030] In a preferred embodiment of the present invention, the feature point comparison module obtains the types of information contained in the filtering conditions through the filtering condition deconstruction module, takes the types of information contained in the filtering conditions as the main types, and takes the types of information not contained in the filtering conditions as the secondary types, and the data set comparison module marks the corresponding single-point database in the final screening database as the main database or the secondary database.

[0031] In a preferred embodiment of the present invention, the method for the data sorting output module to obtain data correlation is as follows:

[0032] The data sorting output module assigns a weight r to the primary database and a weight q to the secondary database. When counting the occurrences of data, the data sorting output module calculates the weighted occurrences of the data using a formula and records them as data relevance. The data sorting output module arranges the data in descending order of relevance and obtains a preset number of filters n, selecting the top n groups of data in the arrangement as the filtering results.

[0033] A storage medium storing a computer program, which, when executed by a processor, implements an internet-based information filtering and matching method.

[0034] Compared with the prior art, the beneficial effects of the present invention are:

[0035] In this invention, when filtering Internet information, the information is deconstructed in multiple dimensions according to the filtering conditions to obtain the types of information in the filtering conditions. Then, the extraction methods of keywords and feature points are applied in a targeted manner according to the types of information to improve the accuracy of the deconstruction of the filtering conditions and provide more accurate conditions for subsequent precise filtering and matching.

[0036] In this invention, during the information filtering process, the database structure is optimized by self-filtering the data source and data type. Then, filtering conditions are introduced, and multi-level searches are performed based on the filtering conditions to filter out data in the database that match the keywords. After the filtering is completed, a dynamic weighted matching degree ranking method is implemented based on the correspondence between the filtering results and the multi-dimensional filtering conditions. This improves the efficiency and accuracy of the correlation ranking in the data filtering process from multiple perspectives. Attached Figure Description

[0037] To facilitate understanding by those skilled in the art, the present invention will be further described below with reference to the accompanying drawings.

[0038] Figure 1 This is a system block diagram of the present invention;

[0039] Figure 2 This is a system flowchart of the present invention. Detailed Implementation

[0040] The technical solution of the present invention will be clearly and completely described below with reference to the embodiments. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0041] Example 1:

[0042] Please see Figure 1 - Figure 2 As shown, the internet-based information filtering and matching method includes the following steps:

[0043] Step 1: Obtain the filtering conditions and deconstruct them to obtain the filtering feature points;

[0044] Step 2: Crawl information from the Internet and record the crawled information as the initial database;

[0045] Step 3: Perform preliminary screening of the data sources and data types in the initial database to obtain the initial screening database;

[0046] Step 4: Based on the feature points selected in Step 1, compare the feature points in the initial screening database to obtain multiple single-point databases;

[0047] Step 5: Select duplicate data from multiple single-point databases and retain the duplicate data as the final screening database;

[0048] Step Six: Compare the data in the final screening database with the screening feature points from Step One to obtain a data correlation ranking, and use the data correlation ranking as the screening matching result.

[0049] Example 2:

[0050] Please see Figure 1 - Figure 2 As shown, the Internet-based information filtering and matching system includes a filtering condition deconstruction module, a data information crawling module, a preliminary filtering module, a feature point comparison module, a data set comparison module, and a data sorting and output module.

[0051] The filtering condition deconstruction module can obtain filtering conditions through manual input. The information types of the filtering conditions include any one or more of text information, image information, or voice information. When deconstructing the filtering conditions, the filtering condition deconstruction module first determines the information types contained in the filtering conditions, separates the different information types, and deconstructs each type of information, including text information, image information, and voice information, separately. Specifically, when deconstructing text information, it first obtains a pre-built custom knowledge base, then performs keyword recognition through an algorithm, and uses the recognized keywords as feature points obtained from the deconstruction of the text information.

[0052] When deconstructing image information, the filtering condition deconstruction module first obtains feature points through spatial transformation and extracts typical feature types from the image information. These typical feature types include: edges, corners, regions, and ridges. After extracting the typical feature types, the total number of documents (N) in the sample set is counted using the chi-square test. Then, the frequency of positive documents (A), the frequency of negative documents (B), the frequency of no positive documents, and the frequency of no negative documents for each word are counted. Next, the chi-square value of each word is calculated. Finally, each word is sorted from largest to smallest according to its chi-square value, and the top k words are selected as features, where k is the feature dimension, thus obtaining k feature points.

[0053] When deconstructing speech information, the screening condition deconstruction module first filters the speech information for noise, then transforms the speech waveform into a parametric representation using linear predictive coding (LPC) or mel-frequency cepstral coefficients (MFCC) methods, and then extracts multidimensional feature vectors for each speech signal and performs recognition to obtain the feature points corresponding to the speech information.

[0054] After obtaining all feature points of a single category, the filtering condition deconstruction module merges the feature points of each single category to obtain the filtering feature points.

[0055] The data crawling module is used to crawl information from the Internet to obtain an initial database;

[0056] The initial screening module accesses the initial database obtained by the data crawling module and judges the data source and data type in the initial database to obtain the data source and data type corresponding to each initial data in the initial database. Data of the same data type will be divided into a data set, which includes text data, image data and audio data. The data sets are text data set, image data set and audio data set.

[0057] The preliminary screening module filters the data sources after the dataset is created, resulting in a preliminary screening database. The specific method is as follows:

[0058] The initial screening module then judges the data source in each dataset, marks the same data source, sorts the marked data sources according to the frequency of occurrence, obtains the data source sort, obtains the preset confidence level M, and selects the top m data from the data source sort to retain based on the preset confidence level.

[0059] After the initial screening module has retained the data for each dataset, it records all datasets as the initial screening database.

[0060] After the feature point comparison module obtains the filter feature points, it selects one feature point from the filter feature points one by one, numbers each individual feature point in the filter feature points and records it as i, and compares feature point i with the data in the initial screening database one by one. If the data in the initial screening database matches the feature point, it is recorded as single point database i.

[0061] After obtaining multiple single-point databases, the data set comparison module selects data from one single-point database in turn and compares it with data from other single-point databases. If the data in the selected single-point database appears at least once in other single-point databases, it is retained; if the data in the selected single-point database does not appear in other single-point databases, it is deleted.

[0062] The data set comparison module iterates through each data in each single-point database to obtain the final screening database, and then selects the data that appears in at least two single-point databases as the final screening database;

[0063] The feature point comparison module obtains the types of information contained in the filtering conditions through the filtering condition deconstruction module, takes the types of information contained in the filtering conditions as the primary types, and takes the types of information not contained in the filtering conditions as the secondary types. The data set comparison module marks the corresponding single-point database in the final screening database as the primary database or the secondary database.

[0064] The data sorting and output module counts the number of times a data point appears in different single-point databases, records this as data correlation, and outputs the data based on the data correlation.

[0065] The data sorting output module obtains data correlation using the following method:

[0066] The data sorting output module assigns a weight *r* to the primary database and a weight *q* to the secondary database, where weight *r* is greater than weight *q*. When calculating the frequency of occurrences of statistical data, the data sorting output module uses a formula... The weighted frequency of occurrence of data is calculated, where X is the weighted frequency of occurrence, a is the frequency of occurrence of data in the primary database, and b is the frequency of occurrence of data in the secondary database. This frequency is recorded as data relevance. The data sorting output module arranges the data in descending order of relevance and obtains the preset number of filters n. The top n groups of data in the sorted data are selected as the filtering results.

[0067] Example 3: A storage medium storing a computer program, which, when executed by a processor, implements an Internet-based information filtering and matching method.

[0068] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods.

[0069] Any references to memory, storage, database or other media used in the embodiments provided in this application may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory.

[0070] By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0071] Thresholds, preset values, preset ranges, etc. are set for result comparison and analysis to determine good or bad. The size of these values ​​is determined by a combination of large-scale model analysis of sample data and human experience. They can also be adjusted appropriately based on seasonal or common-sense influences.

[0072] Furthermore, the settings for weighting ratios, influence factors, etc., are based on the magnitude of each parameter's influence on the results. The specific values ​​are allocated to ultimately reflect the impact on the results. The settings for input and storage are also determined by a combination of large-scale model analysis of sample data and human experience. Appropriate adjustments can also be made based on seasonal or rational influence conditions.

[0073] The preferred embodiments of the present invention disclosed above are merely illustrative of the invention. These preferred embodiments do not exhaustively describe all details, nor do they limit the invention to any specific implementation. Clearly, many modifications and variations can be made based on the content of this specification. This specification selects and specifically describes these embodiments to better explain the principles and practical applications of the invention, thereby enabling those skilled in the art to better understand and utilize the invention. The invention is limited only by the claims and their full scope and equivalents.

Claims

1. An information filtering and matching method based on the Internet, characterized in that, Includes the following steps: Step 1: Obtain the filtering conditions and deconstruct them to obtain filtering feature points. The information types of the filtering conditions include any one or more of text information, image information, or voice information. When the filtering condition deconstruction module deconstructs the filtering conditions, it first determines the information types contained in the filtering conditions and separates the different information types. It then deconstructs each type of information, including text information, image information, and voice information, separately. When deconstructing text information, it first obtains a pre-built custom knowledge base and then performs keyword recognition through an algorithm. The recognized keywords are used as feature points obtained from the deconstruction of the text information. When deconstructing image information, feature points are first obtained through spatial transformation. Typical feature types in the image information are extracted, including edges, corners, regions, and ridges. After the typical feature types are extracted, the first K words are selected as feature points using the chi-square test. When deconstructing speech information, noise is first filtered out, then the speech waveform is converted into a parametric representation, and then a multi-dimensional feature vector is extracted for each speech signal and recognized to obtain the feature points corresponding to the speech information. Step 2: Crawl information from the Internet and record the crawled information as the initial database; Step 3: Perform preliminary screening of the data sources and data types in the initial database to obtain the preliminary screening database. Specifically, determine the data source in each data set, mark the same data sources, sort the marked data sources according to the frequency of occurrence to obtain the data source sorting, obtain the preset confidence level m, select the top m data from the data source sorting according to the preset confidence level and retain them. After retaining the data of each data set, record all data sets as the preliminary screening database. Step 4: Based on the selected feature points in Step 1, compare the feature points in the initial screening database to obtain multiple single-point databases. Specifically, after obtaining the selected feature points, select one feature point from each of the selected feature points, number each individual feature point in the selected feature points and record it as i, and compare feature point i with the data in the initial screening database one by one. If the data in the initial screening database matches the feature point, then record it as single-point database i. Step 5: Select duplicate data from multiple single-point databases and retain the duplicate data as the final screening database; Step Six: Compare the data in the final screening database with the screening feature points from Step One to obtain a data correlation ranking, and use the data correlation ranking as the screening matching result.

2. An internet-based information filtering and matching system, characterized in that, It includes a filter condition deconstruction module, a data information crawling module, a preliminary filtering module, a feature point comparison module, a data set comparison module, and a data sorting and output module; The filtering condition deconstruction module can obtain filtering conditions through manual input. The information types of the filtering conditions include any one or more of text information, image information, or voice information. When deconstructing the filtering conditions, the filtering condition deconstruction module first determines the information types contained in the filtering conditions, separates the different information types, and deconstructs each information type separately to obtain feature points of a single type. Then, it merges the feature points of each single type to obtain the filtering feature points. The data crawling module is used to crawl information via the Internet to obtain an initial database; The preliminary screening module accesses the initial database obtained by the data crawling module and judges the data source and data type in the initial database to obtain the data source and data type corresponding to each initial data in the initial database. The preliminary screening module filters the data type and data source to obtain the preliminary screening database. After obtaining the filtered feature points, the feature point comparison module selects one feature point from the filtered feature points one by one, matches the selected feature point with the feature points of the data in the initial screening database, and records all the matched feature points as a single-point database. The data set comparison module selects duplicate data from multiple single-point databases and retains the duplicate data as the final screening database. The data sorting output module compares the data in the final screening database with the screening feature points to obtain a data correlation ranking, and outputs the data correlation ranking as the screening matching result. When deconstructing text information, the filtering condition deconstruction module first obtains a pre-built custom knowledge base, then performs keyword recognition through an algorithm, and uses the recognized keywords as feature points obtained from the deconstruction of text information. When deconstructing image information, the filtering condition deconstruction module first obtains feature points through spatial transformation and extracts typical feature types in the image information. The typical feature types include: edges, corners, regions and ridges. After the typical feature types are extracted, the first K words are selected as feature points using the chi-square test. When deconstructing speech information, the filtering condition deconstruction module first filters the speech information for noise, then converts the speech waveform into a parameter representation, then extracts multi-dimensional feature vectors for each speech signal and performs recognition to obtain the feature points corresponding to the speech information. After accessing the initial database, the preliminary screening module judges the data types in the database and divides the same data types into a data set. The data types include text data, image data, and sound source data. The data sets are text data set, image data set, and sound source data set. The preliminary screening module then judges the data source in each dataset, marks the same data source, sorts the marked data sources according to the frequency of occurrence, obtains the data source sorting, obtains the preset confidence level m, and selects the top m data in the data source sorting to retain according to the preset confidence level. After the preliminary screening module has retained the data for each dataset, it records all datasets as a preliminary screening database. After the feature point comparison module obtains the selected feature points, it numbers each individual feature point in the selected feature points and records it as i. It then compares feature point i with the data in the initial screening database one by one. If the data in the initial screening database matches the feature point, it records it as single point database i. After obtaining multiple single-point databases, the data set comparison module sequentially selects data from one single-point database and compares it with data from other single-point databases. If the data in the selected single-point database appears at least once in other single-point databases, it is retained; if the data in the selected single-point database does not appear in other single-point databases, it is deleted. The data set comparison module iterates through each data point in each single-point database to obtain the final screening database.

3. The Internet-based information filtering and matching system according to claim 2, characterized in that, The feature point comparison module obtains the types of information contained in the filtering conditions through the filtering condition deconstruction module, takes the types of information contained in the filtering conditions as the most important types, and takes the types of information not contained in the filtering conditions as secondary types. The data set comparison module marks the corresponding single-point database in the final screening database as the most important database or the secondary database.

4. The Internet-based information filtering and matching system according to claim 2, characterized in that, The method by which the data sorting and output module obtains data correlation is as follows: The data sorting output module assigns a weight r to the most important database and a weight q to the secondary database. When counting the occurrences of data, the data sorting output module calculates the weighted occurrences of the data using a formula and records them as data relevance. The data sorting output module arranges the data in descending order of relevance and obtains a preset number of filters n, selecting the top n groups of data in the arrangement as the filtering results.

5. A storage medium, characterized in that, It stores a computer program, which, when executed by a processor, implements the Internet-based information filtering and matching method as described in claim 1.