Price comparison method based on webpage capture and text similarity
By combining the Scrapy framework and the TF-IDF algorithm with the cosine similarity algorithm, this price comparison tool solves the problems of data lag and inaccurate matching in existing tools, and achieves efficient and accurate price comparison and interactive user experience in the Chinese market.
Patent Information
- Application Number
- CN202511010848.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-22
- Publication Date
- 2025-11-07
Smart Images

Figure CN120910334A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of electronic commerce, and particularly relates to a price comparison method based on webpage crawling and text similarity. BACKGROUND
[0002] In recent years, the popularity and rapid growth of e-commerce websites have reached an unprecedented level. A large number of consumers rely on e-commerce websites for shopping instead of physical stores. However, the operators of e-commerce websites sometimes intentionally set prices away from the actual prices to take advantage of people's needs. Therefore, consumers often spend more money to buy products without knowing the actual cost of different products, thus being cheated. Based on the above facts, the importance of price comparison tools is self-evident, which prompts us to develop a price comparison tool specifically for the Chinese market to enable consumers to more easily evaluate prices and make purchase decisions according to their budget.
[0003] The prior art includes the Scraoy framework, Scrapy is a powerful open source web scraping framework for efficiently extracting and processing data from websites. It automates sending requests, parsing responses, and storing data into various formats or databases by defining crawlers. Scrapy supports asynchronous processing, enabling concurrent crawling of multiple web pages, and provides a flexible component system, including crawlers, data models, pipelines, and middleware, to meet different data scraping needs.
[0004] TF-IDF algorithm, the TF-IDF method can generate vector representation, which converts text into vectors. Such vectors not only consider the frequency of word occurrence, but also measure the importance of each word in the entire database. TF (Term Frequency): Calculate the number of times each word appears in a single document. Words that appear more frequently are more important in describing the document. IDF (Inverse Document Frequency): Calculate the rarity of a word across all documents. Words that appear less frequently are considered to represent the uniqueness of the document. Combining the two, we get a score of the importance of each word in the document, which can more accurately represent the content of each document. SUMMARY
[0005] The purpose of the present application is to overcome the shortcomings of the prior art and provide a price comparison method based on webpage crawling and text similarity.
[0006] The technical solution adopted by the present application to solve its technical problems comprises the following steps:
[0007] Step 1, build Scraoy framework tool to obtain information from each website, and build a database.
[0008] Step 2, vectorization of large-scale text data, converting text data into digital form for calculation and analysis;
[0009] Step 3, construct cosine similarity algorithm.
[0010] Step 4, design user interface.
[0011] Further, step 1 uses Scrapy framework to build different crawlers for different websites to collect product information, as follows:
[0012] 1-1. Analyze the differences in page structure, product information is embedded in different positions of different websites, select a dedicated XPath selector for each website;
[0013] 1-2. Store the crawled data in the corresponding "image", "product name", "product price", "price" and "website link" columns of each website in a single CSV file;
[0014] 1-3. Repeat steps 1-1 and 1-2 until all required reference websites are traversed and relevant data is stored.
[0015] Further, since websites will update products after a certain time, an automated script is introduced to dynamically update the data in the database in real time.
[0016] Further, the automated script: through CrawlerRunner, asynchronously start multiple crawlers in the script, and limit the number of crawlers running at the same time, each crawler performs data deduplication and outputs to an independent CSV file, then unify the field names of different websites, and unify the currency format to RMB; Finally, based on the "website" to judge whether the record exists, and combined with the price and update record detection, change the input to the database.
[0017] Further, some websites use multi-layer nested JSON data, which requires the response command to extract.
[0018] Further, step 2 is implemented as follows:
[0019] Convert all text data existing in the database into word groups to preprocess the original data, including the words presented in the corresponding website query, and pair with the word count of each database; that is, first split the long text into words or word groups; Then merge and count all platform text data to generate a unified word frequency table; Finally, use TF-IDF to generate a vector representation of the preprocessed database, achieving the effect of converting unstructured text in the database into structured data.
[0020] Further, step 3 is implemented as follows:
[0021] Create a long list of repeated or less similar products, and introduce a variable to control the length of the long list, that is, the similarity threshold.
[0022] Further, step 4 is specifically operated as follows:
[0023] Create a clear main dashboard to display the search box, price comparison button and result display area; on the main interface, the user inputs the product name or keyword to be compared and clicks the "start comparison" button; the system obtains the price information of different e-commerce websites through the backend service and updates the comparison result in real time; the result page will display the prices of each website and the detailed information of the product in the form of a table or chart.
[0024] Further, the application also provides a price comparison tool based on web crawling and text similarity, which comprises a price comparison method based on web crawling and text similarity.
[0025] The application has the following advantages:
[0026] The application solves the three major pain points of data lag, inaccurate matching and redundant results in traditional price comparison tools through the technical combination of Scrapy dynamic crawling + TF-ID vectorization + cosine similarity matching + interactive UI, and is optimized for the Chinese market. BRIEF DESCRIPTION OF DRAWINGS
[0027] Figure 1 The flowchart of the application.
[0028] Figure 2 The cross-sectional schematic view of the application. DETAILED DESCRIPTION
[0029] The application will be further described below in conjunction with the embodiments and drawings. The application proposes a powerful price comparison tool. The tool uses the Scrapy framework to build a crawler to obtain the information published on the website in the website. The obtained data is stored in the database, and then the title is vectorized, and the cosine similarity algorithm is used to find the similarity between the search text and the dataset text. Finally, the user interface is designed for user-friendly interaction when searching and displaying the results of the corresponding query.
[0030] As shown in Figure 1 and 2 , a price comparison method based on web crawling and text similarity, comprising the following steps:
[0031] Step 1, build Scraoy framework tool to obtain information of each website, and build database.
[0032] Step 2: Vectorization of large-scale text data, converting text data into numerical form for computation and analysis.
[0033] Step 3: Building cosine similarity algorithm.
[0034] Step 4: Designing user interaction interface.
[0035] Further, step 1 is specifically operated as follows:
[0036] First, use Python's powerful Scrapy framework to build different crawlers for different websites to collect product information, and collect necessary data from different websites in a fast and efficient way.
[0037] First, analyze the differences in page structure, and the product information of different websites is embedded in different locations. For each website, select a dedicated XPath selector. Some websites use multi-layer nested JSON data, which requires the use of response commands to extract.
[0038] Secondly, store the crawled data in the corresponding "image", "product name", "product price", "price" and "website link" columns of the single CSV file of each website. Since websites often change products after a certain time, the tool is equipped with a dynamic function to keep up with the pace of change. This dynamic function is implemented by introducing an automated script that runs every 12 hours to update the data in the database.
[0039] The automated script: through CrawlerRunner, multiple crawlers are started asynchronously in the script, and the number of crawlers running simultaneously is limited (such as a maximum of 5 websites parallel crawling), to prevent IP from being banned. Each crawler performs data deduplication and outputs to an independent CSV file, then unifies the field names of different websites to the format in the previous text, and unifies the currency format to RMB. Finally, based on the "website" to judge whether the record exists, and combined with the price and update record detection changes to input into the database.
[0040] Repeat the above process until all the required reference websites are traversed and the relevant data is stored.
[0041] Further, step 2 is specifically operated as follows:
[0042] The process involves preprocessing the raw data by converting all text data in the database into phrases, including words presented in the corresponding website queries, and pairing them with word counts from each database. Specifically, long texts are first split into meaningful words or phrases; for example, "Huawei Mate50 Pro 5G mobile phone obsidian black 512GB" would have word segmentation results in ["Huawei", "Mate50", "Pro", "5G", "mobile phone", "obsidian black", "512GB"]. Then, text data from all platforms are merged and statistically analyzed to generate a unified word frequency table. Finally, TF-IDF is used to generate a vector representation of the preprocessed database, achieving the effect of converting unstructured text in the database into structured data.
[0043] Furthermore, step 3 is performed as follows:
[0044] Design a function to precisely match products searched by a user. Products are presented in ascending order. Therefore, the lowest-priced product is ranked first in this order. To find the consistency between the user's input products and the data existing in the database, construct a cosine similarity algorithm.
[0045] Cosine similarity distinguishes two or more documents based on direction (angle), not size. The similarity between two documents is calculated using the cosine of the angle between the two vectors. Cosine similarity focuses on the frequency of common words between common words. Furthermore, cosine similarity overcomes the limitation that two documents may not match regardless of size. For example, if a and b are two vectors, then cosine similarity uses the following rules:
[0046]
[0047] If the cosine value is 0, then the angle between vectors a and b is 90 degrees, meaning the two vectors are not similar. Therefore, it can be seen that when the cosine value of the angle between two vectors is small, the two vectors are similar.
[0048] Based on the similarity rate, a long list containing duplicate or dissimilar products is created. This certainly increases the difficulty of finding suitable products from a lengthy list. Therefore, this invention introduces a variable to control the length of the list, namely the similarity threshold.
[0049] The similarity threshold controls the number of similar products a user wants to see in the results list. It is controlled by a numbered slider that specifies the minimum percentage of similarity required for a product to appear in the results. This value varies between 0% and 99%.
[0050] Furthermore, step 4 is performed as follows:
[0051] Create a clear main dashboard that displays a search box, price comparison button, and result display area. On the main interface, the user inputs the product name or keyword to be compared and clicks the "Start Comparison" button. The system obtains price information from different e-commerce websites through the backend service and updates the comparison results in real time. The result page will display the prices of each website and the detailed information of the product in the form of a table or chart. Users can also view the price history, set price reminders, and export the comparison results on the result page.
Claims
1. A price comparison method based on web scraping and text similarity, characterized in that, Comprising the following steps: Step 1, building Scraoy framework tool to obtain information of each website, building database; Step 2, vectorizing large-scale text data, converting text data into digital form for calculation and analysis; Step 3, building cosine similarity algorithm; Step 4, designing user interaction interface.
2. The price comparison method based on web scraping and text similarity according to claim 1, characterized in that, Step 1 uses Scrapy framework to build different crawlers for different websites to collect product information, as follows: 1-1. Analyze the differences in page structure, and select different XPath selectors for each website according to the different embedding positions of product information; 1-2. Store the crawled data in the corresponding "image", "product name", "product price", "price" and "website link" columns of each website in a single CSV file; 1-3. Repeat steps 1-1 and 1-2 until all required reference websites are traversed and relevant data is stored.
3. The price comparison method based on web scraping and text similarity according to claim 2, characterized in that, Since websites may change products after a certain time, an automated script is introduced to dynamically update the data in the database in real time.
4. The price comparison method based on web scraping and text similarity according to claim 3, characterized in that, The automated script: through CrawlerRunner, asynchronously start multiple crawlers in the script, and limit the number of crawlers running at the same time, each crawler performs data deduplication and outputs to an independent CSV file, then formats the field names of different websites uniformly, and unifies the currency format to RMB; Finally, based on the "website" to judge whether the record exists, and combined with the price and update record detection, change the input to the database.
5. The price comparison method based on web scraping and text similarity according to claim 3, characterized in that, Some websites use multi-layer nested JSON data, which need to be extracted in combination with the response command.
6. The price comparison method based on web scraping and text similarity according to claim 3, characterized in that, Step 2 is implemented as follows: Convert all text data in the database into word groups to preprocess the original data, including the words presented in the corresponding website query, and pair them with the word count of each database; That is, first split the long text into words or word groups; Then merge and count all platform text data to generate a unified word frequency table; Finally, use TF-IDF to generate a vector representation of the preprocessed database, achieving the effect of converting unstructured text data in the database into structured data.
7. The price comparison method based on web scraping and text similarity according to claim 6, characterized in that, Step 3 is implemented as follows: Create a long list containing duplicate or less similar products, and introduce a variable to control the length of the long list, i.e. the similarity threshold.
8. The price comparison method based on web scraping and text similarity according to claim 6, characterized in that, Step 4 is implemented as follows: Create a clear main dashboard that displays a search box, a price comparison button, and a result display area; On the main interface, the user inputs the product name or keyword to be compared and clicks the "Start Comparison" button; The system obtains price information from different e-commerce websites through the backend service and updates the comparison results in real time; The result page will display the prices of each website and the detailed information of the product in the form of a table or chart.
9. A price comparison tool based on web scraping and text similarity, characterized in that, The tool comprises the method of claim 1.