Method for implementing Amazon commodity information collection based on browser plug-in
By injecting scripts into the browser, simulating internal requests of Amazon websites and avoiding risk control problems, it realizes rapid and comprehensive collection of Amazon product information, and solves the problems of low data collection efficiency and risk control restrictions in traditional methods.
Patent Information
- Application Number
- CN202411992791.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-16
AI Technical Summary
When facing the Amazon platform, traditional data collection methods are difficult to obtain comprehensive and in-depth product information, and are restricted by Amazon's strict risk control mechanism, resulting in low data collection efficiency and high cost.
By injecting scripts into the browser, we simulate internal requests of Amazon websites, avoid risk control issues, and use multiple filtering conditions to collect product information and sub-ASINs.
It realizes rapid and comprehensive collection of Amazon product information, improves the efficiency and integrity of data collection, avoids risk control problems, and ensures the stability and sustainability of data collection.
Smart Images

Figure CN120010945A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of information collection, and more specifically, to a method for collecting Amazon product information based on a browser plug-in. Background Art
[0002] In today's era of rapid development of e-commerce, market competition is increasingly fierce, and enterprises have unprecedented demands for the collection and analysis of product information. As a world-renowned e-commerce giant, Amazon's platform covers many countries and regions, and has a massive and rich product information library. This information covers everything from the basic attributes of the product, brand background, market sales data to detailed feedback from consumers, and is an extremely valuable resource for merchants.
[0003] However, traditional data collection methods have exposed many serious flaws when facing a highly complex platform like Amazon with strict risk control mechanisms. On the one hand, existing tools can only obtain a limited amount of data, which is difficult to meet the urgent needs of merchants to comprehensively and deeply analyze the market. They may only be able to collect some basic information about the goods, but cannot effectively obtain important data such as consumers' in-depth comments and key information in the Q&A session, resulting in merchants being unable to fully understand market dynamics and consumer needs, and lacking sufficient basis when formulating competitive strategies.
[0004] On the other hand, Amazon has implemented a powerful risk control mechanism to protect the security of platform data and user privacy. This mechanism poses a huge challenge to data collection. Traditional collection methods are subject to multiple strict checks on Amazon's interface requests when making data requests, including compatibility checks on browser versions, restrictions on whether it is a cross-domain request, and verification of cookie legitimacy. Although traditional reverse engineering methods attempt to parse these parameter checks, due to the continuous dynamic adjustment of Amazon's risk control strategy, this parsing work has become extremely cumbersome and inefficient, often requiring a lot of time and manpower costs, and it is difficult to ensure the stability and continuity of data collection, which seriously affects the efficiency and quality of data collection.
[0005] Therefore, a method for collecting Amazon product information based on browser plug-ins is proposed. Summary of the invention
[0006] In order to overcome the above-mentioned defects of the prior art, the present invention provides an Amazon product information collection method based on a browser plug-in. The method injects a script into the browser to simulate the internal request of the Amazon website to avoid risk control problems, and adopts a combination of multiple screening conditions to collect product information and sub-ASINs to solve the problems raised in the above-mentioned background technology.
[0007] To achieve the above object, the present invention provides the following technical solution: a method for collecting Amazon product information based on a browser plug-in, comprising the following steps: By injecting browser plug-ins, it simulates the behavior of real users visiting websites, thus avoiding tedious reverse analysis and directly carrying the current browser's cookies to make requests, thus evading Amazon's risk control to the greatest extent possible; Use XPATH technology to locate key fields in the page, and use regular expression parsing technology to extract required data from the page source code; After entering the product details page, the plug-in detects that the current page is a product page and then injects the collection script; The collection script obtains the ASIN code of the current product through the web page identifier as the main ASIN, and stores it in the ASIN collection queue together with other sub-ASINs; Poll the ASIN collection queue and execute a sub-process for each ASIN, wherein the sub-process includes: product information collection, review information collection, and question and answer information collection; After the data collection for each ASIN is completed, data verification is performed to ensure that the collected data is complete and accurate.
[0008] Preferably, the detailed information collected for the commodity information includes, but is not limited to, the following fields: Brand: Get the brand name of the product and provide brand information to merchants; Online time: record the time when the product is put on the shelf to help analyze the market life of the product; Star rating: Collects the overall rating of the product to reflect consumer satisfaction; Title: Get the title of the product to understand the main description of the product; Product images: Download product images for subsequent analysis and display; Price: record the current price of the product and monitor price changes; Number of comments: Count the number of comments on a product and assess its popularity; Total ratings: Summarizes the total number of ratings for a product and is used to calculate the average rating.
[0009] Preferably, the review information collected includes review time, review location, review content, whether the goods have been received, number of likes, review attributes, review language, etc., so as to fully understand consumer feedback.
[0010] Preferably, the process of collecting comment information includes: Read the homepage of the comment page, obtain all the filter conditions of the current comment, arrange and combine the filter conditions, and record the total number of comments; Collect comments based on the screening conditions. Collect the first 100 comments for each screening condition, use the comment ID as the unique identifier, and save them after deduplication. Compare the amount of data after deduplication with the total number of reviews. If the collection has been completed, exit the review collection for this ASIN.
[0011] Preferably, the question and answer information collection includes questions and content to supplement the product and review information, and is used to gain an in-depth understanding of consumers' specific questions and concerns, as well as the responses of merchants or consumers to these questions.
[0012] Preferably, the checking of the ASIN data further includes: Clean and format the collected data for subsequent analysis and application; The collected data is stored in a secure database to provide support for data analysis and mining.
[0013] Preferably, it also includes a user-friendly visual interface for simplifying the product selection and collection confirmation process. Users can easily select the products they want to collect data through this interface. Once the collection button is clicked, the system will automatically trigger the built-in collection script to collect detailed data on the selected products.
[0014] Technical effects and advantages of the present invention: 1. The present invention can quickly and comprehensively collect various types of information about Amazon products through a unique browser plug-in injection method and advanced data location extraction technology. Compared with traditional methods, the efficiency and integrity of data collection are greatly improved. For example, in terms of product information collection, detailed information such as brand, online time, star rating, title, product image, price, number of comments, total number of ratings, etc. can be obtained at one time, while traditional tools may only be able to obtain partial information and the collection speed is slow. In the collection of comment information, through a reasonable combination of screening conditions and a deduplication mechanism, a large amount of valuable consumer feedback can be efficiently obtained, providing merchants with richer market insight data.
[0015] 2. The strategy of simulating real user access behavior and carrying the current browser cookie for request successfully bypassed Amazon's risk control mechanism. This avoids problems such as data collection interruption or account banning due to frequent triggering of risk control, and ensures the stability and continuity of data collection. Compared with the traditional reverse engineering analysis parameter verification method, there is no need to constantly track and adapt to changes in Amazon's risk control strategy, which reduces technical difficulty and maintenance costs.
[0016] 3. Strictly verify, clean and format the collected data and store it in a secure database, providing a high-quality data foundation for subsequent data analysis and mining. Merchants can use this accurate data to analyze market trends and evaluate product advantages and disadvantages, so as to develop more targeted product optimization and marketing strategies and enhance the competitiveness of enterprises in the market. For example, by analyzing consumer comments and Q&A information, enterprises can promptly identify potential problems of products and improve them, or adjust product functions and pricing strategies according to market demand to better meet consumer needs and increase product market share and customer satisfaction.
[0017] 4. The user-friendly visual interface greatly simplifies the product selection and collection confirmation process. Users can easily complete data collection tasks without professional technical knowledge or complex operations. It reduces the tediousness of users switching and operating between different pages, improves the convenience and efficiency of operations, and enables even first-time users to quickly get started, which improves user satisfaction and participation in data collection work. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 It is a schematic diagram of the method flow of the present invention.
[0019] Figure 2 It is a schematic diagram of the process of collecting information through a visual interface according to the present invention.
[0020] Figure 3 This is a practical display diagram of the visual interface display in the operation of the embodiment of the present invention. DETAILED DESCRIPTION
[0021] Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in the field without making any creative work shall fall within the scope of protection of the present invention.
[0022] As attached Figure 1 and Figure 2 The method for collecting Amazon product information based on a browser plug-in shown mainly includes the following steps: 1. Plugin initialization and risk control avoidance First, after the user installs the browser plug-in of the present invention, the plug-in is in standby mode. When the user opens the Amazon website and performs browsing operations, the plug-in automatically detects the page environment. Once the user enters the product details page, the plug-in immediately starts its core functions. In the process of data interaction with the Amazon website, the plug-in simulates the behavior of real users visiting the website by injecting the browser plug-in. It obtains the cookie of the current browser and carries it in the request to the Amazon server. For example, when the user logs in to the Amazon account in the browser and browses the product, the plug-in identifies the cookies related to Amazon stored in the browser, which contain the encrypted form of the user's login information, browsing history and other data. The plug-in uses these cookies to construct the request header, so that the server believes that the request is a legitimate operation from a normal user, thereby circumventing Amazon's risk control mechanism to the greatest extent. This method avoids the cumbersome process of traditional reverse engineering analysis of Amazon website parameter verification, does not need to constantly track and update the changes in Amazon's risk control strategy, and improves the stability and sustainability of data collection.
[0023] 2. Page data location and extraction After successfully circumventing risk control, the plug-in uses a combination of XPATH technology and regular expression parsing technology to collect page data. For key information of products, such as brand, online time, star rating, title, product image, price, number of comments, total number of ratings, etc., the plug-in uses XPATH technology to locate its position in the HTML structure of the page. Taking brand information as an example, the plug-in can accurately find the HTML element node containing the brand name through the pre-set XPATH expression. For some complex or irregular data, such as specific attributes in product descriptions or special format content in user comments, the plug-in uses regular expression parsing technology to extract. For example, when extracting time information in user comments, regular expressions can match and extract in the comment text according to common time format patterns (such as "YYYY-MM-DD" or "MM / DD / YYYY", etc.), ensuring that the required data can be accurately obtained.
[0024] 3. ASIN code processing and queue construction In the product details page, the collection script quickly obtains the ASIN code of the current product as the main ASIN through the web page identifier. This ASIN code is the unique identifier of the Amazon product and is crucial for subsequent data management and association. At the same time, the collection script will also identify other sub-ASINs related to the product and store them in the ASIN collection queue. For example, for some products with multiple accessories or variants, each accessory or variant may have its own sub-ASIN. The plug-in will comprehensively collect this information and manage it in an orderly manner, laying the foundation for subsequent comprehensive data collection.
[0025] 4. Data Collection Process Execution Polling the ASIN collection queue is one of the core links of data collection. For each ASIN in the queue, product information is collected first. The plug-in accurately obtains detailed information of the product according to the predetermined rules. For example, the brand name will be extracted from specific HTML tags or attributes and normalized to remove extra spaces and special characters; the online time will be parsed and converted according to the time format of the Amazon website for subsequent time series analysis; information such as star rating and number of reviews is obtained by analyzing and counting the HTML structure of the corresponding rating and review modules.
[0026] 5. In terms of review information collection, the plug-in first reads the homepage of the review page, uses page parsing technology to obtain all the filter conditions of the current review, such as sorting by time, sorting by star rating, and other filter options, and fully arranges and combines these filter conditions, while recording the total number of reviews. Then, according to each filter condition, the plug-in will collect the first 100 reviews. During the collection process, the comment id is used as the unique identifier, and duplicates are removed and saved by comparing it with the collected data. For example, when a new comment is collected, the plug-in calculates the hash value of its comment id and compares it with the previously stored hash value list. If there is no duplication, the comment data is stored in the local cache. Finally, the plug-in compares the amount of deduplicated data with the total number of comments. If the collection has been completed, that is, the amount of deduplicated data reaches or is close to the total number of comments, the comment collection process for this ASIN will be exited.
[0027] For Q&A information collection, the plug-in deeply analyzes the HTML structure of the product Q&A page to extract questions and content. It identifies the title and detailed description of the question, as well as the corresponding answer content, and organizes and stores this information to supplement product and review information, providing merchants with a more comprehensive perspective on consumer feedback.
[0028] 6. Data Verification and Storage After the data collection for each ASIN is completed, the plug-in will perform strict data verification. First, the collected data is checked for integrity to ensure that all predetermined fields are filled with data without omissions. For example, check whether the product image is successfully downloaded, whether each attribute in the comment information has a value, etc. Then, the data is verified for accuracy by comparing it with known data formats and ranges. For example, whether the price data is within a reasonable price range, whether the star rating is within the range of 1-5 stars, etc.
[0029] After verification, the plug-in will clean and format the collected data. For text data, lexical analysis and standardization will be performed to remove stop words and noise characters, unify the date format, and normalize the number format. Finally, the processed data is stored in a secure database. The plug-in will build a reasonable database table structure based on the type and association of the data, such as storing product information, comment information, and question and answer information in different tables, and linking them through ASIN codes, to provide efficient support for subsequent data analysis and mining.
[0030] 7. Visual interface operation The present invention also provides a user-friendly visual interface. After the user enters a keyword in the Amazon search bar, the search results page will display an additional operation interface provided by the plug-in. This interface displays all the product options listed on the current page in an intuitive list format, and the user can view all available products at a glance. The user only needs to click on the interface to select the product for which data is desired to be collected, and then click the confirmation button, and the system will automatically trigger the built-in collection script to collect detailed data for the selected product. The entire operation process does not require the user to switch between multiple pages or perform complex settings, which greatly simplifies the product selection and collection confirmation process and improves the convenience and efficiency of user operations.
[0031] As attached Figure 3 As shown in the figure, after the user enters a keyword (such as "apple") in the Amazon search bar, the search results page will display an additional operation interface provided by this plug-in. This interface allows users to directly select the products they want to collect data from the search results without leaving the current page or making complicated settings. By clicking the button provided by the plug-in, users can easily trigger the data collection process and achieve comprehensive collection of detailed information, user reviews and Q&A information of the target product.
[0032] This pop-up window contains all the product options listed on the current page, allowing users to view all available products at a glance. Users can select the products they are interested in from this intuitive list and then start the data collection process by clicking the confirmation button.
[0033] This popup is designed to provide the following benefits: User-friendly: It simplifies the product selection process, allowing users to quickly find and select target products.
[0034] Improved efficiency: It reduces the need for users to switch between different pages, and allows users to select products and start collection directly on the search results page.
[0035] Intuitive operation: With the clear interface design, even first-time users can easily understand how to operate.
[0036] Through this design, the present invention not only improves the efficiency of data collection, but also enhances the user experience, making the entire data collection process smoother and more intuitive.
[0037] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention should be included in the protection scope of the present invention.
Claims
1. A method for collecting Amazon product information based on a browser plug-in, characterized in that: The following steps are involved: By injecting browser plug-ins, it simulates the behavior of real users visiting websites, thus avoiding tedious reverse analysis and directly carrying the current browser's cookies to make requests, thus evading Amazon's risk control to the greatest extent possible; Use XPATH technology to locate key fields in the page, and use regular expression parsing technology to extract required data from the page source code; After entering the product details page, the plug-in detects that the current page is a product page and then injects the collection script; The collection script obtains the ASIN code of the current product through the web page identifier as the main ASIN, and stores it in the ASIN collection queue together with other sub-ASINs; Poll the ASIN collection queue and execute a sub-process for each ASIN, wherein the sub-process includes: product information collection, review information collection, and question and answer information collection; After the data collection for each ASIN is completed, data verification is performed to ensure that the collected data is complete and accurate.
2. The method for collecting Amazon product information based on a browser plug-in according to claim 1, characterized in that: The detailed information of the product information collection includes but is not limited to the following fields: Brand: Get the brand name of the product and provide brand information to merchants; Online time: record the time when the product is put on the shelf to help analyze the market life of the product; Star rating: Collects the overall rating of the product to reflect consumer satisfaction; Title: Get the title of the product to understand the main description of the product; Product images: Download product images for subsequent analysis and display; Price: record the current price of the product and monitor price changes; Number of comments: Count the number of comments on a product and assess its popularity; Total ratings: Summarizes the total number of ratings for a product and is used to calculate the average rating.
3. The method for collecting Amazon product information based on a browser plug-in according to claim 1, characterized in that: The review information collected includes review time, review location, review content, whether the goods have been received, number of likes, review attributes, review language, etc., in order to fully understand consumer feedback.
4. The method for collecting Amazon product information based on a browser plug-in according to claim 3, characterized in that: The process of collecting comment information includes: Read the homepage of the comment page, obtain all the filter conditions of the current comment, arrange and combine the filter conditions, and record the total number of comments; Collect comments based on the screening conditions. Collect the first 100 comments for each screening condition, use the comment ID as the unique identifier, and save them after deduplication. Compare the amount of data after deduplication with the total number of reviews. If the collection has been completed, exit the review collection for this ASIN.
5. The method for collecting Amazon product information based on a browser plug-in according to claim 1, characterized in that: The question and answer information collection includes questions and content to supplement product and review information, and is used to gain an in-depth understanding of consumers' specific questions and concerns, as well as the responses of merchants or consumers to these questions.
6. The method for collecting Amazon product information based on a browser plug-in according to claim 1, characterized in that: The verification of the ASIN data also includes: Clean and format the collected data for subsequent analysis and application; The collected data is stored in a secure database to provide support for data analysis and mining.
7. The method for collecting Amazon product information based on a browser plug-in according to claim 1, characterized in that: It also includes a user-friendly visual interface to simplify the product selection and collection confirmation process. Users can easily select the products they want to collect data through this interface. Once the collection button is clicked, the system will automatically trigger the built-in collection script to collect detailed data for the selected products.
Citation Information
Cited By
Digital human live broadcast control method, device, system, equipment and program product
CN121173976A