Data acquisition system

By using an automated data acquisition system to parse and output structured product information, the problem of cumbersome and time-consuming competitor information collection in existing technologies has been solved, achieving efficient and accurate data acquisition and ensuring the real-time nature and integrity of the data.

CN121117293APending Publication Date: 2025-12-12INVENTEC PUDONG TECH CORPOARTION +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410758238.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-06-12
Publication Date
2025-12-12

AI Technical Summary

Technical Problem

The existing technology for collecting competitor information is cumbersome and time-consuming, and is prone to human error, resulting in data that is not real-time and inaccurate, making it impossible to efficiently obtain the latest competitor information.

Method used

An automated data acquisition system is adopted, which executes an automated data acquisition process through a processing unit. This includes connecting to and parsing websites, parsing product display pages, and parsing strings to output structured product information. Automated web crawling technology is used to replace manual operations.

Benefits of technology

It improved the efficiency and accuracy of competitor information collection, ensured the real-time nature and integrity of data, reduced labor costs, and avoided data mis-entry and duplication issues.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121117293A_ABST
    Figure CN121117293A_ABST
Patent Text Reader

Abstract

The invention provides a data acquisition system, which is used for executing an automatic data acquisition method to obtain structured commodity information of competitive products loaded on a website. The automatic data collection method comprises the following steps: connecting and analyzing the website to obtain a commodity display page; connecting and analyzing the commodity display page to obtain specification information of a corresponding commodity; performing character string analysis on the specification information to obtain a character string analysis result; and outputting the structured commodity information according to the character string analysis result.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to a data collection system, in particular, a data collection system capable of collecting data according to an automated crawler technology. BACKGROUND

[0002] In the past, the method for collecting competitive product information in the industry is often performed by manual operation. The process of collecting competitive product information includes manually opening a browser, selecting a webpage, clicking on related links or buttons, searching for page information, and copying competitive product information, until the collection of target information is completed. Finally, the collected information needs to be sorted and stored manually. In addition, the competitive product information to be collected involves multiple manufacturers and multiple specifications of products, and the related content is often distributed on multiple websites and multiple pages. Therefore, the process of collecting competitive product information is often lengthy and tedious.

[0003] In addition to consuming a large amount of manpower, the process of collecting competitive product information by manual operation often takes several working days. In this case, the collected data cannot guarantee to be the latest and most real-time information, and it is difficult to avoid human errors such as data misplacement or duplication in the process of manual operation, which affects the correctness of the data. Therefore, it is urgent to establish an effective method for automatically collecting competitive product information, which can obtain correct and real-time information in a short time to save manpower and reduce costs. SUMMARY

[0004] Therefore, the main purpose of the present application is to provide a data collection system that collects competitive product information published on multiple websites by executing an automated data collection method, thereby improving the shortcomings of the prior art.

[0005] The present application discloses a data collection system, which includes a processing unit and a storage unit. The processing unit is used to execute a program code, and the storage unit coupled to the processing unit is used to store the program code to instruct the processing unit to execute an automated data collection method to obtain a structured product information of a competitive product published on a website. The automated data collection method includes linking and parsing the website to obtain a product display page, linking and parsing the product display page to obtain a specification information of a corresponding product, performing string parsing on the specification information to obtain a string parsing result, and outputting the structured product information according to the string parsing result. BRIEF DESCRIPTION OF DRAWINGS

[0006] Figure 1 A schematic diagram of a data collection system according to an embodiment of the present application is shown.

[0007] Figure 2A schematic diagram of an automatic data collection process according to an embodiment of the present application.

[0008] Figure 3A Figure 3B A schematic diagram of a rolling product display page according to an embodiment of the present application.

[0009] Element Number Description

[0010] 10 data collection system

[0011] 100 processing unit

[0012] 102 storage unit

[0013] 1020 program code

[0014] 104 network interface

[0015] 12 network

[0016] 14 enterprise website

[0017] 20 process

[0018] 200-210 steps

[0019] 30 browser

[0020] 32 roll

[0021] 34 product display page

[0022] P1-P20 product DETAILED DESCRIPTION

[0023] Certain terms are used throughout the description and following claims, which are intended to have the meanings as set forth below. It is to be understood that the terms are used only in a generic and descriptive sense, and not for purposes of limitation.

[0024] In the conventional way of collecting competitive product information, a person is often arranged to collect information about the products of a specific enterprise website, and the collected information is interpreted and arranged manually, which consumes a lot of man-hours and manpower, and may cause data misplacement or duplication due to human error. Therefore, the present application provides a data collection system, which collects competitive product information from multiple websites by executing an automatic data collection method, so as to replace the manual searching process and improve the shortcomings of the prior art. ​

[0025] Referring to Figure 1 , Figure 1 A schematic diagram of a data collection system 10 according to an embodiment of the present application is shown. As shown, the data collection system 10 is connected to a plurality of enterprise websites 14 through a network 12 and performs an automated data collection method according to an automated crawler technique to obtain structured product information of relevant competitors published on the enterprise websites 14. Figure 1

[0026] As shown, the data collection system 10 comprises a processing unit 100, a storage unit 102 and a network interface 104. The processing unit 100 can be a general processor, a microprocessor, an application specific integrated circuit (ASIC) or the like or a combination thereof. The storage unit 102 can be any data storage device for storing a program code 1020 and reading and executing the program code 1020 by the processing unit 100. For example, the storage unit 102 can be a read-only memory (ROM), a flash memory, a random access memory (RAM), a hard disk, an optical data storage device, a non-volatile storage unit or the like or a combination thereof, but is not limited thereto. The network interface 104 can comprise an input / output (I / O) interface, a network hardware device or the like and can be connected to the network 12. The data collection system 10 is used to represent necessary components required for implementing the embodiments of the present application and those skilled in the art can make different modifications and adjustments based on the data collection system 10, but is not limited thereto. Figure 1

[0027] As shown, the automated data collection method according to the embodiments of the present application can be summarized as an automated data collection flow 20 and compiled as the program code 1020, which is read and executed by the processing unit 100. The data collection flow 20 comprises the following steps: Figure 2 Step 200: Start.

[0028] Step 202: Connect and parse the enterprise websites 14 to obtain a product display page.

[0029] Step 204: Connect and parse the product display page to obtain a specification information of a corresponding product.

[0030] Step 206: Perform string parsing on the specification information to obtain a string parsing result.

[0031] Step 208: Output the structured product information according to the string parsing result.

[0032]

[0033] ​​​Step 210: End.

[0034] According to the automated data collection process 20, the data collection system 10 first links the enterprise website 14 through the network interface 104, and obtains a product display page by parsing a product search page of the enterprise website 14 (step 202). Then, the data collection system 10 links and parses the web page source code (HTML source code) of the product display page to obtain the specification information related to the product (step 204). For the specification information of the product, the data collection system 10 further performs string parsing to obtain a string parsing result (step 206). Finally, according to the string parsing result, the data collection system 10 outputs the structured product information related to the competitive product for easy viewing (step 208). In this embodiment, the data collection system 10 can automatically perform the automated data collection process 20 for each of a plurality of enterprise websites 14, and combine and output the structured product information of the competitive product posted on each website. Accordingly, the data collection system 10 can replace the traditional manual collection of competitive product information with an automated crawler technology.

[0035] According to the automated data collection process 20, in step 202, the data collection system 10 links and parses the enterprise website 14 to obtain a product display page. In detail, because the URLs and domains of the various enterprise websites 14 are different, the data collection system 10 automatically opens a web browser to link to the enterprise website according to the URL of the enterprise website 14. Because different enterprise websites 14 include different website design methods and data arrangement methods, etc., the data collection system 10 needs to parse the enterprise website 14 to obtain a website parsing result. According to the website parsing result, the data collection system 10 can obtain the product display page in the enterprise website 14. For example, in addition to posting product information, the enterprise website can also post various all-inclusive information including enterprise information introduction, customer service resources, online store, latest news, frequently asked questions, etc. In addition, according to different enterprises, the product information can also cover a variety of different fields and different types of products. In the case of different layout and design of each website, the data collection system 10 must parse the website to accurately obtain the product display page of the relevant competitive product.

[0036] At step 204, the data collection system 10 links and parses the product display page to obtain the specification information of the corresponding product. In the present embodiment, the data collection system 10 needs to obtain the web page source code of the product display page, parse the obtained web page source code and collect the text data set of the product display frame therefrom, and collect the corresponding specification information from the text data set according to the default specification field of the structured product information. For example, different enterprise websites 14 include different product display modes, and the product display page can be implemented in different ways such as page mode or using pull-down loading mode. Therefore, the data collection system 10 needs to parse the product display page of the enterprise website 14 to obtain the required information. After the data collection system 10 obtains the web page source code of the product display page, it needs to interpret the layout and framework of the web page source code. Accordingly, the content of the product information frame, which is a text data set in the present embodiment, can be obtained. Then, according to the default specification field, the data collection system 10 classifies and organizes the collected text data set. Taking a notebook computer as an example, the data collection system 10 can collect the corresponding specification information from the text data set according to the default specification fields such as display adapter, central processing unit, memory, weight, thickness, etc.

[0037] It should be noted that according to different data presentation modes of the enterprise website 14, the product display content can contain links or other hidden data. Therefore, in the process of collecting the text data set, the data collection system 10 needs to load the hidden data of the product display page by JavaScript. On the other hand, according to different design modes of the enterprise website 14, before the data collection system 10 obtains the web page source code of the product display page, it needs to scroll the product display page to the bottom to load the complete page. Accordingly, it can be avoided that the data collection system 10 only obtains the initial screen preloaded by the website, and misses the complete information content. Please refer to Figure 3A and Figure 3B which show the schematic diagrams of a scrolled product display page according to the embodiments of the present application. As shown in FIGS. 3A and 3B, a web browser 30 includes a scroll bar 32 and a connected product display page 34. The product display page 34 displays a plurality of products P1-P20. However, at the initial stage of loading the product display page 34 by the browser, only 3 products P1-P3 are loaded as shown in FIG. 3A. Only when the user scrolls down the scroll bar 32, the remaining product information will be loaded in batches along with the scroll bar as shown in FIG. 3B. Only when the user scrolls the scroll bar 32 to the bottom, all the products including P1-P20 will be completely loaded into the page as shown in FIG. 3C. Figure 3A Figure 3B Figure 3A ​​The product loading condition shown varies depending on the resolution used by the data collection system 10 and other differences, in addition to the webpage design. Therefore, in order to obtain the original code of the webpage containing complete information, the data collection system 10 needs to scroll the product display page to the bottom to load the complete page. It should be noted that, Figure 3A , Figure 3B Only one possible product display mode and arrangement is shown. During the execution of the automated data collection process 20, the data collection system 10 needs to ensure that the data of each page accessed is completely loaded.

[0038] In step 206, the data collection system 10 performs string parsing on the specification information obtained in step 204 to obtain a string parsing result. In detail, the specification information collected in step 204 is often a long string of characters that cannot be directly recognized, and therefore needs to be filtered and screened, etc. to obtain complete information. For example, the specification information can contain special function characters for webpage writing, such as HTML tags, annotations, etc. Therefore, the data collection system 10 needs to filter the special function characters for webpage writing in the specification information. In addition, the specification information can contain special description boxes for emphasizing the features of the product, and the data collection system 10 needs to obtain the content of the special description boxes. After interpreting the text content of the special description boxes, the data collection system 10 determines whether to include the text content in the structured product information.

[0039] In step 206, the data collection system 10 outputs the structured product information according to the string parsing result. In detail, the data collection system 10 can output the specification information obtained in the previous steps into structured product information according to the default specification field, to facilitate subsequent product analysis and analysis. In this embodiment, the structured product information can be a structured file such as an Excel format, or a data table stored in a database, and is not limited thereto.

[0040] Accordingly, the automated data collection process 20 can replace the traditional manual collection of competitor information, improving the efficiency and accuracy of collecting competitor information.

[0041] It should be noted that some enterprise websites set security protection mechanisms such as anti-crawler programs to prevent various network security threats and improper access to webpages. In this case, when the data collection system 10 executes the automated data collection process 20, it may encounter blocking or offline situations. Therefore, the automated data collection process 20 also includes corresponding exception handling for each website security protection mechanism and the blocking mechanism that may be triggered.

[0042] In the present embodiment, the data collection system 10 performs online error handling when an online error occurs during execution of the automated data collection process 20. The online error handling includes, but is not limited to, changing the Internet Protocol Address (IP Address) and the web user-agent, and re-executing the automated data collection process 20 after changing the IP Address and the user-agent.

[0043] In an embodiment, the automated data collection process 20 can be packaged as a program, and a user can execute the data collection process 20 by opening a specific file, achieving a one-click execution effect. In an embodiment, the data collection system 10 can execute the automated data collection process 20 according to a predetermined time, or can set the automated data collection process 20 to be executed periodically according to needs.

[0044] In summary, the present application provides a data collection system that executes an automated data collection process according to automated crawler technology, replacing the traditional manual collection of competitive product information, thereby improving efficiency and the correctness of data collection. Through an automatic online error handling mechanism, the stability of the system can be improved. In addition, through automatic and efficient information collection, real-time control of network data can be performed, and the latest competitive product information can be obtained.

[0045] The above description is only the preferred embodiment of the present application, and any changes and modifications made within the scope of the patent application of the present application shall be covered.

Claims

1. A data acquisition system, characterized in that, Include: A processing unit is used to execute a piece of program code; as well as A storage unit, coupled to the processing unit, stores the program code to instruct the processing unit to execute an automated data collection method at a predetermined time or periodically to obtain structured product information of competitors posted on a website, the automated data collection method comprising: Link to and parse the website to obtain a product display page; Link to and parse the product display page to obtain the corresponding product's specifications; The specification information is parsed to obtain a string parsing result; as well as Based on the string parsing result, the structured product information is output; Specifically, the structured product information is output and then stored on an disk or in a database.

2. The data acquisition system according to claim 1, characterized in that, The automated data collection method is executed for each of a plurality of websites, and the structured product information of the competitors posted on each website is merged and output.

3. The data acquisition system according to claim 1, characterized in that, The steps of linking to and resolving the website to obtain the product display page include: Automatically open a web browser to link to the website based on a given URL; The website is parsed to obtain a website parsing result; and The product display page is obtained based on the website's parsing results.

4. The data acquisition system according to claim 1, characterized in that, The step of linking to and parsing the product display page to obtain the corresponding product's specification information includes: Obtain the source code of the webpage displaying the product; Parse the webpage source code to collect a text dataset from a product display box; and Based on the default specification field of the structured product information, the corresponding specification information is collected from the text dataset.

5. The data acquisition system according to claim 4, characterized in that, The text dataset also includes hidden data from the product display page loaded in JavaScript.

6. The data acquisition system according to claim 4, characterized in that, Before obtaining the source code of the webpage of the product display page, scroll the product display page to the bottom to load the complete page.

7. The data acquisition system according to claim 1, characterized in that, The step of parsing the specification information to obtain the string parsing result includes: Filter out the special function text used for web page writing in the specification information.

8. The data acquisition system according to claim 1, characterized in that, The automated data acquisition method further includes: Retrieve the content of a special description box; Interpreting the text content of the special description box; and Determine whether to include the text content in the structured product information.

9. The data acquisition system according to claim 1, characterized in that, It also includes performing a connection error handling when a connection error occurs during the execution of the automated data acquisition method.

10. The data acquisition system according to claim 9, characterized in that, The online error handling includes: Change an Internet Protocol address; Change a user agent; and Re-execute the automated data acquisition method.