Multi-source data acquisition method, device and system
By initiating page access requests to web programs, PC applications and mobile applications, and extracting data using the target big model, the problems of limited data sources and high development costs in the prior art are solved, and the flexibility and efficiency of multi-source data acquisition are achieved.
Patent Information
- Application Number
- CN202411999439.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-23
AI Technical Summary
The existing crawler technology has problems such as data source limitation and high development costs in terms of data collection, making it difficult to effectively collect data from PC and mobile applications.
By initiating page access requests to web programs, PC applications and mobile applications, obtaining multiple page documents, and using the target model to extract target data matching the page prompt word from the page document, realizing multi-source data collection.
It realizes data collection of multiple data sources, reduces development costs, improves the flexibility and efficiency of data collection, and can quickly adapt to business changes.
Smart Images

Figure CN120030210A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a multi-source data acquisition method, device and system. Background Art
[0002] Web data scraping, also known as web data collection or web crawler, refers to the process of extracting structured and unstructured information from a large number of web pages by writing programs or using specific tools. After the information is processed according to certain rules and screening criteria, it is saved in a structured database. The uses of web scraping include online price comparison, meteorological data monitoring, web page change detection, scientific research and web data integration.
[0003] Web data crawling can help enterprises and research institutions quickly obtain a large amount of data, providing support for market analysis, customer insights, product decisions, trend forecasting, academic research, etc. In the field of scientific research, web data crawling technology can quickly capture target information in a short period of time and build a large data set to meet the needs of analysis and research. In the field of enterprise production, it can quickly obtain industry information to help enterprises gain an advantage in the competition. By conducting in-depth analysis of the captured data, the correlation between the data can be discovered to provide support for decision-making. For content providers, web data crawling can be used to automatically update website content and improve user experience.
[0004] Existing crawler technology usually uses programming languages (such as Python, Java, etc.) to write programs, simulate browsers to access web pages, parse HTML (Hyper Text Markup Language) / XML (Extensible Markup Language) content, and extract data from web pages. However, this solution has certain defects:
[0005] 1. The supported data sources are limited. It is usually used for web page data collection and does not support data collection of PC (Personal Computer) applications and mobile applications.
[0006] 2. The development cost is high and requires professional developers to implement. Extracting data requires analyzing the page structure, locating the element path, and extracting data through the path. Summary of the invention
[0007] In view of the above problems, embodiments of the present application provide a multi-source data acquisition method, device and system that overcome the above problems or at least partially solve the above problems.
[0008] In a first aspect, an embodiment of the present application provides a multi-source data collection method, which is applied to a collection server, comprising:
[0009] Based on a page access request initiated to a target data source, a plurality of page documents provided by the target data source are obtained, wherein the target data source includes at least two of a Web program, a PC application program, and a mobile application program;
[0010] Obtaining page prompt words corresponding to each page document, and constructing multiple input arrays including the page documents and the corresponding page prompt words;
[0011] The multiple input arrays are input into a target macro model, and the target macro model is used to extract target data matching corresponding page prompt words from the page documents, so as to obtain page data corresponding to the multiple page documents respectively.
[0012] In a second aspect, an embodiment of the present application provides a multi-source data acquisition device, which is applied to an acquisition server, including:
[0013] A first acquisition module, configured to acquire a plurality of page documents provided by a target data source based on a page access request initiated to the target data source, wherein the target data source includes at least two of a Web program, a PC application program, and a mobile application program;
[0014] An acquisition construction module is used to acquire the page prompt words corresponding to each page document, and construct a plurality of input arrays including the page documents and the corresponding page prompt words;
[0015] The processing module is used to input the multiple input arrays into a target large model, and use the target large model to extract target data matching the corresponding page prompt words in the page document to obtain page data corresponding to the multiple page documents respectively.
[0016] In a third aspect, an embodiment of the present application provides a multi-source data acquisition system, including: an acquisition server, a collected device, a data processing device, and a data storage device;
[0017] The collected device provides a page document based on the running mobile application, PC application and Web program;
[0018] The acquisition server runs an acquisition program to acquire page documents provided by at least two of the mobile application, the PC application, and the Web program;
[0019] The data processing device deploys a prompt word output module and a target large model, wherein the prompt word output module is used to output the page prompt word corresponding to the page document, and the target large model is used to extract target data from the page document based on the page prompt word corresponding to the page document;
[0020] The acquisition server, based on the interaction with the device to be acquired and the data processing device, executes the method described in the first aspect to collect page documents and extract page data from the page documents by using the target large model, wherein the extracted page data is stored in the data storage device.
[0021] In a fourth aspect, an embodiment of the present application provides an electronic device, including a processor, a memory, and a computer program stored on the memory and executable on the processor. When the computer program is executed by the processor, the steps of the multi-source data acquisition method described in the first aspect above are implemented.
[0022] In a fifth aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the multi-source data acquisition method described in the first aspect above are implemented.
[0023] The technical solution of the embodiment of the present application can provide relatively rich multi-source data by sending page access requests to at least two target data sources and obtaining multiple page documents provided by the target data sources, and solve the problem of single data source in the traditional data acquisition process. After obtaining the page document and the corresponding page prompt words of the page document, an input array including the page document and the corresponding page prompt words is constructed, and multiple input arrays are input into the target large model. After the target large model obtains the original page content and data acquisition prompt information, it automatically extracts the target data in the page document and outputs the data extraction result, which can realize multi-source data collection with low cost and intelligence, and solve the high-cost disadvantage of data extraction by professional personnel. Description of the Drawings
[0024] Figure 1 A schematic diagram showing the multi-source data acquisition method provided by an embodiment of the present application;
[0025] Figure 2 An interaction flowchart showing the multi-source data acquisition method provided by an embodiment of the present application;
[0026] Figure 3 A schematic diagram showing the multi-source data acquisition device provided by an embodiment of the present application;
[0027] Figure 4 A schematic diagram showing the multi-source data acquisition system provided by an embodiment of the present application;
[0028] Figure 5 A schematic diagram showing the structure of the electronic device provided by an embodiment of the present application. Detailed Embodiments
[0029] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the protection scope of the present application.
[0030] It should be understood that the "one embodiment" or "an embodiment" mentioned throughout the specification means that a specific feature, structure, or characteristic related to the embodiment is included in at least one embodiment of the present application. Therefore, the appearances of "in one embodiment" or "in an embodiment" throughout the specification do not necessarily refer to the same embodiment. In addition, these specific features, structures, or characteristics can be combined in one or more embodiments in any suitable manner. The plurality in the embodiments of the present application may include two and more than two.
[0031] In various embodiments of the present application, it should be understood that the sequence numbers of the following processes do not mean the order of execution, and the execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0032] The embodiments of the present application provide a multi-source data collection method, which is applied to a collection server, such as Figure 1 shown, and the method includes:
[0033] Step 101, based on a page access request sent to a target data source, obtain multiple page documents provided by the target data source, where the target data source includes at least two of a Web program, a PC application program, and a mobile application program.
[0034] The collection server in the embodiments of the present application runs a collection program and uses the collection program to collect page documents from the target data source. When the collection server collects page documents from the target data source, it sends a page access request to the target data source and receives multiple page documents provided by the target data source based on the page access request. Among them, the target data source includes at least two of a Web (World Wide Web) program, a PC (Personal Computer) application program, and a mobile application program. By sending a page access request to at least two programs and receiving the response results fed back by at least two programs, relatively rich multi-source data can be provided, and the problem of single data source in the traditional data collection process can be solved.
[0035] Among them, Web programs are also called Web pages, PC applications refer to applications installed on PC devices, and mobile applications refer to applications installed on mobile devices. The page documents including the original page content provided by the Web programs, PC applications, and mobile applications acquired by the collection server are in HTML format or XML format. Data in HTML format or XML format can be regarded as text in a special format and can be manually recognized.
[0036] Step 102: Obtain the page prompt words corresponding to each page document, and construct multiple input arrays including the page documents and the corresponding page prompt words.
[0037] After acquiring multiple page documents provided by the target data source based on the page access request, the acquisition server acquires the page prompt words corresponding to each page document. The page prompt words corresponding to the page document do not need to be developed using a specific programming language, and can be written by ordinary users, which can reduce development costs.
[0038] In this embodiment, the page prompt words can be provided manually after understanding the page document, or can be provided after the content of the page document is recognized by intelligent means. After obtaining the page document provided by the target data source and the page prompt words corresponding to the page document, the acquisition server constructs an input array based on the page document and the page prompt words corresponding to the page document to determine the input arrays corresponding to the multiple page documents.
[0039] Step 103: Input multiple input arrays into the target macro model, and use the target macro model to extract target data matching the corresponding page prompt words in the page document to obtain page data corresponding to the multiple page documents.
[0040] After determining the input array based on the page document and the page prompt word corresponding to the page document, the multiple input arrays corresponding to the multiple page documents are input into the target large model. The target large model is a large language model. After the multiple input arrays are input into the target large model, the page document in the input array is analyzed by the target large model, and the target data matching the corresponding page prompt word is extracted from the page document, so that the target large model outputs the target data corresponding to the page document, so as to obtain the page data corresponding to the multiple page documents respectively, and realize the collection of the required data for the page document.
[0041] A large language model refers to a deep learning model trained with a large amount of text data, which is composed of an artificial neural network with many parameters, and can generate natural language text or understand the meaning of language text. The large language model can handle a variety of natural language tasks, such as text classification, question answering, dialogue, etc., and is an important way to artificial intelligence. In this embodiment, the target large model belonging to the large language model automatically parses and processes the page content, and intelligently extracts the data in the page. Compared with the traditional solution that requires developers to analyze the page structure, find the target element path (xpath), and extract data according to the target element path when extracting data, the intelligent data collection method based on the large model can realize multi-source data collection at a lower cost and intelligently.
[0042] The above implementation scheme of the present application can provide relatively rich multi-source data by initiating page access requests to at least two target data sources and obtaining multiple page documents provided by the target data sources, thereby solving the problem of single data source in the traditional data collection process; after obtaining the page document and the page prompt word corresponding to the page document, an input array including the page document and the corresponding page prompt word is constructed, and the multiple input arrays are input into the target large model. After obtaining the original page content and data collection prompt information, the target large model automatically extracts the target data in the page document and outputs the data extraction result, thereby realizing multi-source data collection at a relatively low cost and intelligently, thereby solving the high cost disadvantage of requiring professionals to extract data.
[0043] In an optional embodiment of the present application, the method further includes:
[0044] In response to identifying that a first page document among multiple page documents is updated as a second page document, a page prompt word corresponding to the second page document is obtained; based on the second page document and the page prompt word corresponding to the second page document, an associated input array is updated to determine a target input array; the target input array is input into a target large model, and the target large model extracts data from the second page document based on the page prompt word corresponding to the second page document to obtain the latest page data corresponding to the second page document.
[0045] When it is identified that the first page document among the multiple page documents provided by the target data source is updated to the second page document, the second page document is processed in a timely manner. The update of the first page document may be a page revision or a page content update; the update of the page document may be identified based on the comparison of the page documents, for example, by comparing the page documents corresponding to the same page identifier before and after to identify whether the page has been updated, and the page document may be recorded in the form of a screenshot for subsequent page comparison.
[0046] After identifying that the first page document is updated to the second page document, the page prompt word corresponding to the second page document is obtained, and then the associated input array is updated based on the second page document and the page prompt word corresponding to the second page document, and the target input array corresponding to the second page document is determined by updating the input array. The target input array is input into the target large model, and the target large model performs recognition and analysis on the target input array to extract the adapted data in the second page document based on the page prompt word corresponding to the second page document, and output the latest page data corresponding to the second page document.
[0047] Among them, the page prompt words corresponding to the second page document can also be provided manually, or they can be provided after the content of the page document is recognized by intelligent means. When a page change is identified, the updated page and the corresponding page prompt words are input into the target large model in a timely manner, and the target large model is used for data extraction. After fine-tuning the data collection prompt words, intelligent data collection can be realized based on the large language model, which can quickly adapt to business changes. Compared with the existing technology, when a page changes, the developer needs to re-analyze the page structure, locate the target elements, and modify the crawling logic to adapt to the processing method of business changes. Data collection can be flexibly performed, quickly adapt to business changes, and improve the efficiency of data collection.
[0048] In the above solution, when a page update occurs, the data collection strategy is quickly adjusted and a large language model is used for intelligent data collection. There is no need for complex system reconstruction or code modification, which reduces manual participation and greatly improves the speed and efficiency of data collection. It solves the problem of poor data collection flexibility and difficulty in adapting to business changes in existing collection solutions.
[0049] The following describes the process of obtaining a page document provided by a target data source. When obtaining multiple page documents provided by a target data source based on a page access request initiated to the target data source, the process includes:
[0050] Initiate a page access request to at least two of the Web program, the PC application, and the mobile application, where the page access request is a network request; receive page documents fed back by at least two of the Web program, the PC application, and the mobile application;
[0051] Among them, the Web program directly feedbacks the page document based on the page access request initiated by the acquisition server, the PC application receives the page access request through the WinAppDriver service and provides the page document to the acquisition server, and the mobile application receives the page access request through the Web agent program and provides the page document to the acquisition server.
[0052] In this embodiment, the mobile application is deployed on a mobile device (such as a mobile phone, a tablet, etc.), and the Web program and the PC application can be deployed on a PC device or can be directly run on a collection server. The collection server initiates a page access request belonging to a network request to at least two of the Web program, the PC application, and the mobile application through the collection program. For the Web program, after receiving the page access request, it can directly feedback the response result to the collection server, and the feedback response result is a page document provided by the Web program. For the mobile application, it does not support the network request directly initiated by the collection server, and a Web agent program is installed on the mobile device. The mobile application receives the page access request through the Web agent program installed on the mobile device and provides the page document for the collection server. For the PC application, it also does not support the network request directly initiated by the collection server. The PC application receives the page access request through the WinAppDriver service and provides the page document for the collection server. Among them, if the PC application is deployed on the PC device, the WinAppDriver service is also deployed on the PC device. If the PC application runs on the collection server, the WinAppDriver service also runs on the collection server. After the collection server initiates a page access request for the PC application based on the collection program, the PC application provides the page document through the WinAppDriver service.
[0053] By initiating requests to at least two data sources, the collection server can receive page documents provided by different data sources, provide rich and multi-source data for enterprises and scientific research institutions, and use multi-source data to assist enterprise production and scientific research.
[0054] Optionally, when initiating a page access request to a mobile application and obtaining a page document fed back by the mobile application, the following is included:
[0055] Initiate a page access request to the web agent installed in the mobile device;
[0056] Receiving page documents fed back after the automated testing framework of the mobile device is driven by the Web agent program to run and obtain page documents of the mobile application program;
[0057] Among them, the web agent collects page documents of mobile applications based on the automated testing framework.
[0058] The mobile application in the embodiment of the present application does not support directly initiated network requests. A web agent is pre-installed in the mobile device. After receiving a page access request, the web agent drives the automated testing framework to run to obtain the page document of the mobile application. The web agent drives the automated testing framework to run and performs operations such as clicking, sliding, and taking screenshots to obtain the page document of the mobile application.
[0059] During the data collection process, you need to switch pages or turn pages to obtain data. By driving the automated test framework to run, you can use automated testing technology to click and slide to switch pages or turn pages. You can also use automated testing technology to take screenshots, based on OCR (Optical Character Recognition) technology to identify information in the picture to extract data, and through screenshots, you can save the current page status for subsequent detection of page changes. You can also record some logs to facilitate troubleshooting and analysis when anomalies occur.
[0060] It should be noted that the Web agent in this embodiment is a functional concept, and can also be called a test driver. After receiving a page access request, the test driver drives the automated test framework to run to obtain a page document provided by the mobile application.
[0061] After the acquisition server initiates a page access request for the mobile application, the Web agent drives the automated testing framework to run based on the page access request to obtain the page document of the mobile application and feed the page document back to the acquisition server. The acquisition server receives the original page content captured from the mobile device.
[0062] Optionally, when initiating a page access request to the PC application and obtaining a page document fed back by the PC application, the method includes: initiating a page access request to a WinAppDriver service, wherein the acquisition server interacts with the PC application through the WinAppDriver service;
[0063] Receive the page document fed back by WinAppDriver service driving the UI automation framework to run and obtain the page document of the PC application;
[0064] Among them, the WinAppDriver service identifies and operates UI elements in PC applications based on the UI automation framework and collects page documents of PC applications.
[0065] The PC application in the embodiment of the present application does not support directly initiated network requests, and the acquisition server uses the WinAppDriver service to interact with the PC application. Among them, the WinAppDriver service uses the UI (UserInterface, user interface) automation framework to identify and operate the UI elements in the PC application and collect the page documents of the PC application. The page access request initiated by the acquisition server for the PC application is received by the WinAppDriver service, and the WinAppDriver service drives the UI automation framework to run to obtain the page document of the PC application, and the obtained page document is fed back to the acquisition server as a response result. For the WinAppDriver service, after receiving the network request (page access request) sent by the acquisition server based on the acquisition program, it uses the UI automation framework to perform the corresponding operation. After the operation is completed, the result of the operation is sent to the acquisition server in the form of an HTTP response.
[0066] For the collection server, after obtaining the page documents provided by at least two data sources, it inputs the original page content and the matching data collection prompt words into the target big model, and the target big model automatically extracts the target data. It can use the big model to extract multi-source data, provide rich and multi-source data for enterprises and scientific research institutions, and use multi-source data to assist enterprise production and scientific research.
[0067] The following is an introduction to the process of constructing an input array. When obtaining the page prompt words corresponding to each page document and constructing multiple input arrays including page documents and corresponding page prompt words, it includes:
[0068] After displaying multiple page documents provided by the target data source, obtaining a page prompt word corresponding to each page document, which is determined based on analyzing the page content and page structure of the page document;
[0069] For each page document, the page document and the corresponding page prompt word are combined to construct an input array for input into the target large model.
[0070] After receiving the page document provided by the target data source, the acquisition server displays the acquired page document, and manually analyzes the page document and determines the page prompt word based on the page content and page structure. Alternatively, after displaying the page document, the recognition model is used to automatically analyze and determine the page prompt word corresponding to the page document. The acquired page prompt word is the data acquisition prompt word.
[0071] After acquiring the page document and the page prompt word corresponding to the page document, the acquisition server combines the current page document with the corresponding page prompt word for each page document, so as to construct an input array through the combination of the page content and the prompt word.
[0072] After constructing the input arrays corresponding to the respective page documents, when using the target large model to extract the target data matching the page prompt words, it includes:
[0073] The target big model is used to analyze multiple input arrays, and the target big model extracts target data from the page document based on the rules described by the page prompt words corresponding to the page document; wherein the page prompt words are used to assist the target big model in parsing the page document and extracting data from the page document.
[0074] After constructing multiple input arrays, the input arrays are input into the target large model. Since the input array includes pairs of page documents and page prompt words, the target large model can extract target data from the page documents based on the rules described by the page prompt words, and parse the page documents with the assistance of the page prompt words, extract data from the page documents, and output adapted target data.
[0075] After the target large model analyzes and processes the input array and outputs the target data, the output results of the target large model are saved in the database to provide rich and diverse data for enterprises and scientific research institutions based on the valid data extracted from the page document.
[0076] like Figure 2 As shown, it is a specific interactive flow chart of the multi-source data collection method provided by the embodiment of the present application. The collection server initiates page access requests to the Web program, PC application and mobile application, receives the response result directly fed back by the Web program, receives the page document provided by the PC application fed back by the WinAppDriver service, and receives the page document provided by the mobile application fed back by the Web agent program.
[0077] After acquiring the page documents provided by the Web program, PC application and mobile application, the acquisition server receives the page prompt words corresponding to each page document provided manually, combines the page prompt words with the page document into an input array, and then inputs it into the target large model. The target large model extracts the target data from the page document based on the page prompt words, outputs the extraction results, and stores the output extraction results in the database for providing multi-source page content.
[0078] Among them, when the page is updated, by adjusting the page prompt words, the target large model re-extracts the target data in the updated page based on the updated page prompt words, without the need for developers to re-analyze the page structure, locate the target elements, and modify the crawling logic to adapt to business changes.
[0079] The above implementation plan supports data collection in Web programs, PC applications and mobile applications, expands the data collection sources, and can provide rich and multi-source data for enterprises and scientific research institutions; using prompt words, automatically parsing and processing page content based on a large language model, intelligently extracting data from the page, and timely adjusting prompt words based on page updates to extract adaptive data can reduce development costs, flexibly extract data, and quickly respond to business changes. It solves the pain points of high development and maintenance costs, poor flexibility, and difficulty in adapting to business changes in the traditional data collection process. At the same time, it can improve data collection efficiency and help enterprise production and scientific research.
[0080] The present application embodiment provides a multi-source data acquisition device, which is applied to an acquisition server, such as Figure 3 As shown, the device comprises:
[0081] A first acquisition module 301 is used to acquire multiple page documents provided by a target data source based on a page access request initiated to the target data source, wherein the target data source includes at least two of a Web program, a PC application program, and a mobile application program;
[0082] An acquisition and construction module 302 is used to acquire the page prompt words corresponding to each page document, and construct a plurality of input arrays including the page documents and the corresponding page prompt words;
[0083] The processing module 303 is used to input the multiple input arrays into a target macro model, and use the target macro model to extract target data matching the corresponding page prompt words in the page documents to obtain page data corresponding to the multiple page documents respectively.
[0084] Optionally, the device further comprises:
[0085] A second acquisition module, configured to acquire a page prompt word corresponding to the second page document in response to identifying that a first page document among the plurality of page documents is updated to a second page document;
[0086] An update determination module, configured to update the associated input array based on the second page document and the page prompt word corresponding to the second page document, and determine a target input array;
[0087] An input extraction module is used to input the target input array into the target large model, and the target large model extracts data from the second page document based on the page prompt words corresponding to the second page document to obtain the latest page data corresponding to the second page document.
[0088] Optionally, the first acquisition module includes:
[0089] An initiating submodule, used to initiate the page access request to at least two of the Web program, the PC application and the mobile application, wherein the page access request is a network request;
[0090] A receiving submodule, configured to receive page documents respectively fed back by at least two of the Web program, the PC application, and the mobile application;
[0091] Among them, the Web program directly feeds back the page document based on the page access request initiated by the acquisition server, the PC application receives the page access request through the WinAppDriver service and provides the page document to the acquisition server, and the mobile application receives the page access request through the Web agent program and provides the page document to the acquisition server.
[0092] Optionally, when initiating the page access request to the mobile application and acquiring the page document fed back by the mobile application, the first acquisition module is further used to:
[0093] Initiating the page access request to the Web agent program installed in the mobile device;
[0094] Receiving a page document fed back after the Web agent program drives the automated testing framework of the mobile device to run and obtains the page document of the mobile application;
[0095] Wherein, the Web agent program collects page documents of the mobile application based on the automated testing framework.
[0096] Optionally, when initiating the page access request to the PC application and obtaining the page document fed back by the PC application, the first obtaining module is further used to:
[0097] Initiating the page access request to the WinAppDriver service, and the acquisition server interacting with the PC application through the WinAppDriver service;
[0098] Receiving the page document fed back after the WinAppDriver service drives the user interface UI automation framework to run and obtain the page document of the PC application;
[0099] The WinAppDriver service identifies and operates UI elements in the PC application based on a UI automation framework and collects page documents of the PC application.
[0100] Optionally, the acquisition building block includes:
[0101] An acquisition submodule, for acquiring, after displaying a plurality of page documents provided by the target data source, a page prompt word corresponding to each page document and determined based on analyzing the page content and page structure of the page document;
[0102] The construction submodule is used to combine the page document with the corresponding page prompt word for each page document to construct an input array for input into the target large model.
[0103] Optionally, the processing module is further used for:
[0104] Analyzing the multiple input arrays using the target large model, and extracting the target data from the page document based on the rule described by the page prompt word corresponding to the page document by the target large model;
[0105] The page prompt words are used to assist the target large model in parsing the page document and extracting data from the page document.
[0106] As for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0107] The present application also provides a multi-source data acquisition system. Figure 4 As shown, it includes: a collection server 41, a collected device 42, a data processing device 43 and a data storage device 44;
[0108] The collected device 42 provides page documents based on the running mobile application, PC application and Web program;
[0109] The collection server 41 runs a collection program to collect page documents provided by at least two of the mobile application, the PC application, and the Web application;
[0110] The data processing device 43 deploys a prompt word output module 431 and a target large model 432. The prompt word output module 431 is used to output the page prompt word corresponding to the page document, and the target large model 432 is used to extract target data from the page document based on the page prompt word corresponding to the page document.
[0111] The acquisition server 41 executes the above-mentioned multi-source data acquisition method based on the interaction with the acquired device 42 and the data processing device 43 to collect page documents and extract page data in the page documents using the target large model, wherein the extracted page data is stored in the data storage device 44.
[0112] The collected devices 42 may include mobile devices and PC devices. Mobile applications installed on mobile devices may provide page documents, PC applications installed on PC devices may provide page documents, and Web programs deployed on PC devices may also provide page documents. Among them, Web programs and PC applications may also be directly run on the collection server 41, and the PC applications and Web programs running on the collection server 41 provide page documents. The collection server 41 receives page documents provided by at least two of the Web program, PC application, and mobile application by initiating a page access request.
[0113] Among them, on the mobile device side, the Web agent program drives the automated testing framework to run based on the page access request to obtain the page document of the mobile application, and feeds the page document back to the acquisition server 41, so that the acquisition server 41 receives the original page content captured from the mobile device.
[0114] For PC applications, which do not support direct page access requests, the acquisition server 41 uses the WinAppDriver service to interact with the PC application. The WinAppDriver service uses the UI automation framework to identify and operate the UI elements in the PC application, collect the page documents of the PC application, and provide the collected page documents to the acquisition server 41. The Web program can directly feedback the page document based on the page access request initiated by the acquisition server 41.
[0115] The data processing device 43 is deployed with a prompt word output module 431 and a target large model 432, wherein the prompt word output module 431 is used to output the page prompt words corresponding to the page document. The page prompt words can be determined manually. For example, after the acquisition server 41 obtains the page document, it is manually analyzed to obtain the page prompt words, and the page prompt words are provided to the prompt word output module 431, and the prompt word output module 431 outputs the page prompt words.
[0116] The target large model 432 is used to receive an input array based on page prompt words and page documents, analyze the input array, extract target data from the page document based on the page prompt words corresponding to the page document and output it, and the output data is stored in a data storage device 44, such as a database, so as to use the data storage device 44 to store the data output by the target large model 432.
[0117] When a page is updated, the prompt word output module 431 can provide page prompt words again, and the target large model 432 re-processes the updated page document based on the page prompt words, and outputs the latest page data. By adjusting the prompt words, the latest page data can be obtained efficiently and at low cost, without the need for complex system reconstruction or code modification, reducing manual participation and greatly improving the speed and efficiency of data collection.
[0118] The system embodiment is basically similar to the method embodiment and the description is relatively simple. For the relevant parts, please refer to the partial description of the method embodiment.
[0119] The multi-source data acquisition system provided in the above-mentioned embodiments of the present application can provide multi-source data and quickly adjust the data acquisition strategy according to the changes in business needs, thereby solving the pain points of high development and maintenance costs, poor flexibility, and difficulty in adapting to business changes in the traditional data acquisition process, improving data collection efficiency, and being helpful to enterprise production and scientific research.
[0120] An embodiment of the present application also provides an electronic device, including: a processor, a memory, and a computer program stored in the memory and executable on the processor. When the computer program is executed by the processor, each process of the multi-source data acquisition method embodiment described above is implemented, and the same technical effect can be achieved. To avoid repetition, it will not be described here.
[0121] For example, Figure 5 FIG. 1 shows a schematic diagram of the physical structure of an electronic device. Figure 5 As shown, the electronic device may include: a processor 510, a communication interface 520, a memory 530 and a communication bus 540, wherein the processor 510, the communication interface 520 and the memory 530 communicate with each other via the communication bus 540. The processor 510 may call the logic instructions in the memory 530, and the processor 510 is used to execute each process of the multi-source data acquisition method of the embodiment of the present application, which will not be described one by one here.
[0122] In addition, the logic instructions in the above-mentioned memory 530 can be implemented in the form of software functional units and can be stored in a computer-readable storage medium when sold or used as an independent product. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product, which is stored in a storage medium and includes a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present application.
[0123] The embodiment of the present application also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, each process of the above-mentioned multi-source data acquisition method embodiment is implemented, and the same technical effect can be achieved. To avoid repetition, it is not repeated here. The computer-readable storage medium is, for example, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.
[0124] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, method, article or device including the element.
[0125] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, a magnetic disk, or an optical disk), and includes a number of instructions for a terminal (which can be a mobile phone, a computer, a server, an air conditioner, or a network device, etc.) to execute the methods described in each embodiment of the present application.
[0126] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of the present application, ordinary technicians in this field can also make many forms without departing from the purpose of the present application and the scope of protection of the claims, all of which are within the protection of the present application.
Claims
1. A multi-source data collection method, applied to a collection server, characterized in that: include: Based on a page access request initiated to a target data source, a plurality of page documents provided by the target data source are obtained, wherein the target data source includes at least two of a Web program, a PC application program, and a mobile application program; Obtaining page prompt words corresponding to each page document, and constructing multiple input arrays including the page documents and the corresponding page prompt words; The multiple input arrays are input into a target macro model, and the target macro model is used to extract target data matching corresponding page prompt words from the page documents, so as to obtain page data corresponding to the multiple page documents respectively.
2. The method according to claim 1, characterized in that Also includes: In response to identifying that a first page document among the plurality of page documents is updated to a second page document, obtaining a page prompt word corresponding to the second page document; Based on the second page document and the page prompt word corresponding to the second page document, updating the associated input array to determine the target input array; The target input array is input into the target macro model, and the target macro model extracts data from the second page document based on the page prompt words corresponding to the second page document to obtain the latest page data corresponding to the second page document.
3. The method according to claim 1, characterized in that The acquiring of a plurality of page documents provided by the target data source based on a page access request initiated to the target data source includes: Initiate the page access request to at least two of the Web program, the PC application, and the mobile application, wherein the page access request is a network request; Receiving page documents respectively fed back by at least two of the Web program, the PC application, and the mobile application; Among them, the Web program directly feeds back the page document based on the page access request initiated by the acquisition server, the PC application receives the page access request through the WinAppDriver service and provides the page document to the acquisition server, and the mobile application receives the page access request through the Web agent program and provides the page document to the acquisition server.
4. The method according to claim 3, characterized in that When the page access request is initiated to the mobile application and the page document fed back by the mobile application is obtained, it includes: Initiating the page access request to the Web agent program installed in the mobile device; Receiving a page document fed back after the Web agent program drives the automated testing framework of the mobile device to run and obtains the page document of the mobile application; Wherein, the Web agent program collects page documents of the mobile application based on the automated testing framework.
5. The method according to claim 3, characterized in that: When the page access request is initiated to the PC application and the page document fed back by the PC application is obtained, the following steps are included: Initiating the page access request to the WinAppDriver service, wherein the acquisition server interacts with the PC application through the WinAppDriver service; Receiving the page document fed back after the WinAppDriver service drives the user interface UI automation framework to run and obtain the page document of the PC application; The WinAppDriver service identifies and operates UI elements in the PC application based on a UI automation framework and collects page documents of the PC application.
6. The method according to claim 1, characterized in that The step of obtaining the page prompt words corresponding to each page document and constructing a plurality of input arrays including the page documents and the corresponding page prompt words includes: After displaying a plurality of page documents provided by the target data source, obtaining a page prompt word corresponding to each page document, which is determined based on analyzing the page content and page structure of the page document; For each page document, the page document is combined with the corresponding page prompt word to construct an input array for inputting into the target large model.
7. The method according to claim 1, characterized in that The step of inputting the plurality of input arrays into a target macro model and extracting target data matching corresponding page prompt words from the page document using the target macro model comprises: Analyzing the multiple input arrays using the target large model, and extracting the target data from the page document based on the rule described by the page prompt word corresponding to the page document by the target large model; The page prompt words are used to assist the target large model in parsing the page document and extracting data from the page document.
8. A multi-source data acquisition device, applied to an acquisition server, characterized in that: include: A first acquisition module, configured to acquire a plurality of page documents provided by a target data source based on a page access request initiated to the target data source, wherein the target data source includes at least two of a Web program, a PC application program, and a mobile application program; An acquisition construction module is used to acquire the page prompt words corresponding to each page document, and construct a plurality of input arrays including the page documents and the corresponding page prompt words; The processing module is used to input the multiple input arrays into a target large model, and use the target large model to extract target data matching the corresponding page prompt words in the page document to obtain page data corresponding to the multiple page documents respectively.
9. A multi-source data acquisition system, characterized in that: include: Collection server, collected equipment, data processing equipment and data storage equipment; The collected device provides a page document based on the running mobile application, PC application and Web program; The acquisition server runs an acquisition program to acquire page documents provided by at least two of the mobile application, the PC application, and the Web program; The data processing device deploys a prompt word output module and a target large model, wherein the prompt word output module is used to output the page prompt word corresponding to the page document, and the target large model is used to extract target data from the page document based on the page prompt word corresponding to the page document; Based on the interaction with the collected device and the data processing device, the collection server executes the method described in any one of claims 1 to 7 to collect page documents and extract page data in the page documents using the target large model, wherein the extracted page data is stored in the data storage device.
10. An electronic device, characterized in that: The method comprises a processor, a memory and a computer program stored in the memory and executable on the processor, wherein when the computer program is executed by the processor, the steps of the multi-source data acquisition method according to any one of claims 1 to 7 are implemented.