A cross-platform data acquisition system and method
By analyzing the correlation between the data crawling results of a single platform and other cross-platform data sets, the problems of data extraction efficiency and quality in the existing technology are solved, and high-precision cross-platform data acquisition is achieved, which improves the accuracy and crawling efficiency of data results.
Patent Information
- Application Number
- CN202411236460.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-04
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2044-09-04
AI Technical Summary
The existing cross-platform data extraction technology has problems with data crawling speed and quality, which leads to the quality of data results being limited by the accuracy of the correlation calculation process and the crawling efficiency is low.
By analyzing the correlation degree between the data crawling results of a single platform and other cross-platform data sets, using the data structure to sort out and divide the modules to generate a data structure tree, and calculating the correlation degree through the correlation analysis module, achieving high-precision cross-platform data acquisition.
It improves the quality and efficiency of cross-platform data acquisition, ensures the accuracy of data results and improves crawling efficiency.
Smart Images

Figure CN119577222B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data analysis, and particularly to a cross-platform data acquisition system and method. Background Art
[0002] Currently, cross-platform data extraction refers to the process of automatically collecting and extracting specific information required from data sources on multiple different platforms. These platforms can include websites, APIs, Excel files, databases, etc. Cross-platform data extraction technology has been widely applied in many fields, such as business intelligence, marketing, finance, human resources, etc. The commonly used existing technologies for cross-platform data extraction mainly include the following aspects:
[0003] Web Scraping: A web scraping is a program that automatically grabs web page content and can grab information such as text, pictures, links, etc. from websites. Through web scraping technology, the required data can be automatically obtained on multiple platforms.
[0004] API Calls (Application Programming Interface Calls): Many service providers provide APIs that allow developers to access and operate data using code on their platforms. Using APIs can obtain data more conveniently and is usually more efficient than using traditional web scraping.
[0005] Natural Language Processing (NLP): Natural language processing technology is mainly used to process and parse human language text. In the scenario of cross-platform data extraction, NLP technology can be used to analyze unstructured data (such as text messages, emails, etc.) to achieve cross-platform data extraction.
[0006] Data mapping and visualization tools: In order to better understand the data sources and relationships, data mapping and visualization tools can be used to visualize the extracted data, such as charts, timelines, etc.
[0007] In addition, the technical background of cross-platform data extraction also includes: data mining, machine learning, artificial intelligence, etc. These technologies provide strong support for cross-platform data extraction, making the mining and utilization of data more intelligent and efficient.
[0008] However, there are also some deficiencies in cross-platform data extraction, such as data crawling speed and quality issues: Due to the complexity of data source platforms, cross-platform data extraction requires comparing the data of multiple platforms one by one to obtain cross-platform crawling results with a relatively high degree of correlation and meeting the crawling requirements of user input. This may limit the data quality of the data results obtained from cross-platform data acquisition by the accuracy of a large number of correlation calculation processes, and due to the huge amount of cross-platform data requiring a large number of correlation calculation processes, the crawling efficiency is also relatively low.
[0009] Therefore, the present invention proposes a cross-platform data acquisition system and method. Summary of the Invention
[0010] The present invention provides a cross-platform data acquisition system and method for analyzing the degree of association between the data results obtained by crawling on a single platform and the data sets in different data ranges of other cross-platforms based on the data structure sorting results of multiple cross-platforms, and realizing high-precision cross-platform data acquisition based on the analyzed association degree analysis results, improving the quality and efficiency of cross-platform data acquisition.
[0011] The present invention provides a cross-platform data acquisition system, including:
[0012] A single-platform crawling module for performing data crawling on a single platform based on the user's data crawling requirements on the single platform to obtain a single-platform data crawling result;
[0013] A data structure sorting module for sorting the complete data of each cross-platform on the single platform to generate a data structure tree for each cross-platform;
[0014] A data structure division module for dividing the complete data of each cross-platform according to the structure based on the data structure tree of each cross-platform to obtain all data sets at each structural level of each cross-platform;
[0015] An association degree analysis module for performing association degree analysis on the single-platform data crawling result and all data sets at all structural levels of all cross-platforms to obtain an association degree analysis result;
[0016] A cross-platform crawling module for performing cross-platform data crawling in the complete data of each cross-platform based on the association degree analysis result to obtain a cross-platform data crawling result.
[0017] Preferably, the data structure sorting module includes:
[0018] A lowest-level determination sub-module for obtaining all the lowest-level partition addresses of the complete data of each cross-platform on the single platform;
[0019] The upper address generalization sub-module is used to continuously perform upper generalization on each lowest-level partition address based on the hierarchical tags in each lowest-level partition address to obtain all upper addresses of each lowest-level partition address;
[0020] The data chain generation sub-module is used to store all the data stored in each lowest-level partition address and its corresponding all upper addresses, sort and connect them in descending order according to the upper degree of the addresses, and obtain a unidirectional hierarchical data chain for each lowest-level partition address;
[0021] The data chain merging sub-module is used to merge the data nodes with repeated addresses in the unidirectional hierarchical data chains of all the lowest-level partition addresses of each cross-platform to obtain a data structure tree for each cross-platform.
[0022] Preferably, the data structure division module includes:
[0023] The data node determination sub-module is used to determine all the data nodes in the data structure tree of each cross-platform;
[0024] The data set determination sub-module is used to summarize all the data stored in the partition addresses corresponding to each data node as the data set corresponding to the data node, and use the data sets of all the data nodes at each structural level of each cross-platform as all the data sets at the corresponding structural level of the corresponding cross-platform.
[0025] Preferably, the correlation analysis module includes:
[0026] The correlation degree calculation sub-module is used to calculate the correlation degree between the single-platform data crawling result and each data set at the last layer of each cross-platform;
[0027] The correlation degree upper crawling sub-module is used to perform upper crawling calculation based on the data structure tree of each cross-platform and the correlation degree between the single-platform data crawling result and all the data sets at the corresponding last layer to obtain the correlation degree between the single-platform data crawling result and each data set at each structural level of each cross-platform;
[0028] The result summary sub-module is used to regard the correlation degree between the single-platform data crawling result and all the data sets at all the structural levels of all the cross-platforms as the correlation analysis result.
[0029] Preferably, the correlation degree calculation sub-module includes:
[0030] The form standardization processing unit is used to perform form standardization processing on all the data in the single-platform data crawling result to obtain the first form-standardized data. At the same time, it performs form standardization processing on each data set at the last layer of each cross-platform to obtain the second form-standardized data of each data set at the last layer of each cross-platform;
[0031] The first escape error rate calculation unit is configured to use the ratio of the amount of data in each non - standardized form in all data in the single - platform data crawling result to the total amount of data in the single - platform data crawling result as the proportion of the amount of data in each non - standardized form in the single - platform data crawling result, and use the sum of the products of the escape error rates of all non - standardized forms in the single - platform data crawling result and the proportion of the amount of data as the standardized escape error rate of the first - form standardized data;
[0032] The second escape error rate calculation unit is configured to use the ratio of the amount of data in each non - standardized form in all data in each dataset of each cross - platform last layer to the amount of data in the corresponding dataset as the proportion of the amount of data in each non - standardized form in the corresponding dataset, and use the sum of the products of the escape error rates of all non - standardized forms in each dataset of each cross - platform last layer and the proportion of the amount of data as the standardized escape error rate of the corresponding second - form standardized data;
[0033] The first correlation degree calculation unit is configured to calculate the correlation degree between the first - form standardized data and the second - form standardized data of each dataset of each cross - platform last layer;
[0034] The second correlation degree calculation unit is configured to calculate the correlation degree between the single - platform data crawling result and each dataset of each cross - platform last layer based on the correlation degree between the first - form standardized data and the second - form standardized data of each dataset of each cross - platform last layer, the standardized escape error rate of the first - form standardized data, and the standardized escape error rate of the second - form standardized data of each dataset of each cross - platform last layer.
[0035] Preferably, the first correlation degree calculation unit includes:
[0036] The data segmentation subunit is configured to segment the first - form standardized data to obtain a plurality of first - segmented data. At the same time, segment the second - form standardized data of each dataset of each cross - platform last layer to obtain a plurality of second - segmented data;
[0037] The correlation degree analysis subunit is configured to analyze the data correlation degree between each first - segmented data and each second - segmented data based on the paragraph data correlation degree analysis model;
[0038] The correlation degree calculation subunit is used to regard the data correlation degree between each first segmented data in the first-form standardized data and all second segmented data in each dataset of each cross-platform last layer as the available data correlation degree of each first segmented data in the first-form standardized data, and regard the mean value of all available data correlation degrees in the first-form standardized data as the correlation degree between the first-form standardized data and the second-form standardized data of each dataset of each cross-platform last layer.
[0039] Preferably, the second correlation degree calculation unit calculates the correlation degree between the single-platform data crawling result and each dataset of each cross-platform last layer based on the correlation degree between the first-form standardized data and the second-form standardized data of each dataset of each cross-platform last layer, the standardized escape error rate of the first-form standardized data, and the standardized escape error rate of the second-form standardized data of each dataset of each cross-platform last layer, including:
[0040]
[0041] In the formula, ε ’ is the correlation degree between the single-platform data crawling result and the currently calculated dataset of the currently calculated cross-platform last layer, ε is the correlation degree between the first-form standardized data and the second-form standardized data of the currently calculated dataset of the currently calculated cross-platform last layer, σ 1 is the standardized escape error rate of the first-form standardized data, σ 2 is the standardized escape error rate of the second-form standardized data of the currently calculated dataset of the currently calculated cross-platform last layer, min(ε + σ 1 - σ 2 , ε - σ 1 + σ 2 ) is to take the minimum value of ε + σ 1 - σ 2 and ε - σ 1 + σ 2 .
[0042] Preferably, the upper correlation degree sub-module includes:
[0043] The node relationship sorting unit is used to determine all hierarchical adjacent sub-nodes of each data node in each cross-platform data structure tree based on each cross-platform data structure tree;
[0044] The crawling weight determination unit is used to regard the ratio between the data volume of each hierarchical adjacent sub-node of each data node and the sum of the data volumes of all hierarchical adjacent sub-nodes of the corresponding data node as the upper crawling weight of each hierarchical adjacent sub-node of each data node.
[0045] The upper-level crawling calculation unit is used to perform upper-level crawling calculation based on the upper-level crawling weight of each hierarchical adjacent child node of each data node in each cross-platform data structure tree and the correlation degree between the single-platform data crawling result and all data sets in the corresponding last layer, so as to obtain the correlation degree between the single-platform data crawling result and each data set at each structural level of each cross-platform.
[0046] Preferably, the cross-platform crawling module includes:
[0047] The correlation degree calculation sub-module is used to determine the correlation degree between the single-platform data crawling result and all data sets at all structural levels of all cross-platforms based on the correlation degree analysis result;
[0048] The cross-platform crawling sub-module is used to screen out all data sets with a correlation degree not less than the correlation degree threshold from all data sets at all structural levels of all cross-platforms as the cross-platform data crawling result.
[0049] The present invention provides a cross-platform data acquisition method for implementing the cross-platform data acquisition system described in any one of Embodiments 1 to 9, including:
[0050] S1: Based on the data crawling requirements of the user on a single platform, perform data crawling on the single platform to obtain the single-platform data crawling result;
[0051] S2: Sort out the structures of the complete data of each cross-platform on the single platform to generate a data structure tree for each cross-platform;
[0052] S3: Divide the complete data of each cross-platform according to the structure based on the data structure tree of each cross-platform to obtain all data sets at each structural level of each cross-platform;
[0053] S4: Perform correlation degree analysis on the single-platform data crawling result and all data sets at all structural levels of all cross-platforms to obtain the correlation degree analysis result;
[0054] S5: Perform cross-platform data crawling on the complete data of each cross-platform based on the correlation degree analysis result to obtain the cross-platform data crawling result.
[0055] The beneficial effects of the present invention compared with the prior art are as follows: Based on the data structure sorting results of multiple cross-platforms, analyze the correlation degree between the data results crawled on a single platform and the data sets in different data ranges of other cross-platforms, and realize high-precision cross-platform data acquisition based on the analyzed correlation degree analysis result, improving the quality and efficiency of cross-platform data acquisition.
[0056] Other features and advantages of the present invention will be set forth in the following description, and in part will be obvious from the description, or may be learned by practice of the present invention. The objectives and other advantages of the present invention may be realized and attained by the structure particularly pointed out in this application document.
[0057] The technical solution of the present invention will be further described in detail below through the accompanying drawings and embodiments. Description of the Drawings
[0058] The accompanying drawings are used to provide a further understanding of the present invention, and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention, but do not constitute a limitation to the present invention. In the accompanying drawings:
[0059] Figure 1 It is a schematic diagram of the internal functional sub-modules of a cross-platform data acquisition system in an embodiment of the present invention;
[0060] Figure 2 It is a flowchart of a cross-platform data acquisition method in an embodiment of the present invention. Detailed Embodiments
[0061] The following describes the preferred embodiments of the present invention with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are only for explaining and illustrating the present invention, and are not used to limit the present invention.
[0062] Embodiment 1:
[0063] The present invention provides a cross-platform data acquisition system. Refer to Figure 1 , including:
[0064] A single-platform crawling module, which is used to perform data crawling on a single platform based on the user's data crawling requirements on the single platform, and obtain the single-platform data crawling result;
[0065] A data structure sorting module, which is used to sort the structure of each cross-platform complete data on the single platform, and generate a data structure tree for each cross-platform;
[0066] A data structure division module, which is used to divide the complete data of each cross-platform according to the structure based on the data structure tree of each cross-platform, and obtain all data sets at each structural level of each cross-platform;
[0067] A correlation analysis module, which is used to perform correlation analysis on the single-platform data crawling result and all data sets at all structural levels of all cross-platforms, and obtain the correlation analysis result;
[0068] A cross-platform crawling module, which is used to perform cross-platform data crawling on the complete data of each cross-platform based on the correlation analysis result, and obtain the cross-platform data crawling result.
[0069] In this embodiment, the single platform is a data storage platform for directly inputting data crawling requirements by the user, such as a certain search engine website.
[0070] In this embodiment, the data crawling requirement is a data feature representing the data that the user wants to obtain.
[0071] In this embodiment, based on the data crawling requirements of the user on the single platform, data crawling is performed on the single platform, and obtaining the single platform data crawling result can be achieved through existing web crawler technologies.
[0072] In this embodiment, the single platform data crawling result includes all data results that meet the data crawling requirements obtained by crawling from the single platform.
[0073] In this embodiment, the cross-platform is a data storage platform other than the single platform, such as an Excel file, etc.
[0074] In this embodiment, the complete data is all the data stored in a single platform.
[0075] In this embodiment, the data structure tree is a tree structure that includes multiple data nodes belonging to different levels, and the data nodes contain data within different data ranges in the corresponding platform. The higher the level to which the data node belongs, the wider the data range it stores. Vice versa.
[0076] The beneficial effects of the above technologies are as follows: Based on the data structure sorting results of multiple cross-platforms, analyze the degree of association between the data results obtained by crawling on the single platform and the data sets with different data ranges of other cross-platforms, and realize high-precision cross-platform data acquisition based on the analyzed association degree analysis results, improving the quality and efficiency of cross-platform data acquisition.
[0077] Embodiment 2:
[0078] Based on Embodiment 1, the data structure sorting module includes:
[0079] The lowest-level determination sub-module is used to obtain all the lowest-level partition addresses of the complete data of each cross-platform of the single platform;
[0080] The upper address generalization sub-module is used to continuously perform upper generalization on each lowest-level partition address based on the level markers in each lowest-level partition address to obtain all the upper addresses of each lowest-level partition address;
[0081] The data chain generation sub-module is used to store all the data stored in each lowest-level partition address and its corresponding all upper addresses, sort and connect them in descending order according to the upper degree of the addresses, and obtain a one-way hierarchical data chain for each lowest-level partition address;
[0082] A data link merging sub-module, which is used to merge the data nodes with duplicate addresses in the one-way hierarchical data link of all the lowest-level partition addresses of each cross-platform, so as to obtain a data structure tree for each cross-platform.
[0083] In this embodiment, the lowest-level partition address is the smallest unit storage address among all the data storage addresses of the cross-platform, and it is also the address with the most hierarchical markers, that is, the data storage address without a lower-level and more detailed data storage address.
[0084] In this embodiment, the hierarchical marker is the marker used to distinguish the levels of the storage addresses in the storage address, such as " / " in a website address.
[0085] In this embodiment, based on the hierarchical markers in each lowest-level partition address, the lowest-level partition addresses are continuously generalized upwards to obtain all the upper-level addresses of each lowest-level partition address, that is:
[0086] Delete the last slash and the partial address information after the last slash of each lowest-level partition address, which is the first address generalization process, and obtain an upper-level address of the lowest-level partition address;
[0087] Then delete the last slash and the partial address information after the last slash in the upper-level address obtained in the previous step, which is the second address generalization process, and obtain the second upper-level address of the lowest-level partition address;
[0088] And so on, until the upper-level address obtained latest does not contain hierarchical markers, then stop the address generalization process, and obtain all the upper-level addresses of the lowest-level partition address.
[0089] In this embodiment, the upper-level address is the data storage address whose data range includes the data range of the corresponding lowest-level partition address.
[0090] In this embodiment, the larger the data range of the upper-level address is relative to the data range of the corresponding lowest-level partition address, the greater its upper level, and vice versa.
[0091] The beneficial effects of the above technology are as follows: By continuously generalizing the lowest-level partition addresses of all the complete data of each cross-platform of a single platform upwards, all the upper-level addresses of each lowest-level partition address are obtained, realizing the efficient sorting of the data structure of the data stored across platforms by sorting the cross-platform data storage addresses, and clearly representing the data structure of the data stored across platforms by generating a one-way hierarchical data link and a data structure tree.
[0092] Embodiment 3:
[0093] Based on Embodiment 1, the data structure partitioning module includes:
[0094] The data node determination sub-module is used to determine all data nodes in each cross-platform data structure tree;
[0095] The data set determination sub-module is used to summarize all the data stored in the partition addresses corresponding to each data node as the data set corresponding to the data node, and regard the data sets of all data nodes at each structural level of each cross-platform as all data sets at the corresponding structural level of the corresponding cross-platform.
[0096] In this embodiment, the data node is a node in the data structure tree used to represent the corresponding data range.
[0097] The beneficial effect of the above technology is: Based on the structural levels of the data nodes in each cross-platform data structure tree, a clear sorting of all cross-platform data is realized.
[0098] Embodiment 4:
[0099] Based on Embodiment 1, the correlation analysis module includes:
[0100] The correlation degree calculation sub-module is used to calculate the correlation degree between the single-platform data crawling result and each data set at the last layer of each cross-platform;
[0101] The correlation degree upper-level sub-module is used to perform upper-level crawling calculation based on each cross-platform data structure tree and the correlation degree between the single-platform data crawling result and all data sets at the corresponding last layer, and obtain the correlation degree between the single-platform data crawling result and each data set at each structural level of each cross-platform;
[0102] The result summarization sub-module is used to regard the correlation degree between the single-platform data crawling result and all data sets at all structural levels of all cross-platforms as the correlation analysis result.
[0103] In this embodiment, the correlation degree represents the correlation degree between the corresponding two groups of data in terms of data crawling requirements. For example, if the data crawling requirement restricts the content of the crawled data, the correlation degree represents the correlation degree between the data in terms of data content;
[0104] If the data crawling requirement restricts the data form of the crawled data, the correlation degree represents the correlation degree between the data in terms of data form.
[0105] The beneficial effects of the above technology are as follows: By calculating the degree of association between the data crawling results of a single platform and each dataset in the last layer of each cross-platform, and performing upper-level crawling calculations, the acquisition rate of the degree of association between the data crawling results of a single platform and each dataset at each structural level of each cross-platform is greatly improved.
[0106] Embodiment 5:
[0107] Based on Embodiment 4, the degree-of-association calculation sub-module includes:
[0108] A form standardization processing unit for performing form standardization processing on all data in the data crawling results of a single platform to obtain first form-standardized data. At the same time, perform form standardization processing on each dataset in the last layer of each cross-platform to obtain second form-standardized data for each dataset in the last layer of each cross-platform;
[0109] A first escape error rate calculation unit for taking the ratio of the amount of data in each non-standardized form in all data in the data crawling results of a single platform to the total amount of data in the data crawling results of a single platform as the proportion of the amount of data in each non-standardized form in the data crawling results of a single platform, and taking the sum value of the product of the escape error rates of all non-standardized forms in the data crawling results of a single platform and the proportion of the amount of data as the standardized escape error rate of the first form-standardized data;
[0110] A second escape error rate calculation unit for taking the ratio of the amount of data in each non-standardized form in all data in each dataset in the last layer of each cross-platform to the amount of data in the corresponding dataset as the proportion of the amount of data in each non-standardized form in the corresponding dataset, and taking the sum value of the product of the escape error rates of all non-standardized forms in each dataset in the last layer of each cross-platform and the proportion of the amount of data as the standardized escape error rate of the corresponding second form-standardized data;
[0111] A first degree-of-association calculation unit for calculating the degree of association between the first form-standardized data and the second form-standardized data of each dataset in the last layer of each cross-platform;
[0112] A second degree-of-association calculation unit for calculating the degree of association between the data crawling results of a single platform and each dataset in the last layer of each cross-platform based on the degree of association between the first form-standardized data and the second form-standardized data of each dataset in the last layer of each cross-platform, the standardized escape error rate of the first form-standardized data, and the standardized escape error rate of the second form-standardized data of each dataset in the last layer of each cross-platform.
[0113] In this embodiment, the formal standardization process is to convert the data forms of all data within the corresponding data range into data in the standard form. Here, the data in the standard form is in text form, and this process can be completed using existing data format conversion methods.
[0114] In this embodiment, the non-standard form refers to other data forms other than the standard form, such as: image form, voice form, video form, etc.
[0115] In this embodiment, the escape error rate is the error rate generated in the data content when converting the data in the corresponding non-standard form into the data in the standard form.
[0116] In this embodiment, the standardized escape error rate is the error rate generated in the data content when converting each data set in the last layer of each cross-platform into the second-form standardized data.
[0117] The beneficial effects of the above technologies are as follows: By performing formal standardization processing on all data in the single-platform data crawling results and each data set in the last layer of each cross-platform, and calculating the escape error rate generated during the form conversion process, the accuracy of the correlation degree between the calculated single-platform data crawling results and each data set in the last layer of each cross-platform is ensured.
[0118] Embodiment 6:
[0119] Based on Embodiment 5, the first correlation degree calculation unit includes:
[0120] A data segmentation subunit, configured to segment the first-form standardized data to obtain a plurality of first segmented data. At the same time, segment the second-form standardized data of each data set in the last layer of each cross-platform to obtain a plurality of second segmented data;
[0121] A correlation degree analysis subunit, configured to analyze the data correlation degree between each first segmented data and each second segmented data based on the paragraph data correlation degree analysis model;
[0122] A correlation degree calculation subunit, configured to regard the data correlation degree between each first segmented data in the first-form standardized data and all second segmented data in each data set in the last layer of each cross-platform as the available data correlation degree of each first segmented data in the first-form standardized data, and regard the average value of all available data correlation degrees in the first-form standardized data as the correlation degree between the first-form standardized data and the second-form standardized data of each data set in the last layer of each cross-platform.
[0123] In this embodiment, the data segmentation can be performed on the data based on the timestamp of the data or a preset time interval.
[0124] In this embodiment, the paragraph data correlation analysis model is a pre-trained model that can evaluate the data correlation between two input sets of data;
[0125] This model is obtained by training with a large number of two sets of data whose data correlation between them is calculated using the existing technology as training samples.
[0126] In this embodiment, the data correlation represents the correlation between two sets of formally standardized data in terms of data crawling requirements. For example, if the data crawling requirement restricts the content of the crawled data, the degree of correlation represents the correlation between the data in terms of data content;
[0127] If the data crawling requirement restricts the data form of the crawled data, the degree of correlation represents the correlation between the data in terms of data form.
[0128] The beneficial effects of the above technology are as follows: By simultaneously segmenting the first formally standardized data and the second formally standardized data, and introducing a neural network model to accurately calculate the data correlation between the segmented data, the accurate calculation of the correlation degree between the first formally standardized data and the second formally standardized data of each dataset in the last layer of each cross-platform is realized.
[0129] Embodiment 7:
[0130] On the basis of Embodiment 5, the second correlation degree calculation unit calculates the correlation degree between the data crawling result of a single platform and each dataset in the last layer of each cross-platform based on the correlation degree between the first formally standardized data and the second formally standardized data of each dataset in the last layer of each cross-platform, the standardized escape error rate of the first formally standardized data, and the standardized escape error rate of the second formally standardized data of each dataset in the last layer of each cross-platform, including:
[0131]
[0132] In the formula, ε ’ is the correlation degree between the data crawling result of a single platform and the currently calculated dataset in the currently calculated last layer of the cross-platform, ε is the correlation degree between the first formally standardized data and the second formally standardized data of the currently calculated dataset in the currently calculated last layer of the cross-platform, σ 1 is the standardized escape error rate of the first formally standardized data, σ 2 is the standardized escape error rate of the second formally standardized data of the currently calculated dataset in the currently calculated last layer of the cross-platform, min(ε + σ 1 - σ 2 , ε - σ1 +σ 2 ) takes ε + σ 1 -σ 2 and ε - σ 1 +σ 2 as the minimum value.
[0133] The beneficial effects of the above technology are as follows: By introducing the standardized escape error rate of the first-form standardized data and the standardized escape error rate of the second-form standardized data, the accurate calculation of the correlation degree between the data crawling results of a single platform and each dataset of the last layer of each cross-platform is realized.
[0134] Embodiment 8:
[0135] On the basis of Embodiment 4, the upper-level sub-module of the correlation degree includes:
[0136] The node relationship sorting unit is used to determine all hierarchical adjacent sub-nodes of each data node in the data structure tree of each cross-platform based on the data structure tree of each cross-platform;
[0137] The crawling weight determination unit is used to regard the ratio between the data volume of each hierarchical adjacent sub-node of each data node and the sum of the data volumes of all hierarchical adjacent sub-nodes of the corresponding data node as the upper-level crawling weight of each hierarchical adjacent sub-node of each data node;
[0138] The upper-level crawling calculation unit is used to perform upper-level crawling calculation based on the upper-level crawling weights of each hierarchical adjacent sub-node of each data node in the data structure tree of each cross-platform and the correlation degree between the data crawling results of a single platform and all datasets of the corresponding last layer, and obtain the correlation degree between the data crawling results of a single platform and each dataset of each structural level of each cross-platform.
[0139] In this embodiment, all hierarchical adjacent sub-nodes of a data node are the sub-nodes whose levels belong to the adjacent lower level of the level to which the data node belongs.
[0140] In this embodiment, the upper-level crawling weight represents the correlation degree between the hierarchical adjacent sub-node and the data crawling results of a single platform, and the proportion in the correlation degree value between the data node and the data crawling results of a single platform.
[0141] In this embodiment, performing upper-level crawling calculation based on the upper-level crawling weights of each hierarchical adjacent sub-node of each data node in the data structure tree of each cross-platform and the correlation degree between the data crawling results of a single platform and all datasets of the corresponding last layer, and obtaining the correlation degree between the data crawling results of a single platform and each dataset of each structural level of each cross-platform includes:
[0142] The sum of the products of the degree of association between the datasets of all child nodes (which are also one or more data nodes in the last layer) of each data node belonging to the penultimate layer and the single-platform data crawling result and the upper crawling weights of the corresponding child nodes is regarded as the degree of association between the dataset of the corresponding data node in the penultimate layer and the single-platform data crawling result;
[0143] And in the above manner, based on the degree of association between the datasets of all child nodes of each data node belonging to the third-to-last layer and the single-platform data crawling result and the upper crawling weights of the corresponding child nodes, continue to calculate the degree of association between the dataset of the corresponding data node in the third-to-last layer and the single-platform data crawling result;
[0144] And so on, until the degree of association between the dataset of the data node in the top layer and the single-platform data crawling result is determined, and the degree of association between the single-platform data crawling result and each dataset at each structural level of each cross-platform is obtained.
[0145] The beneficial effects of the above technology are as follows: By introducing the upper crawling weights of each level of adjacent child nodes of each data node, continuous upper crawling calculations are performed on the basis of the degree of association between the single-platform data crawling result and all datasets in the corresponding last layer. Compared with directly calculating the degree of association between the single-platform data crawling result and each dataset at each structural level of each cross-platform, the calculation amount of the method in this embodiment is much less. Therefore, the acquisition efficiency of the degree of association between the single-platform data crawling result and each dataset at each structural level of each cross-platform is higher.
[0146] Embodiment 9:
[0147] On the basis of Embodiment 1, the cross-platform crawling module includes:
[0148] The degree-of-association calculation sub-module is used to determine the degree of association between the single-platform data crawling result and all datasets at all structural levels of all cross-platforms based on the association analysis result;
[0149] The cross-platform crawling sub-module is used to screen out all datasets with a degree of association not less than the degree-of-association threshold from all datasets at all structural levels of all cross-platforms as the cross-platform data crawling result.
[0150] In this embodiment, the degree-of-association threshold is a value that the degree of association between all datasets in the cross-platform data crawling result and the single-platform data crawling result must not be less than.
[0151] The beneficial effect of the above technology is: based on the degree of correlation between the data crawling results of a single platform and all data sets of all structural levels across all platforms, cross-platform data that meets the data crawling requirements input by the user is screened out.
[0152] Embodiment 10:
[0153] The present invention provides a cross-platform data acquisition method for executing the cross-platform data acquisition system described in any one of embodiments 1 to 9, referring to Figure 2 ,include:
[0154] S1: Based on the user's data crawling requirements on a single platform, data is crawled on a single platform to obtain data crawling results on a single platform;
[0155] S2: Structuring the complete data of each cross-platform on a single platform to generate a data structure tree for each cross-platform;
[0156] S3: Based on the data structure tree of each cross-platform, the complete data of each cross-platform is divided by structure to obtain all data sets of each structural level of each cross-platform;
[0157] S4: Perform correlation analysis on the data crawling results of a single platform and all data sets of all structural levels across all platforms to obtain correlation analysis results;
[0158] S5: Based on the correlation analysis results, cross-platform data crawling is performed in the complete data of each cross-platform to obtain the cross-platform data crawling results.
[0159] The beneficial effects of the above technology are: based on the results of combing data structures across multiple platforms, the degree of correlation between the data results obtained by crawling on a single platform and data sets of different data ranges across other platforms is analyzed, and high-precision cross-platform data acquisition is achieved based on the analyzed correlation analysis results, thereby improving the quality and efficiency of cross-platform data acquisition.
[0160] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalents, the present invention is also intended to include these modifications and variations.
Claims
1. A cross-platform data acquisition system, characterized in that: include: A single platform crawling module is used to crawl data on a single platform based on the user's data crawling requirements on a single platform and obtain the single platform data crawling results; The data structure combing module is used to comb the structure of the complete data of each cross-platform of a single platform and generate a data structure tree for each cross-platform; A data structure partitioning module is used to partition the complete data of each cross-platform by structure based on the data structure tree of each cross-platform, and obtain all data sets of each structural level of each cross-platform; The correlation analysis module is used to perform correlation analysis on the data crawling results of a single platform and all data sets of all structural levels across all platforms to obtain correlation analysis results; A cross-platform crawling module is used to crawl cross-platform data in the complete data of each cross-platform based on the correlation analysis results to obtain cross-platform data crawling results; Among them, the data structure combing module includes: A lowest-level determination submodule, used to obtain all lowest-level partition addresses of complete data of each cross-platform of a single platform; The upper address summarization submodule is used to continuously perform upper summarization on each lowest partition address based on the hierarchical mark in each lowest partition address to obtain all upper addresses of each lowest partition address; The data link generation submodule is used to sort and connect all the data stored in each lowest partition address and all the corresponding upper addresses according to the upper degree of the address from large to small, so as to obtain a unidirectional hierarchical data link of each lowest partition address; The data link merging submodule is used to merge the data nodes with repeated addresses in the unidirectional hierarchical data links of all the lowest partition addresses of each cross-platform to obtain the data structure tree of each cross-platform.
2. The cross-platform data acquisition system according to claim 1, characterized in that: Data structure partitioning module, including: A data node determination submodule is used to determine all data nodes in each cross-platform data structure tree; The data set determination submodule is used to aggregate all data stored in the partition address corresponding to each data node as the data set of the corresponding data node, and to aggregate the data sets of all data nodes at each structural level of each cross-platform as all data sets of the corresponding structural level of the cross-platform.
3. The cross-platform data acquisition system according to claim 1, characterized in that: Relevance analysis module, including: The correlation degree calculation submodule is used to calculate the correlation degree between the data crawling results of a single platform and each data set in the last layer of each cross-platform; The correlation degree upper submodule is used to perform upper crawling calculation based on the correlation degree between each cross-platform data structure tree and the single platform data crawling result and all corresponding data sets in the last layer, and obtain the correlation degree between the single platform data crawling result and each data set of each structural level of each cross-platform; The result summary submodule is used to take the correlation degree between the data crawling results of a single platform and all data sets of all structural levels across all platforms as the correlation analysis result.
4. The cross-platform data acquisition system according to claim 3, characterized in that: The correlation degree calculation submodule includes: A form standardization processing unit is used to perform form standardization processing on all data in the data crawling results of a single platform to obtain first form standardized data, and at the same time, perform form standardization processing on each data set in the last layer of each cross-platform to obtain second form standardized data of each data set in the last layer of each cross-platform; A first escape error rate calculation unit is used to take the ratio of the amount of data in each non-standardized form in all data in the single platform data crawling result to the amount of all data in the single platform data crawling result as the proportion of the amount of data in each non-standardized form in the single platform data crawling result, and take the sum of the products of the escape error rates of all non-standardized forms in the single platform data crawling result and the data volume proportion as the standardized escape error rate of the first form standardized data; A second escape error rate calculation unit is used to take the ratio of the amount of data in each non-standardized form in all data in each data set in the last layer of each cross-platform to the amount of data in the corresponding data set as the proportion of the amount of data in each non-standardized form in the corresponding data set, and take the sum of the products of the escape error rates of all non-standardized forms of each data set in the last layer of each cross-platform and the proportion of the data amount as the standardized escape error rate of the corresponding second-form standardized data; A first correlation degree calculation unit is used to calculate the correlation degree between the first form of standardized data and the second form of standardized data of each data set in the last layer of each cross-platform; A second correlation degree calculation unit is used to calculate the correlation degree between the data crawling results of a single platform and each data set in the last layer of each cross-platform based on the correlation degree between the first-form standardized data and the second-form standardized data of each data set in the last layer of each cross-platform and the standardized escape error rate of the first-form standardized data and the standardized escape error rate of the second-form standardized data of each data set in the last layer of each cross-platform.
5. The cross-platform data acquisition system according to claim 4, characterized in that: The first correlation degree calculation unit includes: The data segmentation subunit is used to segment the first form of standardized data to obtain a plurality of first segmented data, and at the same time, segment the second form of standardized data of each data set of the last layer of each cross-platform to obtain a plurality of second segmented data; A relevance analysis subunit, used to analyze the data relevance between each first segmented data and each second segmented data based on the paragraph data relevance analysis model; The correlation degree calculation subunit is used to regard the data correlation degree between each first segmented data in the first form standardized data and all the second segmented data in each data set of each cross-platform last layer as the available data correlation degree of each first segmented data in the first form standardized data, and to regard the average of all available data correlation degrees in the first form standardized data as the correlation degree between the first form standardized data and the second form standardized data of each data set of each cross-platform last layer.
6. The cross-platform data acquisition system according to claim 4, characterized in that: The second correlation degree calculation unit calculates the correlation degree between the single platform data crawling result and each data set of the last layer of each cross-platform based on the correlation degree between the first form standardized data and the second form standardized data of each data set of the last layer of each cross-platform, the standardized escape error rate of the first form standardized data, and the standardized escape error rate of the second form standardized data of each data set of the last layer of each cross-platform, including: In the formula, ε ’ is the correlation degree between the data crawling results of a single platform and the currently calculated data set of the last layer of the cross-platform currently calculated, ε is the correlation degree between the first form standardized data and the second form standardized data of the currently calculated data set of the last layer of the cross-platform currently calculated, σ1 is the standardized escape error rate of the first form standardized data, σ2 is the standardized escape error rate of the second form standardized data of the currently calculated data set of the last layer of the cross-platform currently calculated, and min(ε+σ1-σ2,ε-σ1+σ2) is the minimum value of ε+σ1-σ2 and ε-σ1+σ2.
7. The cross-platform data acquisition system according to claim 3, characterized in that: The submodules of the relevant level include: A node relationship combing unit, used to determine all hierarchical adjacent child nodes of each data node in each cross-platform data structure tree based on each cross-platform data structure tree; A crawling weight determination unit, used to take the ratio between the amount of data of each level adjacent child node of each data node and the sum of the amount of data of all level adjacent child nodes of the corresponding data node as the upper crawling weight of each level adjacent child node of each data node; The upper-level crawling calculation unit is used to perform upper-level crawling calculation based on the upper-level crawling weight of each adjacent child node of each level of each data node in each cross-platform data structure tree and the degree of association between the single platform data crawling results and all data sets of the corresponding last layer, so as to obtain the degree of association between the single platform data crawling results and each data set of each structural level of each cross-platform.
8. The cross-platform data acquisition system according to claim 1, characterized in that: Cross-platform crawling modules, including: The correlation degree calculation submodule is used to determine the correlation degree between the data crawling results of a single platform and all data sets of all structural levels across all platforms based on the correlation degree analysis results; The cross-platform crawling submodule is used to screen out all data sets with a correlation degree not less than a correlation degree threshold from all data sets at all structural levels across all platforms as cross-platform data crawling results.
9. A cross-platform data acquisition method, characterized in that: A cross-platform data acquisition system for executing any one of claims 1 to 8, comprising: S1: Based on the user's data crawling requirements on a single platform, data is crawled on a single platform to obtain data crawling results on a single platform; S2: Structuring the complete data of each cross-platform on a single platform to generate a data structure tree for each cross-platform; S3: Based on the data structure tree of each cross-platform, the complete data of each cross-platform is divided by structure to obtain all data sets of each structural level of each cross-platform; S4: Perform correlation analysis on the data crawling results of a single platform and all data sets of all structural levels across all platforms to obtain correlation analysis results; S5: Based on the correlation analysis results, cross-platform data crawling is performed in the complete data of each cross-platform to obtain the cross-platform data crawling results.
Citation Information
Patent Citations
Cross-platform-based content recommendation method and device, equipment and storage medium
CN113742576A
Cross-platform rendering method and device and electronic equipment
CN116661790A