A multi-dimensional correlation analysis and mining system for Internet data

By positioning the target domain name in the Internet and collecting Internet data sets of multiple data dimensions, analyzing the characteristics of each data dimension, the problem of in-depth mining of Internet data relationships in the existing technology is solved, and more accurate data behavior habit analysis is achieved.

CN120144644BActive Publication Date: 2025-08-12SHENZHEN SEA AREA INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510624734.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-15
Publication Date
2025-08-12
Estimated Expiration
2045-05-15

AI Technical Summary

Technical Problem

In the prior art, Internet data analysis and mining are mainly based on the same dimension, resulting in the inability to mine deep-level correlation information of the data, the inability to effectively understand the correlation relationship between Internet data in different dimensions, and the analysis results are inaccurate.

Method used

By obtaining the mining purpose, positioning the target domain name, collecting Internet data sets in multiple data dimensions, analyzing the first and second data characteristics of each data dimension, performing correlation analysis and mining, and determining data behavior habits.

Benefits of technology

It improves the comprehensiveness and accuracy of data mining and can accurately mine valuable regular characteristics in Internet data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120144644B_ABST
    Figure CN120144644B_ABST
Patent Text Reader

Abstract

The present invention provides a multi-dimensional association analysis and mining system for Internet data, comprising: a positioning module for obtaining a mining purpose and locating a target domain name on the Internet according to the mining purpose; an Internet data acquisition module for obtaining an Internet data set under the target domain name and determining multiple data dimensions in the Internet data set; a data analysis module for analyzing first data features of a sub-Internet data set under each data dimension and analyzing second data features between each data dimension; a result output module for performing association analysis and mining on the first data feature and the second data feature, and outputting data behavior habits corresponding to the mining purpose; thereby improving the comprehensiveness and accuracy of data mining and facilitating accurate mining of valuable regular features in Internet data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing technology, and in particular to a multi-dimensional correlation analysis and mining system for Internet data. Background Art

[0002] At present, with the rapid development of the Internet, various network platforms generate massive amounts of data. Analyzing and mining the data generated by various network platforms can effectively obtain the inherent laws and potential value of the data. Therefore, analyzing and mining massive Internet data is particularly important.

[0003] However, current data processing, analysis, and mining often rely on simple data analysis based on the same dimension, resulting in single results and an inability to mine deep-level correlation information from the data, nor to effectively understand the correlation between Internet data in different dimensions. This results in poor Internet data analysis and inaccurate analysis results.

[0004] Therefore, in order to overcome the above technical problems, the present invention provides a multi-dimensional correlation analysis and mining system for Internet data. Summary of the Invention

[0005] The present invention provides a multi-dimensional association analysis and mining system for Internet data, which is used to obtain a mining purpose and effectively locate a target domain name on the Internet according to the mining purpose, thereby facilitating the accurate extraction of an Internet data set under the target domain name. By collecting multiple data dimensions corresponding to the Internet data set and performing analysis, a first data feature and a second data feature are effectively obtained, and then the data behavior habits corresponding to the mining purpose are accurately determined through association analysis and mining, thereby improving the comprehensiveness and accuracy of data mining and facilitating the accurate mining of valuable regular features in Internet data.

[0006] A multi-dimensional correlation analysis and mining system for Internet data, comprising:

[0007] A positioning module is used to obtain the mining purpose and locate the target domain name on the Internet according to the mining purpose;

[0008] An Internet data acquisition module is used to obtain an Internet data set under a target domain name and determine multiple data dimensions in the Internet data set;

[0009] A data analysis module, configured to analyze the first data features of the sub-Internet data sets under each data dimension, and to analyze the second data features between the data dimensions;

[0010] The result output module is used to perform correlation analysis and mining on the first data feature and the second data feature, and output data behavior habits corresponding to the mining purpose.

[0011] Preferably, a multi-dimensional correlation analysis and mining system for Internet data, including a positioning module, includes:

[0012] a request sending unit, configured to send a data mining request to a server based on a user terminal, and identify the data mining request based on the first data receiving node to obtain core keywords of the data mining request;

[0013] a mining purpose determining unit, configured to determine a mining purpose based on the core keyword and input the mining purpose to the second data receiving node;

[0014] an analyzing unit, configured to analyze the data mining purpose according to the second data receiving node, and determine the data content and mining information type corresponding to the data mining purpose;

[0015] The positioning unit is used to retrieve a domain name set in the Internet in the second data receiving node, and locate the domain name set according to the mined information type and data content to obtain the target domain name.

[0016] Preferably, a multi-dimensional correlation analysis and mining system for Internet data, including a positioning unit, comprises:

[0017] A domain name set reading subunit is used to read the domain name set and determine the information distribution interval corresponding to each domain name in the domain name set on the Internet;

[0018] A first positioning subunit is used to perform a first positioning in the information distribution state according to the mined information type, and obtain a key domain name that meets the mined information type;

[0019] A second positioning subunit is used to extract the information distribution interval of the key domain name and perform a second positioning in the information distribution interval of the key domain name according to the data content to obtain the key information distribution interval;

[0020] The target domain name extraction subunit is used to obtain the subsidiary domain name corresponding to the key information distribution interval and extract the subsidiary domain name as the target domain name.

[0021] Preferably, a multi-dimensional correlation analysis and mining system for Internet data, an Internet data acquisition module, includes:

[0022] A data extraction unit, configured to retrieve an Internet dataset from the Internet based on a target domain name;

[0023] The data parsing unit is used to parse the Internet data set to obtain the data dimensions in the Internet data set.

[0024] Preferably, a multi-dimensional correlation analysis and mining system for Internet data, wherein the data extraction unit comprises:

[0025] The page information determination subunit is used to locate the corresponding target page on the Internet according to the target domain name, read the target page, and determine the page information in the target page;

[0026] The valid information determination subunit is used to determine valid information in the page information according to the mining purpose;

[0027] An information data determination subunit is used to extract valid information and determine the information data of the valid information;

[0028] The preprocessing subunit is used to preprocess the information data of valid information to obtain an Internet data set.

[0029] Preferably, a multi-dimensional correlation analysis and mining system for Internet data, a data parsing unit, includes:

[0030] The reading subunit is used to read the Internet data set, determine the data content of the Internet data set, and obtain the core points of the data content;

[0031] The data processing sub-unit is used to obtain the similarity between each content core point, and if the similarity between the target content core points is equal to or greater than a preset similarity threshold, the content core points are merged;

[0032] The data dimension determination subunit is used to use the content core point as the data dimension in the Internet data set according to the merging result.

[0033] Preferably, a multi-dimensional correlation analysis and mining system for Internet data, the data analysis module includes:

[0034] An Internet data splitting unit is used to split the Internet data set based on the data dimension to obtain a sub-Internet data set under each data dimension;

[0035] Internet data analysis unit for:

[0036] Preprocessing the sub-Internet dataset under each data dimension, classifying the sub-Internet dataset under each data dimension based on the preprocessing results, and extracting category features of each category;

[0037] Generate a category identifier for each category based on the category features, and annotate the corresponding classification results based on the category identifier;

[0038] A first data feature determination unit is configured to:

[0039] Based on the annotation results, the text content set and multimedia data set included in the sub-Internet data set under each data dimension are obtained;

[0040] Segment the text content set, and determine the part of speech and frequency of occurrence of different word elements based on the segmentation results, and use the part of speech and frequency of occurrence as word features;

[0041] quantifying the value range of the text content set based on the word segmentation result, and determining the numerical features of the text content set based on the quantification result;

[0042] Input the multimedia dataset included in the sub-Internet dataset under each dimension into the pre-trained convolutional neural network for processing, determine the image composition, color ratio composition and texture distribution corresponding to the multimedia dataset, and use the image composition, color ratio composition and texture distribution as image features;

[0043] The word features, numerical features, and image features under each data dimension are aggregated to obtain the first data features of the sub-Internet dataset under each data dimension.

[0044] Preferably, in a multi-dimensional correlation analysis and mining system for Internet data, in a data analysis module, analyzing the second data features between the data dimensions includes:

[0045] A quantitative analysis unit is used to perform quantitative analysis on the sub-Internet dataset under any data dimension and observe the value change characteristics of the sub-Internet dataset under other data dimensions;

[0046] The second data feature determination unit is used to determine the relative value change trend of the sub-Internet data set under different data dimensions based on the value change characteristics, and determine the correlation relationship between different data dimensions based on the relative change trend, and determine the second data feature between each data dimension based on the correlation relationship.

[0047] Preferably, a multi-dimensional correlation analysis and mining system for Internet data, with a result output module, includes:

[0048] The result acquisition unit is used to:

[0049] Acquire the first data feature and the second data feature, and at the same time, acquire the mining purpose, and determine a business branch related to the mining purpose from the business knowledge system according to the mining purpose;

[0050] Split the mining purpose into specific business analysis goals based on the functional attributes of the business branch;

[0051] Association analysis and mining unit, used for

[0052] Determine the topological structure between data dimensions under different business analysis objectives based on the business knowledge system and the second data feature, and divide the data dimensions with associated nodes into groups based on the topological structure;

[0053] Based on the business analysis goal, the first data feature, and the second data feature, correlation analysis and mining analysis are performed on the sub-Internet data sets to obtain target analysis results corresponding to the business analysis goal under the data dimension within each group;

[0054] The data behavior habit determination unit is used to summarize the target analysis results corresponding to different groups based on the business analysis goals, and obtain the data behavior habits corresponding to the mining purpose based on the summary results.

[0055] Preferably, a multi-dimensional correlation analysis and mining system for Internet data, wherein the data behavior habit determination unit comprises:

[0056] The data acquisition sub-unit is used to acquire data behavior habits corresponding to different mining purposes, and to bundle the mining purposes with the corresponding data behavior habits based on the acquisition results;

[0057] Record subunit, used to:

[0058] Construct a data behavior record table and divide the data behavior record table into table area;

[0059] Based on the results of the table item area division and the bundling results, the mining purpose and the corresponding data behavior habits are filled in the corresponding table item areas in the data behavior record table, and the mining purpose and the corresponding data behavior habits are recorded and archived based on the filling results.

[0060] Compared with the prior art, the present invention has the following beneficial effects:

[0061] By obtaining the mining purpose and effectively locating the target domain name on the Internet based on the mining purpose, it is helpful to accurately extract the Internet data set under the target domain name. By collecting multiple data dimensions corresponding to the Internet data set and analyzing them, the first data feature and the second data feature are effectively obtained. Then, the data behavior habits corresponding to the mining purpose are accurately determined through association analysis and mining, thereby improving the comprehensiveness and accuracy of data mining and facilitating the accurate mining of valuable regular features in Internet data.

[0062] Other features and advantages of the present invention will be described in the following description, and in part will become apparent from the description, or will be understood by practicing the present invention. The purpose and other advantages of the present invention can be achieved and obtained through the structures specifically pointed out in this application document.

[0063] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0064] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings:

[0065] Figure 1 This is a structural diagram of a multi-dimensional correlation analysis and mining system for Internet data in an embodiment of the present invention;

[0066] Figure 2 This is a structural diagram of a positioning module in a multi-dimensional correlation analysis and mining system for Internet data in an embodiment of the present invention;

[0067] Figure 3 This is a structural diagram of a positioning unit in a multi-dimensional association analysis and mining system for Internet data in an embodiment of the present invention. DETAILED DESCRIPTION

[0068] The preferred embodiments of the present invention are described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present invention, and are not used to limit the present invention.

[0069] Example 1

[0070] This embodiment provides a multi-dimensional correlation analysis and mining system for Internet data, including:

[0071] A positioning module is used to obtain the mining purpose and locate the target domain name on the Internet according to the mining purpose;

[0072] An Internet data acquisition module is used to obtain an Internet data set under a target domain name and determine multiple data dimensions in the Internet data set;

[0073] A data analysis module, configured to analyze the first data features of the sub-Internet data sets under each data dimension, and to analyze the second data features between the data dimensions;

[0074] The result output module is used to perform correlation analysis and mining on the first data feature and the second data feature, and output data behavior habits corresponding to the mining purpose.

[0075] In this embodiment, the purpose of mining is to obtain the results required based on user needs. The results are obtained by performing multi-dimensional correlation analysis and mining on Internet data, and are used to characterize the laws in the entity operation process of Internet data records (the laws are data behavior habits), such as change trends.

[0076] In this embodiment, the target domain name is determined according to the mining purpose and is a source identifier for obtaining Internet data from the Internet.

[0077] In this embodiment, the data dimension is determined based on the Internet dataset and is used to represent different categories of attributes presented by the Internet dataset, wherein one data dimension corresponds to one sub-Internet dataset.

[0078] In this embodiment, the first data feature refers to a data transformation trend or characteristic of the sub-Internet dataset, for example, it may be a value transformation situation of the sub-Internet dataset.

[0079] In this embodiment, the second data feature refers to the interaction relationship between various data dimensions.

[0080] The working principle and beneficial effects of the above technical solution are: by obtaining the mining purpose and effectively locating the target domain name on the Internet according to the mining purpose, it is conducive to accurately extracting the Internet data set under the target domain name, and by collecting multiple data dimensions corresponding to the Internet data set and analyzing them, the first data feature and the second data feature are effectively obtained, and then the data behavior habits corresponding to the mining purpose are accurately determined through association analysis and mining, thereby improving the comprehensiveness and accuracy of data mining and facilitating the accurate mining of valuable regular features in Internet data.

[0081] Example 2

[0082] Based on Example 1, this example provides a multi-dimensional correlation analysis and mining system for Internet data, such as Figure 2 As shown, the positioning module includes:

[0083] a request sending unit, configured to send a data mining request to a server based on a user terminal, and identify the data mining request based on the first data receiving node to obtain core keywords of the data mining request;

[0084] a mining purpose determining unit, configured to determine a mining purpose based on the core keyword and input the mining purpose to the second data receiving node;

[0085] an analyzing unit, configured to analyze the data mining purpose according to the second data receiving node, and determine the data content and mining information type corresponding to the data mining purpose;

[0086] The positioning unit is used to retrieve a domain name set in the Internet in the second data receiving node, and locate the domain name set according to the mined information type and data content to obtain the target domain name.

[0087] In this embodiment, the output end of the user terminal is connected to the input end of the server, wherein the server includes: a first data receiving node and a second data receiving node, the input end of the first data receiving node is connected, and the output end of the first data receiving node is connected to the input end of the second data receiving node.

[0088] In this embodiment, the first data receiving node is used to identify the data mining request, thereby effectively obtaining the core keywords of the data mining request, wherein the core keywords are data segments of the object purpose etc. to be mined in the data mining request.

[0089] In this embodiment, the mining purpose is the final mining result determined based on the core keywords.

[0090] In this embodiment, the second data receiving node pre-stores a set of domain names on the Internet, and the second data receiving node effectively determines the data content and the type of mined information according to the mining purpose.

[0091] The working principle and beneficial effects of the above technical solution are: the first data receiving node is used to effectively analyze the data mining request, thereby effectively obtaining the core keywords. By analyzing the core keywords, the data mining purpose can be effectively obtained, and then the target domain name is determined based on the mining purpose through the second data analysis node, thereby improving the accuracy and effectiveness of obtaining the target domain name.

[0092] Example 3

[0093] Based on Example 2, this example provides a multi-dimensional correlation analysis and mining system for Internet data, such as Figure 3 As shown, the positioning unit includes:

[0094] A domain name set reading subunit is used to read the domain name set and determine the information distribution interval corresponding to each domain name in the domain name set on the Internet;

[0095] A first positioning subunit is used to perform a first positioning in the information distribution state according to the mined information type, and obtain a key domain name that meets the mined information type;

[0096] A second positioning subunit is used to extract the information distribution interval of the key domain name and perform a second positioning in the information distribution interval of the key domain name according to the data content to obtain the key information distribution interval;

[0097] The target domain name extraction subunit is used to obtain the subsidiary domain name corresponding to the key information distribution interval and extract the subsidiary domain name as the target domain name.

[0098] In this embodiment, the information distribution interval refers to the information type interval and information data range interval of the web page corresponding to each domain name in the domain name set.

[0099] In this embodiment, the subsidiary domain name refers to a subordinate domain name under the key domain name, and is a domain name corresponding to the key information distribution interval.

[0100] The working principle and beneficial effect of the above technical solution are: by determining the information distribution range of each domain name in the domain name set on the Internet, it is possible to effectively locate and extract the target domain name under the key domain name, effectively ensuring the accuracy of the target domain name extraction.

[0101] Example 4

[0102] Based on Example 1, this embodiment provides a multi-dimensional correlation analysis and mining system for Internet data, and an Internet data acquisition module, including:

[0103] A data extraction unit, configured to retrieve an Internet dataset from the Internet based on a target domain name;

[0104] The data parsing unit is used to parse the Internet data set to obtain the data dimensions in the Internet data set.

[0105] The working principle and beneficial effects of the above technical solution are: by retrieving the corresponding Internet data set from the Internet according to the target domain name and parsing the Internet data set, the data dimensions in the Internet data set can be determined, providing reliable data support for multi-dimensional correlation analysis and mining of Internet data.

[0106] Example 5

[0107] Based on Example 4, this embodiment provides a multi-dimensional correlation analysis and mining system for Internet data, wherein the data extraction unit includes:

[0108] The page information determination subunit is used to locate the corresponding target page on the Internet according to the target domain name, read the target page, and determine the page information in the target page;

[0109] The valid information determination subunit is used to determine valid information in the page information according to the mining purpose;

[0110] An information data determination subunit is used to extract valid information and determine the information data of the valid information;

[0111] The preprocessing subunit is used to preprocess the information data of valid information to obtain an Internet data set.

[0112] In this embodiment, the target page refers to a web page determined from the Internet according to the target domain name, that is, the main body of the Internet data set.

[0113] In this embodiment, the valid information refers to data related to the mining purpose obtained by screening the page information in the target page according to the mining purpose.

[0114] In this embodiment, preprocessing refers to operations such as cleaning and deduplication of information data of valid information.

[0115] The working principle and beneficial effects of the above technical solution are: by locating the corresponding target page on the Internet according to the target domain name, and screening the page information of the located target page, the information data of the valid information is determined, and finally, the information data of the obtained valid information is preprocessed to achieve effective acquisition of the Internet data set, providing reliable data support for multi-dimensional correlation analysis and mining of Internet data.

[0116] Example 6

[0117] Based on Example 4, this embodiment provides a multi-dimensional correlation analysis and mining system for Internet data, wherein the data parsing unit includes:

[0118] The reading subunit is used to read the Internet data set, determine the data content of the Internet data set, and obtain the core points of the data content;

[0119] The data processing sub-unit is used to obtain the similarity between each content core point, and if the similarity between the target content core points is equal to or greater than a preset similarity threshold, the content core points are merged;

[0120] The data dimension determination subunit is used to use the content core point as the data dimension in the Internet data set according to the merging result.

[0121] In this embodiment, the core point of the content refers to a data parameter that can represent the key core or main idea of the data content.

[0122] In this embodiment, the preset similarity threshold is set in advance and is the minimum standard for measuring whether to merge the core points of the content, and can be adjusted.

[0123] The working principle and beneficial effects of the above technical solution are: by determining the data content of the obtained Internet data set and determining the content core points of the data content, secondly, determining the similarity between different content core points, and merging the content core points whose similarity is equal to or greater than the preset similarity threshold, and finally determining the data dimension in the Internet data set based on the merging result, which provides convenience for multi-dimensional correlation analysis and mining of Internet data.

[0124] Example 7

[0125] Based on Example 1, this embodiment provides a multi-dimensional correlation analysis and mining system for Internet data, wherein the data analysis module includes:

[0126] An Internet data splitting unit is used to split the Internet data set based on the data dimension to obtain a sub-Internet data set under each data dimension;

[0127] Internet data analysis unit for:

[0128] Preprocessing the sub-Internet dataset under each data dimension, classifying the sub-Internet dataset under each data dimension based on the preprocessing results, and extracting category features of each category;

[0129] Generate a category identifier for each category based on the category features, and annotate the corresponding classification results based on the category identifier;

[0130] The first data feature determination unit is configured to:

[0131] Based on the annotation results, the text content set and multimedia data set included in the sub-Internet data set under each data dimension are obtained;

[0132] Segment the text content set, and determine the part of speech and frequency of occurrence of different word elements based on the segmentation results, and use the part of speech and frequency of occurrence as word features;

[0133] quantifying the value range of the text content set based on the word segmentation result, and determining the numerical features of the text content set based on the quantification result;

[0134] Input the multimedia dataset included in the sub-Internet dataset under each dimension into the pre-trained convolutional neural network for processing, determine the image composition, color ratio composition and texture distribution corresponding to the multimedia dataset, and use the image composition, color ratio composition and texture distribution as image features;

[0135] The word features, numerical features, and image features under each data dimension are aggregated to obtain the first data features of the sub-Internet dataset under each data dimension.

[0136] In this embodiment, dimension splitting refers to splitting the Internet dataset into data sets corresponding to each dimension according to the data dimension, and the result obtained by the splitting is the sub-Internet dataset.

[0137] In this embodiment, preprocessing refers to operations such as cleaning and data supplementation of the sub-Internet dataset.

[0138] In this embodiment, the category feature refers to the characteristics of each category of data obtained after classifying the sub-Internet data set under each data dimension, for example, it can be the data type and data value characteristics in each category of data.

[0139] In this embodiment, the category identifier refers to a symbol generated according to the category characteristics and used to mark the sub-Internet dataset of each category.

[0140] In this embodiment, the text content refers to the text content contained in the sub-Internet data set.

[0141] In this embodiment, the multimedia data set refers to the pictures and audio data contained in the sub-Internet data set.

[0142] In this embodiment, word segmentation refers to splitting the words contained in the text content.

[0143] In this embodiment, word elements refer to the characters contained in each word obtained after word segmentation of the text content.

[0144] In this embodiment, word features refer to parameter information that can represent the part of speech and frequency of occurrence of each word obtained after word segmentation.

[0145] In this embodiment, the numerical feature refers to the value range and specific value distribution corresponding to the text content set.

[0146] In this embodiment, the pre-trained convolutional neural network is a tool that is trained in advance and is used to process and analyze multimedia data sets, and is used to extract features such as the color composition of images.

[0147] In this embodiment, the image feature refers to a data parameter obtained after analyzing the multimedia data set and capable of representing the core of the multimedia data set.

[0148] The working principle and beneficial effects of the above technical solution are: by dimensionally splitting the Internet data set according to the data dimension, the sub-Internet data set under each data dimension can be accurately and effectively determined, and the obtained sub-Internet data set is classified after preprocessing, and then the text content set and multimedia data set included in the sub-Internet data set under each data dimension are determined according to the classification results. Secondly, the obtained text content set and multimedia data set are processed separately to determine the word features and data features corresponding to the text content set. At the same time, the image features of the multimedia data set are also determined. Finally, the word features, numerical features and image features under each data dimension are summarized to achieve accurate and effective determination of the first data features of the sub-Internet data set under each data dimension, which provides convenience and guarantee for multi-dimensional correlation analysis and mining of Internet data.

[0149] Example 8

[0150] Based on Example 1, this embodiment provides a multi-dimensional correlation analysis and mining system for Internet data. In the data analysis module, the second data feature between each data dimension is analyzed, including:

[0151] A quantitative analysis unit is used to perform quantitative analysis on the sub-Internet dataset under any data dimension and observe the value change characteristics of the sub-Internet dataset under other data dimensions;

[0152] The second data feature determination unit is used to determine the relative value change trend of the sub-Internet data set under different data dimensions based on the value change characteristics, and determine the correlation relationship between different data dimensions based on the relative change trend, and determine the second data feature between each data dimension based on the correlation relationship.

[0153] In this embodiment, the quantitative analysis refers to adjusting the values of the sub-Internet data sets under one of the data dimensions in sequence, while the values of the sub-Internet data sets under the remaining data dimensions remain unchanged.

[0154] In this embodiment, the value change feature refers to the situation where, when the value of a sub-Internet dataset in one data dimension changes, the values of the sub-Internet datasets in other data dimensions also change accordingly, as well as the changing trend of their values when the changes occur.

[0155] In this embodiment, the relative value change trend refers to the relative change of values between different sub-Internet data sets.

[0156] The working principle and beneficial effects of the above technical solution are: by performing quantitative analysis on the sub-Internet data set under any data dimension, the relative value change trend of the sub-Internet data set under different data dimensions is determined according to the analysis results, and then the correlation relationship between different data dimensions is locked according to the relative value change trend. Finally, the correlation relationship is realized to determine the second data feature between each data dimension, which provides a reference basis for multi-dimensional correlation analysis and mining of Internet data, and also ensures the reliability of multi-dimensional correlation analysis and mining of Internet data.

[0157] Example 9

[0158] Based on Example 1, this embodiment provides a multi-dimensional correlation analysis and mining system for Internet data, and the result output module includes:

[0159] The result acquisition unit is used to:

[0160] Acquire the first data feature and the second data feature, and at the same time, acquire the mining purpose, and determine a business branch related to the mining purpose from the business knowledge system according to the mining purpose;

[0161] Split the mining purpose into specific business analysis goals based on the functional attributes of the business branch;

[0162] Association analysis and mining unit, used for:

[0163] Determine the topological structure between data dimensions under different business analysis objectives based on the business knowledge system and the second data feature, and divide the data dimensions with associated nodes into groups based on the topological structure;

[0164] Extracting sub-internet datasets under the data dimension within each group based on the group division results, and determining the mutual restriction relationships between different sub-internet datasets under different feature categories based on the first data feature of the sub-internet datasets;

[0165] The sub-Internet datasets within each group are divided into sets based on the mutual restriction relationships under different feature categories, and the sub-Internet datasets under each feature category are aligned based on the set division results;

[0166] Based on the data alignment results, autocorrelation analysis is performed on the sub-Internet datasets under each feature category according to the business analysis objectives to obtain the analysis results under each feature category. Based on the autocorrelation analysis results, cross-correlation analysis is performed on the analysis results between the sub-Internet datasets of different feature categories within each group to obtain the target analysis results corresponding to the business analysis objectives under the data dimensions within each group;

[0167] The data behavior habit determination unit is used to summarize the target analysis results corresponding to different groups based on the business analysis goals, and obtain the data behavior habits corresponding to the mining purpose based on the summary results.

[0168] In this embodiment, the business knowledge system is prepared in advance and is used to record the association relationship between different business branches and the structural relationship between different business branches.

[0169] In this embodiment, the business branch refers to business information determined from the business knowledge system according to the mining purpose.

[0170] In this embodiment, the functional attribute refers to the business purpose or business result that can be achieved by the business branch.

[0171] In this embodiment, the business analysis target refers to a specific business analysis object that can represent the mining target and is obtained after splitting the mining purpose, that is, each small target included in the mining purpose.

[0172] In this embodiment, the associated nodes refer to nodes with an associated relationship determined according to the topological structure between data dimensions.

[0173] In this embodiment, the mutual restriction relationship refers to the interaction relationship between different categories of sub-Internet data.

[0174] In this embodiment, data alignment refers to aligning the starting points of the sub-Internet data sets after the set is divided, so as to facilitate autocorrelation analysis of the sub-Internet data in each set.

[0175] In this embodiment, the autocorrelation analysis refers to analyzing the correlation between sub-Internet data within the same feature category in each set.

[0176] In this embodiment, the cross-correlation analysis refers to performing correlation analysis between categories on the analysis results of the sub-Internet data sets of different feature categories within each group.

[0177] The working principle and beneficial effects of the above technical solution are: by determining the business branches related to the mining purpose from the business knowledge system according to the mining purpose, and splitting the mining purpose according to the functional attributes of the business analysis, the business analysis goal is determined; secondly, the sub-Internet data is analyzed according to the obtained business analysis goal and the first data feature and the second data feature, so as to accurately and reliably determine the target analysis results corresponding to the business analysis goal under the data dimension in each group; finally, the target analysis results corresponding to different groups are summarized according to the business analysis goal, and finally the data behavior habits corresponding to the mining purpose are locked, ensuring the accuracy and reliability of the data behavior habits finally obtained.

[0178] Example 10

[0179] Based on Example 9, this embodiment provides a multi-dimensional association analysis and mining system for Internet data, wherein the data behavior habit determination unit includes:

[0180] The data acquisition sub-unit is used to acquire data behavior habits corresponding to different mining purposes, and to bundle the mining purposes with the corresponding data behavior habits based on the acquisition results;

[0181] Record subunit, used to:

[0182] Construct a data behavior record table and divide the data behavior record table into table area;

[0183] Based on the results of the table item area division and the bundling results, the mining purpose and the corresponding data behavior habits are filled in the corresponding table item areas in the data behavior record table, and the mining purpose and the corresponding data behavior habits are recorded and archived based on the filling results.

[0184] In this embodiment, the data behavior record table is a table used to record data behavior habits corresponding to different mining purposes.

[0185] In this embodiment, the table item area division refers to dividing the data behavior record table into areas according to the categories of data items to be recorded in the record table.

[0186] The working principle and beneficial effects of the above technical solution are: by bundling the obtained mining purpose with the corresponding data behavior habits, at the same time, constructing a data behavior record table, and recording the mining purpose and data behavior habits in the data behavior record table according to the bundling results, the obtained data behavior habits can be effectively recorded and archived, which is also convenient for tracing the analysis results.

[0187] Obviously, those skilled in the art may make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if such changes and modifications fall within the scope of the claims and their equivalents, the present invention is intended to include such changes and modifications.

Claims

1. A multi-dimensional correlation analysis and mining system for Internet data, characterized in that: include: A positioning module is used to obtain the mining purpose and locate the target domain name on the Internet according to the mining purpose; An Internet data acquisition module is used to obtain an Internet data set under a target domain name and determine multiple data dimensions in the Internet data set; A data analysis module, configured to analyze the first data features of the sub-Internet data sets under each data dimension, and to analyze the second data features between the data dimensions; A result output module is used to perform correlation analysis and mining on the first data feature and the second data feature, and output data behavior habits corresponding to the mining purpose; Positioning module, including: a request sending unit, configured to send a data mining request to a server based on a user terminal, and identify the data mining request based on the first data receiving node to obtain core keywords of the data mining request; a mining purpose determining unit, configured to determine a mining purpose based on the core keyword and input the mining purpose to the second data receiving node; an analyzing unit, configured to analyze the data mining purpose according to the second data receiving node, and determine the data content and mining information type corresponding to the data mining purpose; A positioning unit is configured to retrieve a domain name set on the Internet in the second data receiving node, and locate the target domain name in the domain name set according to the mined information type and data content; Positioning unit, comprising: A domain name set reading subunit is used to read the domain name set and determine the information distribution interval corresponding to each domain name in the domain name set on the Internet; A first positioning subunit is used to perform a first positioning in the information distribution state according to the mined information type, and obtain a key domain name that meets the mined information type; A second positioning subunit is used to extract the information distribution interval of the key domain name and perform a second positioning in the information distribution interval of the key domain name according to the data content to obtain the key information distribution interval; The target domain name extraction subunit is used to obtain the subsidiary domain name corresponding to the key information distribution interval and extract the subsidiary domain name as the target domain name; Result output module, including: The result acquisition unit is used to: Acquire the first data feature and the second data feature, and at the same time, acquire the mining purpose, and determine a business branch related to the mining purpose from the business knowledge system according to the mining purpose; Split the mining purpose into specific business analysis goals based on the functional attributes of the business branch; Association analysis and mining unit, used for Based on the business knowledge system and the second data characteristics, the topological structure between the data dimensions under different business analysis objectives is determined, and the related The data dimension of the join node is divided into groups; Based on the business analysis goal, the first data feature, and the second data feature, correlation analysis and mining analysis are performed on the sub-Internet data sets to obtain target analysis results corresponding to the business analysis goals under the data dimensions within each group; The data behavior habit determination unit is used to summarize the target analysis results corresponding to different groups based on the business analysis goals, and obtain the data behavior habits corresponding to the mining purpose based on the summary results.

2. A multi-dimensional correlation analysis and mining system for Internet data according to claim 1, characterized in that: Internet data acquisition module, including: A data extraction unit, configured to retrieve an Internet dataset from the Internet based on a target domain name; The data parsing unit is used to parse the Internet data set to obtain the data dimensions in the Internet data set.

3. A multi-dimensional correlation analysis and mining system for Internet data according to claim 2, characterized in that: Data extraction unit, including: The page information determination subunit is used to locate the corresponding target page on the Internet according to the target domain name, read the target page, and determine the page information in the target page; The valid information determination subunit is used to determine valid information in the page information according to the mining purpose; An information data determination subunit is used to extract valid information and determine the information data of the valid information; The preprocessing subunit is used to preprocess the information data of valid information to obtain an Internet data set.

4. A multi-dimensional correlation analysis and mining system for Internet data according to claim 2, characterized in that: Data parsing unit, including: The reading subunit is used to read the Internet data set, determine the data content of the Internet data set, and obtain the core points of the data content; The data processing sub-unit is used to obtain the similarity between each content core point, and if the similarity between the target content core points is equal to or greater than a preset similarity threshold, the content core points are merged; The data dimension determination subunit is used to use the content core point as the data dimension in the Internet data set according to the merging result.

5. The multi-dimensional correlation analysis and mining system for Internet data according to claim 1, characterized in that: Data analysis module, including: An Internet data splitting unit is used to split the Internet data set based on the data dimension to obtain a sub-Internet data set under each data dimension; Internet data analysis unit for: Preprocessing the sub-Internet dataset under each data dimension, classifying the sub-Internet dataset under each data dimension based on the preprocessing results, and extracting category features of each category; Generate a category identifier for each category based on the category features, and annotate the corresponding classification results based on the category identifier; A first data feature determination unit is configured to: Based on the annotation results, the text content set and multimedia data set included in the sub-Internet data set under each data dimension are obtained; Segment the text content set, and determine the part of speech and frequency of occurrence of different word elements based on the segmentation results, and use the part of speech and frequency of occurrence as word features; quantifying the value range of the text content set based on the word segmentation result, and determining the numerical features of the text content set based on the quantification result; Input the multimedia dataset included in the sub-Internet dataset under each dimension into the pre-trained convolutional neural network for processing, determine the image composition, color ratio composition and texture distribution corresponding to the multimedia dataset, and use the image composition, color ratio composition and texture distribution as image features; The word features, numerical features, and image features under each data dimension are aggregated to obtain the first data features of the sub-Internet dataset under each data dimension.

6. The multi-dimensional correlation analysis and mining system for Internet data according to claim 1, characterized in that: In the data analysis module, the second data features between the data dimensions are analyzed, including: A quantitative analysis unit is used to perform quantitative analysis on the sub-Internet dataset under any data dimension and observe the value change characteristics of the sub-Internet dataset under other data dimensions; The second data feature determination unit is used to determine the relative value change trend of the sub-Internet data set under different data dimensions based on the value change characteristics, and determine the correlation relationship between different data dimensions based on the relative change trend, and determine the second data feature between each data dimension based on the correlation relationship.

7. The multi-dimensional correlation analysis and mining system for Internet data according to claim 1, characterized in that: Data behavior habit determination unit, including: The data acquisition sub-unit is used to acquire data behavior habits corresponding to different mining purposes, and to bundle the mining purposes with the corresponding data behavior habits based on the acquisition results; Record subunit, used to: Construct a data behavior record table and divide the data behavior record table into table area; Based on the results of the table item area division and the bundling results, the mining purpose and the corresponding data behavior habits are filled in the corresponding table item areas in the data behavior record table, and the mining purpose and the corresponding data behavior habits are recorded and archived based on the filling results.

Citation Information

Patent Citations

  • Data processing method and device, electronic equipment and storage medium

    CN110704751A

  • Universal internet information data mining method

    US20170046434A1