Multi-dimensional correlation analysis and mining system for Internet data
Through multi-dimensional correlation analysis and mining systems, multiple dimensional characteristics of Internet data are acquired and analyzed, and the problem of inability to effectively mine deep-level data correlation information in the existing technology is solved, and a more comprehensive and accurate data mining effect is achieved.
Patent Information
- Application Number
- CN202510624734.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-15
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-05-15
AI Technical Summary
The existing technology mainly relies on simple data analysis of the same dimension in Internet data analysis and mining, and cannot effectively mine the deep-level correlation information of the data and the correlation relationship between different dimensions, resulting in poor analysis results and inaccurate results.
It provides a multi-dimensional correlation analysis and mining system for Internet data. By obtaining mining purposes, positioning the target domain name, collecting multiple data dimensions, analyzing the first and second data characteristics, and determining data behavior habits through correlation analysis and mining, improving the comprehensiveness and accuracy of data mining.
It realizes the accurate mining of valuable regular characteristics in Internet data, improves the comprehensiveness and accuracy of data mining, and can effectively understand the correlation between Internet data in different dimensions.
Smart Images

Figure CN120144644A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and particularly relates to a multi-dimensional correlation analysis and mining system for Internet data. Background Art
[0002] At present, with the rapid development of the Internet, various network platforms generate a large amount of data. Analyzing and mining the data generated by various network platforms can effectively obtain the internal laws and potential values of the data. Therefore, it is particularly important to analyze and mine the massive Internet data; However, current data processing, analysis, and mining often perform simple data analysis based on the same dimension, resulting in a single result, being unable to mine the deep correlation information of the data, and also being unable to effectively understand the correlation relationship between Internet data under different dimensions, thus leading to poor analysis effects and inaccurate analysis results for Internet data; Therefore, in order to overcome the above technical problems, the present invention provides a multi-dimensional correlation analysis and mining system for Internet data. Summary of the Invention
[0003] The present invention provides a multi-dimensional correlation analysis and mining system for Internet data, which is used to obtain a mining purpose, and effectively locate a target domain name in the Internet according to the mining purpose, thereby facilitating accurately extracting the Internet data set under the target domain name, collecting multiple data dimensions corresponding to the Internet data set, and analyzing to effectively obtain a first data feature and a second data feature, and then accurately determining the data behavior habit corresponding to the mining purpose through correlation analysis and mining, improving the comprehensiveness and accuracy of data mining, and being conducive to accurately mining valuable regular features in Internet data.
[0004] A multi-dimensional correlation analysis and mining system for Internet data, comprising: A positioning module, configured to obtain a mining purpose and locate a target domain name in the Internet according to the mining purpose; An Internet data acquisition module, configured to obtain the Internet data set under the target domain name and determine multiple data dimensions in the Internet data set; A data analysis module, configured to analyze the first data feature of the sub-Internet data set under each data dimension and analyze the second data feature between each data dimension; A result output module, configured to perform correlation analysis and mining on the first data feature and the second data feature and output the data behavior habit corresponding to the mining purpose.
[0005] Preferably, a multi-dimensional correlation analysis and mining system for Internet data, the positioning module, comprising: A request sending unit, configured to send a data mining request to a server based on a user terminal, and identify the data mining request based on a first data receiving node to obtain a core keyword of the data mining request; A mining purpose determining unit, configured to determine a mining purpose according to the core keyword and input the mining purpose into a second data receiving node; An analysis unit, configured to analyze the data mining purpose according to the second data receiving node to determine data content corresponding to the data mining purpose and a mining information type; A positioning unit, configured to retrieve a domain name set in the Internet in the second data receiving node, and perform positioning in the domain name set according to the mining information type and the data content to obtain a target domain name.
[0006] Preferably, a multi-dimensional association analysis and mining system for Internet data, the positioning unit includes: A domain name set reading subunit, configured to read the domain name set to determine an information distribution interval corresponding to each domain name in the Internet in the domain name set; A first positioning subunit, configured to perform first positioning in the information distribution state according to the mining information type to obtain a key domain name that conforms to the mining information type; A second positioning subunit, configured to extract the information distribution interval of the key domain name, and perform second positioning in the information distribution interval of the key domain name according to the data content to obtain a key information distribution interval; A target domain name extraction subunit, configured to obtain an affiliated domain name corresponding to the key information distribution interval and extract the affiliated domain name as the target domain name.
[0007] Preferably, a multi-dimensional association analysis and mining system for Internet data, the Internet data acquisition module includes: A data extraction unit, configured to retrieve an Internet data set in the Internet based on the target domain name; A data parsing unit, configured to parse the Internet data set to obtain data dimensions in the Internet data set.
[0008] Preferably, a multi-dimensional association analysis and mining system for Internet data, the data extraction unit includes: A page information determining subunit, configured to locate a corresponding target page in the Internet according to the target domain name, read the target page, and determine page information in the target page; A valid information determining subunit, configured to determine valid information in the page information according to the mining purpose; An information data determining subunit, configured to extract the valid information and determine information data of the valid information; A preprocessing subunit, configured to preprocess the information data of the valid information to obtain an Internet data set.
[0009] Preferably, a multi-dimensional correlation analysis and mining system for Internet data, the data parsing unit includes: A reading subunit for reading an Internet data set, determining the data content of the Internet data set, and at the same time, obtaining the content core points of the data content; A data processing subunit for obtaining the similarity between each content core point. If the similarity between the target content core points is equal to or greater than the preset similarity threshold, the content core points are merged; A data dimension determination subunit for using the content core points as the data dimensions in the Internet data set according to the merging result.
[0010] Preferably, a multi-dimensional correlation analysis and mining system for Internet data, the data analysis module includes: An Internet data splitting unit for splitting the Internet data set based on the data dimensions to obtain sub-Internet data sets under each data dimension; An Internet data analysis unit for: Preprocessing the sub-Internet data sets under each data dimension, classifying the sub-Internet data sets under each data dimension based on the preprocessing results, and extracting the category features of each category; Generating category identifiers for each category based on the category features, and labeling the corresponding classification results based on the category identifiers; A first data feature determination unit for: Obtaining the text content set and the multimedia data set included in the sub-Internet data sets under each data dimension based on the labeling results; Performing word segmentation on the text content set, determining the part-of-speech and occurrence frequency of different word elements based on the word segmentation results, and using the part-of-speech and occurrence frequency as word features; Quantifying the value range of the text content set based on the word segmentation results, and determining the numerical features of the text content set based on the quantification results; Inputting the multimedia data sets included in the sub-Internet data sets under each dimension into a pre-trained convolutional neural network for processing, determining the image composition, color ratio composition and texture distribution corresponding to the multimedia data sets, and using the image composition, color ratio composition and texture distribution as image features; Summarizing the word features, numerical features and image features under each data dimension to obtain the first data features of the sub-Internet data sets under each data dimension.
[0011] Preferably, in the data analysis module of a multi-dimensional correlation analysis and mining system for Internet data, analyzing the second data features between each data dimension, including: A quantitative analysis unit for quantitatively analyzing the sub-internet data sets under any data dimension and observing the value change characteristics of the sub-internet data sets under the remaining data dimensions; A second data feature determination unit for determining the relative value change trend of the sub-internet data sets under different data dimensions according to the value change characteristics, determining the association relationship between different data dimensions based on the relative change trend, and determining the second data feature between each data dimension according to the association relationship.
[0012] Preferably, a multi-dimensional association analysis and mining system for Internet data, the result output module includes: A result acquisition unit for: Acquiring the obtained first data feature and second data feature, and at the same time, acquiring the obtained mining purpose, and determining the business branch related to the mining purpose from the business knowledge system according to the mining purpose; Splitting the mining purpose into specific business analysis objectives based on the functional attributes of the business branch; An association analysis and mining unit for Determining the topological structure between each data dimension under different business analysis objectives based on the business knowledge system and the second data feature, and grouping the data dimensions with associated nodes based on the topological structure; Performing association analysis and mining analysis on the sub-internet data sets under the business analysis objectives, the first data feature, and the second data feature to obtain the target analysis results corresponding to the business analysis objectives under each data dimension within each group; A data behavior habit determination unit for summarizing the target analysis results corresponding to different groups based on the business analysis objectives and obtaining the data behavior habits corresponding to the mining purpose based on the summary results.
[0013] Preferably, a multi-dimensional association analysis and mining system for Internet data, the data behavior habit determination unit includes: A data acquisition sub-unit for acquiring the data behavior habits corresponding to different mining purposes and bundling the mining purpose with the corresponding data behavior habits based on the acquisition results; A recording sub-unit for: Constructing a data behavior record table and dividing the table item area of the data behavior record table; Based on the table item area division result, filling the corresponding table item areas in the data behavior record table with the mining purpose and the corresponding data behavior habits according to the bundling results, and completing the record and filing of the mining purpose and the corresponding data behavior habits based on the filling results.
[0014] Compared with the prior art, the beneficial effects of the present invention are as follows: By obtaining the mining objective and effectively locating the target domain name in the Internet according to the mining objective, it is beneficial to accurately extract the Internet data set under the target domain name. By collecting multiple data dimensions corresponding to the Internet data set and analyzing them, the first data feature and the second data feature can be effectively obtained. Furthermore, by accurately performing correlation analysis and mining, the data behavior habits corresponding to the mining objective can be determined, improving the comprehensiveness and accuracy of data mining and facilitating the accurate mining of valuable regular features in Internet data.
[0015] Other features and advantages of the present invention will be described in the following specification, and, in part, will be obvious from the specification or learned by implementing the present invention. The objectives and other advantages of the present invention can be achieved and obtained through the structure specifically pointed out in this application document.
[0016] The technical solution of the present invention will be further described in detail below through the accompanying drawings and embodiments. Description of the Drawings
[0017] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation to the present invention. In the accompanying drawings: Figure 1 It is a structural diagram of a multi-dimensional correlation analysis and mining system for Internet data in an embodiment of the present invention; Figure 2 It is a structural diagram of a positioning module in a multi-dimensional correlation analysis and mining system for Internet data in an embodiment of the present invention; Figure 3 It is a structural diagram of a positioning unit in a multi-dimensional correlation analysis and mining system for Internet data in an embodiment of the present invention. Detailed Embodiment
[0018] The following describes the preferred embodiments of the present invention with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present invention and are not used to limit the present invention.
[0019] Embodiment 1 This embodiment provides a multi-dimensional correlation analysis and mining system for Internet data, including: A positioning module, configured to obtain the mining objective and locate the target domain name in the Internet according to the mining objective; An Internet data acquisition module, configured to acquire the Internet data set under the target domain name and determine multiple data dimensions in the Internet data set; A data analysis module, configured to analyze the first data feature of the sub-Internet data set under each data dimension and analyze the second data feature between each data dimension; A result output module for performing correlation analysis and mining on the first data feature and the second data feature, and outputting the data behavior habits corresponding to the mining objective.
[0020] In this embodiment, the mining objective is the result that needs to be obtained determined according to user requirements. This result is obtained through multi-dimensional correlation analysis and mining of Internet data, and is used to characterize the rules (the rules are the data behavior habits) in the entity operation process recorded by Internet data. For example, the change trend, etc.
[0021] In this embodiment, the target domain name is the source identifier for obtaining Internet data in the Internet determined according to the mining objective.
[0022] In this embodiment, the data dimension is an attribute used to represent different categories presented by the Internet data set determined according to the Internet data set. Among them, one data dimension corresponds to one sub-Internet data set.
[0023] In this embodiment, the first data feature refers to the data transformation trend or characteristics of the sub-Internet data set. For example, it can be the value transformation situation of the sub-Internet data set.
[0024] In this embodiment, the second data feature refers to the interaction relationship between data dimensions.
[0025] The working principle and beneficial effects of the above technical solution are: By obtaining the mining objective and effectively locating the target domain name in the Internet according to the mining objective, it is beneficial to accurately extract the Internet data set under the target domain name. By collecting multiple data dimensions corresponding to the Internet data set and performing analysis, the first data feature and the second data feature can be effectively obtained, and then the data behavior habits corresponding to the mining objective can be accurately determined through correlation analysis and mining, improving the comprehensiveness and accuracy of data mining and facilitating the accurate mining of valuable rule features in Internet data.
[0026] Embodiment 2 Based on Embodiment 1, this embodiment provides a multi-dimensional correlation analysis and mining system for Internet data, as Figure 2 shown, the positioning module includes: A request sending unit for sending a data mining request to the server based on the user terminal, and identifying the data mining request based on the first data receiving node to obtain the core keyword of the data mining request; A mining objective determination unit for determining the mining objective according to the core keyword and inputting the mining objective into the second data receiving node; An analysis unit for analyzing the data mining objective according to the second data receiving node to determine the data content corresponding to the data mining objective and the mining information type; A positioning unit is used to retrieve a domain name set in the Internet in the second data receiving node, and perform positioning in the domain name set according to the mining information type and data content to obtain a target domain name.
[0027] In this embodiment, the output end of the user terminal is connected to the input end of the server. Among them, the server includes: a first data receiving node and a second data receiving node. The input end of the first data receiving node is connected, and the output end of the first data receiving node is connected to the input end of the second data receiving node.
[0028] In this embodiment, the first data receiving node is used to identify a data mining request, so as to effectively obtain the core keyword of the data mining request. Among them, the core keyword is a data segment such as the object purpose to be mined in the data mining request.
[0029] In this embodiment, the mining purpose is the final mining result determined based on the core keyword.
[0030] In this embodiment, a domain name set in the Internet is pre-stored in the second data receiving node. The second data receiving node effectively determines the data content and the mining information type according to the mining purpose.
[0031] The working principle and beneficial effects of the above technical solution are: through the first data receiving node, the analysis of the data mining request is effectively realized, so as to effectively obtain the core keyword. By analyzing the core keyword, the data mining purpose can be effectively obtained. Furthermore, through the second data analysis node, the target domain name is determined based on the mining purpose, improving the accuracy and effectiveness of obtaining the target domain name.
[0032] Embodiment 3 Based on Embodiment 2, this embodiment provides a multi-dimensional association analysis and mining system for Internet data, as Figure 3 shown. The positioning unit includes: A domain name set reading subunit is used to read the domain name set and determine the information distribution interval corresponding to each domain name in the Internet in the domain name set; A first positioning subunit is used to perform a first positioning in the information distribution state according to the mining information type to obtain a key domain name that meets the mining information type; A second positioning subunit is used to extract the information distribution interval of the key domain name and perform a second positioning in the information distribution interval of the key domain name according to the data content to obtain a key information distribution interval; A target domain name extraction subunit is used to obtain the affiliated domain name corresponding to the key information distribution interval and extract the affiliated domain name as the target domain name.
[0033] In this embodiment, the information distribution range refers to the information type range and the information data range of the web page corresponding to each domain name in the domain name set.
[0034] In this embodiment, the subordinate domain name refers to the subordinate domain name under the key domain name, which is the domain name corresponding to the key information distribution range.
[0035] The working principle and beneficial effects of the above technical solution are as follows: By determining the information distribution range of each domain name in the domain name set in the Internet, the target domain name can be effectively located and extracted under the key domain name, effectively ensuring the accuracy of the target domain name extraction.
[0036] Embodiment 4 Based on Embodiment 1, this embodiment provides a multi-dimensional association analysis and mining system for Internet data. The Internet data acquisition module includes: The data extraction unit is used to retrieve the Internet data set based on the target domain name in the Internet; The data parsing unit is used to parse the Internet data set to obtain the data dimensions in the Internet data set.
[0037] The working principle and beneficial effects of the above technical solution are as follows: By retrieving the corresponding Internet data set from the Internet according to the target domain name and parsing the Internet data set, the data dimensions in the Internet data set are determined, providing reliable data support for the multi-dimensional association analysis and mining of Internet data.
[0038] Embodiment 5 Based on Embodiment 4, this embodiment provides a multi-dimensional association analysis and mining system for Internet data. The data extraction unit includes: The page information determination subunit is used to locate the corresponding target page in the Internet according to the target domain name, read the target page, and determine the page information in the target page; The valid information determination subunit is used to determine the valid information in the page information according to the mining purpose; The information data determination subunit is used to extract the valid information and determine the information data of the valid information; The preprocessing subunit is used to preprocess the information data of the valid information to obtain the Internet data set.
[0039] In this embodiment, the target page refers to the web page determined from the Internet according to the target domain name, that is, the main body recording the Internet data set.
[0040] In this embodiment, the valid information refers to the data related to the mining purpose obtained by screening the page information in the target page according to the mining purpose.
[0041] In this embodiment, preprocessing refers to operations such as cleaning and deduplicating the information data of valid information.
[0042] The working principle and beneficial effects of the above technical solution are as follows: By locating the corresponding target page in the Internet according to the target domain name and screening the page information of the located target page, the information data of valid information is determined. Finally, the information data of the obtained valid information is preprocessed to effectively obtain the Internet data set, providing reliable data support for multi-dimensional correlation analysis and mining of Internet data.
[0043] Embodiment 6 Based on Embodiment 4, this embodiment provides a multi-dimensional correlation analysis and mining system for Internet data. The data parsing unit includes: A reading subunit for reading the Internet data set, determining the data content of the Internet data set, and at the same time, obtaining the content core points of the data content; A data processing subunit for obtaining the similarity between each content core point. If the similarity between the target content core points is equal to or greater than the preset similarity threshold, the content core points are merged; A data dimension determination subunit for taking the content core points as the data dimensions in the Internet data set according to the merging result.
[0044] In this embodiment, the content core point refers to a data parameter that can represent the key core or main idea of the data content.
[0045] In this embodiment, the preset similarity threshold is set in advance and is used as the minimum standard for measuring whether to merge the content core points and can be adjusted.
[0046] The working principle and beneficial effects of the above technical solution are as follows: By determining the data content of the obtained Internet data set and determining the content core points of the data content, secondly, determining the similarity between different content core points, realizing the merging of content core points whose similarity is equal to or greater than the preset similarity threshold, and finally realizing the determination of the data dimensions in the Internet data set according to the merging result, providing convenience for multi-dimensional correlation analysis and mining of Internet data.
[0047] Embodiment 7 Based on Embodiment 1, this embodiment provides a multi-dimensional correlation analysis and mining system for Internet data. The data analysis module includes: An Internet data splitting unit for dimensionally splitting the Internet data set based on the data dimensions to obtain sub-Internet data sets under each data dimension; An Internet data analysis unit, configured to: Preprocess the sub-Internet data sets under each data dimension, classify the sub-Internet data sets under each data dimension based on the preprocessing results, and extract the category features of each category; Generate a category identifier for each category based on the category features, and label the corresponding classification results based on the category identifier; A first data feature determination unit, configured to: Obtain the text content set and the multimedia data set included in the sub-Internet data sets under each data dimension based on the labeling results; Segment the text content set, determine the part-of-speech and occurrence frequency of different word elements based on the segmentation results, and use the part-of-speech and occurrence frequency as word features; Quantify the value range of the text content set based on the segmentation results, and determine the numerical features of the text content set based on the quantification results; Input the multimedia data sets included in the sub-Internet data sets under each dimension into a pre-trained convolutional neural network for processing, determine the image composition, color ratio composition, and texture distribution corresponding to the multimedia data sets, and use the image composition, color ratio composition, and texture distribution as image features; Summarize the word features, numerical features, and image features under each data dimension to obtain the first data features of the sub-Internet data sets under each data dimension.
[0048] In this embodiment, dimension splitting refers to splitting the Internet data set into data sets corresponding to each dimension according to the data dimension, and the obtained result is the sub-Internet data set.
[0049] In this embodiment, preprocessing refers to operations such as cleaning and data supplementation on the sub-Internet data sets.
[0050] In this embodiment, the category feature refers to the characteristics presented by each category of data obtained by classifying the sub-Internet data sets under each data dimension, such as the data type and data value characteristics in each category of data, etc.
[0051] In this embodiment, the category identifier refers to a symbol generated based on the category features for marking the sub-Internet data sets of each category.
[0052] In this embodiment, the text content refers to the text content included in the sub-Internet data sets.
[0053] In this embodiment, the multimedia data set refers to the picture and audio data, etc. included in the sub-Internet data sets.
[0054] In this embodiment, word segmentation refers to splitting the vocabulary included in the text content.
[0055] In this embodiment, a word element refers to the characters contained in each vocabulary obtained after word segmentation of the text content.
[0056] In this embodiment, a word feature refers to the parameter information that can characterize the part-of-speech and frequency of occurrence of each word obtained after word segmentation.
[0057] In this embodiment, a numerical feature refers to the value range and specific value distribution corresponding to the text content set.
[0058] In this embodiment, the pre-trained convolutional neural network is a tool that has been pre-trained for processing and analyzing multimedia data sets, and is used to extract features such as the color composition of images.
[0059] In this embodiment, an image feature refers to the data parameters that can characterize the core of the multimedia data set obtained after analyzing the multimedia data set.
[0060] The working principle and beneficial effects of the above technical solution are as follows: By splitting the Internet data set according to the data dimension, the sub-Internet data sets under each data dimension can be accurately and effectively determined. After preprocessing the obtained sub-Internet data sets, classification processing is performed, and then the text content sets and multimedia data sets included in the sub-Internet data sets under each data dimension can be determined according to the classification results. Secondly, the obtained text content sets and multimedia data sets are processed respectively to determine the word features and data features corresponding to the text content sets. At the same time, the image features of the multimedia data sets are also determined. Finally, the word features, numerical features, and image features under each data dimension are summarized to accurately and effectively determine the first data features of the sub-Internet data sets under each data dimension, providing convenience and guarantee for the multi-dimensional correlation analysis and mining of Internet data.
[0061] Embodiment 8 On the basis of Embodiment 1, this embodiment provides a multi-dimensional correlation analysis and mining system for Internet data. In the data analysis module, the second data features between each data dimension are analyzed, including: A quantitative analysis unit for quantitatively analyzing the sub-Internet data set under any data dimension and observing the value change characteristics of the sub-Internet data sets under the remaining data dimensions; A second data feature determination unit for determining the relative value change trend of the sub-Internet data sets under different data dimensions according to the value change characteristics, determining the association relationship between different data dimensions based on the relative change trend, and determining the second data features between each data dimension according to the association relationship.
[0062] In this embodiment, quantitative analysis refers to adjusting the values of the sub-internet data sets under one data dimension in sequence while keeping the values of the sub-internet data sets under the remaining data dimensions unchanged.
[0063] In this embodiment, the value change feature refers to the situation where when the value of the sub-internet data set of one data dimension changes, the values of the sub-internet data sets under other data dimensions change accordingly, and the change trend of its value when the change occurs.
[0064] In this embodiment, the relative value change trend refers to the relative value change situation between different sub-internet data sets.
[0065] The working principle and beneficial effects of the above technical solution are as follows: By performing quantitative analysis on the sub-internet data sets under any data dimension, the relative value change trend of the sub-internet data sets under different data dimensions is determined according to the analysis results. Furthermore, the association relationship between different data dimensions is locked according to the relative value change trend. Finally, the second data feature between each data dimension is determined through the association relationship, providing a reference basis for the multi-dimensional association analysis and mining of internet data and ensuring the reliability of the multi-dimensional association analysis and mining of internet data.
[0066] Embodiment 9 Based on Embodiment 1, this embodiment provides a multi-dimensional association analysis and mining system for internet data, and a result output module, including: A result acquisition unit, used for: Acquiring the obtained first data feature and second data feature, and at the same time, acquiring the obtained mining purpose, and determining the business branch related to the mining purpose from the business knowledge system according to the mining purpose; Splitting the mining purpose into specific business analysis objectives based on the functional attributes of the business branch; An association analysis and mining unit, used for: Determining the topological structure between each data dimension under different business analysis objectives based on the business knowledge system and the second data feature, and grouping the data dimensions with associated nodes based on the topological structure; Extracting the sub-internet data sets under the data dimensions within each group based on the group division result, and determining the mutual restriction relationship between different sub-internet data sets under different feature categories based on the first data feature of the sub-internet data sets; Performing set division on the sub-internet data sets within each group based on the mutual restriction relationship under different feature categories, and aligning the sub-internet data sets under each feature category based on the set division result; Based on the data alignment results, perform autocorrelation analysis on the sub-Internet datasets under each feature category according to the business analysis objectives, obtain the analysis results under each feature category, and based on the autocorrelation analysis results, perform cross-correlation analysis on the analysis results between the sub-Internet datasets of different feature categories within each group to obtain the target analysis results corresponding to the business analysis objectives under the data dimension within each group; The data behavior habit determination unit is used to summarize the target analysis results corresponding to different groups based on the business analysis objectives, and obtain the data behavior habits corresponding to the mining purposes based on the summary results.
[0067] In this embodiment, the business knowledge system is prepared in advance and is used to record the association relationships between different business branches and the structural relationships between different business branches.
[0068] In this embodiment, the business branch refers to the business information determined from the business knowledge system according to the mining purpose.
[0069] In this embodiment, the functional attribute refers to the business purpose that the business branch can achieve or the business result that can be reached.
[0070] In this embodiment, the business analysis objective refers to the specific business analysis object that can represent the mining objective obtained after splitting the mining purpose, that is, each small objective included in the mining purpose.
[0071] In this embodiment, the associated node refers to the node with an association relationship determined according to the topological structure between data dimensions.
[0072] In this embodiment, the mutual restriction relationship refers to the interaction relationship between sub-Internet data of different categories.
[0073] In this embodiment, data alignment refers to aligning the starting points of the sub-Internet datasets after set partitioning, so as to facilitate autocorrelation analysis of the sub-Internet data within each set.
[0074] In this embodiment, autocorrelation analysis refers to analyzing the association relationship between the sub-Internet data within the same feature category under each set.
[0075] In this embodiment, cross-correlation analysis refers to analyzing the association relationship between categories of the analysis results between the sub-Internet datasets of different feature categories within each group.
[0076] The working principle and beneficial effects of the above technical solution are as follows: By determining the business branches related to the mining purpose from the business knowledge system according to the mining purpose, and splitting the mining purpose according to the functional attributes of business analysis, the business analysis target is determined. Secondly, the sub-internet data is analyzed according to the obtained business analysis target, the first data feature, and the second data feature, so as to accurately and reliably determine the target analysis result corresponding to the business analysis target under the data dimension within each group. Finally, the target analysis results corresponding to different groups are summarized according to the business analysis target, and finally the data behavior habits corresponding to the mining purpose are locked, ensuring the accuracy and reliability of the finally obtained data behavior habits.
[0077] Embodiment 10 Based on Embodiment 9, this embodiment provides a multi-dimensional correlation analysis and mining system for Internet data. The data behavior habit determination unit includes: The data acquisition sub-unit is used to acquire the data behavior habits corresponding to different mining purposes, and bundle the mining purpose with the corresponding data behavior habits based on the acquisition result; The recording sub-unit is used for: Construct a data behavior record table and divide the table item area of the data behavior record table; Based on the table item area division result, fill the corresponding table item area in the data behavior record table with the mining purpose and the corresponding data behavior habits according to the bundling result, and complete the record and filing of the mining purpose and the corresponding data behavior habits based on the filling result.
[0078] In this embodiment, the data behavior record table is a table for recording the data behavior habits corresponding to different mining purposes.
[0079] In this embodiment, the table item area division refers to splitting the data behavior record table according to the data item categories to be recorded in the record table as needed.
[0080] The working principle and beneficial effects of the above technical solution are as follows: By bundling the obtained mining purpose with the corresponding data behavior habits, and at the same time, constructing a data behavior record table and recording the mining purpose and the data behavior habits in the data behavior record table according to the bundling result, the obtained data behavior habits are effectively recorded and filed, and it is also convenient to trace the analysis results.
[0081] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these changes and modifications.
Claims
1. A multi-dimensional correlation analysis and mining system for Internet data, characterized in that: include: A positioning module is used to obtain the mining purpose and locate the target domain name on the Internet according to the mining purpose; An Internet data acquisition module is used to acquire an Internet data set under a target domain name and determine multiple data dimensions in the Internet data set; A data analysis module, used to analyze the first data features of the Internet sub-data set under each data dimension, and to analyze the second data features between the data dimensions; The result output module is used to perform correlation analysis and mining on the first data feature and the second data feature, and output data behavior habits corresponding to the mining purpose.
2. The multi-dimensional correlation analysis and mining system for Internet data according to claim 1, characterized in that: Positioning module, including: A request sending unit, configured to send a data mining request to a server based on a user terminal, and identify the data mining request based on a first data receiving node to obtain core keywords of the data mining request; A mining purpose determining unit, used to determine the mining purpose according to the core keyword, and input the mining purpose to the second data receiving node; An analyzing unit, configured to analyze the data mining purpose according to the second data receiving node, and determine the data content and mining information type corresponding to the data mining purpose; The positioning unit is used to retrieve a domain name set in the Internet in the second data receiving node, and locate the domain name set according to the mining information type and data content to obtain the target domain name.
3. The multi-dimensional correlation analysis and mining system for Internet data according to claim 2, characterized in that: Positioning unit, comprising: A domain name set reading subunit, used to read the domain name set and determine the information distribution interval corresponding to each domain name in the domain name set in the Internet; A first positioning subunit, configured to perform a first positioning in the information distribution state according to the mined information type, and obtain a key domain name that matches the mined information type; A second positioning subunit is used to extract the information distribution interval of the key domain name, and perform a second positioning in the information distribution interval of the key domain name according to the data content to obtain the key information distribution interval; The target domain name extraction subunit is used to obtain the subsidiary domain name corresponding to the key information distribution interval, and extract the subsidiary domain name as the target domain name.
4. The multi-dimensional correlation analysis and mining system for Internet data according to claim 1, characterized in that: Internet data acquisition module, including: A data extraction unit, used for retrieving an Internet data set from the Internet based on a target domain name; The data parsing unit is used to parse the Internet data set to obtain the data dimensions in the Internet data set.
5. The multi-dimensional correlation analysis and mining system for Internet data according to claim 4, characterized in that: Data extraction unit, comprising: The page information determination subunit is used to locate the corresponding target page on the Internet according to the target domain name, read the target page, and determine the page information in the target page; A valid information determination subunit, used to determine valid information in the page information according to the mining purpose; An information data determination subunit is used to extract valid information and determine the information data of the valid information; The preprocessing subunit is used to preprocess the information data of the valid information to obtain the Internet data set.
6. The multi-dimensional correlation analysis and mining system for Internet data according to claim 4, characterized in that: Data analysis unit, including: The reading subunit is used to read the Internet data set, determine the data content of the Internet data set, and obtain the core points of the data content; The data processing subunit is used to obtain the similarity between each content core point, and if the similarity between the target content core points is equal to or greater than a preset similarity threshold, the content core points are merged; The data dimension determination subunit is used to use the content core point as the data dimension in the Internet data set according to the merging result.
7. The multi-dimensional correlation analysis and mining system for Internet data according to claim 1, characterized in that: Data analysis modules, including: An Internet data splitting unit is used to split the Internet data set into dimensions based on the data dimension to obtain a sub-Internet data set under each data dimension; Internet data analysis unit for: Preprocessing the sub-Internet data sets under each data dimension, and classifying the sub-Internet data sets under each data dimension based on the preprocessing results, and extracting the category features of each category; Generate a category identifier for each category based on the category features, and annotate the corresponding classification results based on the category identifier; The first data feature determination unit is configured to: Based on the annotation results, a text content set and a multimedia data set included in the sub-Internet data set under each data dimension are obtained; Segment the text content set, and determine the parts of speech and occurrence frequencies of different word elements based on the segmentation results, and use the parts of speech and occurrence frequencies as word features; Quantifying the value range of the text content set based on the word segmentation result, and determining the numerical features of the text content set based on the quantification result; Input the multimedia data set included in the sub-Internet data set under each dimension into the pre-trained convolutional neural network for processing, determine the image composition, color ratio composition and texture distribution corresponding to the multimedia data set, and use the image composition, color ratio composition and texture distribution as image features; The word features, numerical features and image features under each data dimension are summarized to obtain the first data features of the sub-Internet data set under each data dimension.
8. The multi-dimensional correlation analysis and mining system for Internet data according to claim 1, characterized in that: In the data analysis module, the second data features between the data dimensions are analyzed, including: A quantitative analysis unit, used to perform quantitative analysis on a sub-Internet data set under any data dimension, and observe the value change characteristics of the sub-Internet data sets under other data dimensions; The second data feature determination unit is used to determine the relative value change trend of the sub-Internet data set under different data dimensions according to the value change characteristics, and determine the correlation relationship between different data dimensions based on the relative change trend, and determine the second data feature between each data dimension according to the correlation relationship.
9. The multi-dimensional correlation analysis and mining system for Internet data according to claim 1, characterized in that: The result output module includes: The result acquisition unit is used to: Acquire the first data feature and the second data feature, and at the same time, acquire the mining purpose, and determine a business branch related to the mining purpose from the business knowledge system according to the mining purpose; Split the mining purpose into specific business analysis goals based on the functional attributes of the business branch; Association analysis and mining unit for Determine the topological structure between data dimensions under different business analysis objectives based on the business knowledge system and the second data feature, and divide the data dimensions with associated nodes into groups based on the topological structure; Based on the business analysis target, the first data feature and the second data feature, the sub-Internet data set is subjected to association analysis and mining analysis to obtain the target analysis results corresponding to the business analysis target under the data dimension in each group; The data behavior habit determination unit is used to summarize the target analysis results corresponding to different groups based on the business analysis goals, and obtain the data behavior habits corresponding to the mining purpose based on the summary results.
10. A multi-dimensional correlation analysis and mining system for Internet data according to claim 9, characterized in that: Data behavior habit determination unit, including: The data acquisition subunit is used to acquire data behavior habits corresponding to different mining purposes, and to bundle the mining purpose with the corresponding data behavior habits based on the acquisition results; Record subunit, used to: Construct a data behavior record table and divide the data behavior record table into table item areas; Based on the results of the table item area division and according to the bundling results, the mining purpose and the corresponding data behavior habits are filled in the corresponding table item areas in the data behavior record table, and the mining purpose and the corresponding data behavior habits are recorded and archived based on the filling results.
Citation Information
Patent Citations
Data processing method and device, electronic equipment and storage medium
CN110704751A
Implementation method of alarm data association traceability system based on density clustering
CN118861720A
Large-scale data mining method based on multi-tuple data optimization
CN119089404A
Root cause positioning method and device
CN119806883A
Universal internet information data mining method
US20170046434A1