Web-based data mining method and system

By scoring information value and dividing data modalities on Web data sources, creating a multimodal data fusion framework, extracting cross-modal features and mapping them to common feature space, the problem of neglecting modal differences in the existing technology is solved, and a more accurate and efficient data mining effect is achieved.

CN119988937AInactive Publication Date: 2025-05-13GUANGZHOU YUEZHENG NETWORK INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510081618.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-17
Publication Date
2025-05-13
Estimated Expiration
Not applicable · inactive patent

Smart Images

  • Figure CN119988937A_ABST
    Figure CN119988937A_ABST
Patent Text Reader

Abstract

The invention discloses a Web-based data mining method and system, and relates to the technical field of data processing, and the method comprises the steps: carrying out the information value scoring of each Web data source included in a Web data source set, so as to recognize and select a reference Web data source of which the information value score is greater than a value score threshold; dividing the data in the reference Web data source into a plurality of data block sets according to the data modality of the data in the reference Web data source; dividing the data in the plurality of data block sets into a plurality of structure sub-blocks according to content types, structures and attributes of the data in the plurality of data block sets; creating a multi-modal data fusion framework based on the plurality of structure sub-blocks corresponding to each data block set; based on a multi-modal data fusion framework, fusing the plurality of data block sets to extract cross-modal features; the cross-modal features obtained through extraction are mapped to a common feature space; and executing a data mining task based on the cross-modal features in the common feature space.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing technology, and in particular to a Web-based data mining method and system. Background Art

[0002] Data mining, also known as data mining or knowledge discovery, is the process of extracting valuable information and knowledge from large amounts of data through algorithms and statistical analysis methods. It involves identifying patterns, trends or correlations from raw data to help decision makers make more informed decisions. Data mining is often used in business intelligence, market research, financial analysis, medical diagnosis, scientific research and many other fields.

[0003] Web-based data mining refers to the process of using data mining technology to discover data of different modalities from Web documents and Web servers and extract information or knowledge that people are interested in. By analyzing data of different modalities, hidden patterns, trends, and correlations can be discovered to support decision-making and management. However, when processing data of different modalities, related technologies often ignore the differences in content types and attributes of different modalities, resulting in the model's over-reliance on certain modalities and ignoring other modalities during the fusion process.

[0004] To address the above-mentioned problems, no effective solution has been proposed yet. Summary of the invention

[0005] The embodiments of the present application provide a Web-based data mining method and system to solve the above technical problems.

[0006] The present application provides a Web-based data mining method, comprising: based on the update frequency, user interaction volume and historical information value of a Web data source set obtained from a Web document and a Web server, scoring the information value of each Web data source included in the Web data source set to identify and select a reference Web data source whose information value score is greater than a value scoring threshold; dividing the data in the reference Web data source into a plurality of data block sets according to the data modality of the data in the reference Web data source; dividing the data in the plurality of data block sets into a plurality of structural sub-blocks according to the content type, structure and attribute of the data in the plurality of data block sets; creating a multimodal data fusion framework based on the plurality of structural sub-blocks corresponding to each data block set; based on the multimodal data fusion framework, fusing the plurality of data block sets to extract cross-modal features; mapping the extracted cross-modal features to a common feature space; and performing a data mining task based on the cross-modal features in the common feature space; wherein the data mining task includes a web page content classification task, a web page content clustering task, a user behavior pattern analysis task and a link structure analysis task.

[0007] The present application provides a Web-based data mining system, including: a data source identification and selection module, which is used to score the information value of each Web data source included in the Web data source set based on the update frequency, user interaction volume and historical information value of the Web data source set obtained from Web documents and Web servers, so as to identify and select a reference Web data source whose information value score is greater than a value score threshold; a data partitioning module, which is used to partition the data in the reference Web data source into multiple data block sets according to the data modality of the data in the reference Web data source; and a multimodal data fusion module, which is used to divide the data in the multiple data block sets into multiple data block sets according to the data modality of the data in the reference Web data source. The invention discloses a method for mining cross-modal data mining, which is a method for mining cross-modal data. The method comprises the following steps: dividing the data in a plurality of data block sets into a plurality of structural sub-blocks based on the content type, structure and attributes of the data; creating a multimodal data fusion framework based on the plurality of structural sub-blocks corresponding to each data block set; fusing the plurality of data block sets based on the multimodal data fusion framework to extract cross-modal features; a result aggregation and task execution module for mapping the extracted cross-modal features to a common feature space; and executing data mining tasks based on the cross-modal features in the common feature space; wherein the data mining tasks include the classification tasks of web page content, the clustering tasks of web page content, the user behavior pattern analysis tasks and the link structure analysis tasks.

[0008] Furthermore, the plurality of data block sets include a text data block set, an image data block set, and an audio and video data block set; the data partitioning module partitions the data in the reference Web data source into the plurality of data block sets according to the data modality of the data in the reference Web data source, including:

[0009] Preprocessing the data in the reference Web data source; wherein the preprocessing includes data cleaning, data conversion and feature processing;

[0010] Identifying the data type of the data in the reference Web data source by analyzing the binary features and file header information of the data in the reference Web data source;

[0011] Analyzing statistical characteristics of data in the reference Web data source;

[0012] Performing preliminary modality allocation on the data in the reference Web data source according to the data type and statistical characteristics of the data in the identified reference Web data source;

[0013] For the data initially allocated to the image data block set, dividing the data into macroblocks of a first preset size;

[0014] For the data initially allocated to the text data block set, dividing the data into patches of a second preset size;

[0015] For the data initially allocated to the audio and video data block set, the audio stream part and the video stream part included in the data are separated and processed; wherein, for the audio stream part, the audio stream part is divided into macroblocks of a third preset size, and each macroblock is used as an audio and video data block; for the video stream part, the video stream part is divided into patches of a fourth preset size, and each patch is used as an audio and video data block;

[0016] The divided data blocks are stored in a corresponding data block set; wherein each data block is marked with information about the data in the reference Web data source and the index position of the data block in the data block set.

[0017] Furthermore, the structure of the data includes structured, semi-structured and unstructured; the multimodal data fusion module divides the data in the multiple data block sets into multiple structural sub-blocks according to the content type, structure and attribute of the data in the multiple data block sets; creates a multimodal data fusion framework based on the multiple structural sub-blocks corresponding to each data block set; based on the multimodal data fusion framework, fuses the multiple data block sets to extract cross-modal features, including:

[0018] According to the structure of the data in the text data block set, the image data block set and the audio and video data block set, the data is divided into structured data blocks, semi-structured data blocks and unstructured data blocks;

[0019] Based on the Bayesian network, the nodes of the Bayesian network and the dependency relationships between the nodes are constructed according to the attributes of the data in the structured data block, the semi-structured data block and the unstructured data block; wherein the attributes of the data in the structured data block include field names, data types and constraints; the attributes of the data in the semi-structured data block include labels and nested structures; the attributes of the data in the unstructured data block include the sentiment of the text, the color and texture of the image and the frequency characteristics of the audio;

[0020] Referring to the nodes of the constructed Bayesian network and the dependencies between the nodes, based on the content types of the data in the structured data block, the semi-structured data block and the unstructured data block, the nodes of the Bayesian network are divided into a plurality of structural sub-blocks; wherein the hierarchical relationship of the plurality of structural sub-blocks is determined based on the dependencies between the nodes of the Bayesian network;

[0021] Creating the multimodal data fusion framework based on a plurality of structural sub-blocks corresponding to the structured data block, the semi-structured data block and the unstructured data block;

[0022] Based on the created multimodal data fusion framework, a first structural sub-block among the multiple structural sub-blocks corresponding to the structured data block, a second structural sub-block among the multiple structural sub-blocks corresponding to the semi-structured data block, and a third structural sub-block among the multiple structural sub-blocks corresponding to the unstructured data block are fused multiple times; wherein, in each fusion process, the first structural sub-block, the second structural sub-block, and the third structural sub-block all need to be re-determined;

[0023] According to the results of multiple fusions, the cross-modal features are extracted.

[0024] Based on the embodiments provided by the present application, based on the update frequency, user interaction volume and historical information value of the Web data source set obtained from the Web documents and the Web server, the information value score of each Web data source included in the Web data source set is scored to identify and select a reference Web data source whose information value score is greater than a value score threshold; according to the data modality of the data in the reference Web data source, the data in the reference Web data source is divided into multiple data block sets; according to the content type, structure and attribute of the data in the multiple data block sets, the data in the multiple data block sets is divided into multiple structural sub-blocks; based on the multiple structural sub-blocks corresponding to each data block set, a multimodal data fusion framework is created; based on the multimodal data fusion framework, the multiple data block sets are fused to extract cross-modal features; the extracted cross-modal features are mapped to a common feature space; and based on the cross-modal features in the common feature space, a data mining task is performed. This achieves: Improving the accuracy and efficiency of data mining: By scoring the information value based on the update frequency, user interaction volume and historical information value of Web data sources, this method can identify and select the most valuable data sources, thereby improving the accuracy and efficiency of data mining; Optimizing multimodal data processing: Dividing data into different sets of data blocks according to modality, and further dividing them into structural sub-blocks according to content type, structure and attributes, this method helps to process multimodal data more finely and reduce excessive reliance on a single modality; Enhancing cross-modal feature extraction: By creating a multimodal data fusion framework, this method can effectively extract cross-modal features and enhance the model's ability to perform well on different modalities. It improves the comprehensive analysis capability of cross-modal data and the comprehensiveness and depth of data mining results; improves the consistency of feature space: maps cross-modal features to a common feature space, ensures the consistency of different modal data at the feature level, and provides a unified processing basis for subsequent data mining tasks; supports diversified data mining tasks: this method is suitable for a variety of data mining tasks such as web page content classification, clustering, user behavior pattern analysis, and link structure analysis, and has broad application prospects and practical value; improves the quality of decision support: by more accurately mining and analyzing Web data, this method can provide decision makers with more valuable information, thereby improving the quality and efficiency of decision making. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] The drawings described herein are used to provide a further understanding of the embodiments of the present invention and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0026] Figure 1 is a flowchart of an optional Web-based data mining method according to an embodiment of the present application;

[0027] Figure 2 The structure diagram of an optional Web-based data mining system according to an embodiment of the present application.

[0028] The realization of the purpose, functional features and advantages of the present invention will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION

[0029] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0030] Alternatively, if Figure 1 As shown, the present application provides a Web-based data mining method, comprising:

[0031] S101, based on the update frequency, user interaction volume and historical information value of the Web data source set obtained from the Web document and the Web server, scoring the information value of each Web data source included in the Web data source set, so as to identify and select a reference Web data source whose information value score is greater than a value score threshold;

[0032] S102, dividing the data in the reference Web data source into a plurality of data block sets according to the data modality of the data in the reference Web data source;

[0033] S103, dividing the data in the multiple data block sets into multiple structural sub-blocks according to the content type, structure and attribute of the data in the multiple data block sets; creating a multimodal data fusion framework based on the multiple structural sub-blocks corresponding to each data block set; and fusing the multiple data block sets based on the multimodal data fusion framework to extract cross-modal features;

[0034] S104, mapping the extracted cross-modal features to a common feature space; and performing data mining tasks based on the cross-modal features in the common feature space; wherein the data mining tasks include web page content classification tasks, web page content clustering tasks, user behavior pattern analysis tasks, and link structure analysis tasks.

[0035] Based on the embodiments provided by the present application, based on the update frequency, user interaction volume and historical information value of the Web data source set obtained from the Web documents and the Web server, the information value score of each Web data source included in the Web data source set is scored to identify and select a reference Web data source whose information value score is greater than a value score threshold; according to the data modality of the data in the reference Web data source, the data in the reference Web data source is divided into multiple data block sets; according to the content type, structure and attribute of the data in the multiple data block sets, the data in the multiple data block sets is divided into multiple structural sub-blocks; based on the multiple structural sub-blocks corresponding to each data block set, a multimodal data fusion framework is created; based on the multimodal data fusion framework, the multiple data block sets are fused to extract cross-modal features; the extracted cross-modal features are mapped to a common feature space; and based on the cross-modal features in the common feature space, a data mining task is performed. This achieves: Improving the accuracy and efficiency of data mining: By scoring the information value based on the update frequency, user interaction volume and historical information value of Web data sources, this method can identify and select the most valuable data sources, thereby improving the accuracy and efficiency of data mining; Optimizing multimodal data processing: Dividing data into different sets of data blocks according to modality, and further dividing them into structural sub-blocks according to content type, structure and attributes, this method helps to process multimodal data more finely and reduce excessive reliance on a single modality; Enhancing cross-modal feature extraction: By creating a multimodal data fusion framework, this method can effectively extract cross-modal features and enhance the model's ability to perform well on different modalities. It improves the comprehensive analysis capability of cross-modal data and the comprehensiveness and depth of data mining results; improves the consistency of feature space: maps cross-modal features to a common feature space, ensures the consistency of different modal data at the feature level, and provides a unified processing basis for subsequent data mining tasks; supports diversified data mining tasks: this method is suitable for a variety of data mining tasks such as web page content classification, clustering, user behavior pattern analysis, and link structure analysis, and has broad application prospects and practical value; improves the quality of decision support: by more accurately mining and analyzing Web data, this method can provide decision makers with more valuable information, thereby improving the quality and efficiency of decision making.

[0036] Alternatively, if Figure 2 As shown, the present application provides a Web-based data mining system, comprising:

[0037] The data source identification and selection module 201 is used to score the information value of each Web data source included in the Web data source set based on the update frequency, user interaction volume and historical information value of the Web data source set obtained from the Web document and the Web server, so as to identify and select a reference Web data source whose information value score is greater than a value score threshold;

[0038] A data partitioning module 202, for partitioning the data in the reference Web data source into a plurality of data block sets according to the data modality of the data in the reference Web data source;

[0039] The multimodal data fusion module 203 is used to divide the data in the multiple data block sets into multiple structural sub-blocks according to the content type, structure and attribute of the data in the multiple data block sets; create a multimodal data fusion framework based on the multiple structural sub-blocks corresponding to each data block set; and fuse the multiple data block sets based on the multimodal data fusion framework to extract cross-modal features;

[0040] The result aggregation and task execution module 204 is used to map the extracted cross-modal features to a common feature space; and to perform data mining tasks based on the cross-modal features in the common feature space; wherein the data mining tasks include web page content classification tasks, web page content clustering tasks, user behavior pattern analysis tasks, and link structure analysis tasks.

[0041] Furthermore, the plurality of data block sets include a text data block set, an image data block set, and an audio and video data block set; and the data partitioning module partitions the data in the reference Web data source into the plurality of data block sets according to the data modality of the data in the reference Web data source, including:

[0042] Preprocess the data in the reference Web data source; the preprocessing includes data cleaning, data conversion and feature processing;

[0043] Data cleaning includes processing missing values, outliers, and duplicate values; data conversion includes standardization, normalization, unique hot encoding, etc.

[0044] Identify the data type of the data in the reference Web data source by analyzing the binary features and file header information of the data in the reference Web data source; for example, a text file usually starts with a specific character encoding, an image file has a specific image format identifier, and an audio and video file has a corresponding encoding format header;

[0045] Analyze the statistical characteristics of data in the reference web data source; for example, text data usually has a high vocabulary density, while image and audio and video data have different signal characteristics;

[0046] According to the data types and statistical characteristics of the data in the reference Web data sources identified, preliminary modality allocation is performed on the data in the reference Web data sources;

[0047] For the data initially allocated to the image data block set, the data is divided into macroblocks of a first preset size; wherein the first preset size is determined based on the following formula:

[0048] Block_Size = 2 n ×Base_Size; wherein Block_Size is a first preset size; n is an integer index; Base_Size is a basic block size determined based on data characteristics;

[0049] For the data initially allocated to the text data block set, the data is divided into patches of a second preset size; wherein the second preset size is determined based on the following formula:

[0050] Wherein, Patch_Size is the second preset size; Total_Data_Length is the total length of the data; Number_of_Patches is the expected number of patches;

[0051] For the data initially allocated to the audio and video data block set, the audio stream part and the video stream part included in the data are processed separately; this ensures the alignment of the audio and video data in time series while maintaining spatial independence.

[0052] For the audio stream part, the audio stream part is divided into macroblocks of a third preset size, and each macroblock is used as an audio and video data block; for the video stream part, the video stream part is divided into patches of a fourth preset size, and each patch is used as an audio and video data block;

[0053] The divided data blocks are stored in the corresponding data block set; wherein each data block is marked with information of the data in the reference Web data source and the index position of the data block in the data block set.

[0054] Furthermore, the structure of the data includes structured, semi-structured and unstructured; the multimodal data fusion module divides the data in the multiple data block sets into multiple structural sub-blocks according to the content type, structure and attribute of the data in the multiple data block sets; based on the multiple structural sub-blocks corresponding to each data block set, a multimodal data fusion framework is created; based on the multimodal data fusion framework, the multiple data block sets are fused to extract cross-modal features, including:

[0055] According to the structure of data in the text data block set, the image data block set and the audio and video data block set, the data is divided into structured data blocks, semi-structured data blocks and unstructured data blocks;

[0056] Based on the Bayesian network, the nodes of the Bayesian network and the dependency relationships between the nodes are constructed according to the attributes of the data in the structured data block, the semi-structured data block and the unstructured data block; the attributes of the data in the structured data block include the field name, data type and constraint; the attributes of the data in the semi-structured data block include the label and the nested structure; the attributes of the data in the unstructured data block include the sentiment of the text, the color and texture of the image and the frequency characteristics of the audio;

[0057] Referring to the nodes of the constructed Bayesian network and the dependencies between the nodes, the nodes of the Bayesian network are divided into a plurality of structural sub-blocks based on the content types of the data in the structured data blocks, the semi-structured data blocks and the unstructured data blocks; wherein the hierarchical relationships of the plurality of structural sub-blocks are determined based on the dependencies between the nodes of the Bayesian network;

[0058] Content type: for text data, it includes paragraphs, titles, lists, news, blogs, comments, etc.; for image data, it includes visual features such as color, texture, shape, and content such as faces, scenery, and objects; for audio and video data, it is divided according to its audio characteristics such as frequency and beats, and video content such as actions and scenes.

[0059] A multimodal data fusion framework is created based on multiple structural sub-blocks corresponding to structured data blocks, semi-structured data blocks and unstructured data blocks;

[0060] Based on the created multimodal data fusion framework, a first structural sub-block among the multiple structural sub-blocks corresponding to the structured data block, a second structural sub-block among the multiple structural sub-blocks corresponding to the semi-structured data block, and a third structural sub-block among the multiple structural sub-blocks corresponding to the unstructured data block are fused multiple times; wherein, in each fusion process, the first structural sub-block, the second structural sub-block, and the third structural sub-block all need to be re-determined;

[0061] Based on the results of multiple fusions, cross-modal features are extracted.

[0062] Furthermore, based on the Bayesian network, according to the attributes of the data in the structured data block, the semi-structured data block and the unstructured data block, the nodes of the Bayesian network and the dependency relationships between the nodes are constructed, including:

[0063] Construct the nodes of the Bayesian network according to the field names, data types and constraints of the data in the structured data block, the labels and nested structures of the data in the semi-structured data block, the sentiment of the text, the color and texture of the image, and the frequency characteristics of the audio in the unstructured data block;

[0064] Among them, each node in the Bayesian network represents an attribute or a set of related attributes. For example: Structured nodes: can be fields such as "age" and "gender". Semi-structured nodes: can be labels such as "author" and "publication date". Unstructured nodes: can be features such as "positive emotions" and "cats in images"; For structured data, mutual information or conditional mutual information is used to evaluate the interdependence between variables. For semi-structured data, the co-occurrence matrix (Co-occurrenceMatrix) can be used to evaluate the correlation between labels. For unstructured data, correlation analysis (CorrelationAnalysis) can be used to evaluate the correlation between features.

[0065] Determine the dependency between nodes based on the correlation and mutual influence between the field name, data type and constraint of the data in the structured data block, the label and nested structure of the data in the semi-structured data block, the sentiment of the text of the data in the unstructured data block, the color and texture of the image and the frequency characteristics of the audio, including the following steps:

[0066] Creating an initial dependency network structure corresponding to the dependency relationship between nodes; wherein the initial dependency network structure is a network structure pre-selected based on the maximum information coefficient;

[0067] Use the Bayesian Information Criterion and the Akaike Information Criterion to define a scoring function to measure the quality of the network structure; the Akaike Information Criterion, also known as the Akaike Information Criterion; for structured data, the Bayesian Information Criterion (BIC) can be used; for semi-structured and unstructured data, a more general scoring function such as the Akaike Information Criterion (AIC) can be considered;

[0068] Based on the initial dependency network structure, the neighborhood structure is generated by adding, deleting or reversing edges; the neighborhood structure refers to all possible new structures obtained by making small modifications to the current network structure. These modifications include adding an edge, deleting an edge or reversing the direction of an edge. The neighborhood structure is not part of the initial dependency network structure, but a potential new structure obtained by performing these operations on the initial structure.

[0069] For each neighborhood structure, the Bayes factor is used to calculate its score; wherein, the Bayes factor is used to calculate its score, including: for each dependency between a pair of nodes in the current network structure, the likelihood ratio of two models including and excluding the dependency is calculated, and the Bayes factor is the logarithm of the likelihood ratio of the two models; the Bayes factor is a scoring function that measures the degree of fit between the network structure and the data, and the higher the score, the more likely the network structure is to correctly reflect the dependency in the data;

[0070] If the score of the neighborhood structure is higher than the current network structure, the current network structure is updated to the neighborhood structure with the highest score. This step is repeated until no structure with a higher score can be found.

[0071] After the above iterative steps are repeated, the target network result is obtained;

[0072] According to the target network results, the dependencies between nodes are determined.

[0073] Further, based on the created multimodal data fusion framework, a first structural sub-block among the multiple structural sub-blocks corresponding to the structured data block, a second structural sub-block among the multiple structural sub-blocks corresponding to the semi-structured data block, and a third structural sub-block among the multiple structural sub-blocks corresponding to the unstructured data block are fused multiple times; and cross-modal features are extracted according to the results of the multiple fusions, including:

[0074] The definition of multiple structure sub-blocks corresponding to the structured data block is defined as A = {A1, A2, ..., A k1}; define multiple structure sub-blocks corresponding to the semi-structured data block as B = {B1, B2, ..., B k2}; define multiple structured sub-blocks corresponding to the unstructured data block as C = {C1, C2, ..., C k3};

[0075] Determine a selection strategy; wherein the selection strategy is used to determine how to select sub-blocks from sets A, B, and C each time when merging; the selection strategy is any one of the following strategies: selecting a sub-block from each set each time; selecting a fixed number of sub-blocks from each set each time; selecting a different number of sub-blocks from each set each time;

[0076] Set the number of fusion times k;

[0077] First fusion: select the first x sub-blocks from A; select the first y sub-blocks from set B; select the first z sub-blocks from set C;

[0078] Subsequent fusion: For the g-th fusion, g = 2, 3, ..., k, update the selected sub-blocks: select the g×x-th sub-block to the (g+1)×x-1-th sub-block from A; select the g×y-th sub-block to the (g+1)×y-1-th sub-block from B; select the g×z-th sub-block to the (g+1)×z-1-th sub-block from C;

[0079] According to the results of multiple fusions, the cross-modal features are extracted.

[0080] Furthermore, the result aggregation and task execution module maps the extracted cross-modal features to a common feature space; and executes data mining tasks based on the cross-modal features in the common feature space, including:

[0081] Through cross-modal embedding, data from different modalities are mapped to a common feature space;

[0082] Based on the cross-modal features in a common feature space, data mining tasks are performed.

[0083] In the process of performing data mining tasks, semantically enhanced entity linking: In the text data block, based on entity recognition technology, the data is further divided into sub-blocks containing specific entities. For each sub-block containing an entity, the semantically enhanced entity linking method is used to match the entity in the text with the corresponding entity in the knowledge base, enhance the semantic information of the entity, and build a relationship network between entities.

[0084] Behavior pattern mining: In the user behavior data block, the data is divided into sub-blocks representing different behavior patterns according to the type and frequency of user behavior.

[0085] Sentiment trend analysis: In the social media data block, the data is divided into positive, negative, and neutral sentiment sub-blocks based on the sentiment polarity.

[0086] Furthermore, the Web-based data mining system also includes:

[0087] The privacy-protected data desensitization module is used to formulate personalized desensitization strategies based on data sensitivity and business needs during the execution of data mining tasks; personalized desensitization strategies include data generalization, data perturbation, and data replacement; create a desensitization engine; the desensitization engine is used to process the identified sensitive data according to the personalized desensitization strategy; after processing the identified sensitive data, use the data quality assessment algorithm to evaluate the analytical value of the desensitized data; record logs in the process of processing sensitive data for auditing and problem tracking.

[0088] Furthermore, the Web data source set includes text, images, audio, video and metadata; the Web-based data mining system also includes:

[0089] An adaptive learning and optimization module, which is used to adjust the cross-modal features in the common feature space according to the accuracy scores of the mining results of the data mining task and user feedback;

[0090] The user-customized report module is used to generate user-customized data mining reports based on the data source, mining depth and result display method selected by the user.

[0091] Further,

[0092]

[0093] Among them, BFij is the Bayes factor of the dependency between node i and node j; It is a network structure containing dependency (i,j) The likelihood probability of the next data D; It is a network structure that does not contain dependency (i,j) The likelihood probability of the next data D;

[0094] If the Bayes factor is greater than a certain threshold, the dependency is determined to be statistically significant; based on the result of the Bayes factor, the statistically significant dependencies are retained and the insignificant dependencies are removed. A certain threshold, for example, 1 or 3.

[0095] The Bayes factor is the logarithm of the likelihood ratio of two models and can be used to measure the strength of evidence for one model relative to another.

[0096] It should be noted that in the present application, the embodiments implemented on the Web-based data mining system side can be referenced to each other with the embodiments implemented on the Web-based data mining method side, and the present application will not describe them one by one.

[0097] The above are only preferred embodiments of the present invention, and are not intended to limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made using the contents of the present invention specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.

Claims

1. A Web-based data mining method, characterized in that: include: Based on the update frequency, user interaction volume and historical information value of the Web data source set obtained from the Web documents and the Web server, score the information value of each Web data source included in the Web data source set to identify and select a reference Web data source whose information value score is greater than a value score threshold; Dividing the data in the reference Web data source into a plurality of data block sets according to the data modality of the data in the reference Web data source; According to the content type, structure and attribute of the data in the multiple data block sets, the data in the multiple data block sets are divided into multiple structural sub-blocks; based on the multiple structural sub-blocks corresponding to each data block set, a multimodal data fusion framework is created; Based on the multimodal data fusion framework, multiple data block sets are fused to extract cross-modal features; Map the extracted cross-modal features into a common feature space; And based on the cross-modal features in the common feature space, data mining tasks are performed; wherein the data mining tasks include web page content classification tasks, web page content clustering tasks, user behavior pattern analysis tasks and link structure analysis tasks.

2. A Web-based data mining system, the system implementing the method according to claim 1, characterized in that: include: A data source identification and selection module, configured to score the information value of each Web data source included in the Web data source set based on the update frequency, user interaction volume and historical information value of the Web data source set obtained from the Web document and the Web server, so as to identify and select a reference Web data source whose information value score is greater than a value score threshold; A data partitioning module, configured to partition the data in the reference Web data source into a plurality of data block sets according to the data modality of the data in the reference Web data source; A multimodal data fusion module is used to divide the data in the multiple data block sets into multiple structural sub-blocks according to the content type, structure and attribute of the data in the multiple data block sets; and to create a multimodal data fusion framework based on the multiple structural sub-blocks corresponding to each data block set; Based on the multimodal data fusion framework, multiple data block sets are fused to extract cross-modal features; The result aggregation and task execution module is used to map the extracted cross-modal features into a common feature space; and to execute data mining tasks based on the cross-modal features in the common feature space; wherein the data mining tasks include web page content classification tasks, web page content clustering tasks, user behavior pattern analysis tasks, and link structure analysis tasks.

3. The Web-based data mining system according to claim 2, characterized in that: The plurality of data block sets include a text data block set, an image data block set, and an audio and video data block set; The data partitioning module partitions the data in the reference Web data source into a plurality of data block sets according to the data modality of the data in the reference Web data source, including: Preprocessing the data in the reference Web data source; wherein the preprocessing includes data cleaning, data conversion and feature processing; Identifying the data type of the data in the reference Web data source by analyzing the binary features and file header information of the data in the reference Web data source; Analyze the statistical characteristics of the data in the reference Web data source.

4. The Web-based data mining system according to claim 3, characterized in that: The plurality of data block sets include a text data block set, an image data block set, and an audio and video data block set; The data partitioning module partitions the data in the reference Web data source into a plurality of data block sets according to the data modality of the data in the reference Web data source, including: Performing preliminary modality allocation on the data in the reference Web data source according to the data type and statistical characteristics of the data in the identified reference Web data source; For the data initially allocated to the image data block set, dividing the data into macroblocks of a first preset size; For the data initially allocated to the text data block set, the data is divided into patches of a second preset size.

5. The Web-based data mining system according to claim 4, characterized in that: The plurality of data block sets include a text data block set, an image data block set, and an audio and video data block set; The data partitioning module partitions the data in the reference Web data source into a plurality of data block sets according to the data modality of the data in the reference Web data source, including: For the data initially allocated to the audio and video data block set, the audio stream part and the video stream part included in the data are separated and processed; wherein, for the audio stream part, the audio stream part is divided into macroblocks of a third preset size, and each macroblock is used as an audio and video data block; for the video stream part, the video stream part is divided into patches of a fourth preset size, and each patch is used as an audio and video data block; The divided data blocks are stored in a corresponding data block set; wherein each data block is marked with information about the data in the reference Web data source and the index position of the data block in the data block set.

6. The Web-based data mining system according to claim 5, characterized in that: The structure of the data includes structured, semi-structured and unstructured; the multimodal data fusion module divides the data in the multiple data block sets into multiple structural sub-blocks according to the content type, structure and attribute of the data in the multiple data block sets; Based on multiple structural sub-blocks corresponding to each data block set, a multimodal data fusion framework is created; Based on the multimodal data fusion framework, multiple data block sets are fused to extract cross-modal features, including: According to the structure of the data in the text data block set, the image data block set and the audio and video data block set, the data is divided into structured data blocks, semi-structured data blocks and unstructured data blocks; Based on a Bayesian network, the nodes of the Bayesian network and the dependencies between the nodes are constructed according to the attributes of the data in the structured data block, the semi-structured data block and the unstructured data block; wherein the attributes of the data in the structured data block include field names, data types and constraints; the attributes of the data in the semi-structured data block include labels and nested structures; the attributes of the data in the unstructured data block include the sentiment of the text, the color and texture of the image, and the frequency characteristics of the audio.

7. The Web-based data mining system according to claim 6, characterized in that: The structure of the data includes structured, semi-structured and unstructured; the multimodal data fusion module divides the data in the multiple data block sets into multiple structural sub-blocks according to the content type, structure and attribute of the data in the multiple data block sets; Based on multiple structural sub-blocks corresponding to each data block set, a multimodal data fusion framework is created; Based on the multimodal data fusion framework, multiple data block sets are fused to extract cross-modal features, including: Referring to the nodes of the constructed Bayesian network and the dependencies between the nodes, based on the content types of the data in the structured data block, the semi-structured data block and the unstructured data block, the nodes of the Bayesian network are divided into a plurality of structural sub-blocks; wherein the hierarchical relationship of the plurality of structural sub-blocks is determined based on the dependencies between the nodes of the Bayesian network; The multimodal data fusion framework is created based on a plurality of structural sub-blocks corresponding to the structured data block, the semi-structured data block and the unstructured data block.

8. The Web-based data mining system according to claim 7, characterized in that: The structure of the data includes structured, semi-structured and unstructured; the multimodal data fusion module divides the data in the multiple data block sets into multiple structural sub-blocks according to the content type, structure and attribute of the data in the multiple data block sets; Based on multiple structural sub-blocks corresponding to each data block set, a multimodal data fusion framework is created; Based on the multimodal data fusion framework, multiple data block sets are fused to extract cross-modal features, including: Based on the created multimodal data fusion framework, a first structural sub-block among the multiple structural sub-blocks corresponding to the structured data block, a second structural sub-block among the multiple structural sub-blocks corresponding to the semi-structured data block, and a third structural sub-block among the multiple structural sub-blocks corresponding to the unstructured data block are fused multiple times; wherein, in each fusion process, the first structural sub-block, the second structural sub-block, and the third structural sub-block all need to be re-determined; According to the results of multiple fusions, the cross-modal features are extracted.

9. The Web-based data mining system according to claim 8, characterized in that: The web-based data mining system also includes: The privacy-protected data desensitization module is used to formulate personalized desensitization strategies according to data sensitivity and business needs during the execution of the data mining task; the personalized desensitization strategies include data generalization, data perturbation and data replacement; create a desensitization engine; wherein, the desensitization engine is used to process the identified sensitive data according to the personalized desensitization strategy; after processing the identified sensitive data, use the data quality assessment algorithm to evaluate the analytical value of the desensitized data; and record logs in the process of processing sensitive data for auditing and problem tracking.

10. The Web-based data mining system according to claim 9, characterized in that: The Web data source set includes text, images, audio, video and metadata; the Web-based data mining system also includes: An adaptive learning and optimization module, configured to adjust the cross-modal features in the common feature space according to the accuracy scores of the mining results of the data mining task and user feedback; The user-customized report module is used to generate user-customized data mining reports based on the data source, mining depth and result display method selected by the user.