A data mining system based on AI big model
By adopting AI large models in the data mining system, combining data type consistency coefficient, string matching degree and fuzzy matching similarity, the problem of inaccurate data screening and matching in the existing technology is solved, and more efficient data mining and output is achieved.
Patent Information
- Application Number
- CN202510259510.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-06
- Publication Date
- 2025-05-02
- Estimated Expiration
- 2045-03-06
AI Technical Summary
It is difficult for the prior art to achieve accurate screening of target data during data mining and accurate matching of data in the database with target data.
A data mining system based on AI large model is adopted, including data preprocessing module, feature extraction module and mining topic analysis module. By calculating the data type consistency coefficient, string matching degree, and fuzzy matching similarity, we determine the adaptability of the target topic to the data in the underlying database, and give priority to output data with the highest adaptability.
Improve the accuracy of data mining and output data, ensuring that the selected data is more consistent with the target topic.
Smart Images

Figure CN119739764B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data mining technology, and in particular to a data mining system based on an AI big model. Background Art
[0002] The data mining system of AI large models is usually based on machine learning, deep learning and other technologies, and processes large data sets to identify and extract valuable information. The goal of data mining is to extract potential patterns, trends, associations and knowledge from large amounts of complex data.
[0003] The existing patent discloses a big data mining system (CN107577771A), which includes a big data storage module, a data extraction module, a data inspection module, a data mining module, a result verification module, a data reporting module and a log module. The data extraction module extracts a data set that meets the user's needs from the big data storage module and sends the data set to the data inspection module. In the technology disclosed in the patent, it is difficult to accurately screen the target data and accurately match the data in the database with the target data during the data mining process. Summary of the invention
[0004] The main technical problem solved by the present invention is to provide a data mining system based on an AI big model, which solves the problems in the above-mentioned background technology.
[0005] To solve the above technical problems, according to one aspect of the present invention, more specifically, a data mining system based on an AI big model includes a data preprocessing module, a feature extraction module, and a mining topic analysis module;
[0006] The data preprocessing module is used to preprocess the original data to generate a basic database;
[0007] The feature extraction module is used to extract and sort features according to the degree of fit between the target subject and the data in the basic database;
[0008] The mining topic analysis module is used to determine the target topic and target keywords;
[0009] Wherein, the degree of adaptation is calculated by at least one of a type consistency coefficient, a string matching degree or a fuzzy matching similarity;
[0010] The feature extraction module determines the degree of adaptability between the selected target topic and the mined data in the basic database based on the data type consistency coefficient, string matching degree, and fuzzy matching similarity between the data in the basic database and the target topic data provided by the mining topic analysis module. , The specific steps are as follows: In the formula, Represents the type consistency coefficient between the target subject and the retrieved data in the basic database, Indicates the string matching degree between the target subject and the search data in the basic database. Represents the similarity of fuzzy matching between the target topic and the retrieved data in the base database.
[0011] Furthermore, the degree of adaptation The larger the value is, the more consistent the data information filtered and processed from the basic database is with the target subject, and the mining data with the largest adaptation coefficient between the target subject and the mining data in the basic database is output preferentially.
[0012] Furthermore, in the feature extraction module, the type consistency coefficient between the target subject and the search data in the basic database is The specific steps for obtaining are as follows: In the formula, A represents the label and classification number of the target subject data, and B represents the label and classification number of the search data in the basic database.
[0013] Furthermore, in the feature extraction module, the degree of string matching between the target subject and the search data in the basic database is The specific steps for obtaining are as follows: Where C represents all character strings of the target subject data, and D represents all character strings of the search data in the basic database.
[0014] Furthermore, in the feature extraction module, the similarity of fuzzy matching between the target subject and the search data in the basic database is The specific steps for obtaining are as follows: In the formula, represents the number of overlapping strings in C and D. Indicates the number of swaps that occur while matching character pairs.
[0015] Furthermore, the system also includes a rule definition module, which is used to control the data mining scope of the feature extraction module by defining screening conditions for mining data, and the screening conditions can be displayed through a visual interface.
[0016] Furthermore, the visualization interface is used to display the screening conditions set by the rule definition module and the specific feature contents arranged by the feature extraction module.
[0017] Furthermore, the mining topic analysis module is used to determine the target topic and target keyword of the filtered data according to the filtering conditions set by the rule definition module.
[0018] The present invention provides a data mining system based on an AI large model. Compared with the prior art, the present method has the following effects:
[0019] 1. The present invention determines the degree of adaptability between the screened target subject and the mined data in the basic database through the data type consistency coefficient between the data in the basic database and the target subject data provided by the mining subject parsing module, the degree of string matching, and the similarity of fuzzy matching between the two data, and gives priority to outputting the mining data with the highest adaptability as the output target, which can improve the accuracy of mining and outputting data.
[0020] 2. The present invention controls the data mining scope of the feature extraction module by defining the screening conditions of the mining data, and the screening conditions can be displayed through a visual interface. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 It is a schematic diagram of the structure of the present invention. DETAILED DESCRIPTION
[0022] In order to make the technical solution of the present invention clearer, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments. Example
[0023] like Figure 1 As shown, according to one aspect of the present invention, there is provided a data mining system based on an AI big model, including a data preprocessing module, a conversion task metadata parsing tool, a feature extraction module, a rule definition module, and a mining topic parsing module;
[0024] The data preprocessing module preprocesses the original file data and the data in the original database. The preprocessing includes removing duplicate data and processing missing data. The processed data constitute the basic database; the feature extraction module is used to extract features in the basic database according to the extraction objects fed back by the mining subject analysis module, and sort them according to feature similarity to mine the optimal feature data; the rule definition module is used to control the data mining scope of the feature extraction module by defining the filtering conditions of the mining data, and the filtering conditions can be displayed through the visual interface.
[0025] The mining topic analysis module is used to determine the target topic and target keywords of the filtered data according to the filtering conditions set by the rule definition module. The visual interface is used to display the filtering conditions set by the rule definition module and the specific feature content arranged by the feature extraction module. The degree of adaptation of the filtered target topic and the mined data in the basic database is determined by the consistency of data types, string matching degree, and fuzzy matching similarity between the data in the basic database and the target topic data provided by the mining topic analysis module, and the output target is given priority. This design can greatly improve the accuracy of mining and outputting data. Example
[0026] like Figure 1 As shown, the feature extraction module determines the degree of adaptability between the selected target topic and the mined data in the basic database based on the data type consistency coefficient, string matching degree, and fuzzy matching similarity between the data in the basic database and the target topic data provided by the mining topic analysis module. ,in: In the formula, It represents the adaptation coefficient between the selected target topic and the mined data in the basic database. Represents the type consistency coefficient between the target subject and the retrieved data in the basic database, Indicates the string matching degree between the target subject and the search data in the basic database. It represents the similarity of fuzzy matching between the target topic and the retrieved data in the basic database;
[0027] when The larger the value is, the more consistent the data information filtered and processed from the basic database is with the target subject, and the mining data with the largest adaptation coefficient between the target subject and the mining data in the basic database is output preferentially.
[0028] And to prove the relationship between the above formula G and h, d, and m, at least one of the type consistency coefficient h, string matching degree d, and fuzzy matching similarity m will be used to study the degree of consistency between the data screened and processed from the basic database and the target topic. So first:
[0029] 1) Establish a mathematical relationship between the commonality and adaptation degree G of the type consistency coefficient h, string matching degree d, and fuzzy matching similarity m. Then we can have: in, It is used to express the mathematical relationship between the type consistency coefficient h and the degree of adaptation G. Indicates the value of the type consistency coefficient. And by analogy, we can establish a mathematical relationship between the string matching degree d and the adaptation degree G , and the mathematical relationship between the fuzzy matching similarity m and the adaptation degree G .
[0030] 2) After combining the type consistency coefficient h, string matching degree d, and fuzzy matching similarity m, a mathematical relationship is constructed with the adaptation degree G. It can be found that: Among them, X and Y represent two numerical values of type consistency coefficient h, string matching degree d, and fuzzy matching similarity m that are combined in pairs.
[0031] 3) Then, mathematical relationships are established between the three variables of type consistency coefficient h, string matching degree d, and fuzzy matching similarity m and the adaptation degree G, and we can know that: Then there will be: Therefore, the adaptation degree G can be calculated by at least one of the type consistency coefficient h, the string matching degree d, and the fuzzy matching similarity m.
[0032] The type consistency coefficient between the target subject and the retrieved data in the basic database in the feature extraction module The search for: In the formula, It represents the type consistency coefficient between the target subject and the search data in the basic database. A represents the number of labels and classification numbers owned by the target subject data, and B represents the number of labels and classification numbers owned by the search data in the basic database.
[0033] The degree of string matching between the target subject and the search data in the basic database in the feature extraction module , where: In the formula, It indicates the degree of string matching between the target subject and the search data in the basic database. C indicates all the strings of the target subject data, and D indicates all the strings of the search data in the basic database.
[0034] The similarity of fuzzy matching between the target subject and the search data in the basic database in the feature extraction module , where: In the formula, Represents the similarity of fuzzy matching between the target subject and the retrieved data in the basic database, represents the number of overlapping strings in C and D. Indicates the number of swaps that occur while matching character pairs.
[0035] When matching the target subject X data with the data Y mined from the basic database, the type consistency coefficient between the target subject and the retrieved data in the basic database is The degree of string matching between the target subject and the search data in the basic database is The similarity between the target topic and the search data in the basic database is obtained by fuzzy matching. , then we have: From the above calculation, we can know that the adaptation coefficient between the target subject X and the data Y is .
[0036] Then, to match multiple sets of data with the target topic X, we have:
[0037] Table 1 Relationship between parameters and adaptation coefficients of some embodiments
[0038] Type consistency coefficient h String matching degree d Fuzzy matching similarity m Adaptation coefficient G Implementation data 1 64% 63% 55% 2.2 Implementation data 2 50% 50% 50% 1.8 Implementation data 3 40% 40% 40% 1.5 Implementation Data 4 30% 30% 30% 1.3
[0039] It can be seen from Table 1 above that the mining data with a larger adaptation coefficient G has a higher degree of consistency with the target subject X data, and can be mined and outputted first.
[0040] Example 3
[0041] Duplicate data can be removed based on similarity. Duplicate data based on similarity is determined by calculating the similarity between data items to determine whether they are duplicate records. Similarity measures include Jaccard similarity, Cosine similarity, and Jaro-Winkler distance. Selecting the similarity measure and threshold according to your needs can effectively remove duplicate items in the data.
[0042] Jaccard similarity: Jaccard similarity measures the ratio of the intersection to the union of two sets. Its value is between 0 and 1. The larger the value, the higher the similarity. It is suitable for deduplication of sets or keyword matching, such as processing short text, tags, and keywords. You can use TfidfVectorizer or CountVectorizer in sklearn to calculate Jaccard similarity.
[0043] Cosine similarity: Cosine similarity measures the cosine value of the angle between two vectors. It is used to measure the similarity between texts and is particularly suitable for processing large amounts of text data. It is suitable for similarity comparison of long texts, articles, and sentences.
[0044] Jaro-Winkler distance: Jaro-Winkler distance is an improved string similarity measure specifically for short strings (such as names). It considers not only character matches but also the order of characters. You can use the jellyfish library to calculate the Jaro-Winkler distance. Example
[0045] The features extracted by the feature extraction module include data type consistency between data, string matching degree, and fuzzy matching similarity, edit distance, Jaccard similarity, Cosine similarity and Jaro-Winkler distance between two data. Example
[0046] Methods for controlling the scope of data mining: In data mining, the size of the search space usually determines the complexity of the algorithm. Too large a search space may lead to inefficient algorithms, waste of computing resources, or even failure to obtain effective mining results. Therefore, it is necessary to limit the search space:
[0047] Reduce the scope through data screening: Filter the data set according to time, region, specific conditions, etc. Limitation based on domain knowledge: With the help of the knowledge of domain experts, limit the scope of data mining, focus only on specific fields or problems, and avoid involving irrelevant data and features.
[0048] Scenario 1: Data mining does not have to start from the global perspective, but can gradually converge to valuable information through local exploration. This method can be achieved in the following ways:
[0049] Clustering and grouping: Divide the data into several clusters and conduct in-depth mining in only one cluster at a time. For example, when segmenting customers, analyze them separately according to different customer types to avoid irrelevant data across fields or types that affect the mining effect.
[0050] Multi-stage mining: First, a rough screening is performed in a larger range to find potentially valuable data, and then further analysis is performed on the finely screened data set. This phased local exploration can avoid the high complexity of global calculations.
[0051] Scenario 2: When faced with extremely large-scale data, full calculation is often not practical. Data sampling methods can replace the entire data with a smaller sample set for analysis without losing information:
[0052] Simple random sampling: randomly select samples from the data set so that each sample has an equal probability of being selected.
[0053] Stratified sampling: Divide the data into different levels according to certain characteristics of the data (such as category labels, regions, etc.), and extract samples within each level, so as to ensure that samples of different categories or intervals are representative.
[0054] Scenario 3: In data mining, certain thresholds are usually set to filter out unimportant or irrelevant information. Threshold settings can help control the scope of data mining and avoid the generation of irrelevant patterns:
[0055] Minimum support: In association rule mining, support indicates the frequency of an item set in a data set. By setting a minimum support threshold, you can avoid mining some rarely occurring item sets, thereby narrowing the mining scope.
[0056] Minimum confidence: Confidence indicates the reliability of a rule. Setting a minimum confidence threshold can avoid generating rules with low confidence and little significance.
[0057] Decision threshold in classification: In classification tasks, the range of classification can be controlled by setting a probability threshold of the classification result (such as 0.5). Only when the predicted probability exceeds a certain value, the sample is classified into a certain category.
[0058] Scenario 4: Gradual screening Gradually limit the scope of mining through a refined data filtering process to ensure the accuracy and business relevance of the mining results:
[0059] Rule-based filtering: Limit the scope of data through pre-set rules. For example, in purchase behavior data, only focus on users who have purchased specific products, or in user behavior data, only focus on active users.
[0060] Feature engineering: Through reasonable feature selection or construction, irrelevant features are excluded and the complexity of the feature space is reduced, thereby reducing unnecessary search scope.
[0061] The above-mentioned embodiments only express several implementation methods of the present invention, and the description thereof is relatively specific and detailed, but it cannot be understood as limiting the scope of the patent of the present invention. It should be pointed out that, for ordinary technicians in this field, several variations and improvements can be made without departing from the concept of the present invention, which all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention shall be subject to the attached claims.
Claims
1. A data mining system based on AI big model, characterized in that: It includes data preprocessing module, feature extraction module and mining topic analysis module; The data preprocessing module is used to preprocess the original data to generate a basic database; The feature extraction module is used to extract and sort features according to the degree of fit between the target subject and the data in the basic database; The mining topic analysis module is used to determine the target topic and target keywords; The feature extraction module determines the degree of adaptability between the selected target topic and the mined data in the basic database based on the data type consistency coefficient, string matching degree, and fuzzy matching similarity between the data in the basic database and the target topic data provided by the mining topic analysis module. , the specific steps are as follows: In the formula, Represents the type consistency coefficient between the target subject and the retrieved data in the basic database, Indicates the string matching degree between the target subject and the search data in the basic database. It represents the similarity of fuzzy matching between the target topic and the retrieved data in the basic database; In the feature extraction module, the type consistency coefficient between the target subject and the search data in the basic database The specific steps for obtaining are as follows: In the figure, A represents the label and classification number of the target subject data, and B represents the label and classification number of the search data in the basic database; In the feature extraction module, the degree of string matching between the target subject and the search data in the basic database is The specific steps for obtaining are as follows: In the formula, C represents all the character strings of the target subject data, and D represents all the character strings of the search data in the basic database; In the feature extraction module, the similarity of fuzzy matching between the target subject and the search data in the basic database is The specific steps for obtaining are as follows: In the formula, represents the number of overlapping strings in C and D. Indicates the number of swaps that occur while matching character pairs.
2. The data mining system based on AI big model according to claim 1 is characterized in that: The degree of adaptation The larger the value is, the more consistent the data information filtered and processed from the basic database is with the target subject, and the mining data with the largest adaptation coefficient between the target subject and the mining data in the basic database is output preferentially.
3. The data mining system based on AI big model according to claim 1 is characterized in that: The system also includes a rule definition module, which is used to control the data mining scope of the feature extraction module by defining screening conditions for mining data, and the screening conditions can be displayed through a visual interface.
4. The data mining system based on AI big model according to claim 3 is characterized by: The visualization interface is used to display the screening conditions set by the rule definition module and the specific feature contents arranged by the feature extraction module.
5. The data mining system based on AI big model according to claim 3 is characterized by: The mining topic analysis module is used to determine the target topic and target keyword of the filtered data according to the filtering conditions set by the rule definition module.
Citation Information
Patent Citations
Big data mining system
CN107577771A
Video efficient retrieval system supporting fuzzy comment mining
CN113656641A
Support vector machines processing system
US20050049990A1