Data analysis method, device and system
The data analysis method and system address the challenge of classifying large volumes of heterogeneous data by using data warehouses and semantic analysis to enhance classification accuracy and efficiency.
Patent Information
- Application Number
- GB2025001137
- Authority / Receiving Office
- GB · GB
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-15
- Filing Date
- 2025-01-02
- Publication Date
- 2026-01-28
AI Technical Summary
The challenge of accurately classifying and analyzing large volumes of heterogeneous and dynamic data is exacerbated by the lack of human-like logical thinking in computers, leading to difficulties in extracting useful information efficiently.
A data analysis method and system that utilizes data warehouses, reference data, and semantic analysis to classify data by matching data features, employing word splitting, keyword extraction, and feature vector comparison to enhance accuracy.
Enables precise data classification by leveraging semantic analysis and feature vectors to accurately categorize data, reducing misjudgment and improving the efficiency of data classification processes.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[0001] The present disclosure belongs to the technical field of data processing, and in particular to a data analysis method, a device thereof and a system thereof BACKGROUND
[0002] With the rapid development of Internet technology, the amount of data has exploded. There is a great value contained in these data. How to extract useful information accurately and quickly from massive, heterogeneous and dynamic data is a major challenge in the field of data mining.
[0003] The classification and analysis of data is an important aspect of data mining, which helps people understand the essential features and internal relations of data by dividing data into predefined categories. However, computers do not have the logical thinking ability of human beings, so that it is difficult to accurately analyze and classify data. SUMMARY
[0004] The purpose of the present disclosure is to provide a data analysis method, a device thereof and a system thereof. By matching and analyzing the unclassified data with the reference data of each category, the present disclosure realizes the accurate classification of the data.
[0005] In order to solve the above technical problems, the present disclosure is realized by the following technical scheme.
[0006] The present disclosure provides a data analysis method, including:
[0007] acquiring a category of classifying data,
[0008] acquiring a data warehouse of each category;
[0009] acquiring reference data of each of the data warehouses;
[0010] acquiring unclassified data;
[0011] obtaining data features of the data warehouse according to data features of classified data in the data warehouse and data features of the corresponding reference data;
[0012] acquiring and obtaining the category of each classified data according to the data features of unclassified data and the data features of each of the data warehouses.
[0013] The present disclosure further discloses a data analysis method, including:
[0014] establishing a data warehouse for each category of data;
[0015] receiving the classified data and the category of the classified data;
[0016] storing each classified data into the corresponding data warehouse according to the category.
[0017] The present disclosure further discloses a data analysis device, including:
[0018] a data warehouse reading interface, which is configured to acquire a category of classifying data;
[0019] acquire a data warehouse of each category;
[0020] acquire several reference data of each of the data warehouses,
[0021] an analysis service input interface, which is configured to acquire unclassified data;
[0022] an arithmetic unit, which is configured to obtain data features of the data warehouse according to data features of classified data in the data warehouse and data features of the corresponding reference data;
[0023] acquire and obtain the category’ of each classified data according to the data features of unclassified data and the data features of each of the data warehouses;
[0024] an analysis sendee output interface, which is configured to output the category of each classified data.
[0025] The present disclosure further discloses data analysis system, including:
[0026] a data analysis device, which is configured to output the category of each classified data; and
[0027] a storage unit, which is configured to establish a data warehouse for each category of data;
[0028] receive the classified data and the category of the classified data;
[0029] store each classified data into the corresponding data warehouse according to the category.
[0030] According to the present disclosure, the data features of each data warehouse are obtained by analyzing the reference data and the classified data of each data warehouse, and then are compared with the data features of the unclassified data, so that the accurate classification of the data is realized.
[0031] Of course, it is not necessary for any product implementing the present disclosure to achieve all the advantages mentioned above at the same time. BRIEF DESCRIPTION OF THE DR.W1NGS
[0032] In order to explain the technical scheme of the embodiment of the present disclosure more clearly, the drawings needed for the description of the embodiment wall be briefly introduced hereinafter. Obviously, the drawings in the following description are only some embodiments of the present disclosure. For those skilled in the art, other drawings can be obtained according to these drawings without paying creative labor.
[0033] FIG. 1 is a schematic diagram of a functional unit and an information flow of a data analysis system according to an embodiment of the present disclosure.
[0034] FIG. 2 is a schematic diagram of a step flow of a data analysis device according to an embodiment of the present disclosure.
[0035] FIG. 3 is a schematic diagram of a step flow of a storage unit according to an embodiment of the present disclosure.
[0036] FIG. 4 is a schematic diagram of a step flow of Step S5 according to an embodiment of the present disclosure.
[0037] FIG. 5 is a schematic diagram of a step flow of Step S52 according to an embodiment of the present disclosure.
[0038] FIG. 6 is a schematic diagram of a step flow of Step S528 according to an embodiment of the present disclosure.
[0039] FIG. 7 is a schematic diagram of a step flow of Step S55 according to an embodiment of the present disclosure.
[0040] FIG. 8 is a schematic diagram of a step flow of Step S553 according to an embodiment of the present disclosure. [004.1] FIG. 9 is a schematic diagram of a step flow' of Step S6 according to an embodiment of the present disclosure.
[0042] In the attached drawings, the list of components represented by each reference number is as follows:
[0043] I -data warehouse reading interface, 2-analysis sendee input interface, 3-arithmetic unit, 4-analysis sendee output interface, and 5-storage unit. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0044] In order to make the purpose, the technical scheme and the advantages of the present disclosure more clear, the embodiment of the present disclosure will be further described in detail with reference to the attached drawings.
[0045] It should be noted that the terms "first" and "second" in the present disclosure are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present disclosure described herein can be implemented in other orders than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. Rather, the embodiments are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0046] Computer data analysis refers to the use of computers and related technologies for data processing, analysis and mining. The process of computer data analysis usually involves the use of computer software and tools to collect, clean, transform and analyze a large amount of data. It is difficult to conduct manual review of computer data one by one due to the large amount of computer data. At the same time, due to the limitation of the computer identification technology, automatic computer classification is prone to misjudgment. In order to improve the accuracy of data classification, the present disclosure provides the following scheme.
[0047] As shown in FIG. 1 to FIG. 3, the present disclosure provides a data analysis system, which can provide users with data analysis services, thus realizing accurate judgment of data categories. According to the interactive functions, the data analysis system includes a data warehouse reading interface 1, an analysis service input interface 2, an arithmetic unit 3, an analysis service output interface 4 and a storage unit 5. The data warehouse reading interface 1, the analysis service input interface 2, the arithmetic unit 3 and the analysis service output interface 4 cooperate to output the category of each classified data, and the storage unit 5 is configured to store the classified data.
[0048] In the specific data analysis process, first, Step SI is executed by the data warehouse reading interface 1 to acquire a category'- of classifying data. Next, Step S2 can be executed to acquire a data warehouse of each category-. Next, Step S3 can be executed to acquire several reference data of each of the data warehouses. The reference data here can be an exemplary sample selected by the staff for the purpose of giving the system a sample for classified comparison.
[0049] After that, the analysis service input interface 2 executes Step S4 to acquire unclassified data. The arithmetic unit 3 executes Step S5 to obtain data features of the data warehouse according to data features of classified data in the data warehouse anti data features of the corresponding reference data. Next, Step S6 can be executed to acquire and obtain the category' of each classified data according to the data features of unclassified data and the data features of each of the data warehouses. Finally, the analysis service output interface 4 outputs the categoiy of each classified data. Of course, in actual operation, the classified data itself can also be synchronously output by the analysis service output interface 4.
[0050] In the process allowed by the storage unit 5, first, Step SOI can be executed to establish a data warehouse for each category of data. Next, Step S02 can be executed to receive the classified data and the category of the classified data. Finally, Step S03 can be executed to store each classified data into the corresponding data warehouse according to the category'. The data warehouse here can be an aggregation concept, that is, the classified data with the same category are stored in a centralized way. Alternatively, the data warehouse can be an abstract concept, that is, the existence of data is not differentiated, but each classified data is labeled according to the category. Both of the aggregation concept and the abstract concept are feasible and belong to the scope of protection of this scheme.
[0051] As shown in FIG. 4, the reference data corresponding to the data warehouse is provided with many feature dimensions, but from the perspective that human beings utilize data, the semantics of the data should be taken into account first. Therefore, this scheme classifies the data according to the semantics contained in the data. In other words, the semantics with distinguishing features in the reference data are regarded as the data features corresponding to the data warehouse. In the specific operation, first, Step S51 can be executed to perform word splitting on each of the reference data to obtain split words in each of the reference data and the corresponding number. Next, Step S52 can be executed to obtain keywords in each of the reference data and the word frequencies of the keywords according to the split words in each of the reference data and the corresponding number. Next, Step S53 can be executed to take the keywords of the reference data corresponding to each of the data warehouses as the keywords of the classified data in the data warehouse. Next, Step S54 can be executed to perform word splitting on the classified data in each of the data warehouses to obtain word frequencies of the keywords of each classified data in each of the data warehouses. Finally, Step S55 can be executed to obtain an effective distribution range of word frequencies of each keyword in the classified data in the data warehouse as data features of the data warehouse according to the keywords in each of the reference data and the word frequencies of the keywords, and the word frequencies of the keywords of each classified data in each of the data warehouses. The semantics with distinguishing features are embodied as the effective distribution range of word frequencies of each keyword, which can be operated and compared by a computer.
[0052] As shown in FIG. 5, because the number of split words generated after semantic splitting of the reference data is too large, in order to select the important and representative split words as the data features, first, Step S521 can be executed to acquire each split word in all the reference data. Next, Step S522 can be executed to acquire the frequency of occurrence of each split word in all the reference data. Next, Step S523 can be executed to acquire the cumulative frequency of occurrence of all split words in ail the reference data. Next, Step S524 can be executed to take the ratio of the frequency of occurrence of each split word in all the reference data to the cumulative frequency of occurrence of all split words as a global word frequency of each split word. Next, Step S525 can be executed to acquire the cumulative frequency of occurrence of ail split words in each of the reference data. Next, Step S526 can be executed to acquire the frequency of occurrence of each split word in each of the reference data. Next, Step S527 can be executed to take the ratio of the frequency of occurrence of each split word in each of the reference data to the cumulative frequency of occurrence of all split words in the reference data as an internal word frequency of each split word in each of the reference data. Finally, Step S528 can be executed to obtain keywords in each of the reference data, and the word frequencies of the keywords according to the internal word frequency of each split word in each of the reference data and the corresponding global word frequency.
[0053] In order to supplement the implementation process of the above Step S521 to Step S528, the source codes of some functional modules are provided, and comparative explanations are carried out in the comments. In order to avoid data leakage involving trade secrets, desensitization i s carried out for some data that will not affect the implementation of the scheme. The same applies hereinafter. ^include <iostream> #include <vector> #include <map> #include <string> ^include <sstream> ^include <iomanip> / / used to set the output format / / define the function of splitting the string std::vector<std::string> split(const stdxstring &text, char delim) { std: :vector<std: :string> tokens; std:: string token; std: :i stringstream token Streamttext); while (std::getline(tokenStream, token, delim)) { if ('token.empty()) { tokens.push back(token); i 1 return tokens, / / define the function of calculating the word frequency std::map<std::string, int> calculateWordFrequency(const std::vector<std::string>& words) { std: :map<std:: string, int> wordFreq; for (const std::string& word : words) { wordF req [word]++; return wordFreq; int main() { / / example reference data std::vector<std::string> referenceData = { "apple banana apple chew banana", "banana apple grape banana cherry", "cherry banana apple grape” }; std: :map<std:: string, int> global Frequency: std: :vector<std: :map<std:: string, float» intemalFrequencies; int totalWordCount = 0; / 7 calculate the global word frequency for (const std::string& text: referenceData) { auto words = split(text,''); auto wordFreq = calculateWordFrequency(words); for (const auto& pair : wordFreq) { globalFrequency[pair.firsf] • pair.second, totalWordCount += pair, second; } } / / calculate the internal word frequency of split words in each of the reference data for (const std::string& text • referenceData) { auto words = split(text,''); auto wordFreq = calculateWordFrequency(w'ords); int totalWordsInReference = 0; for (const auto& pair : wordFreq) { totalWordsInReference += pair, second; } std: :map<std::string, float> internalFrequency; for (const auto& pair : wordFreq) { internalFrequency [pai r. fi r st ] = stai i c_cast<float>(pai r second) total Word slnReference; intemalFrequencies.push_back(internalFrequency); / 7 output the keywords of each of the reference data and the word frequencies of the keywords for (sizej i = 0, i <internal Frequencies. size(); ++i) { std::cout « “Reference Data ’’ « i + 1 « ":\n”; for (const auto& pair : internal Frequencies^]) { std::cout« "Keyword: ” « std::setw(10) « pair.first « ", Internal Frequency: "« std::setw(5) « pair.second « ", Global Frequency: "« std::setw(5) « static_cast<float>(globalFrequen^^ total WordCount«'\n'; stdzcout«'\n'; return 0; I
[0054] This code first reads a set of reference data strings, and splits and counts the words in each string. Thereafter, the global word frequency of each word (frequency of occurrence in all the reference data) and the internal word frequency of each word in each of the reference data text (frequency of occurrence in the specific reference data) are calculated. .Finally, the code outputs the internal word frequency and the global word frequency of each keyword in each reference data set. These data can be used to identify keywords and their importance.
[0055] As shown in FIG. 6, because the importance of each split word is different, and the number of keywords in the reference data is limited, the importance of keywords is significantly higher than that of other split words. In view of this, in the process of acquiring the keywords of each of the reference data, first, Step S5281 can be executed to acquire the ratio of the internal word frequency of each split word to the corresponding global word frequency as word frequency coefficients of each split word. Next, Step S5282 can be executed to arrange the word frequency coefficients of each split word according to the magnitude of numerical values to obtain a coefficient list. Next, Step S5283 can be executed to acquire an average value of the difference between each word frequency coefficient and an adjacent word frequency coefficient in the coefficient list as a coefficient average difference. Next, Step S5284 can be executed to start from the word frequency coefficient with a maximum value in the coefficient list and calculate the difference with the adjacent smaller word frequency coefficient in sequence, and judge whether the difference is greater than the coefficient average difference. If so, next, Step S5285 can be executed to stop execution. If not, next, return to execute Step S5284 to continuously execute the step of starting from the word frequency coefficient with a maximum value in the coefficient list and calculating the difference with the adjacent smaller word frequency coefficient in sequence, and judging whether the difference is smaller than the coefficient average difference. Next, Step S5286 can be executed to take the split word corresponding to the word frequency coefficient participating in the calculation as the keyword. Finally, Step S5287 can be executed to summarize and acquire the keywords in each of the reference data and the word frequencies of the keywords. (0056] In order to supplement the implementation process of the above Step S5281 to Step S5287, the source codes of some functional modules are provided, and comparative explanations are carried out in the comments. #include <iostream> #include <vector> ^include <map> ^include <string> ^include <algorithm> / / used for sort and other algorithms / 7 function of splitting the string std::vector<std::string> split(const stdxstring &str, char delim) { std::vector<std::string> elements, std:: stringstream ss(str); std::string item; while (std::getline(ss, item, delim)) { elements.push back(item); i return elements; } H function of calculating the word frequency std::map<std::string, int> calculateFrequency(const std::vector<std::string> &data) { std: :map<std::string, int> freq; for (const auto &word : data) { freq [word]++; return freq; / / the word frequency coefficient structure struct WordCoefficient { std::string word; float coefficient; / / main program int mainO ( / 7 / example reference data std::vector<std::string> referenceData = { "apple banana apple cherry'", "banana apple grape banana cherry", "cherry banana apple grape" std: :map<std:: string, int> globalFrequency; std: :vector<std: :map<std:: string, int» allFrequencies; / / calculate the global word frequency and the word frequency of each of the reference data for (const auto &data : referenceData) { auto words ::: split(data,1'); auto freq = cdculateFrequency(words); allFrequencies.push ^back(freq); for (const auto &entry : freq) { globalFrequency[entiy.first] += entry-', second; / / circulate each of the reference data for (size_t i = 0; i <referenceData.size(); ++i) { std::vector<WordCoefficient> coefficients, for (const auto &entry : allFrequencies[i]) { WordCoefficient wc; wc.word = entry.first; wc. coefficient = static jiast<float>(entry. second) gl obalFrequency [entry, fl r st], coeffi ci ents.push_back(wc); J / 7 sort according to coefficients std::sort(coefficients.begin(), coefficients.end(), [](const WordCoefficient &a, const WordCoefficient &b) { return a.coefficient >b.coefficient; }); / / calculate the coefficient average difference float averageDifference ::: O.Of; for (size^t j = 1; j <coefficients.size(); ++j) f averageDifference += coefficients!] - 1]. coefficient c oeffi ci ents [j ]. coeffi ci ent; averageDifference / = coefficients. sizeQ - 1; / 7 determine keywords std: wector<std: :string> keywords; for (size_tj = 0; j <coefficients.size() - 1; ++j) { if (coefficients!]], coefficient - coefficients!] + 1],coefficient > averageDifference) { key words. push_back( coeffi ci ents [j ]. word); break; / / if exceeding the average difference, stop execution J A / / output results std::cout« "Reference dataset ” « i + 1 « "keywords:\n"; for (const auto &keyword : keywords) { std::cout« keyword « "(Coefficient: "« allFrequencies'"cpp / / include <iostream> / / include <vector> / / include <map> / / include <algorithm> / / include <numeric> / / user-defined type used to store the word frequency and the corresponding global word frequency struct Word Freq ( std:: string word; int local freq; int global freq; }; / / user-defined sorting function used to sort according to the word frequency coefficient bool sortByCoefficient(const WordFreq& a, const WordFreq& b) { double coefa = a.localfreq / static_cast<double>(a.global_freq); double coef b ::: blocal freq i static cast<double>(b.global freq), return coef_a >coef_b; int main() { / 7 data of the internal word frequency of the split words and the corresponding global word frequency / / use static data for simulation here std::vector<WordFreq> referenceData = { {"apple", 2, 10}, {"banana", 3,5}, ("cherry", 1, 8} / 7 ... add more data }; / / calculate the word frequency coefficient of each split word for (auto& wf: referenceData) { wf.coefficient = static cast<double>(wf.local freq) 7 wf.global freq; J / 7 sort according to the word frequency coefficient std::sort(referenceData.begin(), referenceData.end(), sortByCoefficient), / 7 acquire the average value of the difference between each word frequency coefficient in the coefficient list and the adjacent word frequency coefficient / / calculate all differences first std: :vector<double> differences, for (size_t i = 0, i <referenceData.size() -1; ++i) { double diff = referenceData[i],coefficient - referenceDatafi + 1],coefficient; differences. push_back(diff); / / calculate the average value of the differences double meanDifference == std::accumulate(differences begm(), differences. endQ, 0.0) / differences.size(); / / filter keywords std: :vector<std: :string> keywords; for (size t i== 0; i <differences.size(), ~H-i) { if (differences^] >meanDifference) { key words .push back(referenceData[i ]. word); break; / / if the difference is greater than the average difference, stop execution / 7 output keywords and the word frequencies of the key words stdzcout « "Keywords and their local frequencies:" « stdxendl; for (const auto& keyword : keywords) { std::cout « keyword « ": "« referenceDatafi].local freq « std::endl; } return 0; }
[0057] This code first defines a structure WordFreq, which is used to store the text, the local word frequency and the global word frequency of each word. Thereafter, a standard library function std::sort is used to sort the word frequency coefficients, and the average value of the difference between the word frequency coefficients is calculated. The keywords are filtered according to whether the difference is greater than the average value. Finally, each keyword and the local word frequency of the keyword are output. The keywords are determined by comparing the word frequency coefficients and the average gap between these coefficients.
[0058] As shown in FIG. 7 and FIG.8, due to a single amount of reference data, relying solely on reference data to classify data may result in the classification and comparison scope being too narrow, causing a large amount of unclassified data to be unable to be effectively classified and become dirty data. Therefore, the classified data in the data warehouse is also needed for auxiliary' classification, that is, the data features of the classified data are extracted synchronously as the data features of the data warehouse. Specifically, for each of the data warehouses, Step S551 can be executed to acquire the word frequencies of the key words of each classified data in the data warehouse. Next, Step S552 cart be executed to arrange the keywords of each classified data in the same order to obtain a multi-dimensional vector consisting of the numerical values of the word frequencies of the keywords of each classified data as the feature vector of each classified data. The classified data here can be selected by the staff and injected into the data warehouse, or can be classified data after classification.
[0059] Next, Step S553 can be executed to obtain a plurality of distribution ranges of the word frequency of each key word in the classified data in the data warehouse according to the feature vector of each classified data. In this process, first, Step S5531 can be executed to select several feature vectors of a plurality of classified data as target feature vectors. Next, Step S5532 can be executed to calculate and acquire a vector difference modulus length of each target feature vector and each other feature vector. Next, Step S5533 can be executed to form a vector set by each other feature vector and the target feature vector with the smallest vector difference modulus length. Next, Step S5534 can be executed to calculate and acquire a feature vector with the smallest vector difference modulus length from a mean vector of al l the feature vectors in each vector set as an updated target feature vector. Next, Step S5535 can be executed to judge whether the updated target feature vector changes. If so, next, Step S5532 to Step S5535 can be executed to return to continuously update the vector set and the target feature vector. If not, next, Step S5536 can be executed to acquire the distribution range of the word frequency of each keyword of the classified data corresponding to all the feature vectors in each vector set. That is to say, the classification and comparison scope of the data features of the classified data are extended
[0060] Because the classified data in the data warehouse may not strictly meet the target requirements, the classified data needs to be calibrated with reference data. Therefore, finally, Step S554 can be executed io take the distribution range where the keywords in the reference data and the word frequencies of the keywords are located as the effective distribution range of word frequencies of each keyword in the classified data in the data warehouse.
[0061] In order to supplement the implementation process of the above Step S5531 to Step S5536, the source codes of some functional modules are provided, and comparative explanations are carried out in the comments. #include <iostream> ^include <vector> #include <cmath> include <limits> ^include <algorithm> / 7 user-defined type used to store keywords and the word frequencies of the keywords struct KeywordFreq { std::string keyword; std::vector<int> freqs; / 7 the word frequencies of the same keyword in different data i ■ / , / / user-defined type used to store the feature vectors struct Feature Vector ’ std::vector<int> features; H feature vectors }; / / calculate the Euclidean distance between two feature vectors double euciideanDistance(const FeatureVector& a, const FeatureVector& b) f double distance ::: 0.0; for (size_t i = 0; i <a.features.size / ); ++1) { distance += std: :pow<static cast<double>(a.fe^ - b.features[i]), 2); return std:: sqrt(distance); / / calculate the mean vector of the feature vector set Feature Vector calculateMeanVector(const std::vector<FeatureVector>& vectors) { F eature Vector mean Vector; if (' vectors, empty / )) { meanVector.features.resize / vectorsfO].features.size / ), 0); for (const auto& vec : vectors) { for (size_t i = 0, i <vec.features.size / ), ++i) { meanVector.featuresfi] += vec.features[i]; r j } for (size J i = 0; i <mean Vector.features. size / ); ++i) { meanVector.featuresfi] / = vectors, size / ); } return mean Vector; } int main / ) { / 7 keywords and the word frequencies of the key words in different data have been extracted and stored in the KeywordFreq structure std::vector<KeywordFreq> keywordData = { {“apple”, {1, 5, 3}}, {"banana", {2, 1, 0}}, {"cherry”, {3,4,2}} / 7 ... add more data x- i , / / initialize the target feature vector (here, simply select the first few target feature vectors as the initial target feature vectors) std: :vector<FeatureVector> target Vectors; for (size t i==: 0; i <keywordData si.ze() &&i <3, ++i) { / / the number of target vectors is 3 target Vector s. pu sh b ack( {key wordData[i ]. freq s}); bool changed; do { / / classify all feature vectors according to the target feature vectors std::vector<std::vector<FeatureVector» chtsters( target Vectors, size})); for (const auto& keyword : keywordData) { double minDistance ::: std: .numeric limits<double>::max(); size J clusterindex == 0; for (size J i = 0; i <target Vectors. size(); ++i) { double distance == euclideanDistance({keyword. freqs}, target Vectors}!]); if (distance <minDistance) { minDistance == distance; clusterindex = i; } } clusters[clusterlndex], push _back({key word, freqs}); } / / calculate a new target feature vector of each set changed = false; for (size t i = 0; i <clusters.size(); ++i) { FeatureVector newMean Vector == calcul ateMean Vector(clusters}i J); if (1 std: :equal(newMeanVector.features.beginO, newMeanVector.features.end(), targetVectorsfi].features.begin())) { targetVectors[i] = new Mean Vector; changed = true; } white (changed); / 7 acquire the distribution range of word frequencies corresponding to the feature vectors in each set for (size ;? i = 0; i <keywordData.sizei); ++i) { auto& freqs::: keywordDatafi] freqs; std;:cout « "Keyword: "« key wordDatafi], key word « ", Freq Range: "; stdzcout « *std::min element(freqs.begin(), freqs.end()) « "”cpp *std::max_element(freqs.beginO, freqs.endQ) « stdzendl; f return 0; I
[0062] During the above code operation, the data are organized into keywords and the word frequencies of the keywords in different data items. The initial target feature vector is simply selected as the first items in the data set. Thereafter, the code enters a loop, constantly updating the target feature vectors until they no longer change. In the loop process, the Euclidean distance is used to classify feature vectors, and a new mean vector of each classification is calculated. Once the target feature vector stabilizes, the loop ends. Finally, the code calculates and outputs the distribution range of the word frequencies of each keyword.
[0063] As shown in FIG. 9, in the process of comparing the unclassified data, with the data features of the data warehouse, in order to reduce the calculation amount of comparison, first, Step S61 can be executed to perform word splitting on the unclassified data to obtain split words in the unclassified data and the corresponding number. Next, Step S62 can be executed to judge whether the split words in the unclassified data cover the keywords in the data features of the data warehouse. If not, next, Step S63 can be executed to perform no process. If so, next, Step S64 can be executed to take the data warehouse as an alternative data warehouse. This can avoid comparing the data features of each of the data warehouses, and reduce the calculation amount of comparison without losing the comparison and classification accuracy.
[0064] In the process of comparing the unclassified data with each alternative data warehouse, first, Step S65 can be executed to take the keywords of the alternative data warehouse as keywords of the unclassified data, and take each keyword of the unclassified data and the word frequency of the keyword as data features. Next, Step S66 can be executed to judge whether the word frequency of each keyword of the unclassified data falls within the effective distribution range of the word frequencies of each keyword in the classified data in the alternative data warehouse. If so, Step S67 can be executed to label the unclassified data as the classified data, and take the category of the alternative data warehouse as the category of the classified data. If not, next. Step S65 to Step S67 can be executed to compare with the next alternative data warehouse.
[0065] In order to supplement the implementation process of the above Step S61 to Step S67, the source codes of some functional modules are provided, and comparative explanations are carried out in the comments. #include <iostream> #include <string> #i nclude <unordered_m ap> ^include <vector> ^include <algori thm> / / data feature structure struct DataFeature { std::string category; / 7 data category' std::unordered__map<std::string, std::pair<int, int» keywordFreqRange; H effective distribution range of the word frequencies of the keywords i ■ / / split words of unclassified data, and calculate the word frequency std::unordered map<std::string, int> tokenizeAndCountFreq(const std::string& data) { std:: unordered_map<std:: s tri ng, int> wordCounts; / . / splitFunction is a function used to split strings / / std::vector<std::string> words = splitFunction(data); / 7 the example simplified method used here needs to be replaced by the real word splitting method in practical application size__t prev = 0, pos = 0; do { pos := data.fmd(" ”, prev), if (pos == std::string::npos) pos = data.length / ); std::string word = data.substr(prev, pos-prev); if (Iword.emptyO) wordCounts[word]++; prev = pos + 1; } while (pos <data.length() &&prev <data.length()); return wordCounts; } / / check whether all keywords exist in the data warehouse feature bool are AllKey wordsCovered(const std::unordered_map<std:: string, int>& wordCounts, const DataFeature& dataFeature) { for (const auto& wordCount. wordCounts) { if (dataFeature.keywordFreqRange.find(wordCount.first) == dataFeature.keywordFreqRange.endO) { return fal se, r return true; / / judge whether the word frequency falls within the effective distribution range bool isFreqInRange(const std:.unordered map<std::string, int>& wordCounts, const DataFeature& dataFeature) { for (const auto& wordCount: wordCounts) { auto rangelt = dataFeature.keywordFreqRange.find(wordCount.first); if (range h != dataFeature.keywordFreqRange.endQ) { const auto& range = rangelt->second; if (wordCount.second <range.first l| wordCount.second >range.second) { return false; / / the word frequency falls out of range I J .( return true, H all word frequencies fall within the range } int main() { H unclassified data std::string unclassifiedData = "apple banana apple cherry’"; / / data features of the data warehouse std::vector<DataFeature> dataFeatures = { {"Fruit", {{"apple", {1, 3}}, {"banana", {1,2)}, {"cherry", {0,2}}}}, / 7 add more data feature categories and word frequency ranges t-J f H count the word frequencies of the unclassified data auto wordCounts = tokenizeAndCountFreqfunclassifiedData); / 7 traverse each of the data warehouses to search for the matched category std::string category = "Unclassified"; H default category7 for (const auto& dataFeature . dataFeatures) { / 7 check whether all keywords exist in the data warehouse feature if (areAllKeywordsCovered( wordCounts, dataFeature)) { / 7 judge whether the word frequency falls within the effective distribution range if (isFreqInRange(wordCounts, dataFeature)) { category' = dataFeature.category'; ii the matched data category7 break; / / successful matching, exit the loop I J } / 7 output the final classification results stdvcout« "The unclassified data belongs to category: "« category « stdzendl; return 0; r
[0066] The above code first defines a structure DataFeature for storing data features and word frequency ranges. In the main function main, an unclassified data string unclassifiedData and an array of dataFeatures containing various data features are defined. tokenizeAndCountFreq function is used for unclassified data to split words and count the word frequency. Thereafter, the data features of each of the data warehouses are traversed, and areAHKeywordsCovered function is used to check whether the split words of unclassified data cover the keywords in the data features of the data warehouse. If so, the isFreqlnRange function is further used to judge whether the word frequency of each keyword in the unclassified data falls within the effective distribution range of the word frequencies of the corresponding keyword in the data warehouse. If all of the word frequencies fall within the range, the unclassified data is labelled as the classified data, and the category' of the data warehouse is taken as the category of the classified data. If the word frequencies fail out of range, continue to compare with the next data warehouse. If none of the data warehouses match, the data remains unclassified. Finally, the final classification result is output.
[0067] The flow charts and block diagrams in the drawings show the architectures, functions and operations of possible implementations of the device, the system, the method and the computer program product according to various embodiments of the present disclosure. In this regard, each block in the flow chart, or block diagram may represent a module, a program segment or a part of an instruction, which contains one or more executable instructions for implementing specified logical functions. In some alternative implementations, the functions noted in the blocks may also occur in a different order than those noted in the drawings. For example, two consecutive blocks may actually be executed substantially in. parallel, and may sometimes be executed in a reverse order, depending on the functions involved.
[0068] It should also be noted that each block in the block diagram and / or flow chart, and the combination of blocks in the block diagram and / or flow chart, can be realized by hardware that performs the corresponding functions or actions, such as a circuit or an ASIC (Application Specific Integrated Circuit), or by a combination of hardware and software, such as firmware.
[0069] Although the present disclosure has been described herein in connection with various embodiments, in the process of implementing the claimed present disclosure, those skilled in the art can understand and realize other variations of the disclosed embodiments by viewing the drawings, the present disclosure, and the appended claims. In the claims, the word "comprising" does not exclude other components or steps, and "a" or "one" does not exclude a plurality. A single processor or other units can realize several functions recited in the claims. Certain measures are recorded in mutually different dependent claims, but this does not mean that these measures cannot be combined to produce good results.
[0070] Embodiments of the present disclosure have been described above. The above description is exemplary, rather than exhaustive, and is not limited to the disclosed embodiments. Many modifications and changes will be obvious to those skilled in the art without departing from the scope of the illustrated embodiments. The terminology used herein is chosen to best explain the principles of various embodiments, practical application or improvement of the technology in the market, or to enable other ordinary people in the technical field to understand various embodiments disclosed herein.
Claims
WHAT IS CLAIMED IS:
1. Adata analysis method, comprising:acquiring a category of classifying data,acquiring a data warehouse of each category;acquiring reference data of each of the data warehouses;acquiring unclassified data;obtaining data features of the data warehouse according to data features of classified data in the data warehouse and data features of the corresponding reference data;acquiring and obtaining the category of each classified data according to the data features of unclassified data and the data features of each of the data warehouses.
2. The method according to claim I, wherein the step of obtaining data features of the data warehouse according to data features of classified data in the data warehouse and data features of the corresponding reference data comprises:performing word splitting on each of the reference data to obtain split words in each of the reference data and the corresponding number;obtaining keywords in each of the reference data and the word frequencies of the keywords according to the split words in each of the reference data and the corresponding number;taking the keywords of the reference data corresponding to each of the data warehouses as the keywords of the classified data in the data warehouse;performing word splitting on the classified data in each of the data warehouses to obtain word frequencies of the keywords of each classified data in each of the data warehouses;obtaining an effective distribution range of word frequencies of each keyword in the classified data in the data warehouse as data features of the data warehouse according to the keywords in each of the reference data and the word frequencies of the keywords, and the word frequencies of the keywords of each classified data in each of the data warehouses.
3. The method according to claim 2, wherein the step of obtaining keywords in each of the reference data and the word frequencies of the keywords according to the split words in each of the reference data and the corresponding number comprises:acquiring each split word in all the reference data;acquiring the frequency of occurrence of each split word in all the reference data.;acquiring the cumulative frequency of occurrence of all split words in all the reference data;taking the ratio of the frequency of occurrence of each split word in all the reference data tothe cumulative frequency of occurrence of all split words as a global word frequency of each split word;acquiring the cumulative frequency of occurrence of all split words in each of the reference data;acquiring the frequency of occurrence of each split word in each of the reference data.;taking the ratio of the frequency of occurrence of each split word in each of the reference data to the cumulative frequency of occurrence of all split words in the reference data as an internal word frequency of each split word in each of the reference data;obtaining keywords in each of the reference data and the word frequencies of the keywords according to the internal word frequency of each split word in each of the reference data and the corresponding global word frequency.
4. The method according to claim 3, wherein the step of obtaining key words in each of the reference data and the word frequencies of the keywords according to the internal word frequency of each split word in each of the reference data and the corresponding global word frequency comprises:for each of the reference data,acquiring the ratio of the internal word frequency of each split word to the corresponding global word frequency as word frequency coefficients of each split word,arranging the word frequency coefficients of each split word according to the magnitude of numerical values to obtain a coefficient list,acquiring an average value of the difference between each word frequency coefficient and an adjacent word frequency coefficient in the coefficient list as a coefficient average difference,starting from the word frequency coefficient with a maximum value in the coefficient list and calculating the difference with the adjacent smaller word frequency coefficient in sequence, and judging whether the difference is greater than the coefficient average difference,if so, stopping execution,if not, continuously executing the step of starting from the word frequency coefficient with a maximum value in the coefficient list and calculating the difference with the adjacent smaller word frequency coefficient in sequence, and judging whether the difference is smaller than the coefficient average difference,taking the split word corresponding to the word, frequency coefficient participating in the calculation as the keyword;summarizing and acquiring the keywords in each of the reference data and the word frequencies of the keywords.
5. The method according to claim 2, wherein the step of obtaining an effective distribution range of word frequencies of each keyword in the classified data in the data warehouse according to the keywords in each of the reference data and the word frequencies of the keywords, and the word frequencies of the keywords of each classified data in each of the data warehouses comprises:for each of the data warehouses,acquiring the word frequencies of the keywords of each classified data in the data warehouse, arranging the keywords of each classified data in the same order to obtain a multi-dimensional vector consisting of the numerical values of the word frequencies of the keywords of each classified data as the feature vector of each classified data,obtaining a plurality of distribution ranges of the word frequency of each keyword in the classified data in the data warehouse according to the feature vector of each classified data,taking the distribution range where the keywords in the reference data and the word frequencies of the keywords are located as the effective distribution range of word frequencies of each keyword in the classified data in the data warehouse.
6. The method according to claim 5, wherein the step of obtaining a plurality of distribution ranges of the word frequency of each keyw'ord in the classified data in the data warehouse according to the feature vector of each classified data comprises:selecting several feature vectors of a plurality of classified data as target feature vectors;calculating and acquiring a vector difference modulus length of each target feature vector and each other feature vector;forming a vector set by each other feature vector and the target feature vector with the smallest vector difference modulus length;calculating and acquiring a feature vector with the smallest vector difference modulus length from a mean vector of all the feature vectors in each vector set as an updated target feature vector;judging whether the updated target feature vector changes,if so, returning to continuously update the vector set and the target feature vector;if not, acquiring the distribution range of the word frequency of each keyword of the classified data corresponding to all the feature vectors in each vector set.
7. The method according to any one of claims 2 to 6, wherein the step of acquiring and obtaining the category of each classified data according to the data features of unclassified data and the data features of each of the data warehouses comprises:performing word splitting on the unclassified data to obtain split words in the unclassified data and the corresponding number;judging whether the split words in the unclassified data cover the keywords in the datafeatures of the data warehouse;if not, performing no process;if so, taking the data warehouse as an alternative data warehouse;in the process of comparing the unclassified data with each alternative data warehouse,taking the keywords of the alternative data warehouse as keywords of the unclassified data, and taking each keyword of the unclassified data and the word frequency of the keyword as data features,judging whether the word frequency of each keyword of the unclassified data falls within the effective distribution range of the word frequencies of each keyword in the classified data in the alternative data warehouse,if so, labelling the unclassified data as the classified data, and taking the category of the alternative data warehouse as the category of the classified data,if not, comparing with the next alternative data warehouse.
8. A data analysis method, comprising:establishing a data warehouse for each category' of data;receiving the classified data and the category of the classified data in the data analysis method according to any one of claims I to 7;storing each classified data into the corresponding data warehouse according to the category.
9. A data analysis device, comprising:a data warehouse reading interface, which is configured to acquire a category of classifying data;acquire a data warehouse of each category;acquire several reference data of each of the data warehouses;an analysis service input interface, which is configured to acquire unclassified data;an arithmetic unit, which is configured to obtain data features of the data warehouse according to data features of classified data in the data warehouse and data features of the corresponding reference data;acquire and obtain the category' of each classified data according to the data features of unclassified data and the data features of each of the data warehouses,an analysis service output interface, which is configured to output the category of each classified data.
10. A data analysis system, comprising:a data analysis device according to claim 9, which is configured to output the category' of each classified data; anda storage unit, which is configured to establish a data warehouse for each category of data; receive the classified data and the category' of the classified data;store each classified data into the corresponding data warehouse according to the category.
Citation Information
Patent Citations
Construction method and system for topic model category of city-level data warehouse
CN113849639A
New word classification method and device, electronic equipment and storage medium
CN116108180A
Text classification method and electronic equipment
CN116701616A
Data analysis method, device and system
CN117574243A
System and method for classification of spend data
US20230385765A1