Data analysis method, device and system
By obtaining the reference data characteristics and classified data characteristics of the data warehouse and combining the characteristics of the unclassified data, it solves the problem that computers find it difficult to accurately classify massive, heterogeneous, and dynamically changing data, and realizes the accurate classification and efficient utilization of data.
Patent Information
- Application Number
- PCT/CN2025/070183
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-15
- Filing Date
- 2025-01-02
- Publication Date
- 2025-07-24
AI Technical Summary
It is difficult for computers to accurately classify and analyze massive, heterogeneous, and dynamically changing data, and the existing technology cannot effectively utilize useful information in the data.
By obtaining the reference data characteristics and classified data characteristics of the data warehouse, and combining the characteristics of the unclassified data, the accurate classification of unclassified data can be achieved.
Accurate classification of data is achieved, and the accuracy and efficiency of data classification are improved.
Smart Images

Figure CN2025070183_24072025_PF_FP_ABST
Abstract
Description
Data analysis method, device and system Technical Field
[0001] The present invention belongs to the field of data processing technology, and in particular relates to a data analysis method, device and system. Background Art
[0002] With the rapid development of Internet technology, the amount of data has exploded. This data contains enormous value. How to accurately and quickly extract useful information from massive, heterogeneous, and dynamically changing data is a major challenge facing the field of data mining.
[0003] Data classification analysis is a key aspect of data mining. It helps people understand the essential characteristics and inherent relationships of data by classifying it into predefined categories. However, computers lack the logical thinking ability of humans and find it difficult to accurately analyze and classify data. Summary of the Invention
[0004] The object of the present invention is to provide a data analysis method, device and system, which achieve accurate classification of data by matching and analyzing unclassified data with reference data of each general category.
[0005] To solve the above technical problems, the present invention is achieved through the following technical solutions:
[0006] The present invention provides a data analysis method, comprising:
[0007] Get the category for classifying data;
[0008] Get the data bin for each category;
[0009] Obtaining reference data for each of the data bins;
[0010] Get unclassified data;
[0011] Obtaining data features of the data bin according to data features of the classified data in the data bin and data features of the corresponding reference data;
[0012] The type of each classified data is obtained based on the data features of the unclassified data and the data features of each data bin.
[0013] The present invention also discloses a data analysis method, comprising:
[0014] Create a data warehouse for each type of data;
[0015] receiving classified data and its categories;
[0016] Each classified data is stored in the corresponding data warehouse according to its category.
[0017] The present invention also discloses a data analysis device, comprising:
[0018] Data warehouse reading interface, used to obtain the type of data classification;
[0019] Get the data bin for each category;
[0020] Acquiring a plurality of reference data of each of the data bins;
[0021] Analyze service input interface to obtain unclassified data;
[0022] a computing unit, configured to obtain data features of the data bin according to data features of the classified data in the data bin and data features of the corresponding reference data;
[0023] Obtaining and obtaining the type of each classified data according to the data characteristics of the unclassified data and the data characteristics of each of the data bins;
[0024] Analysis Service output interface, used to output the category of each classified data.
[0025] The present invention also discloses a data analysis system, comprising:
[0026] A data analysis device for outputting a category of each classified data; and
[0027] A storage unit, used to establish a data warehouse for each type of data;
[0028] receiving classified data and its categories;
[0029] Each classified data is stored in the corresponding data warehouse according to its category.
[0030] The present invention obtains the data features of each data bin by analyzing the reference data and classified data of each data bin, and then compares the data features of the unclassified data with them, thereby achieving accurate classification of the data.
[0031] Of course, any product implementing the present invention does not necessarily need to achieve all of the advantages described above at the same time. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0033] FIG1 is a schematic diagram of functional units and information flow of a data analysis system according to an embodiment of the present invention;
[0034] FIG2 is a schematic diagram of a process flow of a data analysis device according to an embodiment of the present invention;
[0035] FIG3 is a schematic diagram of a process flow of a storage unit according to an embodiment of the present invention;
[0036] FIG4 is a schematic diagram of a flow chart of step S5 according to an embodiment of the present invention;
[0037] FIG5 is a schematic diagram of a flow chart of step S52 according to an embodiment of the present invention;
[0038] FIG6 is a schematic diagram of a flowchart of step S528 according to an embodiment of the present invention;
[0039] FIG7 is a schematic diagram of a flow chart of step S55 according to an embodiment of the present invention;
[0040] FIG8 is a schematic diagram of a flow chart of step S553 according to an embodiment of the present invention;
[0041] FIG9 is a schematic diagram of a flow chart of step S6 according to an embodiment of the present invention;
[0042] In the accompanying drawings, the components represented by the reference numerals are as follows:
[0043] 1-Data warehouse reading interface, 2-Analysis service input interface, 3-Calculation unit, 4-Analysis service output interface, 5-Storage unit. DETAILED DESCRIPTION
[0044] In order to make the objectives, technical solutions and advantages of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.
[0045] It should be noted that the terms "first," "second," and the like in this application are used to distinguish similar objects and are not necessarily used to describe a specific order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. Instead, they are merely examples of devices and methods consistent with certain aspects of the present application as detailed in the appended claims.
[0046] Computer data analysis refers to the use of computers and related technologies to process, analyze, and mine data. This process typically involves using computer software and tools to collect, clean, convert, and analyze large amounts of data. Due to the large volume of computer data, manual review is difficult. Furthermore, due to the limitations of computer recognition technology, automated classification can easily lead to misclassification. To improve the accuracy of data classification, the present invention provides the following solution.
[0047] Referring to Figures 1 to 3 , the present invention provides a data analysis system that can provide users with data analysis services, thereby enabling accurate judgment of data types. Based on interactive functionality, the system comprises a data warehouse read interface 1, an analysis service input interface 2, a computing unit 3, an analysis service output interface 4, and a storage unit 5. These three interfaces work together to output the type of each categorized data item, while the storage unit 5 is used to store the categorized data.
[0048] During the specific data analysis process, the data bin reading interface 1 first executes step S1 to obtain the data classification category. Next, step S2 can be executed to obtain the data bin for each category. Next, step S3 can be executed to obtain a number of reference data for each data bin. The reference data here can be illustrative samples selected by staff, with the goal of providing the system with a sample for classification comparison.
[0049] The analysis service input interface 2 then executes step S4 to obtain the unclassified data. The computing unit 3 then executes step S5 to obtain the data characteristics of the data bin based on the data characteristics of the classified data within the data bin and the corresponding data characteristics of the reference data. Next, step S6 can be executed to obtain and determine the type of each classified data item based on the data characteristics of the unclassified data and the data characteristics of each data bin. Finally, the analysis service output interface 4 outputs the type of each classified data item. Of course, in actual operations, the analysis service output interface 4 can also simultaneously output the classified data itself.
[0050] In the process allowed by the storage unit 5, step S01 can first be executed to establish a data warehouse for each type of data. Next, step S02 can be executed to receive the classified data and its type. Finally, step S03 can be executed to store each classified data in the corresponding data warehouse according to its type. The data warehouse here can be an aggregate concept, that is, storing classified data of the same type in a centralized manner, or it can be an abstract concept, that is, not distinguishing the existence of data, but only labeling each classified data with a type. Both are feasible and fall within the scope of protection of this solution.
[0051] As shown in Figure 4, the reference data corresponding to the data bin has numerous characteristic dimensions. However, given the nature of how humans use data, the semantics of the data are the primary consideration. Therefore, this solution classifies the data based on the semantics inherent in the data. In other words, the semantics within the reference data that provide distinguishing features serve as the data features of the corresponding data bin. In specific operations, step S51 can first be executed to perform word segmentation on each reference data item, obtaining the segmented words and their corresponding numbers within each reference data item. Next, step S52 can be executed to obtain the keywords and their frequency within each reference data item based on the segmented words and their corresponding numbers within each reference data item. Next, step S53 can be executed to use the keywords of the reference data corresponding to each data bin as the keywords of the classified data within the data bin. Next, step S54 can be executed to perform word segmentation on the classified data within each data bin, obtaining the keyword frequency of each keyword within each classified data item within each data bin. Finally, step S55 can be executed to obtain the effective frequency distribution range of each keyword within the classified data within the data bin based on the keywords and their frequency within each reference data item and the keyword frequency of each keyword within each classified data item within each data bin, as the data features of the data bin. The semantic features that distinguish features are materialized as the effective frequency distribution range of each keyword, which allows computers to perform calculations and comparisons.
[0052] As shown in FIG5 , since the number of segmented words generated after semantic segmentation of the reference data is excessive, in order to select important and representative segmented words as data features, step S521 can be first executed to obtain each segmented word in all the reference data. Next, step S522 can be executed to obtain the number of occurrences of each segmented word in all the reference data. Next, step S523 can be executed to obtain the cumulative number of occurrences of all the segmented words in all the reference data. Next, step S524 can be executed to use the ratio of the number of occurrences of each segmented word in all the reference data to the cumulative number of occurrences of all the segmented words as the global word frequency of each segmented word. Next, step S525 can be executed to obtain the cumulative number of occurrences of all the segmented words in each reference data. Next, step S526 can be executed to obtain the number of occurrences of each segmented word in each reference data. Next, step S527 can be executed to use the ratio of the number of occurrences of each segmented word in each reference data to the cumulative number of occurrences of all the segmented words in the reference data as the internal word frequency of each segmented word in each reference data. Finally, step S528 may be executed to obtain the keywords and their frequencies in each reference data according to the internal frequency of each segmented word in each reference data and the corresponding global frequency.
[0053] To supplement the implementation of steps S521 to S528, we provide the source code for some functional modules, with cross-references and explanations provided in the comments. To prevent the leakage of data involving commercial secrets, data that does not affect the implementation of the solution is desensitized. The same applies below.
[0054] This code first reads a set of reference data strings and segments and counts the words in each string. It then calculates the global frequency of each word (how often it appears across all reference data) and the internal frequency of each word within each reference data text (how often it appears within a specific reference data). Finally, the code outputs the internal and global frequency of each keyword in each reference data set. This data can be used to identify keywords and their importance.
[0055] As shown in FIG6 , since each segmented word has a different degree of importance, the number of keywords in the reference data is limited, and the importance of keywords significantly exceeds that of other segmented words. Therefore, in the process of obtaining keywords for each reference data, step S5281 can be first executed to obtain the ratio of each segmented word's internal frequency to the corresponding global frequency as the frequency coefficient of each segmented word. Next, step S5282 can be executed to arrange the frequency coefficients of each segmented word by numerical value to obtain a coefficient list. Next, step S5283 can be executed to obtain the mean of the difference between each frequency coefficient and its adjacent frequency coefficients in the coefficient list as the mean coefficient difference. Next, step S5284 can be executed to calculate the difference between the frequency coefficient with the largest value and the adjacent smaller frequency coefficients in the coefficient list, and determine whether it is greater than the mean coefficient difference. If so, step S5285 can be executed to stop the process. If not, step S5284 can be executed to continue the process, calculating the difference between the frequency coefficient with the largest value and the adjacent smaller frequency coefficients in the coefficient list, and determining whether it is less than the mean coefficient difference. Next, step S5286 may be executed to use the segmented words corresponding to the word frequency coefficients involved in the calculation as keywords. Finally, step S5287 may be executed to summarize and obtain the keywords and their word frequencies in each reference data.
[0056] In order to supplement the implementation process of the above-mentioned steps S5281 to S5287, the source code of some functional modules is provided, and a comparative explanation is given in the comment section.
[0057] This code first defines a structure called WordFreq to store the text, local word frequency, and global word frequency of each word. It then uses the standard library function std::sort to sort the word frequency coefficients and calculates the mean of the differences between the word frequency coefficients. Keywords are filtered based on whether the difference is greater than the mean. Finally, each keyword and its local word frequency are output. Keywords are identified by comparing the word frequency coefficient with the mean difference between these coefficients.
[0058] Please refer to Figures 7 to 8. Due to the single number of reference data, relying solely on reference data to classify data may result in a classification comparison scope that is too narrow, resulting in a large amount of unclassified data being unable to be effectively classified and becoming dirty data. Therefore, it is also necessary to perform auxiliary classification on the classified data in the data warehouse, that is, to simultaneously extract the data features of the classified data as the data features of the data warehouse. Specifically, for each data warehouse, step S551 can be first executed to obtain the word frequency of the keywords of each classified data in the data warehouse. Next, step S552 can be executed to arrange the keywords of each classified data in the same order to obtain a multidimensional vector composed of the numerical values of the word frequency of the keywords of each classified data as the feature vector of each classified data. The classified data here can be the data injected into the data warehouse after selection by the staff, or it can be the classified data after classification.
[0059] Next, step S553 can be executed to obtain multiple distribution ranges of the word frequency of each keyword in the classified data within the data warehouse based on the feature vector of each classified data. During this process, step S5531 can be first executed to select several feature vectors from the multiple classified data as target feature vectors. Next, step S5532 can be executed to calculate the vector difference modulus between each target feature vector and each other feature vector. Next, step S5533 can be executed to form a vector set by combining each other feature vector with the target feature vector with the smallest vector difference modulus. Next, step S5534 can be executed to calculate the feature vector with the smallest vector difference modulus between each vector set and the mean vector of all feature vectors as the updated target feature vector. Next, step S5535 can be executed to determine whether the updated target feature vector has changed. If so, steps S5532 to S5535 can be executed to return to the continuously updated vector set and the target feature vector. If not, step S5536 can be executed to obtain the distribution range of the word frequency of each keyword in the classified data corresponding to all feature vectors in each vector set. That is to expand the classification comparison scope of the data features of the classified data.
[0060] Since the classified data in the data warehouse may not strictly meet the target requirements, it needs to be calibrated with reference data. Therefore, step S554 can be executed finally to use the distribution range of the keywords and their word frequencies in the reference data as the effective distribution range of the word frequencies of each keyword in the classified data in the data warehouse.
[0061] In order to provide supplementary explanation for the implementation process of the above-mentioned steps 553S1 to S5536, the source code of some functional modules is provided, and a comparative explanation is provided in the comment section.
[0062] As the above code runs, the data is organized into keywords and their frequency across different data items. The initial target feature vectors are simply selected as the first few items in the dataset. The code then enters a loop that continuously updates the target feature vectors until they no longer change. During this loop, the feature vectors are classified using Euclidean distance, and a new mean vector is calculated for each classification. Once the target feature vectors stabilize, the loop ends. Finally, the code calculates and outputs the frequency distribution range for each keyword.
[0063] As shown in FIG9 , in the process of comparing the unclassified data with the data features of the data bin, in order to reduce the computational complexity of the comparison, step S61 can first be executed to segment the unclassified data into words, thereby obtaining the segmented words and the corresponding number within the unclassified data. Next, step S62 can be executed to determine whether the segmented words within the unclassified data cover the keywords in the data features of the data bin. If not, step S63 can be executed to not process the data. If so, step S64 can be executed to select the data bin as an alternative data bin. This avoids comparing the data features of each data bin, reducing the computational complexity of the comparison without sacrificing the accuracy of the comparison and classification.
[0064] In the process of comparing the unclassified data with each candidate data bin, step S65 can be first executed to use the keywords of the candidate data bin as the keywords of the unclassified data, and each keyword and its word frequency of the unclassified data as the data feature. Next, step S66 can be executed to determine whether the word frequency of each keyword in the unclassified data falls within the effective distribution range of the word frequency of each keyword in the classified data in the candidate data bin. If so, step S67 can be executed to mark the unclassified data as classified data, and the type of the candidate data bin can be used as the type of the classified data. If not, steps S65 to S67 can be executed to compare with the next candidate data bin.
[0065] In order to provide supplementary explanation for the implementation process of the above-mentioned steps S61 to S67, source codes of some functional modules are provided, and comparative explanations are provided in the comments.
[0066] The above code first defines a DataFeature structure to store data features and word frequency ranges. The main function, main, defines an unclassified data string, unclassifiedData, and an array, dataFeatures, containing various data features. The tokenizeAndCountFreq function is used to perform word segmentation and frequency counting on the unclassified data. The data features of each data bin are then iterated over, and the areAllKeywordsCovered function is used to check whether the segmented words in the unclassified data cover the keywords in the data features of that data bin. If so, the isFreqInRange function is used to determine whether the frequency of each keyword in the unclassified data falls within the valid frequency distribution range of the corresponding keyword in the data bin. If so, the unclassified data is marked as classified, and the data bin type is used as the classified data type. If not, the comparison continues with the next data bin. If no data bin matches, the data remains unclassified. Finally, the final classification result is output.
[0067] The flow charts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the devices, systems, methods and computer program products according to multiple embodiments of the present application. In this regard, each box in the flow chart or block diagram can represent a part for a module, program segment or instruction, and the part for the module, program segment or instruction comprises one or more executable instructions for realizing the logical function of the specification. In some alternative implementations, the functions marked in the box can also occur in a sequence different from that marked in the accompanying drawings. For example, two continuous boxes can actually be performed substantially in parallel, and they can sometimes also be performed in the opposite order, depending on the function involved.
[0068] It should also be noted that each box in the block diagram and / or flowchart, and the combination of boxes in the block diagram and / or flowchart, can be implemented by hardware that performs the corresponding function or action, such as a circuit or ASIC (Application Specific Integrated Circuit), or can be implemented by a combination of hardware and software, such as firmware.
[0069] Although the present invention is described herein in conjunction with various embodiments, in the process of implementing the claimed invention, those skilled in the art can understand and implement other variations of the disclosed embodiments by reviewing the drawings, the disclosure, and the appended claims. In the claims, the word "comprising" does not exclude other components or steps, and "a" or "an" does not exclude multiple situations. A single processor or other unit can implement several functions listed in the claims. Certain measures are recorded in different dependent claims, but this does not mean that these measures cannot be combined to produce good results.
[0070] The embodiments of the present application have been described above. The above description is illustrative and not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope of the described embodiments. The terminology used herein is selected to best explain the principles of the embodiments, practical applications, or improvements to the technology in the market, or to enable other persons skilled in the art to understand the embodiments disclosed herein.
Claims
1. A data analysis method, characterized in that, including, obtaining the categories for classifying data; obtaining the data warehouses for each category; obtaining the reference data for each of the data warehouses; obtaining the unclassified data; obtaining the data characteristics of the data warehouse according to the data characteristics of the classified data in the data warehouse and the data features of the corresponding reference data; obtaining and determining the category of each classified data according to the data features of the unclassified data and the data characteristics of each data warehouse.
2. The method according to claim 1, wherein The step of obtaining the data characteristics of the data warehouse according to the data characteristics of the classified data in the data warehouse and the data features of the corresponding reference data includes: performing word segmentation on each of the reference data to obtain the segmented words and their corresponding quantities within each of the reference data; obtaining the keywords and their word frequencies within each of the reference data according to the segmented words and their corresponding quantities within each of the reference data; using the keywords of the reference data corresponding to each data warehouse as the keywords of the classified data in the data warehouse; performing word segmentation on the classified data in each data warehouse to obtain the word frequencies of the keywords of each classified data in each data warehouse; obtaining the effective distribution range of the word frequencies of each keyword in the classified data in the data warehouse as the data characteristics of the data warehouse according to the keywords and their word frequencies within each of the reference data and the word frequencies of the keywords of each classified data in each data warehouse.
3. The method according to claim 2, wherein The step of obtaining the keywords and their word frequencies within each of the reference data according to the segmented words and their corresponding quantities within each of the reference data includes: obtaining each segmented word within all the reference data; obtaining the number of occurrences of each segmented word within all the reference data; obtaining the cumulative number of occurrences of all the segmented words within all the reference data; using the ratio of the number of occurrences of each segmented word within all the reference data to the cumulative number of occurrences of all the segmented words as the global word frequency of each segmented word; obtaining the cumulative number of occurrences of all the segmented words within each reference data; obtaining the number of occurrences of each segmented word within each reference data; using the ratio of the number of occurrences of each segmented word within each reference data to the cumulative number of occurrences of all the segmented words within that reference data as the internal word frequency of each segmented word within each reference data; obtaining the keywords and their word frequencies within each of the reference data according to the internal word frequency of each segmented word within each reference data and the corresponding global word frequency.
4. The method according to claim 3, wherein The step of obtaining the keywords and their word frequencies within each of the reference data according to the internal word frequency of each segmented word within each reference data and the corresponding global word frequency includes: for each of the reference data, obtaining the ratio of the internal word frequency of each segmented word to the corresponding global word frequency as the word frequency coefficient of each segmented word; arranging the word frequency coefficients of each segmented word in ascending order of numerical value to obtain a coefficient list; obtaining the average value of the differences between each word frequency coefficient in the coefficient list and the adjacent word frequency coefficient as the average coefficient difference; starting from the word frequency coefficient with the largest numerical value in the coefficient list, calculating the difference with the adjacent smaller word frequency coefficient in turn and determining whether it is greater than the average coefficient difference; if so, stop execution. If not, continuously execute the step of calculating the difference between the highest-frequency coefficient and the adjacent smaller frequency coefficient in the coefficient list in turn, and determine whether it is less than the average coefficient difference. Use the segmentation word corresponding to the frequency coefficient participating in the calculation as the keyword. Summarize and obtain the keywords and their frequencies in each reference data.
5. The method according to claim 2, wherein The step of obtaining the effective frequency distribution range of each keyword in the classified data in the data warehouse according to the keywords and their frequencies in each reference data and the frequencies of the keywords of each classified data in each data warehouse includes: For each data warehouse, Obtain the frequencies of the keywords of each classified data in the data warehouse. Arrange the keywords of each classified data in the same order to obtain a multi-dimensional vector composed of the frequency values of the keywords of each classified data as the feature vector of each classified data. Obtain multiple distribution ranges of the frequencies of each keyword in the classified data in the data warehouse according to the feature vectors of each classified data. Use the distribution range where the keywords and their frequencies in the reference data are located as the effective frequency distribution range of each keyword in the classified data in the data warehouse.
6. The method according to claim 5, wherein The step of obtaining multiple distribution ranges of the frequencies of each keyword in the classified data in the data warehouse according to the feature vectors of each classified data: Includes: Select several from the feature vectors of multiple classified data as target feature vectors. Calculate and obtain the vector difference norm between each target feature vector and each other feature vector. Form a vector set by combining each other feature vector with the target feature vector with the smallest vector difference norm. Calculate and obtain the feature vector with the smallest vector difference norm between each vector set and the mean vector of all feature vectors as the updated target feature vector. Judge whether the updated target feature vector has changed. If so, return to continuously update the vector set and the target feature vector. If not, obtain the distribution range of the frequencies of each keyword of the classified data corresponding to all feature vectors in each vector set.
7. The method according to any one of claims 2 to 6, characterized in that, The step of obtaining and classifying each classified data according to the data characteristics of the unclassified data and the data characteristics of each data warehouse includes: Segment the unclassified data to obtain the segmentation words and corresponding quantities in the unclassified data. Judge whether the segmentation words in the unclassified data cover the keywords in the data characteristics of the data warehouse. If not, do not process. If so, use the data warehouse as an alternative data warehouse. During the comparison between the unclassified data and each alternative data warehouse, Use the keywords of the alternative data warehouse as the keywords of the unclassified data, and use each keyword and its frequency of the unclassified data as data characteristics. Judge whether the frequency of each keyword of the unclassified data falls within the effective frequency distribution range of each keyword in the classified data in the alternative data warehouse. If so, mark the unclassified data as classified data, and use the category of the alternative data warehouse as the category of the classified data. If not, compare with the next alternative data warehouse.
8. A data analysis method, characterized in that, Includes: Establish a data warehouse for each type of data. Receive the classified data and its types in the data analysis method according to any one of claims 1 to 7; Store each classified data into the corresponding data warehouse according to its type.
9. A data analysis device, characterized in that, Including, A data warehouse reading interface for obtaining the types for classifying data; Obtain the data warehouses for each type; Obtain a number of reference data for each of the data warehouses; An analysis service input interface for obtaining unclassified data; An operation unit for obtaining the data characteristics of the data warehouse according to the data characteristics of the classified data in the data warehouse and the data characteristics of the corresponding reference data; Obtain and determine the type of each classified data according to the data characteristics of the unclassified data and the data characteristics of each data warehouse; An analysis service output interface for outputting the types of each classified data.
10. A data analysis system, characterized in that, Including, A data analysis device according to claim 9 for outputting the types of each classified data; And, A storage unit for creating a data warehouse for each type of data; Receive classified data and its types; Store each classified data into the corresponding data warehouse according to its type.
Citation Information
Patent Citations
Construction method and system for topic model category of city-level data warehouse
CN113849639A
New word classification method and device, electronic equipment and storage medium
CN116108180A
Text classification method and electronic equipment
CN116701616A
Data analysis method, device and system
CN117574243A
System and method for classification of spend data
US20230385765A1