Intelligent data asset classification and archiving management system

Through the intelligent data asset classification and archive management system, the dynamic deconstruction of fuzzy semantic features and multi-path parallel inference technology are used to solve the accuracy problem of semantic fuzzy data classification, and more efficient and accurate data asset management is achieved.

CN120067899AInactive Publication Date: 2025-05-30CHANGSHA DIGITAL GROUP CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510552251.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-29
Publication Date
2025-05-30
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing data classification technology is difficult to process data in semantic fuzzy or complex context scenarios, resulting in inaccurate classification results, affecting the efficiency and quality of data asset management.

Method used

An intelligent data asset classification and archiving management system was designed, including a dynamic deconstruction module of fuzzy semantic features, a rule generation module and a fuzzy classification multi-task decision-making module. Through adaptive hierarchical deconstruction, semantic fuzzy point positioning, multi-dimensional feature deconstruction and semantic sensitive rule generation, a classification rule tree is dynamically constructed, and classification credibility is calculated through multi-path parallel inference.

Benefits of technology

The precise expression and classification of semantic fuzzy data is realized, which significantly improves classification accuracy, especially in complex contexts or multi-sense scenarios, and improves the efficiency and quality of data asset management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067899A_ABST
    Figure CN120067899A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data management, in particular to an intelligent data asset classification and archiving management system, which comprises a fuzzy semantic feature dynamic deconstruction module for decomposing and reconstructing semantic fuzzy key features in data by utilizing self-adaptive hierarchical deconstruction; the rule generation module is used for inputting hierarchical features and constructing a classification rule tree based on a semantic hierarchical feature map, and rules comprise semantic sensitive rules and classification generation rules; and the fuzzy classification multi-task decision module is used for mapping the classification rules to a plurality of classification paths, distributing initial weights, calculating the classification credibility of each classification path and outputting an optimal path and a classification result. According to the method, the classification efficiency is improved, path missing or misjudgment possibly existing in a traditional method is avoided, meanwhile, a fuzzy conflict analysis module is canceled, the system process is simplified, and the real-time performance is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data management, and particularly to an intelligent data asset classification and archiving management system. Background Art

[0002] With the continuous advancement of digital transformation, the scale of data assets has grown exponentially. All kinds of data play a core role in storage, management, and analysis. However, how to efficiently classify and archive a large amount of data assets has become a key issue in data governance. Currently, most data classification techniques are based on fixed rules or traditional machine learning algorithms. These methods perform well in dealing with data with clear semantics, but have obvious limitations in dealing with data with ambiguous semantics or complex context scenarios.

[0003] Ambiguous semantic data usually exhibits polysemy, strong context dependence, or complex syntactic structures. Traditional classification methods are difficult to accurately analyze the characteristics of such data. For example, the same keyword may have different meanings in different contexts, and it is difficult to comprehensively capture its semantic features relying only on static rules. In addition, existing classification systems usually lack the ability to dynamically adapt to data characteristics and cannot generate refined classification rules for ambiguous semantic points, resulting in inaccurate classification results, which in turn affects the efficiency and quality of data asset management. Summary of the Invention

[0004] The present invention provides an intelligent data asset classification and archiving management system.

[0005] The intelligent data asset classification and archiving management system includes: A dynamic deconstruction module for fuzzy semantic features: decomposes and reconstructs the key features with ambiguous semantics in the data by using adaptive hierarchical deconstruction, specifically including: An ambiguous semantic point positioning unit: identifies ambiguous semantic points by analyzing keywords, context relationships, and word vector distributions in the data; A multi-dimensional feature deconstruction unit: disassembles ambiguous semantic words into multiple semantic branches based on the semantic features of ambiguous semantic points, combining semantic, syntactic, context, and time series dimensions, to form a semantic hierarchical feature map; A rule generation module: based on the semantic hierarchical feature map, inputs hierarchical features, constructs a classification rule tree, the rules include semantic sensitive rules, generates classification rules, and sends the results to the fuzzy classification multi-task decision module; A fuzzy classification multi-task decision module: receives classification rules and the semantic hierarchical feature map, performs multi-path parallel reasoning, maps the classification rules to multiple classification paths, assigns initial weights, calculates the classification credibility of each classification path, and outputs the optimal path and classification results.

[0006] Optionally, the semantic ambiguity point positioning unit specifically includes: Keyword extraction: Based on TF-IDF (Term Frequency-Inverse Document Frequency), extract high-weight candidate keywords from the input data.

[0007] Context relationship analysis: Use a convolutional neural network to perform semantic analysis on the context where the keyword is located, and capture the dependency relationship between the keyword and its context; Word vector distribution evaluation: Calculate the distribution of keywords in the semantic space through a pre-trained word vector model, and identify the feature points related to different semantic branches; Ambiguity point determination rule: Combine the context relationship weight of the keyword and the result of the word vector distribution evaluation to locate the keyword with polysemy or ambiguity, and mark it as a semantic ambiguity point.

[0008] Optionally, the use of a convolutional neural network to perform semantic analysis on the context where the keyword is located and capture the dependency relationship between the keyword and its context is expressed as: , where is the candidate keyword, is the keyword is the context feature vector of is the keyword is the context window of the keyword (including the previous and subsequent words), is the embedding vector of the context word, is the weight matrix and bias vector of is the convolutional neural network, is the activation function, and the output is the context feature vector of the keyword ; ; In the word vector distribution evaluation, the pre-trained word vector model uses the pre-trained Word2Vec model to calculate the word vector distribution of keywords and generate the distribution matrix of keywords in the semantic space.

[0009] Ambiguity point determination rule: Perform clustering analysis on the context relationship of the candidate keyword and the distribution matrix of the word vector to determine the semantic ambiguity point: , where is the number of context branches of the keyword, is the semantic ambiguity point. When , it is determined that is a semantic ambiguity point, is the ambiguity threshold, set to 0.6, is the ambiguity point context branch The eigenvector of ( a word with context relationship) is defined to be the same as that of , but the specific context is the window of.

[0010] Optionally, the multi-dimensional feature deconstruction unit specifically includes: Semantic dimension decomposition: Based on the word vector distribution of semantic fuzzy points, use principal component analysis to reduce the dimension of its distribution in the semantic space, decompose it into multiple semantic sub-spaces, and generate the eigenvector representation of the fuzzy points in the semantic sub-space ; Syntactic dimension analysis: Adopt dependency parsing technology to extract the syntactic structure of the sentence where the fuzzy point is located, analyze the syntactic dependency relationship, and represent it through a syntax tree to disassemble the association path between the fuzzy point and other elements in the syntactic structure; Context dimension enhancement: Combine the context eigenvector and the syntactic dependency relationship to generate a context feature enhanced representation , capturing its semantic changes in the dynamic context; Time series feature expansion: For time-related data, use a long short-term memory network to capture the evolution features of fuzzy points in the time dimension, forming a time series eigenvector ; Feature integration and hierarchical reconstruction: Combine semantic, syntactic, context, and time series features into a multi-dimensional eigenvector, and use an adaptive hierarchical algorithm to generate a semantic hierarchical feature map, where each layer represents an independent semantic branch of the fuzzy point in a specific dimension.

[0011] Optionally, the combination of semantic, syntactic, context, and time series features into a multi-dimensional eigenvector is expressed as: ; The semantic hierarchical feature map is expressed as: , where represents the feature representation of a certain dimension, including semantic, syntactic, context features, and time series features, represents the index of the feature dimension, represents the set of feature dimensions.

[0012] Optionally, the rule generation module specifically includes: Data reception: Receive the semantic hierarchical feature map , as the basis for constructing the classification rule; Rule node generation: For each feature dimension , by using the semantic branch and context relationship in the features, rule nodes are generated, and the rule nodes are represented as: , where represents the triggering condition based on the semantic-sensitive rule, represents the classification operation on the fuzzy point when the condition is satisfied; Generation of semantic-sensitive rules: According to the context features and the syntax tree dynamically generate context-sensitive rules ; Generation of rule tree: Organize all the generated rule nodes into a classification rule tree according to the feature dimension .

[0013] Optionally, the context-sensitive rule is represented as: ; where is scaled by Min-Max normalization according to the minimum and maximum values of the context feature values, so that has a value range of [0, 1], represents the context correlation threshold, represents the syntactic structure that conforms to the syntactic dependency relationship. When the syntactic structure matches , the condition is satisfied, is the first category label, is the second category label.

[0014] Optionally, the classification rule tree is represented as: .

[0015] Optionally, the fuzzy classification multi-task decision module specifically includes: Input reception: Receive the classification rule tree ; Receive the semantic hierarchical feature map ; Classification path generation: Map the rule nodes of to the feature dimension Feature in to generate multiple classification paths: , where is the classification path index, is the path triggering condition, is the classification operation after the path is triggered, is the initial weight of the rule node, and is generated according to the number of classification rules and the feature distribution, Classification path; Weight assignment: Initial weight calculation: Set the initial weight of each path according to the rule weight and semantic feature weight in the rule tree : , where and are the rule weight and semantic feature weight respectively; Classification path credibility calculation: Calculate the classification credibility of each classification path , comprehensively considering the path weight and feature matching degree: , where represents the matching score between the path trigger condition and the hierarchical feature map, measuring the degree to which the features meet the rules. Sort all paths according to from high to low, and select the path with the highest value as the optimal path , and perform classification operations according to the optimal path , as the classification result.

[0016] Optionally, the matching score between the path trigger condition and the semantic hierarchical feature map is calculated by comparing the feature values of the path trigger condition and the semantic hierarchical feature map ; For each condition , calculate the similarity with the corresponding feature value in the semantic hierarchical feature map; Matching condition: If the path trigger condition is fully satisfied , the similarity , if not satisfied, assign .

[0017] Advantages of the present invention: In the present invention, by constructing a semantic hierarchical feature map, accurate expression and classification of semantically ambiguous data are achieved. The map hierarchically analyzes the semantic, syntactic, context, and time series features of the data, captures the multi-dimensional characteristics of fuzzy points, enables the classification rules to fully utilize the semantic features of the data, significantly improves the classification accuracy, and shows higher adaptability and reliability especially in complex contexts or polysemous scenarios.

[0018] The present invention adopts a multi-path parallel reasoning mechanism based on a classification rule tree, dynamically maps classification rules with a hierarchical feature map, generates multiple classification paths, calculates path credibility based on path weights and matching degrees, selects the optimal path to output classification results, significantly improves classification efficiency by parallelly evaluating the effectiveness of multiple paths, avoids possible path omissions or misjudgments in traditional methods, and at the same time cancels the fuzzy conflict resolution module, simplifies the system process, improves real-time performance, automatically constructs semantic-sensitive rules according to context and grammatical relationships, and dynamically adjusts rule weights to optimize classification decisions. This rule self-adaptation mechanism enhances the classification flexibility and adaptability of the system, enabling it to handle scenarios with complex and variable semantic features, reducing manual intervention, and improving the intelligence level and classification performance of the system. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only those of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0020] Figure 1 Schematic diagram of the functional modules of the management system according to an embodiment of the present invention; Figure 2 Schematic diagram of the fuzzy classification multi-task decision module according to an embodiment of the present invention; Figure 3 Schematic diagram of the rule generation module according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0021] The present invention will be described in detail below in conjunction with the drawings and specific embodiments. At the same time, it should be noted here that in order to make the embodiments more detailed, the following embodiments are the best and preferred embodiments. For some well-known technologies, those skilled in the art can also adopt other alternative methods for implementation; moreover, the drawings are only for more specifically describing the embodiments and are not intended to specifically limit the present invention.

[0022] It should be pointed out that in the specification, it is mentioned that "an embodiment", "embodiments", "exemplary embodiments", "some embodiments", etc. indicate that the described embodiments may include specific features, structures or characteristics, but not necessarily every embodiment includes the specific features, structures or characteristics. In addition, when combining embodiments to describe specific features, structures or characteristics, implementing such features, structures or characteristics in combination with other embodiments (whether explicitly described or not) should be within the knowledge scope of those skilled in the relevant art.

[0023] Generally, terms can be understood at least in part from their use in context. For example, depending at least in part on the context, the term "one or more" as used herein can be used to describe any feature, structure, or property in a singular sense, or can be used to describe a combination of features, structures, or properties in a plural sense. Additionally, the term "based on" can be understood to not necessarily be intended to convey a set of exclusive factors, but rather can alternatively, depending at least in part on the context, allow for the existence of other factors that are not necessarily explicitly described.

[0024] As Figures 1 - 3 shown, the intelligent data asset classification and archiving management system includes: Fuzzy semantic feature dynamic deconstruction module: uses adaptive hierarchical deconstruction to decompose and reconstruct the key features with semantic fuzziness in the data, specifically including: Semantic fuzziness point positioning unit: identifies semantic fuzziness points by analyzing keywords, context relationships, and word vector distributions in the data; Multi-dimensional feature deconstruction unit: based on the semantic features of the semantic fuzziness points, combines semantic, syntactic, context, and time series dimensions to disassemble semantic fuzzy words into multiple semantic branches, forming a semantic hierarchical feature map; Rule generation module: based on the semantic hierarchical feature map, inputs hierarchical features, constructs a classification rule tree, the rules include semantic sensitive rules, generates classification rules, and sends the results to the fuzzy classification multi-task decision module; Fuzzy classification multi-task decision module: receives classification rules and the semantic hierarchical feature map, performs multi-path parallel reasoning, maps the classification rules to multiple classification paths, assigns initial weights, calculates the classification credibility of each classification path, and outputs the optimal path and classification results.

[0025] The semantic fuzziness point positioning unit specifically includes: Keyword extraction: based on TF-IDF (term frequency-inverse document frequency), extracts high-weight candidate keywords from the input data, and extracts keywords through the TF-IDF method: , where is the word, is the document, is the document collection, , represents the term frequency of the word in the document , represents the inverse document frequency, and selects the top words with the highest TF-IDF values as candidate keywords; For short documents (<500 words) set to 5 - 10; For medium-length documents (500 - 2000 words): Set to 10 - 20; For long documents (>2000 words): Set to 20 - 50.

[0026] Context relationship analysis: Use a convolutional neural network to perform semantic analysis on the context where the keyword is located, capturing the dependency relationship between the keyword and its context; Word vector distribution evaluation: Calculate the distribution of keywords in the semantic space through a pre - trained word vector model, identifying feature points related to different semantic branches; Fuzzy point determination rule: Combine the context relationship weight of the keyword and the result of the word vector distribution evaluation to locate keywords with polysemy or ambiguity, and mark them as semantic fuzzy points.

[0027] Using a convolutional neural network to perform semantic analysis on the context where the keyword is located, capturing the dependency relationship between the keyword and its context is expressed as: , where is the candidate keyword, is the keyword is the context feature vector of is the keyword is the context window of the keyword (including the previous and next words), is the embedding vector of the context word, is the weight matrix and bias vector of is the convolutional neural network, is the activation function, and the output is the context feature vector of the keyword ; ; The pre - trained word vector model in the word vector distribution evaluation uses the pre - trained Word2Vec model to calculate the word vector distribution of keywords, expressed as: , where represents two words, is the word 's vector representation, Sim represents the cosine similarity between word vectors; Generate the distribution matrix of the keyword in the semantic space: , represents the vocabulary set; is to calculate the similarity between two specific words and to measure the direct relationship of word vectors in the semantic space, is used to generate the target keyword in the entire vocabulary set Determine the target keyword from the distribution matrix with all words similarity to construct its semantic distribution pattern for subsequent semantic ambiguity point analysis Single similarity calculation for specific word pairs Extended to all vocabulary sets for generating the global semantic distribution matrix of keywords

[0028] Ambiguity point determination rule: Perform clustering analysis on the context relationship of candidate keywords and the distribution matrix of word vectors to determine semantic ambiguity points , where is the number of context branches of the keyword is the semantic ambiguity point. When , it is determined that is a semantic ambiguity point is the ambiguity threshold, set to 0.6 is the ambiguity point context branch (words with a context relationship with ) has a feature vector defined the same as , but the specific context is the window of

[0029] The multi-dimensional feature deconstruction unit specifically includes Semantic dimension decomposition: Based on the word vector distribution of semantic ambiguity points, use principal component analysis to reduce the dimension of its distribution in the semantic space, decompose it into multiple semantic subspaces, and generate the feature vector representation of the ambiguity point in the semantic subspace , and the semantic space representation reduced by PCA (principal component analysis) is , where is the word vector matrix of the ambiguity point , is the principal component matrix, and the multiple semantic subspaces generated by the ambiguity point are , where is the semantic vector, representing the multi-dimensional feature vector of the ambiguity point in the semantic space represents the number of principal components Syntactic dimension analysis: Use dependency parsing technology to extract the syntactic structure of the sentence where the ambiguity point is located, analyze the syntactic dependency relationship, and disassemble the association path between the ambiguity point and other elements in the syntactic structure through the syntax tree represented to disassemble the association path between the ambiguity point and other elements in the syntactic structure ​​​Syntactic Dimension Analysis: Using dependency parsing technology, construct a syntactic dependency graph between the fuzzy points and other sentence elements, expressed as: , where is the syntax tree, representing the dependency relationship of the fuzzy point in the syntactic structure, is the word that has a syntactic relationship with the fuzzy point , representing the syntactic dependency relationship (such as subject, object, etc.), , representing the dependency relationship, indicating the syntactic connection between two words, generated by a dependency parsing tool (SpaCy); Context Dimension Enhancement: Combine the context feature vector and the syntactic dependency relationship to generate a context feature enhanced representation , capturing its semantic changes in the dynamic context; Context Dimension Enhancement: Combine the context feature vector and the syntactic dependency relationship to generate a context feature enhanced representation: , where is the context enhancement vector, representing the syntactic relationship weight, assigned according to the dependency grammar relationship, CtxVec represents the basic context feature vector, used to describe the semantic features of the keyword w in its context window; Time Series Feature Expansion: For time-related data, use the long short-term memory network to capture the evolution features of the fuzzy point in the time dimension, forming a time series feature vector ; Time Series Feature Expansion: For time-related data, use LSTM to calculate the time features: , where is the time series feature, representing the fuzzy point in time pattern of change, represents the time step, represents the hidden state of the previous time step, represents the input of the current time step, represents the weight and bias matrix of LSTM; Feature Integration and Hierarchical Reconstruction: Combine semantic, syntactic, context, and time series features into a multi-dimensional feature vector, and use an adaptive hierarchical algorithm to generate a semantic hierarchical feature map, where each layer represents an independent semantic branch of the fuzzy point in a specific dimension.

[0030] The analysis method of the basic context feature vector CtxVec is as follows: 1. Determine the context window: Select multiple words around the keyword w as the context window, including the words on the left and right, and the specific size is set according to requirements.

[0031] 2. Obtain the word vectors of the context words: Use a pre-trained Word2Vec word vector model to convert each word in the context window into a corresponding word vector, represented as a high-dimensional vector.

[0032] 3. Aggregate the context word vectors: Aggregate the word vectors in the context window, take the average value, and use the aggregated context feature representation as the basic context feature vector of the keyword w for subsequent analysis and processing.

[0033] Combine semantic, syntactic, context, and time series features into a multi-dimensional feature vector represented as: ; The semantic hierarchical feature map is represented as: , where represents the feature representation of a certain dimension, including semantic, syntactic, context features, and time series features, represents the index of the feature dimension, represents the set of feature dimensions, containing the following elements: Semantic dimension (semantic vector, ); Syntactic dimension (syntax tree, ); Context dimension (context-enhanced vector, ); Time series dimension (time series features, ).

[0034] The rule generation module specifically includes: Data reception: Receive the semantic hierarchical feature map , as the basis for constructing classification rules; Rule node generation: For each feature dimension , use the semantic branches and context relationships in the features to generate rule nodes, and the rule nodes are represented as: , where represents the trigger condition based on the semantic-sensitive rule, represents the classification operation on the fuzzy point when the condition is satisfied, for example, assigning a specific classification label to the keyword ; Semantic-sensitive rule generation: Dynamically generate context-sensitive rules according to the context features and the syntax tree ; Rule tree generation: Organize all the generated rule nodes by feature dimension into a classification rule tree .

[0035] Context-sensitive rules Expressed as: ; Wherein, Through Min-Max normalization, it is scaled according to the minimum and maximum values of the context feature values, so that The value range is [0, 1], Represents the context correlation threshold, with a value of 0.7; Represents the grammatical structure that conforms to the grammatical dependency relationship. When the grammatical structure Matches It meets the conditions, Is the first category label, indicating the classification label assigned to the keyword w when it meets the specified context feature and grammatical structure conditions, Is the second category label, indicating the classification label assigned to the keyword w when it does not meet the context feature or grammatical structure conditions; The generation rule is explained as: If the keyword The context enhancement feature value of Is greater than the set context correlation threshold , and The grammatical dependency relationship of Matches a specific pattern , then the keyword Is classified as ; Otherwise, the keyword .

[0036] Classification rule tree Expressed as: ; The generated classification rule tree is sent as output to the fuzzy classification multi-task decision-making module. The decision-making module uses the rule tree for multi-path classification reasoning and selects the optimal path to complete the classification operation.

[0037] The fuzzy classification multi-task decision-making module specifically includes: Input reception: Receive the classification rule tree ; Receive the semantic hierarchical feature map ; Classification path generation: Map the rule nodes of To the feature dimension Feature in To generate multiple classification paths: , wherein, Is the classification path index, Is the path trigger condition, It is the classification operation after path triggering. It is the initial weight of the rule node, which is generated according to the number of classification rules and the feature distribution. classification paths; Weight assignment: Initial weight calculation: According to the rule weights and semantic feature weights in the rule tree, set the initial weight of each path. : , where and are the rule weight and semantic feature weight respectively, both set to 0.5, and the sum of all path weights is 1; Classification path credibility calculation: Calculate the classification credibility of each classification path. , comprehensively considering the path weight and the degree of feature matching: , where represents the matching score between the path trigger condition and the hierarchical feature map, measuring the degree to which the feature satisfies the rule. Sort all paths according to from high to low, and select the path with the highest score as the optimal path. : , according to the classification operation of the optimal path. , as the classification result.

[0038] Each path contains multiple conditions , and the matching score between the path trigger condition and the semantic hierarchical feature map is calculated by comparing the path trigger condition with the feature values in the semantic hierarchical feature map ; For each condition , calculate the similarity with the corresponding feature value in the semantic hierarchical feature map (semantic, syntactic, context features, and time series features): , where Sim is the similarity measure between the condition value and the feature value, is the feature representation corresponding to the condition in the feature map; Matching condition: If the path trigger condition is completely satisfied , the similarity , if not satisfied, then assign ; Finally, is the average matching score between the path condition and the semantic hierarchical feature map, used to measure the path credibility.

[0039] The present invention encompasses any alternatives, modifications, equivalent methods, and solutions that fall within the spirit and scope of the present invention. To enable the public to have a thorough understanding of the present invention, specific details are described in detail in the following preferred embodiments of the present invention. However, those skilled in the art can fully understand the present invention even without the description of these details. Additionally, well-known methods, processes, procedures, components, and circuits are not described in detail to avoid unnecessary confusion with the essence of the present invention.

[0040] The above description is only a preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. Intelligent data asset classification and archiving management system, characterized by: include: Dynamic deconstruction module of fuzzy semantic features: It uses adaptive hierarchical deconstruction to decompose and reconstruct key features with semantic fuzziness in the data, including: Semantic fuzzy point location unit: identifies semantic fuzzy points by analyzing keywords, contextual relationships, and word vector distribution in the data; Multi-dimensional feature deconstruction unit: Based on the semantic features of semantic fuzzy points, combined with semantic, grammatical, contextual and time series dimensions, semantically fuzzy words are decomposed into multiple semantic branches to form a semantic hierarchical feature map; Rule generation module: Based on the semantic hierarchical feature map, input hierarchical features, build a classification rule tree, including semantic sensitive rules, generate classification rules, and send the results to the fuzzy classification multi-task decision module; Fuzzy classification multi-task decision module: receives classification rules and semantic hierarchical feature maps, performs multi-path parallel reasoning, maps classification rules to multiple classification paths, assigns initial weights, calculates the classification credibility of each classification path, and outputs the optimal path and classification results.

2. The intelligent data asset classification and archiving management system according to claim 1 is characterized in that: The semantically ambiguous point positioning unit specifically includes: Keyword extraction: Based on TF-IDF, high-weight candidate keywords are extracted from the input data; Contextual relationship analysis: Use convolutional neural networks to perform semantic analysis on the context in which the keywords are located, capturing the dependency between the keywords and their context; Word vector distribution evaluation: Calculate the distribution of keywords in the semantic space through the pre-trained word vector model and identify feature points related to different semantic branches; Fuzzy point determination rule: Combine the contextual relationship weights of keywords and the word vector distribution evaluation results to locate keywords with polysemous or ambiguous meanings and mark them as semantic fuzzy points.

3. The intelligent data asset classification and archiving management system according to claim 2 is characterized in that: The convolutional neural network is used to perform semantic analysis on the context of the keyword, and the dependency relationship between the keyword and its context is captured as follows: ,in, is a candidate keyword, is the keyword The context feature vector of is the keyword The context window, is the embedding vector of the context word, yes The weight matrix and bias vector of is a convolutional neural network, is the activation function, the output is the keyword The context feature vector ; The pre-trained word vector model in the word vector distribution evaluation uses the pre-trained Word2Vec model to calculate the word vector distribution of keywords and generate keywords. Distribution matrix in semantic space; Fuzzy point determination rule: cluster analysis is performed on the contextual relationship of candidate keywords and the distribution matrix of word vectors to determine semantic fuzzy points: ,in, is the number of context branches of the keyword, Is a semantic ambiguity point, when When is a semantic fuzzy point. is the fuzziness threshold, It's a blur Context branch The feature vector of .

4. The intelligent data asset classification and archiving management system according to claim 3 is characterized in that: The multi-dimensional feature deconstruction unit specifically includes: Semantic dimension decomposition: Based on the word vector distribution of semantic fuzzy points, principal component analysis is used to reduce the dimension of its distribution in the semantic space, decompose it into multiple semantic subspaces, and generate fuzzy points. Feature vector representation in semantic subspace ; Grammatical dimension analysis: Dependency parsing technology is used to extract the grammatical structure of the sentence where the ambiguity point is located, analyze the grammatical dependency relationship, and use the grammatical tree Representation, dismantling the association path between the fuzzy point and other elements in the syntactic construction; Context dimension enhancement: Combine context feature vectors and grammatical dependencies to generate context feature enhanced representations , capturing its semantic changes in dynamic contexts; Time series feature expansion: For time-related data, the long short-term memory network is used to capture the evolution characteristics of fuzzy points in the time dimension to form a time series feature vector ; Feature integration and hierarchical reconstruction: Semantic, grammatical, contextual and time series features are combined into a multi-dimensional feature vector, and an adaptive hierarchical algorithm is used to generate a semantic hierarchical feature map, where each layer represents an independent semantic branch of the fuzzy point in a specific dimension.

5. The intelligent data asset classification and archiving management system according to claim 4 is characterized in that: The combination of semantic, grammatical, contextual and time series features into a multidimensional feature vector is expressed as: ; The semantic hierarchical feature map is expressed as: ,in, Represents the feature representation of a certain dimension, including semantic, grammatical, contextual features and time series features. represents the index of the feature dimension, Represents a collection of feature dimensions.

6. The intelligent data asset classification and archiving management system according to claim 5 is characterized in that: The rule generation module specifically includes: Data reception: receiving semantic layered feature maps , as the basis for the construction of classification rules; Rule node generation: for each feature dimension , using the semantic branches and contextual relationships in the features, rule nodes are generated. The rule nodes are represented as: ,in, Indicates the triggering conditions based on semantically sensitive rules, Indicates that when the condition is met, the fuzzy point Classification operation; Semantic-sensitive rule generation: based on contextual features and syntax tree Dynamically generate context-sensitive rules ; Rule tree generation: All generated rule nodes are organized into classification rule trees according to feature dimensions. .

7. The intelligent data asset classification and archiving management system according to claim 6 is characterized in that: The context-sensitive rules It is expressed as: ; in, Through Min-Max normalization, The value range is [0, 1]. represents the context relevance threshold, Indicates a grammatical structure that complies with grammatical dependencies. match When the condition is met, is the first category label, is the second category label.

8. The intelligent data asset classification and archiving management system according to claim 7 is characterized in that: The classification rule tree It is expressed as: 。 9. The intelligent data asset classification and archiving management system according to claim 1, characterized in that: The fuzzy classification multi-task decision module specifically includes: Input reception: receiving classification rule tree ; Receive semantic hierarchical feature maps ; Classification path generation: The rule nodes and Feature dimension in Mapping, generating multiple classification paths: ,in, is the classification path index, is the path trigger condition, It is the classification operation after the path is triggered. is the initial weight of the rule node, which is generated according to the number of classification rules and feature distribution. Category paths; Weight allocation: Initial weight calculation: According to the rule weight and semantic feature weight in the rule tree, set the initial weight of each path : ,in, , They are rule weight and semantic feature weight respectively; Classification path credibility calculation: Calculate the classification credibility of each classification path , comprehensively considering the path weight and feature matching degree: ,in, Indicates the matching score between the path trigger condition and the hierarchical feature map, which measures the degree to which the feature satisfies the rule. Sort from high to low, select The highest path is taken as the optimal path , according to the classification operation of the optimal path , as the classification result.

10. The intelligent data asset classification and archiving management system according to claim 9, characterized in that: The matching score between the path trigger condition and the semantic hierarchical feature map Through path trigger conditions and semantic layered feature maps The characteristic values ​​of are compared and calculated; For each condition , and calculate the similarity with the corresponding feature value in the semantic hierarchical feature map: Matching condition: If the path triggers the condition Completely satisfied , similarity If not satisfied, then assign .

Citation Information

Cited By

  • Data resource classification method and device based on rule and semantic approximation in combination with AI

    CN121456535A

  • Intelligent patent document classification method and system based on semantic understanding

    CN121833954A