Data classification management method and system based on multilevel semantic network
By conducting basic feature testing and differential evaluation of the data flow, establishing a semantic analytical composite model, and constructing a semantic relationship diagram, the problem of low efficiency and accuracy of multi-level and multi-field data classification is solved, and efficient and accurate data classification management is achieved.
Patent Information
- Application Number
- CN202510337846.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-07-04
AI Technical Summary
The lack of accurate semantic relationship diagrams in the prior art makes it difficult to effectively classify and process multi-level and multi-field complex data flows, resulting in poor data classification efficiency and accuracy.
By conducting basic feature testing and differential evaluation of the data flow to be processed, integrating data field labels, establishing a semantic analytical composite model, constructing multiple data semantic vectors and disambiguation compensation, and finally constructing a semantic relationship diagram for classification management.
It improves the efficiency and accuracy of data classification management, and ensures the accuracy of data classification and efficient management.
Smart Images

Figure CN120256698A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data processing, and specifically to a data classification management method and system based on a multi-level semantic network. Background Art
[0002] With the rapid development of information technology, the amount of data has grown explosively. How to efficiently manage and classify massive data has become a key problem to be solved urgently. Traditional data classification management methods are unable to cope when faced with complex, diverse, and semantically rich data. Especially when dealing with complex data that crosses multiple fields, it is impossible to perform efficient and accurate data classification. For example, in the databases of some large enterprises, data from different sources and in different formats are mixed together. Relying solely on simple rules or shallow feature analysis, it is difficult to accurately classify the data, resulting in low data retrieval efficiency and the inability to fully exploit the data value, thus affecting the accuracy and management efficiency of diverse data classification.
[0003] Therefore, in the current related technologies, there is a technical problem that there is a lack of an accurate semantic relationship graph, making it difficult to effectively classify and process multi-level and multi-field complex data streams, resulting in poor data classification efficiency and accuracy. Summary of the Invention
[0004] This application provides a data classification management method and system based on a multi-level semantic network, which solves the technical problem in the prior art of lacking an accurate semantic relationship graph and being difficult to effectively classify and process multi-level and multi-field complex data streams, resulting in poor data classification efficiency and accuracy, and achieves the technical effect of improving the efficiency and accuracy of data classification management.
[0005] This application provides a data classification management method based on a multi-level semantic network. The method includes: performing a basic feature check on a first data stream to be processed to obtain a basic feature check result, and optimizing the first data stream to be processed according to the basic feature check result to obtain a second data stream to be processed; evaluating the difference degree of multiple data domain labels corresponding to the second data stream to be processed to obtain multiple domain difference coefficients; if all the multiple domain difference coefficients are less than a domain difference threshold, fusing the multiple data domain labels to obtain a fused domain label; based on a semantic parsing loss function, adaptively adjusting and learning training resources for multi-level semantic factors according to the fused domain label to establish a semantic parsing composite model; inputting the second data stream to be processed into the semantic parsing composite model to establish multiple data semantic vectors; performing disambiguation compensation according to the multiple data semantic vectors to obtain multiple semantic optimization vectors, and constructing a semantic relationship graph according to the multiple semantic optimization vectors; classifying and managing the second data stream to be processed according to the semantic relationship graph.
[0006] In a possible implementation, the data classification and management method based on a multi-level semantic network further performs the following processing: retrieving a semantic parsing sample set corresponding to a data sample set in the same domain according to the fused domain label; classifying the semantic parsing sample set according to the multi-level semantic factors to obtain an element semantic parsing sample set, a direct semantic parsing sample set, an indirect semantic parsing sample set, and a structural semantic parsing sample set; performing training resource adaptive adjustment learning on the data sample set in the same domain and the element semantic parsing sample set based on the semantic parsing loss function to obtain an element semantic parsing network; performing training resource adaptive adjustment learning on the data sample set in the same domain and the direct semantic parsing sample set based on the semantic parsing loss function to obtain a direct semantic parsing network; performing training resource adaptive adjustment learning on the data sample set in the same domain and the indirect semantic parsing sample set based on the semantic parsing loss function to obtain an indirect semantic parsing network; performing training resource adaptive adjustment learning on the data sample set in the same domain and the structural semantic parsing sample set based on the semantic parsing loss function to obtain a structural semantic parsing sample set parsing network; connecting the element semantic parsing network, the direct semantic parsing network, the indirect semantic parsing network, and the structural semantic parsing sample set parsing network to generate the semantic parsing composite model.
[0007] In a possible implementation, the data classification and management method based on a multi-level semantic network further performs the following processing: aligning the data sample set in the same domain and the element semantic parsing sample set to obtain an element semantic parsing construction set; equally allocating training resources to the element semantic parsing construction set to obtain a first result of training resource allocation; performing supervised training on a semantic parsing predetermined network according to the element semantic parsing construction set based on the first result of training resource allocation to obtain an initial element semantic parsing network; testing the initial element semantic parsing network according to the element semantic parsing construction set based on the semantic parsing loss function to obtain a semantic parsing loss sequence; if any semantic parsing loss coefficient in the semantic parsing loss sequence is greater than or equal to a semantic parsing loss threshold, optimizing and adjusting the first result of training resource allocation according to the semantic parsing loss sequence to obtain a second result of training resource allocation; using a value less than the semantic parsing loss threshold as a training target, and performing optimized training on the initial element semantic parsing network according to the second result of training resource allocation to obtain the element semantic parsing network.
[0008] In a possible implementation, the data classification and management method based on a multi-level semantic network further performs the following processing: The semantic parsing loss function is:
[0009] SPL = log f SIM(SPY, SPO);
[0010] Among them, SPL represents the semantic parsing loss coefficient, f represents the predetermined factor of the semantic parsing loss, 0 < f < 1, SPY represents the semantic parsing prediction data, SPO represents the semantic parsing sample data corresponding to the semantic parsing prediction data, and SIM(SPY, SPO) represents the similarity between the semantic parsing prediction data and the semantic parsing sample data.
[0011] In a possible implementation, the data classification management method based on the multi-level semantic network further performs the following processing: If any one of the multiple domain difference coefficients is greater than or equal to the domain difference threshold, cluster the multiple data domain labels according to the multiple domain difference coefficients to obtain N domain label clusters, where N is a positive integer greater than 1; Based on the semantic parsing loss function, adaptively adjust and learn the training resources for the multi-level semantic factors according to the N domain label clusters, and establish N domain semantic parsing models; Perform semantic parsing on the second data stream to be processed according to the N domain semantic parsing models.
[0012] In a possible implementation, the data classification management method based on the multi-level semantic network further performs the following processing: Traverse the multiple data semantic vectors, and extract the first data semantic vector; Perform fuzzy detection on the first data semantic vector to obtain a first semantic fuzzy detection result; Evaluate the degree of ambiguity of the first semantic fuzzy detection result to obtain a first ambiguity coefficient; If the first ambiguity coefficient is greater than or equal to the predetermined ambiguity coefficient, perform semantic association backtracking on the second data stream to be processed according to the first semantic fuzzy detection result to obtain a first semantic association feature; Disambiguate and optimize the first data semantic vector according to the first semantic association feature to obtain a first semantic optimization vector, and add the first semantic optimization vector to the multiple semantic optimization vectors.
[0013] In a possible implementation, the data classification management method based on the multi-level semantic network further performs the following processing: Perform error detection on the first data stream to be processed to obtain each data error detection result; Perform integrity detection on each data in the first data stream to be processed to obtain each data integrity detection result; Output the each data error detection result and the each data integrity detection result as the basic feature test result.
[0014] In a possible implementation manner, the data classification and management method based on a multi-level semantic network further performs the following processing: collecting the language format information and storage format information of the first data stream to be processed to obtain first format feature information; determining whether the first format feature information meets a predetermined format constraint; if the first format feature information does not meet the predetermined format constraint, using the format difference feature between the first format feature information and the predetermined format constraint as a format optimization factor; and optimizing the format of the first data stream to be processed according to the format optimization factor.
[0015] In a possible implementation manner, the data classification and management method based on a multi-level semantic network further performs the following processing: the multi-level semantic factors include element semantics, direct semantics, indirect semantics, and structural semantics.
[0016] The present application further provides a data classification and management system based on a multi-level semantic network, including: a basic feature verification module, configured to perform basic feature verification on a first data stream to be processed, obtain a basic feature verification result, and optimize the first data stream to be processed according to the basic feature verification result to obtain a second data stream to be processed; a difference degree evaluation module, configured to evaluate the difference degree of a plurality of data domain labels corresponding to the second data stream to be processed to obtain a plurality of domain difference coefficients; a fused domain label obtaining module, configured to, if all the plurality of domain difference coefficients are less than a domain difference threshold, fuse the plurality of data domain labels to obtain a fused domain label; a semantic parsing composite model building module, configured to perform self-adaptive adjustment learning of training resources on multi-level semantic factors based on a semantic parsing loss function according to the fused domain label, and build a semantic parsing composite model; a data semantic vector building module, configured to input the second data stream to be processed into the semantic parsing composite model to build a plurality of data semantic vectors; a semantic relationship graph construction module, configured to perform disambiguation compensation according to the plurality of data semantic vectors to obtain a plurality of semantic optimization vectors, and construct a semantic relationship graph according to the plurality of semantic optimization vectors; and a classification and management module, configured to perform classification and management on the second data stream to be processed according to the semantic relationship graph.
[0017] The data classification management method and system based on a multi-level semantic network proposed in this application perform basic feature verification on the first data stream to be processed to obtain the basic feature verification result; evaluate the difference degree of multiple data domain labels corresponding to the second data stream to be processed; if the multiple domain difference coefficients are all less than the domain difference threshold, fuse the multiple data domain labels; establish a semantic parsing composite model based on the semantic parsing loss function; establish multiple data semantic vectors; perform disambiguation compensation to obtain multiple semantic optimization vectors, and construct a semantic relationship graph; perform classification management according to the semantic relationship graph. This solves the technical problem in the prior art that there is a lack of an accurate semantic relationship graph, making it difficult to effectively classify and process complex data streams at multiple levels and in multiple domains, resulting in poor data classification efficiency and accuracy, and achieves the technical effect of improving the efficiency and accuracy of data classification management. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] To more clearly illustrate the technical solutions of the embodiments of the present disclosure, the drawings of the embodiments of the present disclosure will be briefly introduced below. Flowcharts are used in this application to illustrate the operations performed by the system according to the embodiments of the application. It should be understood that the operations before or below do not necessarily need to be executed precisely in sequence. Instead, according to the need, they can be executed in reverse order or simultaneously. At the same time, other operations can also be added to these processes, or one or several operations can be removed from these processes.
[0019] Figure 1 It is a schematic flowchart of the data classification management method based on a multi-level semantic network provided by an embodiment of the present application.
[0020] Figure 2 It is a schematic structural diagram of the data classification management system based on a multi-level semantic network provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0021] The above description is only an overview of the technical solutions of this application. In order to be able to understand the technical means of this application more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features, and advantages of this application more obvious and understandable, the specific embodiments of this application are specifically given below.
[0022] In order to make the purpose, technical solutions, and advantages of this application clearer, the present application will be further described in detail below in conjunction with the drawings. The described embodiments should not be regarded as limitations of the present application. All other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of this application.
[0023] In the following description, reference is made to "some embodiments", which describe a subset of all possible embodiments. However, it is understood that "some embodiments" may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict. The terms "first" and "second" are only used to distinguish similar objects and do not represent a specific order for the objects. The terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or server that comprises a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or modules not clearly listed or inherent to these processes, methods, products or devices. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application.
[0024] Embodiments of this application provide a data classification and management method based on a multi-level semantic network, as Figure 1 shown. The method includes:
[0025] Step S100, perform a basic feature check on the first data stream to be processed, obtain a basic feature check result, and optimize the first data stream to be processed according to the basic feature check result to obtain a second data stream to be processed.
[0026] Preferably, the first data stream to be processed refers to the original data set, which is data from various different data sources. For example, transaction data collected in real time from an enterprise's business system, including order information, customer information, product information, etc.; it may also be monitoring data obtained from a sensor network, such as environmental monitoring data, industrial equipment operation status data, etc.; it may also be user-generated content data collected from social media platforms, such as text, pictures, videos, etc. Then, a basic feature check is performed on the first data stream to be processed, that is, each piece of data in the first data stream to be processed is checked and its basic features are analyzed, which may include the type of data (such as numeric, character, image, etc.), the format of the data (such as file format, data record format, etc.), the integrity of the data (whether there are missing values, error values, etc.), the value range of the data (the maximum and minimum values of numeric data, etc.), and the time stamp of the data and other relevant feature information. By checking these basic features, the basic situation of the data in the data stream can be comprehensively understood, the basic feature check result can be obtained, and the blockchain technology is used to record the basic feature check result and optimization operation of each step, including recording the error detection result, integrity detection result of the data, and key information before and after formatting on the blockchain, which not only ensures the traceability of the data processing process. Moreover, the immutable feature of the blockchain ensures the authenticity and integrity of the data. When it is necessary to verify the accuracy of data processing, relevant records can be directly obtained from the blockchain. For example, in a data stream containing sales data, the features of fields such as product name (character type), sales quantity (numeric type), and sales date (time type) in each sales record may be checked to see if there are situations where the product name is empty (data integrity problem) or the sales quantity is negative (value range problem), etc.
[0027] Preferably, according to the problems found in the basic feature check, the data is optimized to obtain the second data stream to be processed. Specifically, if missing values are found in the data, a filling method may be used, such as using the mean, median, or specific estimated values to fill in the missing data; for error values, the corresponding data records may be corrected or deleted; if the data format does not meet the requirements, operations such as format conversion are performed. For example, for the record with an empty product name in the above sales data, the product name can be filled with other relevant information (such as product information associated with the sales order number); for the record with a negative sales quantity, it can be determined as an input error and corrected according to the actual situation, or these error records can be directly deleted, and finally the optimized second data stream to be processed is obtained. By preprocessing the original data stream, the data quality is improved.
[0028] Further, step S100 further includes step S101, collecting the language format information and storage format information of the first data stream to be processed to obtain first format feature information; step S102, determining whether the first format feature information meets a predetermined format constraint; step S103, if the first format feature information does not meet the predetermined format constraint, using the format difference feature between the first format feature information and the predetermined format constraint as a format optimization factor; step S104, performing format optimization on the first data stream to be processed according to the format optimization factor.
[0029] Preferably, for the first data stream to be processed, collect its language format information and storage format information. Specifically, the language format information may include the programming language used for the data, character encoding (such as UTF-8, GBK, etc.), natural language type (such as Chinese, English, Japanese, etc.), etc. The storage format information covers the specific form of data storage, such as whether it is stored in a relational database (such as MySQL, Oracle), or in the form of a file (such as CSV, JSON, XML), as well as the version and structure of the file; and then form the first format feature information corresponding to the first data stream to be processed; the predetermined format constraint is the data format preset according to the standard data. Compare the collected first format feature information with the predetermined format constraint. For example, the predetermined format constraint stipulates that the data must use UTF-8 character encoding and be stored in JSON format, and the JSON data structure must conform to a specific schema. Then judge whether the character encoding, storage format, and data structure in the first format feature information all conform; if the first format feature information does not meet the predetermined format constraint, find the differences between the two. For example, the first format feature information shows that the data uses GBK character encoding, while the predetermined format constraint requires UTF-8; or the storage format is CSV, while the predetermined requirement is JSON, etc.; these differential features are used as format optimization factors. Among them, the format optimization factor clarifies the direction and specific content of the data format that needs to be improved; finally, perform corresponding format adjustment on the first data stream to be processed according to the determined format optimization factor. If the format optimization factor is that the character encodings are inconsistent, then convert the character encoding of the data from GBK to UTF-8; if it is a storage format problem, such as converting from CSV to JSON, use a data conversion tool or write corresponding program code to reorganize the data into a form that meets the JSON format requirements; thus realizing the format optimization of the first data stream to be processed.
[0030] Further, step S100 further includes step S110 of performing error detection on the to-be-processed first data stream to obtain respective data error detection results; step S120 of performing integrity detection on each data in the to-be-processed first data stream to obtain respective data integrity detection results; and step S130 of outputting the respective data error detection results and the respective data integrity detection results as the basic feature inspection results.
[0031] Preferably, each item of data in the to-be-processed first data stream is detected to discover possible errors. Specifically, there may be various error types, such as a numeric data exceeding a reasonable value range (for example, a negative number or an unreasonable value exceeding 150 appears in the age field), data type mismatch (such as a field that should be numeric stores text content), logical error (for example, in an order data, the quantity of goods is 0 but the total order amount is not 0), etc. Each data is detected one by one to obtain the error detection result of each data, clearly indicating whether each data has an error and the specific type of the error; then the integrity of each data in the to-be-processed first data stream is detected, that is, it is judged whether there is a missing situation of the data, which may include checking whether each field in the data record has a value, or whether some key information is missing. For example, in the employee information data stream, fields such as the employee's name, age, and position should all have corresponding records. If the position field of an employee is empty, it means that there is an integrity problem with this data. By detailed inspection of each data, the integrity detection results of each data are obtained, clearly indicating the integrity degree of each data, such as being completely complete, partially missing, or key information missing, etc.; finally, the respective data error detection results and the respective data integrity detection results are integrated and output as the basic feature inspection results of the to-be-processed first data stream to comprehensively reflect the quality status of the data in this data stream, including the error situation of the data and the integrity degree of the data.
[0032] Step S200, performing a difference evaluation on multiple data domain labels corresponding to the to-be-processed second data stream to obtain multiple domain difference coefficients.
[0033] Preferably, the degree of difference between different data domain labels in the second data stream to be processed after processing is quantitatively evaluated. Here, the data domain label is an identifier for classifying or describing data, used to represent the specific domain or theme to which the data belongs. For example, in a data stream containing multiple types of data, there may be different domain labels such as "sales data", "customer data", "product data", etc., and each label represents a specific data domain. Specifically, a difference degree evaluation is performed on multiple data domain labels in the second data stream to be processed, such as calculating the similarity or distance of data features, that is, calculating the cosine similarity or edit distance based on the vector space model, etc., to analyze and quantify the degree of difference between them. Then, the evaluation process and results are recorded on the blockchain, which helps to accurately trace the basis of the difference degree evaluation during subsequent label fusion or clustering operations. At the same time, the consensus mechanism of the blockchain can ensure the consistency and immutability of the evaluation results. If the data corresponding to two domain labels are significantly different in features, then their difference degree is high; conversely, if the difference is small, the difference degree is low. It may include comparing multiple aspects of data (such as the content, structure, semantics, etc. of the data). For example, "sales data" may include information such as sales amount, sales quantity, and sales time, while "customer data" includes information such as customer name, age, and contact information. By comparing the differences in this information, the difference degree between the two domain labels is evaluated.
[0034] Preferably, through the difference degree evaluation, multiple domain difference coefficients will be finally obtained, which are quantitative representations of the degree of difference between different data domain labels. Usually, the value range of the domain difference coefficient can be determined according to the specific evaluation method and application scenario. For example, when using cosine similarity calculation, the similarity value range is between 0 and 1, then the difference coefficient can be defined as 1 minus the similarity, so the value range of the difference coefficient is 0 to 1. 0 indicates that the two domain labels are exactly the same, and 1 indicates that the two domain labels are completely different. And each domain difference coefficient corresponds to a pair of data domain labels, representing the degree of difference between these two domains. Through these domain difference coefficients, the relationship between different data domains in the second data stream to be processed can be better understood, which helps to improve the efficiency and accuracy of data classification management.
[0035] Step S300, if all the multiple domain difference coefficients are less than the domain difference threshold, fuse the multiple data domain labels to obtain a fused domain label.
[0036] Preferably, the domain difference threshold is a preset standard value based on specific business requirements and data characteristics, and is used to determine whether the difference degree between data domain labels is small enough to decide whether to fuse them. When multiple domain difference coefficients are all smaller than the domain difference threshold, it indicates that the differences between the data domains represented by these data domain labels are small, and they have high similarity or relevance. Then, these data domain labels are fused, that is, merged into a fused domain label. When clustering data domain labels according to the domain difference coefficient, if it involves label data from different data sources, privacy computing can be used to perform clustering analysis on the premise of protecting data privacy. For example, using federated learning technology, different data sources can collaborate to complete the training of the clustering model without sharing the original label data, and determine N domain label clusters, which not only protects data privacy but also realizes effective data classification management. Among them, whether it is the original multiple data domain labels or the results after operations such as difference evaluation and clustering, they can be protected through encrypted storage. When constructing a semantic parsing model or performing other data processing based on these labels subsequently, the security of the data is guaranteed, preventing chaos in data classification management caused by label data leakage. For example, assume there are three data domain labels: "Customer Basic Information", "Customer Contact Information", and "Customer Preference Information", and their respective corresponding domain difference coefficients are all smaller than the set domain difference threshold, indicating that these three data domains have many similarities in terms of content, structure, or semantics, and may all be different aspects of information centered around the customer. Then, these three labels can be fused into a fused domain label of "Customer Comprehensive Information" to facilitate the unified management, analysis, and processing of these related data, improving the utilization efficiency of the data and the convenience of management.
[0037] Step S400: Based on the semantic parsing loss function, adaptively adjust the learning of training resources for multi-level semantic factors according to the fused domain label, and establish a semantic parsing composite model.
[0038] Step S400 further includes that the semantic parsing loss function is:
[0039] SPL = log f SIM(SPY, SPO);
[0040] Among them, SPL represents the semantic parsing loss coefficient, f represents the semantic parsing loss predetermined factor, 0 < f < 1, SPY represents the semantic parsing prediction data, SPO represents the semantic parsing sample data corresponding to the semantic parsing prediction data, and SIM(SPY, SPO) represents the similarity between the semantic parsing prediction data and the semantic parsing sample data.
[0041] Step S400 further includes that the multi-level semantic factors include element semantics, direct semantics, indirect semantics, and structural semantics.
[0042] Preferably, the semantic parsing loss function SPL = log f SIM(SPY, SPO), which is a function used to measure the difference between the semantic parsing result predicted by the model and the true semantics. In the semantic parsing task, the model attempts to convert the input data into a specific semantic representation, and the role of the loss function is to quantify the accuracy of this conversion. By minimizing the value of the loss function, the prediction result of the model can be made as close as possible to the true semantics, thereby improving the performance of the model; during the calculation process, sensitive data may be involved. Using privacy computing technologies such as secure multi-party computing, the loss function can be calculated without exposing the original data. For example, SPY and SPO come from different data sources and contain sensitive information. Through the secure multi-party computing protocol, all parties jointly calculate SIM(SPY, SPO) without revealing their own data, and then obtain the semantic parsing loss coefficient SPL, which not only completes the calculation of the loss function but also protects data privacy. Among them, the smaller the value of the semantic parsing loss function, the closer the prediction result is to the sample data, and the better the model performance; the semantic parsing loss predetermined factor, with a value range between 0 and 1, is used to adjust the calculation method of the loss function and can be set according to specific business scenarios and data characteristics; the semantic parsing prediction data is the result obtained by the model after semantic parsing of the input data; the semantic parsing sample data is the pre-determined and considered correct semantic parsing result, which is used as a reference standard for evaluating the model; the greater the similarity value between the semantic parsing prediction data and the semantic parsing sample data, the more similar the two are.
[0043] Preferably, the multi-level semantic factors are factors that describe the data semantics from different levels and perspectives. The semantic information of data has a multi-level structure. For example, at the lexical level, each word has its specific meaning; at the sentence level, the combination and grammatical structure of words form more complex semantics; at the discourse level, the logical relationship and overall theme between sentences constitute higher-level semantics; the multi-level semantic factors are the abstraction and representation of these different-level semantic information, which are used to assist the model in understanding and parsing the semantics of data; the multi-level semantic factors include element semantics (the semantics of a single element in the data, such as the basic semantics contained in a single word in the text, a single object in the image, etc.), direct semantics (the semantics formed by the direct combination of data elements, such as the meaning expressed by a phrase or a simple sentence), indirect semantics (the implicit semantics mined through reasoning, association, etc., which is not as intuitive as direct semantics), and structural semantics (the semantics conveyed by the organizational structure of the data, such as the semantic information reflected by the chapter structure of a document, the relationship between database tables, etc.). For example, as shown in Table 1:
[0044] Table 1 Semantic Type Interaction Relationship Table
[0045]
[0046]
[0047] Preferably, according to the fusion domain label, the semantic parsing loss function is used to perform adaptive adjustment learning of training resources for multi-level semantic factors, that is, with the semantic parsing loss function as the optimization target, according to the characteristics of the data under the fusion domain label, the resource allocation in the training process is dynamically adjusted, including parameters such as the selection of training data, the number of training iterations, and the learning rate. For example, if the data domain represented by the fusion domain label has a high complexity, the model may increase the usage of training data or adjust the learning rate to learn more slowly, so as to more accurately capture the semantic information of the data. For instance, if it is found that the element semantics are complex and diverse, the relevant training data volume may be increased or the training iteration times may be adjusted; if the structure semantics are difficult to learn, parameters such as the learning rate are adjusted. When constructing a semantic parsing composite model based on the fusion domain label, the update of training data and model parameters can be recorded and shared through the blockchain. Different nodes can participate in the model training process and verify the source and authenticity of the training data through the blockchain. For example, when performing adaptive adjustment learning of training resources for the same-domain data sample set and each semantic parsing sample set, the allocation of training resources and the update of model parameters each time can be recorded on the blockchain to ensure the transparency and credibility of the training process. In the process of performing adaptive adjustment learning of training resources, the multi-level semantic factors are continuously optimized so that they can better reflect the semantic characteristics of the data represented by the fusion domain label. By continuously trying and adjusting the model parameters, the multi-level semantic factors can better adapt to the semantic characteristics of the data, reduce the value of the semantic parsing loss function, and finally obtain a semantic parsing composite model with good performance, which can accurately perform semantic parsing on the data covered by the fusion domain label and convert the input data into a meaningful semantic representation.
[0048] Further, step S400 further includes step S410 of retrieving a semantic parsing sample set corresponding to the same-domain data sample set according to the fusion domain label; step S420 of classifying the semantic parsing sample set according to the multi-level semantic factors to obtain an element semantic parsing sample set, a direct semantic parsing sample set, an indirect semantic parsing sample set, and a structural semantic parsing sample set; step S430 of performing training resource adaptive adjustment learning on the same-domain data sample set and the element semantic parsing sample set based on the semantic parsing loss function to obtain an element semantic parsing network; step S440 of performing training resource adaptive adjustment learning on the same-domain data sample set and the direct semantic parsing sample set based on the semantic parsing loss function to obtain a direct semantic parsing network; step S450 of performing training resource adaptive adjustment learning on the same-domain data sample set and the indirect semantic parsing sample set based on the semantic parsing loss function to obtain an indirect semantic parsing network; step S460 of performing training resource adaptive adjustment learning on the same-domain data sample set and the structural semantic parsing sample set based on the semantic parsing loss function to obtain a structural semantic parsing sample set parsing network; step S470 of connecting the element semantic parsing network, the direct semantic parsing network, the indirect semantic parsing network, and the structural semantic parsing sample set parsing network to generate the semantic parsing composite model.
[0049] Preferably, according to the fusion domain label, retrieve the data sample set of the same domain and the corresponding semantic parsing sample set from the stored data. For example, if the fusion domain label is "e-commerce product information", then retrieve the data samples related to e-commerce products and the correct semantic parsing results corresponding to these samples; classify the semantic parsing sample set according to the multi-level semantic factors (element semantics, direct semantics, indirect semantics, and structural semantics) to obtain an element semantic parsing sample set, a direct semantic parsing sample set, an indirect semantic parsing sample set, and a structural semantic parsing sample set. For example, classify the semantic parsing results of single-attribute words describing products into the element semantic parsing sample set; classify the semantic parsing results of simple introduction statements of products into the direct semantic parsing sample set; classify the parsing results of potential semantics inferred from product evaluations into the indirect semantic parsing sample set; classify the semantic parsing results related to the layout structure of product information on the web page into the structural semantic parsing sample set.
[0050] Preferably, taking the semantic parsing loss function as the optimization objective, perform adaptive adjustment learning of training resources on the same-domain data sample set and the element semantic parsing sample set, continuously optimize the model parameters, and obtain an element semantic parsing network specialized for processing element semantics. Similarly, based on the semantic parsing loss function, perform adaptive adjustment learning of training resources on the same-domain data sample set, the direct semantic parsing sample set, the indirect semantic parsing sample set, and the structural semantic parsing sample set, and respectively obtain a direct semantic parsing network, an indirect semantic parsing network, and a structural semantic parsing network. Finally, connect the trained element semantic parsing network, direct semantic parsing network, indirect semantic parsing network, and structural semantic parsing network to form a semantic parsing composite model capable of comprehensively processing different levels of semantics, that is, input the data into each semantic parsing network simultaneously. Each network independently processes and outputs its own semantic parsing result, and then these results are integrated through a fusion layer (such as a fully connected layer, an attention mechanism layer, etc.). For example, the output feature vectors of the element semantic parsing network, the direct semantic parsing network, etc. are concatenated, and then weighted fusion is performed through a fully connected layer to obtain the final semantic representation. Through the semantic parsing composite model, the input same-domain data can be comprehensively parsed from multiple semantic levels such as element, direct, indirect, and structure, improving the efficiency and accuracy of data classification management.
[0051] Further, step S430 further includes step S431, aligning the same-domain data sample set and the element semantic parsing sample set to obtain an element semantic parsing construction set; step S432, equally allocating training resources to the element semantic parsing construction set to obtain a first result of training resource allocation; step S433, based on the first result of training resource allocation, supervisedly train a semantic parsing predetermined network according to the element semantic parsing construction set to obtain an initial element semantic parsing network; step S434, based on the semantic parsing loss function, test the initial element semantic parsing network according to the element semantic parsing construction set to obtain a semantic parsing loss sequence; step S435, if any semantic parsing loss coefficient in the semantic parsing loss sequence is greater than or equal to the semantic parsing loss threshold, optimize and adjust the first result of training resource allocation according to the semantic parsing loss sequence to obtain a second result of training resource allocation; step S436, taking less than the semantic parsing loss threshold as the training objective, optimize and train the initial element semantic parsing network according to the second result of training resource allocation to obtain the element semantic parsing network.
[0052] Preferably, align the data sample set in the same field and the element semantic parsing sample set, that is, make the data samples and the corresponding element semantic parsing results correspond one by one to form an element semantic parsing construction set. For example, in e-commerce review data, the data sample set in the same field is individual user reviews, and the element semantic parsing sample set is the correct annotation of the semantics of individual elements (such as words) in these reviews. After alignment, a construction set for training can be obtained; then, equal allocation of training resources is performed on the element semantic parsing construction set. Among them, the training resources can be computing resources, the amount of training data, training time, the number of training times, etc. Equal allocation means evenly distributing these samples and other resources to each link in the training process to obtain the first result of training resource allocation. For example, evenly dividing 100 samples into 5 groups, with 20 samples in each group for different training iterations, is a simple equal allocation method; then, based on the first result of training resource allocation, use the element semantic parsing construction set to perform supervised training on the predetermined semantic parsing network. Among them, the predetermined semantic parsing network is an initial network model with a preset structure, which may be a BP neural network or a fully connected neural network. A BP neural network is a multi-layer feedforward network trained according to the error backpropagation algorithm, and continuously adjusts the weights and thresholds of the network through backpropagation to minimize the sum of squared errors of the network; a fully connected neural network means that all neurons in adjacent layers of the network are connected. It includes using sample data and its corresponding correct labels (element semantic parsing sample set) to guide network learning and adjust network parameters to obtain the initial element semantic parsing network. For example, in a simple neural network, input user review data, and through comparison with the corresponding element semantic annotation, continuously adjust the network weights to complete the initial training.
[0053] Preferably, according to the first result of training resource allocation, corresponding data samples are selected from the element semantic parsing construction set and input into the predetermined semantic parsing network. For example, if the data is evenly divided into several groups, each group of data is sent into the network in sequence according to the grouping order. For the data in the element semantic parsing construction set of text type, preprocessing (such as word vector conversion) may be required first to make it acceptable to the network; the data propagates forward in the network according to the predetermined network structure. Taking a fully connected neural network as an example, after the neurons in the input layer receive the data, the signal is transmitted to the neurons in the hidden layer through the weight connection. The neurons in the hidden layer perform weighted summation on the input signal and process it through the activation function, and then transmit the processed signal to the next layer until the output layer generates the output result. This output result is the element semantic parsing prediction of the network for the input data; the output result of the network is compared with the correct label corresponding in the element semantic parsing construction set, and the semantic parsing loss function is used to calculate the error between the prediction result and the correct result; if it is a BP neural network, the calculated error will be propagated along the reverse path of the network connection. During the backpropagation process, the weights and thresholds between the layers in the network are adjusted according to the error, so that the error can change in the direction of decrease; all data samples in the element semantic parsing construction set are iteratively trained, and the network parameters are further adjusted according to the current error. As the training progresses, the prediction result of the network will be closer and closer to the correct element semantic parsing result. When the predetermined number of training rounds is completed or a certain training stop condition is met (such as the error is reduced to a certain extent), the training process ends, and the initial element semantic parsing network is obtained.
[0054] Preferably, the initial element semantic parsing network is tested with the element semantic parsing construction set based on the semantic parsing loss function, that is, a semantic parsing loss coefficient is calculated for each test sample and a semantic parsing loss sequence is formed; then the semantic parsing loss sequence is checked. If any one of the semantic parsing loss coefficients is greater than or equal to the semantic parsing loss threshold (a pre-set standard for measuring the acceptable degree of loss, such as 0.25), the first result of training resource allocation is optimized and adjusted according to the loss sequence to obtain the second result of training resource allocation; for example, if it is found that the loss coefficient of a certain sample is 0.3, which is greater than the threshold of 0.25, the training times of this sample or similar samples may be increased or more computing resources may be allocated, thus changing the resource allocation plan; finally, with the training target of being less than the semantic parsing loss threshold, the initial element semantic parsing network is retrained according to the second result of training resource allocation, continuously optimizing the network parameters, and finally the element semantic parsing network is obtained. For example, after adjusting the training resources and training multiple times, the loss coefficients of all test samples are made less than 0.25. At this time, the obtained network is the element semantic parsing network that meets the requirements.
[0055] Preferably, assume that in a text semantic parsing task, the in-domain data sample set is 1000 news titles, and the element semantic parsing sample set is the correct annotation of the semantics of each word in these titles. Divide these 1000 pieces of data into 10 groups evenly, with 100 pieces of data in each group, as the first result of training resource allocation for initial training. Use 200 pieces of data for testing to obtain a semantic parsing loss sequence, such as [0.22, 0.18, 0.26, 0.19,...]. Set the semantic parsing loss threshold to 0.2. It is found that the loss coefficient of the 3rd sample, 0.26, is greater than the threshold. Then increase the training iteration times of the training group containing this sample by 50%, adjust to obtain the second result of training resource allocation, and continue training according to the new resource allocation scheme until the loss coefficients of all test samples are less than 0.2, thus obtaining an element semantic parsing network.
[0056] Further, step S400 further includes step S480. If any one of the multiple domain difference coefficients is greater than or equal to the domain difference threshold, cluster the multiple data domain labels according to the multiple domain difference coefficients to obtain N domain label clusters, where N is a positive integer greater than 1; step S490, based on the semantic parsing loss function, adaptively adjust and learn the training resources for the multi-level semantic factors respectively according to the N domain label clusters, and establish N domain semantic parsing models; step S4100, perform semantic parsing on the to-be-processed second data stream according to the N domain semantic parsing models.
[0057] Preferably, when any one of the multiple domain difference coefficients is greater than or equal to the domain difference threshold, it indicates that there are significant differences between the data domains represented by these data domain labels. Cluster the multiple data domain labels according to these domain difference coefficients. Specifically, group the data domain labels with similarity (smaller domain difference coefficients) into one cluster. Through clustering algorithms (such as K-Means, hierarchical clustering, etc.), finally obtain N domain label clusters (N is a positive integer greater than 1). For example, in an e-commerce data processing scenario, there may be multiple data domain labels such as "product sales data", "user evaluation data", "inventory management data", "logistics distribution data", etc. If the domain difference coefficient between "product sales data" and "inventory management data" is small, while their domain difference coefficients with "user evaluation data" and "logistics distribution data" are large, after clustering, there may be two domain label clusters, one containing "product sales data" and "inventory management data", and the other containing "user evaluation data" and "logistics distribution data".
[0058] Preferably, the semantic parsing loss function is a function used to measure the difference between the predicted result and the true result in the semantic parsing process. By minimizing this loss function, the performance of the semantic parsing model can be optimized. That is, according to the obtained N domain label clusters, adaptive adjustment learning of training resources is performed on multi-level semantic factors respectively. Specifically, for each domain label cluster, adaptive adjustment learning of training resources is performed on multi-level semantic factors (including element semantics, direct semantics, indirect semantics, and structural semantics, etc.). The training resources include parameters such as the selection of training data, the number of training iterations, and the learning rate. Adaptive adjustment means dynamically adjusting these training resources according to the specific situation of the current domain label cluster. For example, if the data of a certain domain label cluster has high complexity, the number of training data may be increased, or the learning rate may be adjusted to learn more slowly, so as to better capture the semantic information of this domain. Through this process of adaptive adjustment learning of training resources, a dedicated domain semantic parsing model is established for each domain label cluster, and then N domain semantic parsing models are obtained. Each model is optimized for a specific domain label cluster and can better understand and parse the semantics of data within this domain.
[0059] Preferably, the data in the to-be-processed second data stream are respectively input into the corresponding domain semantic parsing models according to the domain label clusters they belong to. Each domain semantic parsing model will perform semantic parsing on the input data according to its own training results and semantic understanding ability. For example, for the data belonging to the "commodity sales data" domain label cluster, the corresponding domain semantic parsing model is used for parsing, and semantic information such as the sales quantity, sales amount, and sales time of the commodity may be identified. For the data belonging to the "user evaluation data" domain label cluster, the corresponding model is used for parsing, and semantic information such as the emotional tendency of the user and the theme of the evaluation may be extracted. In this way, comprehensive and in-depth semantic parsing can be performed on the entire to-be-processed second data stream, thereby ensuring the efficiency and accuracy of data classification management.
[0060] Step S500, input the to-be-processed second data stream into the semantic parsing composite model to establish multiple data semantic vectors.
[0061] Preferably, the second data stream to be processed obtained after performing basic feature inspection (such as error detection, integrity detection, etc.) and optimization (such as format optimization, data repair, etc. according to the results of basic feature inspection) on the original first data stream to be processed is input into the semantic parsing composite model for parsing. Specifically, the semantic parsing composite model analyzes and processes these data, that is, comprehensively considers multi-level semantic factors (element semantics, direct semantics, indirect semantics, structural semantics, etc.), conducts in-depth semantic parsing on the input data, and extracts the features of the data at different semantic levels. For example, for a piece of text data, the semantic parsing composite model will analyze the element semantics of each word in it (the meaning of a single word), the direct semantics formed by the combination of words (the meaning of a phrase or a simple sentence), the indirect semantics obtained through reasoning, etc. (such as implicit emotional tendency, potential theme, etc.) and the structural semantics of the text (such as the relationship between paragraphs, the grammatical structure of sentences, etc.); then integrate and transform these extracted semantic features to generate multiple data semantic vectors. Among them, each data semantic vector contains information about the data in a specific semantic aspect. For example, a data semantic vector may represent the emotional tendency of the text (positive, negative, or neutral), and another vector may represent the theme category involved in the text (such as technology, entertainment, sports, etc.), thereby ensuring the processing accuracy of tasks such as data classification, data retrieval, and data similarity calculation, ensuring that the semantic information of the data is accurately reflected, and thus ensuring the efficiency of data classification management.
[0062] Step S600, perform disambiguation compensation according to the multiple data semantic vectors to obtain multiple semantic optimization vectors, and construct a semantic relationship graph according to the multiple semantic optimization vectors.
[0063] Preferably, there may be ambiguity or uncertainty in representing data semantics by data semantic vectors, that is, one semantic vector may correspond to multiple actual semantics. For example, in natural language processing, the semantic vector corresponding to the word "apple" may represent both the fruit apple and the Apple Inc. This ambiguity and uncertainty affect the accurate understanding and processing of data semantics. Therefore, disambiguation compensation is performed based on multiple data semantic vectors, that is, some additional information sources, such as context information, domain knowledge, corpus statistical information, etc., are used to clarify the specific meaning represented by the semantic vector. Taking "apple" as an example, if the context mentions related words such as "eat" and "orchard", then it can be determined that "apple" here refers to the fruit; if the context mentions "mobile phone" and "Jobs", then it can be judged that it refers to Apple Inc. From the first data stream to be processed to the finally constructed semantic relationship graph, a large amount of data generated throughout the process needs to be securely stored, and encryption storage technology is used to protect this data. Specifically, the second data stream to be processed obtained after basic feature inspection and optimization of the first data stream to be processed, as well as the generated multiple data semantic vectors, semantic optimization vectors, etc., can all be encrypted and stored using encryption algorithms (such as AES encryption). Even if the data storage medium is illegally obtained, attackers cannot easily read and understand the data content, ensuring the security of the data during storage and providing a secure basis for classification management. In this way, the original semantic vectors are adjusted and optimized to eliminate ambiguity, and semantic optimization vectors that more accurately reflect the true semantics of the data are obtained.
[0064] Preferably, a semantic relationship graph is a structure that graphically represents the semantic relationships between data, capable of intuitively showing the semantic associations between different data elements, helping people better understand the overall semantic structure of the data, and discovering potential connections and patterns between the data. Specifically, when constructing a semantic relationship graph based on multiple semantic optimization vectors, each semantic optimization vector can be regarded as a node in the graph, and the edges between the nodes represent the semantic relationships between them. These relationships can be similarity relationships, inclusion relationships, causal relationships, etc. For example, if the data represented by two semantic optimization vectors are very similar semantically, then a similarity edge can be established between them; if the data represented by one semantic optimization vector includes the semantics of the data represented by another semantic optimization vector, then an inclusion edge can be established. By analyzing various semantic relationships between semantic optimization vectors and representing them in graphical form, a semantic relationship graph is finally constructed to comprehensively and accurately reflect the semantic structure of the entire data set. For example, in a data set of news articles, multiple semantic optimization vectors are obtained after processing, representing different news topics such as "technology development", "sports events", "political events", etc. If a technology news article mentions both artificial intelligence and 5G technology at the same time, then there will be an edge between the semantic optimization vectors representing "artificial intelligence" and "5G technology", indicating that there is a semantic association in this article, and they may be related technologies under the large theme of "technology development".
[0065] Further, step S600 further includes step S610 of traversing the multiple data semantic vectors to extract a first data semantic vector; step S620 of performing fuzzy detection on the first data semantic vector to obtain a first semantic fuzzy detection result; step S630 of evaluating the ambiguity degree of the first semantic fuzzy detection result to obtain a first ambiguity coefficient; step S640 of, if the first ambiguity coefficient is greater than or equal to a predetermined ambiguity coefficient, performing semantic association backtracking on the second data stream to be processed according to the first semantic fuzzy detection result to obtain a first semantic association feature; step S650 of performing disambiguation optimization on the first data semantic vector according to the first semantic association feature to obtain a first semantic optimization vector, and adding the first semantic optimization vector to the multiple semantic optimization vectors.
[0066] Preferably, each of the established multiple data semantic vectors is viewed (traversed) one by one, and a data semantic vector is randomly selected from them as the first data semantic vector. The selected first data semantic vector is analyzed to determine whether its semantics is ambiguous. For example, in the scenario of natural language processing, if this vector represents the semantics of a text, there may be a word or phrase with an unclear meaning. This kind of ambiguity is detected through rules to obtain the first semantic ambiguity detection result. For example, it is detected that the word "apple" in the text may refer to both a fruit and an electronic brand in the text represented by the current semantic vector, which is a semantic ambiguity situation and is recorded to form the first semantic ambiguity detection result; then, according to the first semantic ambiguity detection result, the ambiguity situation therein is quantitatively evaluated, such as by calculating the number of fuzzy semantics, the degree of difference between fuzzy semantics, etc., to obtain the first ambiguity coefficient. For example, if there are two fuzzy semantic items in a semantic vector and the difference between these two semantic items is large, then a relatively high first ambiguity coefficient may be obtained; conversely, if the number of fuzzy semantic items is small and the difference is small, the coefficient is lower.
[0067] Preferably, when the first ambiguity coefficient is greater than or equal to a predetermined ambiguity coefficient (a pre-set standard value used to determine whether the ambiguity degree is serious enough to be processed), semantic association backtracking is performed, that is, information related to the first data semantic vector is searched from the to-be-processed second data stream to determine its accurate semantics. For example, in e-commerce data, for the first data semantic vector representing the description of a certain product with ambiguity, it is backtracked to other relevant data of this product in the to-be-processed second data stream, such as user evaluations, other descriptions on the product details page, etc., to obtain more information to clarify the semantics, and finally the first semantic association feature is obtained; finally, according to the obtained first semantic association feature, the first data semantic vector is adjusted and optimized to eliminate the ambiguity therein, and a more accurate first semantic optimization vector is obtained. For example, according to the feature of "can be eaten", the semantic representation of "apple" in the first data semantic vector is adjusted to clearly be the apple of the fruit category. Then, this first semantic optimization vector after disambiguation and optimization is added to the existing set of multiple semantic optimization vectors to enrich the content of the semantic optimization vectors, thereby ensuring the accuracy and reliability of the semantic optimization vectors.
[0068] Step S700, classify and manage the to-be-processed second data stream according to the semantic relationship graph.
[0069] Preferably, by using the semantic association information between the data presented in the semantic relationship graph, the data is classified in an organized and logical manner, thereby achieving more efficient data management. Among them, the semantic relationship graph intuitively shows the semantic connections between different data elements in the second data stream to be processed. Based on the semantic relationships reflected in the semantic relationship graph, corresponding classification criteria are formulated. Specifically, the data can be divided into different categories according to the similarity between nodes, and the data with high similarity is grouped into one category; or according to the inclusion relationship, the data with inclusion relationship can be organized into a hierarchical category; or according to the causal relationship, the data with causal association can be grouped into related categories. For example, in an academic literature dataset, the literature with similar topics can be grouped into the same topic category, and the literature with citation relationships can be organized into different hierarchical categories according to the citation logical relationship; for each data element in the second data stream to be processed, according to its position in the semantic relationship graph and its relationship with other nodes, it is classified into the corresponding category; finally, targeted management is carried out on the data of different categories. For example, the data of the same category can be stored, retrieved, and analyzed uniformly to improve the efficiency of data processing; data mining can be carried out according to the classification results to discover the potential laws and trends between different categories of data; more targeted services can also be provided for users, such as quickly locating and recommending relevant categories of data according to the user's needs. And when classifying and managing the second data stream to be processed according to the semantic relationship graph, special encryption technologies such as homomorphic encryption can be used to directly perform some classification-related calculation operations on the encrypted data. For example, homomorphic encryption allows comparison and clustering operations on encrypted semantic vectors, thereby realizing the classification of encrypted data, avoiding the security risks brought by frequent decryption of data, and at the same time ensuring the efficiency and security of data classification management. Generally speaking, by classifying and managing the second data stream to be processed according to the semantic relationship graph, efficient data management based on data semantic understanding is realized, ensuring the accuracy, reliability, and efficiency of data classification management.
[0070] In the above text, with reference to Figure 1 the data classification management method based on the multi-level semantic network according to the embodiments of the present invention is described in detail. Next, with reference to Figure 2 the data classification management system based on the multi-level semantic network according to the embodiments of the present invention will be described.
[0071] The data classification management system based on the multi-level semantic network according to the embodiments of the present invention is used to solve the technical problems existing in the prior art, such as the lack of accurate semantic relationship graphs, which makes it difficult to effectively classify and process complex data streams in multiple levels and multiple fields, resulting in poor data classification efficiency and accuracy, and achieves the technical effect of improving the efficiency and accuracy of data classification management. As Figure 2As shown in the figure, the data classification and management system based on the multi-level semantic network includes: a basic feature inspection module 10, a difference evaluation module 20, a fused domain label acquisition module 30, a semantic parsing composite model establishment module 40, a data semantic vector establishment module 50, a semantic relationship graph construction module 60, and a classification management module 70.
[0072] The basic feature inspection module 10 is used to perform basic feature inspection on the first data stream to be processed, obtain the basic feature inspection result, and optimize the first data stream to be processed according to the basic feature inspection result to obtain the second data stream to be processed; the difference evaluation module 20 is used to perform difference evaluation on multiple data domain labels corresponding to the second data stream to be processed to obtain multiple domain difference coefficients; the fused domain label acquisition module 30 is used to fuse the multiple data domain labels to obtain a fused domain label if the multiple domain difference coefficients are all less than the domain difference threshold; the semantic parsing composite model establishment module 40 is used to perform self-adaptive adjustment learning of training resources on multi-level semantic factors based on a semantic parsing loss function according to the fused domain label to establish a semantic parsing composite model; the data semantic vector establishment module 50 is used to input the second data stream to be processed into the semantic parsing composite model to establish multiple data semantic vectors; the semantic relationship graph construction module 60 is used to perform disambiguation compensation according to the multiple data semantic vectors to obtain multiple semantic optimization vectors, and construct a semantic relationship graph according to the multiple semantic optimization vectors; the classification management module 70 is used to perform classification management on the second data stream to be processed according to the semantic relationship graph.
[0073] Next, the specific configuration of the semantic parsing composite model building module 40 will be described in detail. The semantic parsing composite model building module 40 further includes: retrieving a semantic parsing sample set corresponding to the same-domain data sample set according to the fused domain label; classifying the semantic parsing sample set according to the multi-level semantic factors to obtain an element semantic parsing sample set, a direct semantic parsing sample set, an indirect semantic parsing sample set, and a structural semantic parsing sample set; based on the semantic parsing loss function, performing training resource adaptive adjustment learning on the same-domain data sample set and the element semantic parsing sample set to obtain an element semantic parsing network; based on the semantic parsing loss function, performing training resource adaptive adjustment learning on the same-domain data sample set and the direct semantic parsing sample set to obtain a direct semantic parsing network; based on the semantic parsing loss function, performing training resource adaptive adjustment learning on the same-domain data sample set and the indirect semantic parsing sample set to obtain an indirect semantic parsing network; based on the semantic parsing loss function, performing training resource adaptive adjustment learning on the same-domain data sample set and the structural semantic parsing sample set to obtain a structural semantic parsing sample set parsing network; connecting the element semantic parsing network, the direct semantic parsing network, the indirect semantic parsing network, and the structural semantic parsing sample set parsing network to generate the semantic parsing composite model.
[0074] Next, the specific configuration of the semantic parsing composite model building module 40 will be continued to be described in detail. The semantic parsing composite model building module 40 further includes: aligning the same-domain data sample set and the element semantic parsing sample set to obtain an element semantic parsing construction set; equally allocating training resources to the element semantic parsing construction set to obtain a first result of training resource allocation; based on the first result of training resource allocation, supervising and training a semantic parsing predetermined network according to the element semantic parsing construction set to obtain an initial element semantic parsing network; based on the semantic parsing loss function, testing the initial element semantic parsing network according to the element semantic parsing construction set to obtain a semantic parsing loss sequence; if any semantic parsing loss coefficient in the semantic parsing loss sequence is greater than or equal to the semantic parsing loss threshold, optimizing and adjusting the first result of training resource allocation according to the semantic parsing loss sequence to obtain a second result of training resource allocation; taking less than the semantic parsing loss threshold as the training target, optimizing and training the initial element semantic parsing network according to the second result of training resource allocation to obtain the element semantic parsing network.
[0075] Next, the specific configuration of the semantic parsing composite model building module 40 will be continued to be described in detail. The semantic parsing composite model building module 40 further includes: The semantic parsing loss function is:
[0076] SPL = logf SIM(SPY, SPO);
[0077] Among them, SPL represents the semantic parsing loss coefficient, f represents the semantic parsing loss predetermined factor, 0 < f < 1, SPY represents the semantic parsing prediction data, SPO represents the semantic parsing sample data corresponding to the semantic parsing prediction data, and SIM(SPY, SPO) represents the similarity between the semantic parsing prediction data and the semantic parsing sample data.
[0078] Next, the specific configuration of the semantic parsing composite model building module 40 will be further described in detail. The semantic parsing composite model building module 40 further includes: if any one of the multiple domain difference coefficients is greater than or equal to the domain difference threshold, clustering the multiple data domain labels according to the multiple domain difference coefficients to obtain N domain label clusters, where N is a positive integer greater than 1; based on the semantic parsing loss function, respectively performing adaptive adjustment learning of training resources on the multi-level semantic factors according to the N domain label clusters to establish N domain semantic parsing models; performing semantic parsing on the second data stream to be processed according to the N domain semantic parsing models.
[0079] Next, the specific configuration of the semantic relationship graph construction module 60 will be described in detail. The semantic relationship graph construction module 60 further includes: traversing the multiple data semantic vectors to extract the first data semantic vector; performing fuzzy detection on the first data semantic vector to obtain a first semantic fuzzy detection result; evaluating the ambiguity degree of the first semantic fuzzy detection result to obtain a first ambiguity coefficient; if the first ambiguity coefficient is greater than or equal to the predetermined ambiguity coefficient, performing semantic association backtracking on the second data stream to be processed according to the first semantic fuzzy detection result to obtain a first semantic association feature; performing disambiguation optimization on the first data semantic vector according to the first semantic association feature to obtain a first semantic optimization vector, and adding the first semantic optimization vector to the multiple semantic optimization vectors.
[0080] Next, the specific configuration of the basic feature inspection module 10 will be described in detail. The basic feature inspection module 10 further includes: performing error detection on the first data stream to be processed to obtain each data error detection result; performing integrity detection on each data in the first data stream to be processed to obtain each data integrity detection result; outputting the each data error detection result and the each data integrity detection result as the basic feature inspection result.
[0081] Next, the specific configuration of the basic feature verification module 10 will be further described in detail. The basic feature verification module 10 further includes: collecting the language format information and storage format information of the first data stream to be processed to obtain first format feature information; determining whether the first format feature information meets a predetermined format constraint; if the first format feature information does not meet the predetermined format constraint, using the format difference feature between the first format feature information and the predetermined format constraint as a format optimization factor; and performing format optimization on the first data stream to be processed according to the format optimization factor.
[0082] Next, the specific configuration of the semantic parsing composite model establishment module 40 will be further described in detail. The semantic parsing composite model establishment module 40 further includes: the multi-level semantic factors include element semantics, direct semantics, indirect semantics, and structural semantics.
[0083] The data classification management system based on a multi-level semantic network provided by the embodiments of the present invention can execute the data classification management method based on a multi-level semantic network provided by any embodiment of the present invention, and has functional modules and beneficial effects corresponding to the execution of the method.
[0084] Although the present application makes various references to certain modules in the system according to the embodiments of the present application, however, any number of different modules can be used and run on the user terminal and / or the server. The included individual units and modules are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific names of the functional units are only for the convenience of mutual distinction and do not limit the protection scope of the present invention.
[0085] The above specific implementation manners do not constitute a limitation on the protection scope of the present application. Those skilled in the art should understand that various modifications, combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principle of the present application shall be included within the protection scope of the present application.
Claims
1. A data classification and management method based on a multi-level semantic network, characterized in that Including: Conduct a basic feature inspection on the first data stream to be processed, obtain the basic feature inspection result, and optimize the first data stream to be processed according to the basic feature inspection result to obtain the second data stream to be processed; Evaluate the difference degree of multiple data domain labels corresponding to the second data stream to be processed to obtain multiple domain difference coefficients; If all the multiple domain difference coefficients are less than the domain difference threshold, fuse the multiple data domain labels to obtain a fused domain label; Based on the semantic parsing loss function, perform training resource adaptive adjustment learning on the multi-level semantic factors according to the fused domain label to establish a semantic parsing composite model; Input the second data stream to be processed into the semantic parsing composite model to establish multiple data semantic vectors; Perform disambiguation compensation according to the multiple data semantic vectors to obtain multiple semantic optimization vectors, and construct a semantic relationship graph according to the multiple semantic optimization vectors; Classify and manage the second data stream to be processed according to the semantic relationship graph.
2. The data classification management method based on a multi-level semantic network according to claim 1, characterized in that Based on the semantic parsing loss function, perform training resource adaptive adjustment learning on the multi-level semantic factors according to the fused domain label to establish a semantic parsing composite model, including: According to the fused domain label, retrieve the semantic parsing sample set corresponding to the data sample set in the same domain; Classify the semantic parsing sample set according to the multi-level semantic factors to obtain an element semantic parsing sample set, a direct semantic parsing sample set, an indirect semantic parsing sample set, and a structural semantic parsing sample set; Based on the semantic parsing loss function, perform training resource adaptive adjustment learning on the data sample set in the same domain and the element semantic parsing sample set to obtain an element semantic parsing network; Based on the semantic parsing loss function, perform training resource adaptive adjustment learning on the data sample set in the same domain and the direct semantic parsing sample set to obtain a direct semantic parsing network; Based on the semantic parsing loss function, perform training resource adaptive adjustment learning on the data sample set in the same domain and the indirect semantic parsing sample set to obtain an indirect semantic parsing network; Based on the semantic parsing loss function, perform training resource adaptive adjustment learning on the data sample set in the same domain and the structural semantic parsing sample set to obtain a structural semantic parsing sample set parsing network; Connect the element semantic parsing network, the direct semantic parsing network, the indirect semantic parsing network, and the structural semantic parsing sample set parsing network to generate the semantic parsing composite model.
3. The data classification management method based on a multi-level semantic network according to claim 2, wherein Based on the semantic parsing loss function, perform training resource adaptive adjustment learning on the data sample set in the same domain and the element semantic parsing sample set to obtain an element semantic parsing network, including: Align the data sample set in the same domain and the element semantic parsing sample set to obtain an element semantic parsing construction set; Equally allocate the training resources for the element semantic parsing construction set to obtain the first result of training resource allocation; Based on the first result of training resource allocation, perform supervised training on the semantic parsing predetermined network according to the element semantic parsing construction set to obtain an initial element semantic parsing network; Based on the semantic parsing loss function, the initial network for element semantic parsing is tested according to the constructed set of element semantic parsing, and a semantic parsing loss sequence is obtained; If any semantic parsing loss coefficient in the semantic parsing loss sequence is greater than or equal to the semantic parsing loss threshold, the first result of training resource allocation is optimized and adjusted according to the semantic parsing loss sequence to obtain the second result of training resource allocation; Taking less than the semantic parsing loss threshold as the training target, the initial network for element semantic parsing is optimized and trained according to the second result of training resource allocation to obtain the element semantic parsing network.
4. The data classification management method based on a multi-level semantic network according to claim 1, wherein, The semantic parsing loss function is: SPL = log f SIM(SPY, SPO); where SPL represents the semantic parsing loss coefficient, f represents the predetermined factor of semantic parsing loss, 0 < f < 1, SPY represents the semantic parsing prediction data, SPO represents the semantic parsing sample data corresponding to the semantic parsing prediction data, and SIM(SPY, SPO) represents the similarity between the semantic parsing prediction data and the semantic parsing sample data.
5. The data classification and management method based on a multi-level semantic network according to claim 1, characterized in that It also includes: If any one of the multiple domain difference coefficients is greater than or equal to the domain difference threshold, the multiple data domain labels are clustered according to the multiple domain difference coefficients to obtain N domain label clusters, where N is a positive integer greater than 1; Based on the semantic parsing loss function, according to the N domain label clusters, the training resources of the multi-level semantic factors are adaptively adjusted and learned respectively, and N domain semantic parsing models are established; The second data stream to be processed is semantically parsed according to the N domain semantic parsing models.
6. The data classification management method based on a multi-level semantic network according to claim 1, characterized in that Disambiguation compensation is performed on the multiple data semantic vectors to obtain multiple semantic optimization vectors, including: Traverse the multiple data semantic vectors and extract the first data semantic vector; Perform fuzzy detection on the first data semantic vector to obtain the first semantic fuzzy detection result; Evaluate the ambiguity degree of the first semantic fuzzy detection result to obtain the first ambiguity coefficient; If the first ambiguity coefficient is greater than or equal to the predetermined ambiguity coefficient, semantic association backtracking is performed on the second data stream to be processed according to the first semantic fuzzy detection result to obtain the first semantic association feature; The first data semantic vector is disambiguated and optimized according to the first semantic association feature to obtain the first semantic optimization vector, and the first semantic optimization vector is added to the multiple semantic optimization vectors.
7. The data classification management method based on a multi-level semantic network according to claim 1, characterized in that Perform basic feature inspection on the first data stream to be processed to obtain the basic feature inspection result, including: Perform error detection on the first data stream to be processed to obtain the error detection results of each data; Perform integrity detection on each data in the first data stream to be processed to obtain the integrity detection results of each data; Output the error detection results of each data and the integrity detection results of each data as the basic feature inspection result.
8. The data classification management method based on a multi-level semantic network according to claim 1, characterized in that Before performing basic feature inspection on the first data stream to be processed, it includes: Collect the language format information and storage format information of the first data stream to be processed to obtain the first format feature information; Judge whether the first format feature information meets the predetermined format constraint; If the first format feature information does not meet the predetermined format constraint, use the format difference feature between the first format feature information and the predetermined format constraint as the format optimization factor; Optimize the format of the first data stream to be processed according to the format optimization factor.
9. The data classification management method based on a multi-level semantic network according to claim 1, characterized in that The multi-level semantic factors include element semantics, direct semantics, indirect semantics, and structural semantics.
10. A data classification and management system based on a multi-level semantic network, characterized in that, The system is used to implement the data classification management method based on the multi-level semantic network according to any one of claims 1 to 9. The system includes: A basic feature verification module, configured to perform basic feature verification on the first data stream to be processed, obtain a basic feature verification result, and optimize the first data stream to be processed according to the basic feature verification result to obtain a second data stream to be processed; A difference evaluation module, configured to perform difference evaluation on multiple data domain labels corresponding to the second data stream to be processed to obtain multiple domain difference coefficients; A fused domain label obtaining module, configured to fuse the multiple data domain labels to obtain a fused domain label if the multiple domain difference coefficients are all smaller than the domain difference threshold; A semantic parsing composite model building module, configured to perform adaptive adjustment learning of training resources on the multi-level semantic factors based on a semantic parsing loss function according to the fused domain label, and build a semantic parsing composite model; A data semantic vector building module, configured to input the second data stream to be processed into the semantic parsing composite model to build multiple data semantic vectors; A semantic relationship graph construction module, configured to perform disambiguation compensation according to the multiple data semantic vectors to obtain multiple semantic optimization vectors, and construct a semantic relationship graph according to the multiple semantic optimization vectors; A classification management module, configured to perform classification management on the second data stream to be processed according to the semantic relationship graph.