Airline company data grading method, device, equipment, medium and product
By working together with metadata lineage, feature grading models, and natural language recognition models, the problem of low efficiency in manual identification of airline data grading has been solved, achieving automatic grading and improved accuracy of data sensitivity levels, and promoting the open sharing of data assets.
Patent Information
- Application Number
- CN202511157940.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-19
- Publication Date
- 2025-12-16
AI Technical Summary
Existing technologies for classifying airline data suffer from low efficiency, long cycles, and a high risk of omissions due to manual identification, making it difficult to improve the accuracy and efficiency of data sensitivity classification and hindering the sharing and openness of data assets.
The system employs multiple collaborative methods, including metadata lineage, feature grading models, and natural language recognition models, to automatically grade metadata. This includes lineage tracing grading, feature model grading, sensitive word recognition grading, and data recognition grading. It comprehensively processes airline flight data, airport data, and personal information, and achieves automatic grading through grading models and verification.
It enables automatic metadata classification, reduces manual costs, improves the efficiency of data sensitivity level confirmation, and promotes open sharing of enterprise data.
Smart Images

Figure CN121145845A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of data management, in particular to an airline data grading method, device, equipment, medium and product. BACKGROUND
[0002] In large airline data governance, sensitive grading of data is a very important link, which needs to identify the sensitive level of data for security protection, so as to facilitate data asset sharing and open management, and improve enterprise data governance level. At present, the commonly used data classification and grading technology mainly includes: (1) Establishment of data classification and grading rule library: this is the basic ability of automatic identification, and the rule library adopts technologies including keyword, regular expression, file attribute recognition, and metadata information-based custom recognition; (2) Data scanning, identification and secret level marking: through scanning structured / semi-structured / non-structured data, automatically discovering the category, level and other attribute information of sensitive data and the storage location, and forming a sensitive data asset distribution map; (3) Text processing technology: based on a series of natural language processing capabilities provided by HanLP, including but not limited to word segmentation, part-of-speech tagging, named entity recognition, dependency syntax analysis, and semantic role labeling technologies; (4) Data security technology: the development of data security technology, such as data encryption, data desensitization, data leakage prevention, data tracking and tracing, and database security prevention, provides technical support for data classification and grading; (5) Data security regulations and standards: regulations and standards such as the Network Security Law of the People's Republic of China and the GDPR of the European Union require enterprises to classify and manage data to meet compliance requirements; (6) Data bloodline tracking technology: by establishing a data bloodline map, the data flow process is recorded in real time, the cause of the data security event is tracked and analyzed, and the security risk is reduced.
[0003] Through the above technologies and standards, the aviation industry can effectively classify and grade the shared directory data, but the manual sorting work of data grading is large, the traditional manual identification and confirmation method is relatively low in efficiency, the implementation cycle is relatively long, and it is easy to miss due to negligence. Therefore, it is urgent to solve how to improve the accuracy of data grading and improve the grading efficiency, so as to better promote the open sharing of data assets. SUMMARY
[0004] The present application provides an airline data grading method, device, equipment, medium and product, which can realize batch processing of data sets to be graded, realize automatic metadata grading, reduce manual cost, improve the efficiency of data sensitive level confirmation, and promote enterprise data open sharing through the cooperation of metadata blood relationship, feature grading model, natural semantic recognition model and other grading methods in the aviation industry.
[0005] To achieve the above objectives, embodiments of the present invention provide an airline data classification method, including:
[0006] Obtain metadata information from airlines; wherein, the metadata information includes flight data, airport data, operations data, and personal information;
[0007] The metadata information is subjected to data classification processing to obtain the classification recommendation results of the metadata information; wherein, the data classification processing includes lineage tracing classification, feature model classification, sensitive word recognition classification, and data recognition classification;
[0008] The tiered recommendation results are reviewed and confirmed to obtain the final tiered recommendation results of the metadata information, thereby updating the airline's metadata tiered result database.
[0009] As an improvement to the above scheme, the step of performing data classification processing on the metadata information to obtain the classification recommendation result of the metadata information includes:
[0010] The first classification result of the metadata information is determined by tracing its lineage.
[0011] The metadata information is input into a trained feature classification model for classification processing. After matching by the feature classification model, a second classification result of the metadata information is obtained.
[0012] The metadata information is subjected to sensitive word identification and classification processing to obtain the third classification result of the metadata information;
[0013] The metadata information is subjected to data identification and classification processing to obtain the fourth classification result of the metadata information;
[0014] By comprehensively comparing the results of the first, second, third, and fourth levels, the hierarchical recommendation results of the metadata information are obtained.
[0015] As an improvement to the above scheme, the step of determining the first hierarchical result of the metadata information by tracing its lineage includes:
[0016] The metadata information is used as a character set to be classified, and the metadata lineage of each field to be classified in the character set is traced.
[0017] If each field to be classified has a superior lineage relationship and the superior lineage relationship contains classification data, then each field to be classified is classified according to the field classification corresponding to the superior lineage relationship, until all fields to be classified in the character set to be classified are traversed to obtain the first classification result of the metadata information.
[0018] As an improvement to the above scheme, the step of performing sensitive word identification and classification processing on the metadata information to obtain the third classification result of the metadata information includes:
[0019] Sensitive word identification and classification processing is performed on the metadata information. If there is preset sample data in the metadata information, data sensitivity rules are identified on the fields to be classified corresponding to the metadata information according to the preset sample data to obtain the third classification result of the metadata information.
[0020] As an improvement to the above scheme, the step of performing data identification and classification processing on the metadata information to obtain the fourth classification result of the metadata information includes:
[0021] The metadata information is classified according to the regular expression with defined sensitivity levels, and the fourth classification result of the metadata information is output.
[0022] As an improvement to the above scheme, the step of comprehensively comparing the first, second, third, and fourth level results to obtain the level recommendation result of the metadata information includes:
[0023] The first, second, third, and fourth classification results are sorted and analyzed, and the classification result with the highest ranking is taken as the classification recommendation result of the metadata information.
[0024] To achieve the above objectives, embodiments of the present invention provide an airline data classification device, comprising:
[0025] The data acquisition module is used to acquire metadata information from airlines; wherein, the metadata information includes flight data, airport data, operations data, and personal information.
[0026] The data information classification module is used to perform data classification processing on the metadata information to obtain the classification recommendation results of the metadata information; wherein, the data classification processing includes lineage tracing classification, feature model classification, sensitive word recognition classification, and data recognition classification;
[0027] The grading result review module is used to review and confirm the grading recommendation results, obtain the final grading recommendation results of the metadata information, and update the metadata grading result database of the airline.
[0028] To achieve the above objectives, embodiments of the present invention provide an airline data classification device, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements the above-described airline data classification method.
[0029] To achieve the above objectives, embodiments of the present invention also provide a computer-readable storage medium, the computer-readable storage medium including a stored computer program, wherein, when the computer program is executed, it controls the device where the computer-readable storage medium is located to execute the above-described airline data classification method.
[0030] To achieve the above objectives, embodiments of the present invention also provide a computer program product, which is stored in a storage medium and executed by at least one processor to implement the steps of the above-described airline data classification method.
[0031] Compared with existing technologies, this invention discloses a method, apparatus, device, medium, and product for classifying airline data. This involves acquiring metadata information from airlines, including flight data, airport data, operational control data, and personal information. The metadata information is then processed for classification to obtain classification recommendation results. These data classification processes include lineage-based classification, feature model classification, sensitive word recognition classification, and data recognition classification. The classification recommendation results are then reviewed and confirmed to obtain the final classification recommendation results for the metadata information, thereby updating the airline's metadata classification result database. By employing multiple classification methods such as metadata lineage, feature classification models, and natural language processing models in a collaborative manner, this invention enables batch processing of datasets to be classified in the aviation industry when data sharing and openness are implemented. This achieves automatic metadata classification, reduces manual costs, improves the efficiency of data sensitivity level confirmation, and promotes open data sharing among enterprises. Attached Figure Description
[0032] Figure 1 This is a flowchart illustrating an airline data classification method provided in an embodiment of the present invention;
[0033] Figure 2 This is a schematic diagram of the structure of an airline data classification device provided in an embodiment of the present invention;
[0034] Figure 3 This is a structural block diagram of an airline data classification device provided in an embodiment of the present invention. Detailed Implementation
[0035] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0036] It should be noted that the terms "comprising" and "specific" in this invention, and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not necessarily limited to those steps or units that are explicitly listed, but may include other steps or units that are not explicitly listed or that are inherent to such process, method, product, or device.
[0037] Please see Figure 1 , Figure 1 This is a flowchart illustrating an airline data classification method provided in an embodiment of the present invention. The airline data classification method includes:
[0038] S1, Obtain the airline's metadata information; wherein, the metadata information includes flight data, airport data, operation and control data, and personal information;
[0039] S2, perform data classification processing on the metadata information to obtain the classification recommendation results of the metadata information; wherein, the data classification processing includes lineage tracing classification, feature model classification, sensitive word recognition classification, and data recognition classification;
[0040] S3, review and confirm the hierarchical recommendation results to obtain the final hierarchical recommendation results of the metadata information, so as to update the metadata hierarchical result database of the airline.
[0041] For example, the airline data classification method described in this embodiment of the invention can be implemented by a data classification server capable of interacting with target users. The data classification server acquires metadata information from the airline; wherein the metadata information includes flight data, airport data, operational control data, and personal information; it performs data classification processing on the metadata information to obtain classification recommendation results; wherein the data classification processing includes lineage tracing classification, feature model classification, sensitive word recognition classification, and data recognition classification; the classification recommendation results are reviewed and confirmed to obtain the final classification recommendation results for the metadata information, thereby updating the airline's metadata classification result database. By using multiple classification methods such as metadata lineage, feature classification models, and natural language recognition models in a collaborative manner, the system can handle batch processing of datasets to be classified when conducting data sharing and openness in the aviation industry, achieving automatic metadata classification, reducing manual costs, improving the efficiency of data sensitivity level confirmation, and promoting enterprise data openness and sharing.
[0042] Specifically, step S2 includes:
[0043] S21, perform lineage tracing on the metadata information to determine the first classification result of the metadata information;
[0044] S22, The metadata information is input into the trained feature classification model for classification processing. After matching by the feature classification model, the second classification result of the metadata information is obtained.
[0045] S23, perform sensitive word identification and classification processing on the metadata information to obtain the third classification result of the metadata information;
[0046] S24, perform data identification and classification processing on the metadata information to obtain the fourth classification result of the metadata information;
[0047] S25. By comprehensively comparing the first, second, third, and fourth level results, the level recommendation result of the metadata information is obtained.
[0048] For example, in step S22, the feature classification model is first trained, and the classified sensitive classification data is used as a corpus. Basic standard data and field data covering all sensitive classifications are collected and divided into different folders or tag files according to level (level 1, level 2, level 3, and level 4). The data is cleaned using HanLP, including removing irrelevant characters, punctuation marks, stop words, etc. Word segmentation and part-of-speech tagging are then performed. Using the Naive Bayes classifier training interface provided by HanLP, the model learns how to distinguish different levels based on text features. It can be understood that the Naive Bayes classifier is a classification method based on Bayes' theorem and the assumption of conditional independence of features. It assumes that each feature is unrelated to other features, and then uses Bayes' theorem to calculate the probability of a given sample belonging to each category, selecting the category with the highest probability as the prediction result. Bayesian principle: Given a category (y) and a feature vector (x1, x2, ..., x...), ... n ), representing the conditional probability (P(y|x1,x2,…,x)). n The following methods can be used to calculate: Where P(y) is the prior probability of category (y); P(x1,x2,…,x n |y) is the feature vector (x1, x2, ..., xy) under a given category (y). n The conditional probability of ) P(x1,x2,…,x n ) is the eigenvector (x1, x2, ..., x) n The prior probability of a feature is usually considered constant because all samples in a given dataset have been observed. Naive Bayes assumes that features are conditionally independent. This assumption greatly simplifies computation because we can calculate the conditional probability of each feature individually without considering combinations of features. That is:
[0049] P(x1,x2,…,x n|y)=P(x1|y)P(x2|y)…P(x n |y).
[0050] For a new sample, the Naive Bayes classifier calculates the posterior probability P(y|x1,x2,…,x) that it belongs to each class. n Then, the class with the highest posterior probability is selected as the predicted class. A separate dataset (test set) is used to evaluate the model's performance. Common evaluation metrics include accuracy, precision, and recall. Model parameters are adjusted based on the evaluation results to improve performance.
[0051] Accuracy represents the proportion of samples correctly classified by a classifier out of the total number of samples. The formula is as follows: Accuracy = (TP + TN) / (TP + FP + TN + FN), where TP represents true positives (the number of samples correctly classified as positive), TN represents true negatives (the number of samples correctly classified as negative), FP represents false positives (the number of samples incorrectly classified as positive), and FN represents false negatives (the number of samples incorrectly classified as negative).
[0052] Precision represents the proportion of samples correctly identified as positive out of all samples predicted as positive. It is typically used to focus on false positives. The formula is as follows: Precision = TP / (TP + FP).
[0053] Recall represents the proportion of samples correctly identified as positive out of all true positive samples. It refers to the proportion of samples correctly predicted as positive by the classifier out of all samples that are actually positive. It is usually used to focus on the case of false negatives. The formula is as follows: Recall = TP / (TP + FN).
[0054] For example, metadata information is used as the character set T1 to be classified. Based on the optimal feature model method, the fields to be classified in the character set T1 are input into the model. The model will automatically classify them into the corresponding level and category according to the learned knowledge, and store the classification results in the character set W2. The status variable M2 of the recognition result is updated to 1, and the status variable M2 of the classification result that was not successfully recognized is set to 0.
[0055] Specifically, step S21 includes:
[0056] The metadata information is used as a character set to be classified, and the metadata lineage of each field to be classified in the character set is traced.
[0057] If each field to be classified has a superior lineage relationship and the superior lineage relationship contains classification data, then each field to be classified is classified according to the field classification corresponding to the superior lineage relationship, until all fields to be classified in the character set to be classified are traversed to obtain the first classification result of the metadata information.
[0058] For example, metadata lineage tracing is performed on the field to be classified in character set T1. The field to be classified is checked for whether there is a superior lineage relationship. The search is traversed. If there is a superior lineage relationship and the superior has sensitive classification data, the superior classification result W1 is obtained. If there is no sensitive classification data in the superior, the tracing continues upward until the superior data classification result is obtained. If not found, the character set W1 is set to empty. Define a character variable M1 for the status of the metadata lineage tracing result. Determine whether W1 is NULL. If it is empty, record the classification status as 0, that is, M1 = 0. If W1 has a value, record the classification status as 1, that is, M1 = 1.
[0059] Specifically, step S23 includes:
[0060] Sensitive word identification and classification processing is performed on the metadata information. If there is preset sample data in the metadata information, data sensitivity rules are identified on the fields to be classified corresponding to the metadata information according to the preset sample data to obtain the third classification result of the metadata information.
[0061] For example, determine whether sample data exists. If so, continue with the identification of data sensitivity rules. Input the character set T1 to be identified. Based on the sample data, determine whether the data to be classified is sensitive data. Store the identification result in the character set W3. At the same time, update the identification result status character variable M3 = 1. If the classification result is not successfully identified, set the classification result status variable M3 to 0.
[0062] Specifically, step S24 includes:
[0063] The metadata information is classified according to the regular expression with defined sensitivity levels, and the fourth classification result of the metadata information is output.
[0064] For example, regular expressions are used to define and confirm the sensitivity levels of sensitive personal information in the aviation industry, such as mobile phone numbers, names, ID cards, and addresses, as well as sensitive corporate data such as flight data, airport data, and operational control data. Inputting the character set T1 to be classified, the system identifies the fields to be classified based on the defined sensitivity levels using regular expressions. The identification results are stored in character set W4, and the status variable M4 is updated to 1 for identification results. If the classification result is unsuccessful, the status variable M4 is set to 0.
[0065] It's worth noting that the core of regular expression processing is the character set and metacharacters. The character set defines the range of characters to be matched, while metacharacters provide rich matching control. When we write regular expressions in a program, we are actually defining a rule that tells the computer how to recognize and match strings. First, let's understand the matching process of regular expressions. When a regular expression matches a string, the computer checks the string character by character in a specific order. Examples of metacharacters are shown in Table 1.
[0066] Table 1 Examples of Metacharacters
[0067] Meta-Character Name Meta-Character Description Line Feed \n Matches the line feed character, which is used to find the end of a line in text files. Windows generally uses \r\n Carriage Return \r See above Tab \t Matches the tab character, which is used to indent text in many programs Space \s Matches the space character, which is used to separate words in text Digit \d Matches any digit, 0-9 Matches any letter, a-z and A-Z \w Equivalent to [a-Za-z0-9_] Any Character . Matches any character
[0068] Regular expression example: (<[Aa]\s+[^>]+>\s*)? <[Ii][Mm][Gg]\s+[^>]+>(?(1)\s*)< / [Aa]> The matching process of a regular expression: Preprocessing stage: Before matching begins, the computer performs preprocessing operations on the regular expression, such as removing leading and trailing spaces and escaping special characters, to ensure the correct format. Matching stage: The computer compares each character of the regular expression from left to right. When a matching character is encountered, the matching process continues; otherwise, other paths are tried. This stage uses the character set and metacharacters mentioned earlier, such as "." representing any character, and "*" indicating that the preceding character can be repeated zero or more times. Post-processing stage: After matching is complete, the computer performs post-processing operations on the results, such as extracting the matched text and processing capture groups.
[0069] Specifically, step S25 includes:
[0070] The first, second, third, and fourth classification results are sorted and analyzed, and the classification result with the highest ranking is taken as the classification recommendation result of the metadata information.
[0071] For example, the status values M1, M2, M3, and M4 of the classification results for the four classification methods are determined. All classification results with a result of 1 are placed into character set W5. The W5 character set is sorted and analyzed to obtain the classification result value with the most occurrences. The final classification result variable H1 is defined, and the classification recommendation result is output to H1 to obtain the classification recommendation result of the metadata information.
[0072] This invention discloses a method for classifying airline data. The method involves acquiring metadata information from airlines, including flight data, airport data, operational control data, and personal information. The metadata information is then processed for classification to obtain classification recommendation results. These data classification processes include lineage-based classification, feature model classification, sensitive word recognition classification, and data identification classification. The classification recommendation results are then reviewed and confirmed to obtain the final classification recommendation results for the metadata information, thereby updating the airline's metadata classification result database. By employing multiple classification methods such as metadata lineage, feature-based classification models, and natural language processing models in collaboration, this method enables batch processing of datasets to be classified in the aviation industry when data sharing and openness are implemented. This achieves automatic metadata classification, reduces manual costs, improves the efficiency of data sensitivity level confirmation, and promotes open data sharing among enterprises.
[0073] See Figure 2 , Figure 2 This is a schematic diagram of the structure of an airline data classification device 10 provided in an embodiment of the present invention. The airline data classification device 10 includes:
[0074] The data information acquisition module 11 is used to acquire the airline's metadata information; wherein, the metadata information includes flight data, airport data, operation and control data, and personal information;
[0075] The data information classification module 12 is used to perform data classification processing on the metadata information to obtain the classification recommendation result of the metadata information; wherein, the data classification processing includes lineage tracing classification, feature model classification, sensitive word recognition classification, and data recognition classification.
[0076] The grading result review module 13 is used to review and confirm the grading recommendation results to obtain the final grading recommendation results of the metadata information, so as to update the metadata grading result database of the airline.
[0077] The airline data classification device 10 provided in this embodiment of the invention can realize all the processes of the airline data classification method in the above embodiment. The functions and technical effects of each module in the device are the same as the functions and technical effects of the airline data classification method in the above embodiment, and will not be repeated here.
[0078] See Figure 3 , Figure 3This is a schematic diagram of the structure of an airline data classification device 20 provided in an embodiment of the present invention. The airline data classification device 20 of this embodiment includes: a processor 21, a memory 22, and a computer program stored in the memory 22 and executable on the processor 21. When the processor 21 executes the computer program, it implements the steps in the above-described airline data classification method embodiment. Alternatively, when the processor 21 executes the computer program, it implements the functions of each module in the above-described airline data classification device embodiment.
[0079] For example, the computer program may be divided into one or more modules, which are stored in the memory 22 and executed by the processor 21 to complete the present invention. The one or more modules may be a series of computer program instruction segments capable of performing specific functions, which describe the execution process of the computer program in the airline data classification device 20.
[0080] The airline data classification device 20 can be a computing device such as a desktop computer, laptop, handheld computer, or cloud server. The airline data classification device 20 may include, but is not limited to, a processor 21 and a memory 22. Those skilled in the art will understand that the schematic diagram is merely an example of the airline data classification device 20 and does not constitute a limitation on the airline data classification device 20. It may include more or fewer components than illustrated, or combine certain components, or use different components. For example, the airline data classification device 20 may also include input / output devices, network access devices, buses, etc.
[0081] The processor 21 may be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor. The processor 21 is the control center of the airline data classification equipment 20, connecting all parts of the airline data classification equipment 20 via various interfaces and lines.
[0082] The memory 22 can be used to store the computer programs and / or modules. The processor 21 implements various functions of the airline data classification device 20 by running or executing the computer programs and / or modules stored in the memory 22 and calling the data stored in the memory 22. The memory 22 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the mobile phone (such as audio data, phonebook, etc.). In addition, the memory 22 may include high-speed random access memory, and may also include non-volatile memory, such as hard disk, memory, plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, at least one disk storage device, flash memory device, or other volatile solid-state storage device.
[0083] If the modules integrated into the airline data classification device 20 are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by the processor 21, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording media, USB flash drives, portable hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content contained in the computer-readable medium may be appropriately added to or subtracted from the content as required by the legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium may not include electrical carrier signals and telecommunication signals.
[0084] It should be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Furthermore, in the accompanying drawings of the device embodiments provided by this invention, the connection relationships between modules indicate that they have communication connections, which can be specifically implemented as one or more communication buses or signal lines. Those skilled in the art can understand and implement this without any creative effort.
[0085] This invention also provides a computer-readable storage medium comprising a stored computer program, wherein the computer program, when running, controls the device containing the computer-readable storage medium to execute the airline data classification method as described above.
[0086] Furthermore, embodiments of the present invention also provide a computer program product, which is stored in a storage medium and executed by at least one processor to implement the steps of the airline data classification method described above.
[0087] The above description represents the preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications are also considered to be within the scope of protection of the present invention.
Claims
1. A method for classifying airline data, characterized in that, include: Obtain metadata information from airlines; wherein, the metadata information includes flight data, airport data, operations data, and personal information; The metadata information is subjected to data classification processing to obtain the classification recommendation results of the metadata information; wherein, the data classification processing includes lineage tracing classification, feature model classification, sensitive word recognition classification, and data recognition classification; The tiered recommendation results are reviewed and confirmed to obtain the final tiered recommendation results of the metadata information, thereby updating the airline's metadata tiered result database.
2. The airline data classification method as described in claim 1, characterized in that, The step of performing data classification processing on the metadata information to obtain the classification recommendation result of the metadata information includes: The first classification result of the metadata information is determined by tracing its lineage. The metadata information is input into a trained feature classification model for classification processing. After matching by the feature classification model, a second classification result of the metadata information is obtained. The metadata information is subjected to sensitive word identification and classification processing to obtain the third classification result of the metadata information; The metadata information is subjected to data identification and classification processing to obtain the fourth classification result of the metadata information; By comprehensively comparing the results of the first, second, third, and fourth levels, the hierarchical recommendation results of the metadata information are obtained.
3. The airline data classification method as described in claim 1, characterized in that, The process of determining the first hierarchical result of the metadata information by tracing its lineage includes: The metadata information is used as a character set to be classified, and the metadata lineage of each field to be classified in the character set is traced. If each field to be classified has a superior lineage relationship and the superior lineage relationship contains classification data, then each field to be classified is classified according to the field classification corresponding to the superior lineage relationship, until all fields to be classified in the character set to be classified are traversed to obtain the first classification result of the metadata information.
4. The airline data classification method as described in claim 1, characterized in that, The process of performing sensitive word identification and classification on the metadata information to obtain the third classification result of the metadata information includes: Sensitive word identification and classification processing is performed on the metadata information. If there is preset sample data in the metadata information, data sensitivity rules are identified on the fields to be classified corresponding to the metadata information according to the preset sample data to obtain the third classification result of the metadata information.
5. The airline data classification method as described in claim 4, characterized in that, The step of performing data identification and classification processing on the metadata information to obtain the fourth classification result of the metadata information includes: The metadata information is classified according to the regular expression with defined sensitivity levels, and the fourth classification result of the metadata information is output.
6. The airline data classification method as described in claim 1, characterized in that, The hierarchical recommendation result of the metadata information is obtained by comprehensively comparing the first, second, third, and fourth level results, including: The first, second, third, and fourth classification results are sorted and analyzed, and the classification result with the highest ranking is taken as the classification recommendation result of the metadata information.
7. An airline data classification device, characterized in that, include: The data acquisition module is used to acquire metadata information from airlines; wherein, the metadata information includes flight data, airport data, operations data, and personal information. The data information classification module is used to perform data classification processing on the metadata information to obtain the classification recommendation results of the metadata information; wherein, the data classification processing includes lineage tracing classification, feature model classification, sensitive word recognition classification, and data recognition classification; The grading result review module is used to review and confirm the grading recommendation results, obtain the final grading recommendation results of the metadata information, and update the metadata grading result database of the airline.
8. An airline data classification device, characterized in that, It includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein the processor, when executing the computer program, implements the airline data classification method as described in any one of claims 1-6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored computer program, wherein, when the computer program is executed, it controls the device on which the computer-readable storage medium is located to perform the airline data classification method as described in any one of claims 1-6.
10. A computer program product, characterized in that, The computer program product is stored in a storage medium, and the program product is executed by at least one processor to implement the steps of the airline data classification method as described in any one of claims 1-6.