An Automatic Data Classification and Grading Method and System Based on NLP Algorithm Model
Through the automatic data classification and grading method based on the NLP algorithm model, the high cost and inefficiency problems caused by manual operations in the traditional method are solved, and more efficient and accurate data classification and grading are achieved.
Patent Information
- Application Number
- CN202211254591.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-13
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2042-10-13
AI Technical Summary
Traditional data classification and grading methods rely on manual operations, resulting in high cost, low efficiency and low accuracy, especially when the data is complex and huge.
An automatic data classification and grading method based on the NLP algorithm model is adopted to realize automated data classification and grading by determining standard elements, adding identification rules, training NLP models and labeling classification tags.
It improves the accuracy and efficiency of data classification and grading, reduces the consumption of manpower and financial resources, and achieves efficient, flexible and intelligent data classification and grading.
Smart Images

Figure CN115544256B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data security technology, and specifically relates to an automatic data classification and grading method and system based on an NLP algorithm model. Background Art
[0002] Article 21 of the Data Security Law clearly stipulates that the state shall establish a data classification and grading protection system, and classify and grade the protection of data security according to the importance of the data in economic and social development, and the degree of harm caused to national security, public interests, or the legitimate rights and interests of individuals and organizations in the event of tampering, destruction, leakage, illegal acquisition, or illegal use. Implementing data classification and grading is a prerequisite for ensuring data security and an extremely important part of the data security governance process.
[0003] The traditional approach is for business personnel to sort out the classification and grading according to the classification and grading guidelines promulgated by the state or local governments. Since the implementation of classification and grading requires business personnel to have experience in implementing classification and grading, and there is a lack of relevant standards in various industries, the accuracy of classification and grading is relatively low. Moreover, the data of some governments or enterprises is very complex and huge, and solely relying on manual classification and grading requires a large amount of manpower and financial resources, with high costs and very low efficiency.
[0004] In view of this, the present invention proposes an automatic data classification and grading method and system based on an NLP algorithm model, which can automatically classify and grade, improving the sorting efficiency and the accuracy of classification and grading. Summary of the Invention
[0005] In order to solve the problems of relying on manual classification and grading, which requires a large amount of manpower and financial resources, high costs and very low efficiency, etc., this application provides an automatic data classification and grading method and system based on an NLP algorithm model to solve the above technical defect problems.
[0006] According to one aspect of the present invention, an automatic data classification and grading method based on an NLP algorithm model is proposed, and the method includes the following steps:
[0007] S1. Determine standard elements, and configure the classification catalogs to which the standard elements belong according to the classification and grading standards;
[0008] S2. Add recognition rules to the standard elements, set the credibility and priority of each recognition rule, and the recognition rules include recognition rules of traditional algorithms and recognition rules of NLP algorithms;
[0009] S3. Train the NLP model based on the recognition rules, execute the matching logic in descending order of the priority of the recognition rules, perform NLP algorithm matching or traditional algorithm matching on the preprocessed data, and obtain multiple matching results; and
[0010] S4. Find the standard element corresponding to the result with the highest matching degree from multiple matching results, and mark the classification label for the standard element corresponding to the result with the highest matching degree.
[0011] In a specific embodiment, in step S3, training the NLP model based on the recognition rules specifically includes the following sub-steps:
[0012] S31. Convert all the field names of the preprocessed data into lowercase letters, and remove special symbols and numbers;
[0013] S32. Split the data preprocessed in step S31 according to spaces and punctuation marks;
[0014] S33. Determine whether the data split in step S32 is pinyin or English through the language model and the norms of the initials and finals of Chinese pinyin;
[0015] S34. Perform word segmentation on the data processed in step S32 to obtain multiple combination results;
[0016] S35. Infer and complete the combination results obtained in step S34 according to the content of the English word library or the writing norms of Chinese pinyin to finally obtain the matching results.
[0017] In a specific embodiment, in step S2, the recognition rules of the traditional algorithm include: content-based matching, field annotation matching, exact matching of field names, fuzzy matching of field names, prefix matching, suffix matching, and regular matching.
[0018] In a specific embodiment, in step S2, set the credibility of each recognition rule. The credibility is set as a numerical value, which represents the degree of dependence of the recognition rule and affects the matching degree of the recognition rule.
[0019] In a specific embodiment, in step S2, set the priority of each recognition rule. The priority is set as a numerical value, and the larger the value, the higher the priority. The matching is executed in order from high to low according to the priority.
[0020] In a specific embodiment, in step S2, add recognition rules to the standard element, set the credibility and priority of each recognition rule. The recognition rules include the recognition rules of the traditional algorithm and the NLP algorithm, including:
[0021] S211. Add the recognition rule R1 of the traditional algorithm and the recognition rule R2 of the NLP algorithm;
[0022] S212. Set the matching type of the recognition rule R1 of the traditional algorithm to "field name matching", and set the matching method of the recognition rule R1 of the traditional algorithm to "regular matching"; set the matching method of the recognition rule R2 of the NLP algorithm to "NLP matching"; and
[0023] S213. Respectively set the credibility and priority of the recognition rule R1 of the traditional algorithm and the recognition rule R2 of the NLP algorithm, and the priority of the recognition rule R1 of the traditional algorithm is greater than the priority of the recognition rule R2 of the NLP algorithm.
[0024] In a specific embodiment, in step S2, add recognition rules to the standard elements, set the credibility and priority of each recognition rule, and the recognition rules include the recognition rules of the traditional algorithm and the recognition rules of the NLP algorithm, including:
[0025] S221. Add the recognition rule R1 of the traditional algorithm and the recognition rule R2 of the NLP algorithm;
[0026] S222. Set the matching type of the recognition rule R1 of the traditional algorithm to "field name matching", and set the matching method of the recognition rule R1 of the traditional algorithm to "regular matching"; set the matching method of the recognition rule R2 of the NLP algorithm to "NLP matching"; and
[0027] S223. Respectively set the credibility and priority of the recognition rule R1 of the traditional algorithm and the recognition rule R2 of the NLP algorithm, and the priority of the recognition rule R1 of the traditional algorithm is less than the priority of the recognition rule R2 of the NLP algorithm.
[0028] In a second aspect, the present application proposes an automatic data classification and grading system based on an NLP algorithm model, and the system includes:
[0029] A standard element determination module, configured to determine a standard element and configure a classification directory to which the standard element belongs according to the classification and grading standard;
[0030] An identification rule adding module, configured to add identification rules to the standard elements, set the credibility and priority of each identification rule, and the identification rules include the identification rules of the traditional algorithm and the identification rules of the NLP algorithm;
[0031] A matching module, training an NLP model based on the identification rules, sequentially executing matching logic from high to low according to the priority of the identification rules, performing NLP algorithm matching or traditional algorithm matching on the preprocessed data, and obtaining multiple matching results; and
[0032] A marking module, finding the standard element corresponding to the result with the highest matching degree from multiple matching results, and marking a classification label for the standard element corresponding to the result with the highest matching degree.
[0033] In a specific embodiment, in the matching module, training the NLP model based on the recognition rules specifically includes the following sub-steps:
[0034] S31. Convert all field names of the preprocessed data to lowercase letters, and remove special symbols and numbers;
[0035] S32. Split the data preprocessed in step S31 according to spaces and punctuation marks;
[0036] S33. Determine whether the data split in step S32 is pinyin or English through a language model and the norms of initials and finals of Chinese pinyin;
[0037] S34. Perform word segmentation on the data processed in step S32 to obtain multiple combination results;
[0038] S35. Infer and complete the combination results obtained in step S34 according to the content of the English word library or the writing norms of Chinese pinyin to finally obtain the matching results.
[0039] In a third aspect, the present application provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the method described in any one of the above is implemented.
[0040] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0041] For the NLP model training method provided by the present invention, the government or enterprise can improve the matching degree of the program by training its own specific data set, thereby improving the accuracy of classification and grading; through the NLP algorithm and traditional algorithms, automatic classification and grading can be performed, replacing manual classification, and realizing efficient, flexible, and intelligent classification and grading; for the classification and grading results, create an NLP learning model training task and establish a knowledge algorithm, which can realize continuous and rapid data classification and grading capabilities. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] By reading the detailed description of the non-limiting embodiments with reference to the following drawings, other features, objects, and advantages of the present application will become more obvious:
[0043] Figure 1 is a flowchart of the automatic data classification and grading method based on the NLP algorithm model according to the present application;
[0044] Figure 2 is a schematic diagram of the main framework of the automatic data classification and grading method based on the NLP algorithm model according to the present application;
[0045] Figure 3 is a flowchart of the training of the NLP algorithm model according to the present application;
[0046] Figure 4 It is a schematic diagram of an automatic data classification and grading system based on an NLP algorithm model according to the present application;
[0047] Figure 5 It is a schematic structural diagram of a computer system of an electronic device suitable for implementing the embodiments of the present application. Detailed implementation manners
[0048] The present application will be further described in detail below with reference to the accompanying drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the related invention, rather than limiting the invention. Additionally, it should be noted that for the sake of description, only parts related to the relevant invention are shown in the drawings.
[0049] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other. The present application will be described in detail below with reference to the drawings and embodiments.
[0050] Figure 1 shows a flowchart of an automatic data classification and grading method based on an NLP algorithm model according to the present application, Figure 2 shows a schematic diagram of the main framework of an automatic data classification and grading method based on an NLP algorithm model according to the present application. With reference to Figure 1 and Figure 2 , the method includes the following steps:
[0051] S1. Determine standard elements, and configure the classification catalogs to which the standard elements belong according to the classification and grading criteria.
[0052] In this embodiment, the classification catalogs to which the standard elements belong can be configured according to the classification and grading criteria of the country or local government.
[0053] S2. Add recognition rules to the standard elements, set the credibility and priority of each recognition rule. The recognition rules include recognition rules of traditional algorithms and recognition rules of NLP algorithms.
[0054] In this embodiment, each recognition rule has a corresponding credibility and priority. Credibility refers to the degree to which this rule can be relied on. It is set as a numerical value, which will ultimately affect the matching degree of this rule. Priority refers to the execution order of this rule. It is set as a numerical value. The larger the numerical value, the higher the priority. The program will match and execute in order from high to low according to the priority.
[0055] The recognition rules of traditional algorithms include: content-based matching, field annotation matching, exact matching of field names, fuzzy matching of field names, prefix matching, suffix matching, regular matching.
[0056] In one embodiment, in step S2, recognition rules are added to the standard elements, and the confidence level and priority of each recognition rule are set. The recognition rules include the recognition rules of traditional algorithms and the recognition rules of NLP algorithms, including:
[0057] S211. Add the recognition rule R1 of the traditional algorithm and the recognition rule R2 of the NLP algorithm;
[0058] S212. Set the matching type of the recognition rule R1 of the traditional algorithm as "field name matching", and set the matching method of the recognition rule R1 of the traditional algorithm as "regular matching"; set the matching method of the recognition rule R2 of the NLP algorithm as "NLP matching"; and
[0059] S213. Respectively set the confidence level and priority of the recognition rule R1 of the traditional algorithm and the recognition rule R2 of the NLP algorithm, and the priority of the recognition rule R1 of the traditional algorithm is greater than the priority of the recognition rule R2 of the NLP algorithm.
[0060] In another embodiment, in step S2, recognition rules are added to the standard elements, and the confidence level and priority of each recognition rule are set. The recognition rules include the recognition rules of traditional algorithms and the recognition rules of NLP algorithms, including:
[0061] S221. Add the recognition rule R1 of the traditional algorithm and the recognition rule R2 of the NLP algorithm;
[0062] S222. Set the matching type of the recognition rule R1 of the traditional algorithm as "field name matching", and set the matching method of the recognition rule R1 of the traditional algorithm as "regular matching"; set the matching method of the recognition rule R2 of the NLP algorithm as "NLP matching"; and
[0063] S223. Respectively set the confidence level and priority of the recognition rule R1 of the traditional algorithm and the recognition rule R2 of the NLP algorithm, and the priority of the recognition rule R1 of the traditional algorithm is less than the priority of the recognition rule R2 of the NLP algorithm.
[0064] S3. Train the NLP model based on the recognition rules, and sequentially execute the matching logic from high to low according to the priority of the recognition rules, and perform NLP algorithm matching or traditional algorithm matching on the preprocessed data to obtain multiple matching results.
[0065] Figure 3 The flowchart of the training of the NLP algorithm model according to the present application is shown. Refer to Figure 3 , the training of the NLP model based on the recognition rules in the present application specifically includes the following sub-steps:
[0066] S31. Convert all the field names of the preprocessed data into lowercase letters, and remove special symbols and numbers;
[0067] S32. Split according to spaces and punctuation marks, and split the data preprocessed in step S31.
[0068] S33. Determine whether the data split in step S32 is pinyin or English through the language model and the norms of initials and finals of Chinese pinyin.
[0069] S34. Segment the data processed in step S32 to obtain multiple combination results.
[0070] S35. Infer and complete the combination results obtained in step S34 according to the content of the English word library or the writing norms of Chinese pinyin, and finally obtain the matching results.
[0071] S4. Find the standard element corresponding to the result with the highest matching degree from multiple matching results, and mark the classification label for the standard element corresponding to the result with the highest matching degree.
[0072] Taking the matching of "identity card" as an example, the solution of this application will be elaborated in detail below.
[0073] (1) Training of the NLP model for recognition rules
[0074] First step, first convert all the uppercase and lowercase letters of English or Chinese pinyin in "identity_cardNo" into lowercase letters, and the resulting result is "identity_cardno".
[0075] Second step, split "identity_cardno". It can be split into multiple component strings through the underscore, and then output the multiple component strings to the next step, that is, split into identity and cardno.
[0076] Third step, perform language recognition on identity and cardno according to the language model and the norms of initials and finals of Chinese pinyin. Through detection and judgment, it can be known that identity and cardno do not conform to the writing norms of Chinese pinyin, so it can be judged that these words are English words or English abbreviations, and proceed to the next step.
[0077] Fourth step, segment the longer English words, and then arrange and combine the multiple individual strings to form multiple combination results. Here, cardno is segmented into two words, card and no, so after segmentation, it is identity, card, no.
[0078] Fifth step, based on the third step, it has been inferred that the word type is English. According to the content of the English word library, identity, card, and no are English words, representing identity, certificate, and number respectively.
[0079] Step 6: Obtain the result. Based on the result inferred and completed in the previous step, the field name is obtained as ID number.
[0080] (2) Create a new standard element STD01 with the element code "identity_card" and the element name "ID card". Guided by the "Practice Guide for Network Security Standards - Guidelines for Classifying and Grading Network Data", the classification directory to which this standard element belongs is "Personal Identity Information / Personal Information".
[0081] (3) Add recognition rules to this standard element. First, add a traditional algorithm rule R1. Set the matching type of R1 to "field name matching", set the matching method of R1 to "regular expression matching", set the content of the regular expression matching to "\s*?_card", set the credibility of R1 to "80", and the priority to "1". Second, add an NLP algorithm rule R2. Set the matching type of R2 to "field comment matching", set the matching method of R2 to "NLP matching", set the matching content of R2 to "ID card", set the credibility of R2 to "90", the priority to "2", and the matching degree threshold to "0.8".
[0082] (4) Assume the asset A01 to be matched, with the field name "id_card" and the field comment "ID number". As Figure 2 shown, the steps of the matching process of the program are as follows: First, sort the priorities of the recognition rules of each standard element. In this implementation example, the priority of R2 is higher than that of R1. Second, the program will execute the matching logic in order from high to low according to the priority of the recognition rules. First, match the R2 rule and determine whether the recognition method is "NLP matching". If it is "NLP matching", it will parse the data to be matched, and then perform matching through the extraction pattern obtained by training, and finally obtain a matching result. In this implementation example, NLP will match the similarity between "ID number" and "ID card", and the reason for the high similarity depends on the accuracy of the training set. If the similarity in the result of the NLP match is greater than the matching threshold of 0.8 (the highest similarity is 1), then the binding relationship between this asset and the standard element will be added. If it is not "NLP matching", then the traditional algorithm matching logic will be executed. When performing traditional algorithm matching, different matching logics will be executed according to different matching methods. In this implementation example, it will be determined whether "id_card" conforms to the regular rule "\s*?_card" of R1. If it conforms, then the binding relationship between this asset and the standard element will be added.
[0083] (5) The asset A01 to be matched matches the standard element STD01 according to step (4), and the program will automatically assign the classification directory configured by the standard element STD01 to the asset A01 ("Personal Identity Information / Personal Information"), that is, the classification directory to which the asset A01 belongs is "Personal Identity Information / Personal Information".
[0084] The NLP model training method provided by the present invention enables the government or enterprises to improve the matching degree of the program by training their own specific data sets, thereby improving the accuracy of classification and grading; through the NLP algorithm and traditional algorithms, automatic classification and grading can be achieved, replacing manual classification, and realizing efficient, flexible, and intelligent classification and grading; for the classification and grading results, an NLP learning model training task is created, and a knowledge algorithm is established, which can realize continuous and rapid data classification and grading capabilities.
[0085] Further reference Figure 4 , as an implementation of the above method, the present application provides an embodiment of an automatic data classification and grading system based on the NLP algorithm model. This system embodiment corresponds to Figure 1 the method embodiment shown, and this system can be specifically applied to various electronic devices. The system 400 includes the following modules:
[0086] The standard element determination module 410 is used to determine the standard element and configure the classification directory to which the standard element belongs according to the classification and grading standard;
[0087] The recognition rule addition module 420 is used to add recognition rules to the standard element, set the credibility and priority of each recognition rule, and the recognition rules include the recognition rules of traditional algorithms and the recognition rules of NLP algorithms;
[0088] The matching module 430 trains the NLP model based on the recognition rules, executes the matching logic in order from high to low according to the priority of the recognition rules, and performs NLP algorithm matching or traditional algorithm matching on the preprocessed data to obtain multiple matching results; and
[0089] The marking module 440 finds the standard element corresponding to the result with the highest matching degree from multiple matching results, and marks the classification label for the standard element corresponding to the result with the highest matching degree.
[0090] In the matching module 430, training the NLP model based on the recognition rules specifically includes the following sub-steps:
[0091] S31. Convert all the field names of the preprocessed data to lowercase letters, and remove special symbols and numbers;
[0092] S32. Split the data preprocessed in step S31 according to spaces and punctuation marks;
[0093] S33. Determine whether the data split in step S32 is a pinyin or an English word through the language model and the norms of initials and finals of Chinese pinyin;
[0094] S34. Segment the data processed in step S32 to obtain multiple combination results; S35. Infer and complete the combination results obtained in step S34 according to the content of the English word dictionary or the writing norms of Chinese pinyin to finally obtain the matching results.
[0095] This application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements any one of the above methods.
[0096] Next, refer to Figure 5 which shows a schematic structural diagram of a computer system 500 of an electronic device suitable for implementing the embodiments of this application. Figure 5 The electronic device shown is only an example and should not impose any limitations on the functions and usage scope of the embodiments of this application.
[0097] As Figure 5 shown, the computer system 500 includes a central processing unit (CPU) 501, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 502 or the program loaded from the storage section 508 into the random access memory (RAM) 503. In the RAM 503, various programs and data required for the operation of the system 500 are also stored. The CPU 501, ROM 502, and RAM 503 are connected to each other through a bus 504. The input / output (I / O) interface 505 is also connected to the bus 504.
[0098] The following components are connected to the I / O interface 505: an input section 506 including a keyboard, a mouse, etc.; an output section 507 including a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 508 including a hard disk, etc.; and a communication section 509 including a network interface card such as a LAN card, a modem, etc. The communication section 509 performs communication processing via a network such as the Internet. The drive 510 is also connected to the I / O interface 505 as needed. A removable medium 511, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 510 as needed so that the computer program read from it can be installed into the storage section 508 as needed.
[0099] In particular, according to an embodiment of the present disclosure, the processes described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product that includes a computer program carried on a computer-readable medium, and the computer program includes program code for performing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication section 509, and / or installed from the removable medium 511. When the computer program is executed by the central processing unit (CPU) 501, the above functions defined in the method of the present application are executed.
[0100] It should be noted that the computer-readable storage medium described in the present application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, the computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present application, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries the computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable storage medium other than the computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable storage medium can be transmitted by any suitable medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.
[0101] Computer program code for performing the operations of this application can be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0102] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a portion of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks can occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0103] The units involved in the embodiments of the present application can be implemented in software or in hardware. The described units can also be provided in a processor. For example, it can be described as: a processor includes a first determination unit, a second determination unit, a generation unit, a first extraction unit, and a first storage unit. Among them, the names of these units do not constitute a limitation to the unit itself in some cases. For example, the first determination unit can also be described as "a unit for determining whether there is new event information in a preset event information list". On the other hand, the present application also provides a computer-readable storage medium, which can be included in the electronic device described in the above embodiments; or it can exist alone without being assembled into the electronic device. The above computer-readable storage medium carries one or more programs. When the above one or more programs are executed by the electronic device, the electronic device is caused to: determine whether there is new event information in a preset event information list, where each event information in the event information list includes event description information; in response to determining the existence, determine the new event information as target event information; identify the event description information of the target event information and generate a label for the target event information; extract a set of element information from the target event information; and store the target event information, the set of element information, and the label in association with each other in a preset event information library.
[0104] The above description is only a preferred embodiment of the present application and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of the invention involved in the present application is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above inventive concept. For example, the technical solutions formed by mutually replacing the above features with the (but not limited to) technical features with similar functions disclosed in the present application.
Claims
1. An automatic data classification and grading method based on an NLP algorithm model, characterized in that, It includes the following steps: S1. Determine the standard elements and configure the classification catalogs to which the standard elements belong according to the classification and grading criteria; S2. Add recognition rules to the standard elements, set the credibility and priority of each recognition rule. The recognition rules include the recognition rules of traditional algorithms and the recognition rules of NLP algorithms, specifically including: S211. Add the recognition rule R1 of the traditional algorithm and the recognition rule R2 of the NLP algorithm; S212. Set the matching type of the recognition rule R1 of the traditional algorithm as "field name matching", and set the matching method of the recognition rule R1 of the traditional algorithm as "regular matching"; set the matching method of the recognition rule R2 of the NLP algorithm as "NLP matching"; and S213. Set the credibility and priority of the recognition rule R1 of the traditional algorithm and the recognition rule R2 of the NLP algorithm respectively; S3. Train the NLP model based on the recognition rules, execute the matching logic in descending order of the priority of the recognition rules, perform NLP algorithm matching or traditional algorithm matching on the preprocessed data, and obtain multiple matching results. Training the NLP model based on the recognition rules specifically includes the following sub-steps: S31. Convert all field names of the preprocessed data into lowercase letters, and remove special symbols and numbers; S32. Split the data preprocessed in step S31 according to spaces and punctuation marks; S33. Determine whether the data split in step S32 is pinyin or English through the language model and the norms of the initials and finals of Chinese pinyin; S34. Perform word segmentation on the data processed in step S32 to obtain multiple combination results; S35. Infer and complete the combination results obtained in step S34 according to the content of the English word library or the writing norms of Chinese pinyin, and finally obtain the matching results; and S4. Find the standard element corresponding to the result with the highest matching degree from multiple matching results, and mark the classification label for the standard element corresponding to the result with the highest matching degree.
2. The automatic data classification and grading method based on an NLP algorithm model according to claim 1, characterized in that, In step S2, the recognition rules of the traditional algorithm include: content-based matching, field annotation matching, exact matching of field names, fuzzy matching of field names, prefix matching, suffix matching, regular matching.
3. The automatic data classification and grading method based on an NLP algorithm model according to claim 1, characterized in that, In step S2, set the credibility of each recognition rule. The credibility is set as a numerical value, and the credibility represents the degree of dependence of the recognition rule, which affects the matching degree of the recognition rule.
4. The automatic data classification and grading method based on an NLP algorithm model according to claim 1, characterized in that, In step S2, set the priority of each recognition rule. The priority is set as a numerical value, and the larger the numerical value, the higher the priority. The matching is executed in descending order of priority.
5. The automatic data classification and grading method based on an NLP algorithm model according to claim 1, characterized in that, In step S2, add recognition rules to the standard elements, set the credibility and priority of each recognition rule. The recognition rules include the recognition rules of traditional algorithms and the recognition rules of NLP algorithms, including: The priority of the recognition rule R1 of the traditional algorithm is greater than the priority of the recognition rule R2 of the NLP algorithm.
6. The automatic data classification and grading method based on an NLP algorithm model according to claim 1, characterized in that, In step S2, recognition rules are added to the standard elements, and the credibility and priority of each recognition rule are set. The recognition rules include the recognition rules of traditional algorithms and the recognition rules of NLP algorithms, including: The priority of the recognition rule R1 of the traditional algorithm is lower than the priority of the recognition rule R2 of the NLP algorithm.
7. An automatic data classification and grading system based on an NLP algorithm model, characterized in that, The system includes: A standard element determination module for determining standard elements and configuring the classification catalog to which the standard elements belong according to the classification and grading criteria; A recognition rule addition module for adding recognition rules to the standard elements, setting the credibility and priority of each recognition rule. The recognition rules include the recognition rules of traditional algorithms and the recognition rules of NLP algorithms, specifically including: S211. Add the recognition rule R1 of the traditional algorithm and the recognition rule R2 of the NLP algorithm; S212. Set the matching type of the recognition rule R1 of the traditional algorithm as "field name matching", and set the matching method of the recognition rule R1 of the traditional algorithm as "regular matching"; set the matching method of the recognition rule R2 of the NLP algorithm as "NLP matching"; and S213. Respectively set the credibility and priority of the recognition rule R1 of the traditional algorithm and the recognition rule R2 of the NLP algorithm; A matching module that trains an NLP model based on the recognition rules, executes the matching logic in order from high to low according to the priority of the recognition rules, performs NLP algorithm matching or traditional algorithm matching on the preprocessed data, and obtains multiple matching results. Training the NLP model based on the recognition rules specifically includes the following sub-steps: S31. Convert all field names of the preprocessed data to lowercase letters, and remove special symbols and numbers; S32. Split the data preprocessed in step S31 according to spaces and punctuation marks; S33. Determine whether the data split in step S32 is pinyin or English through a language model and the norms of the initials and finals of Chinese pinyin; S34. Perform word segmentation on the data processed in step S32 to obtain multiple combination results; S35. Infer and complete the combination results obtained in step S34 according to the content of the English word library or the writing norms of Chinese pinyin to finally obtain matching results; and A marking module that finds the standard element corresponding to the result with the highest matching degree from multiple matching results, and marks the classification label for the standard element corresponding to the result with the highest matching degree.
8. A computer-readable storage medium storing a computer program which, when executed by a processor, implements the method according to any one of claims 1-6.
Citation Information
Patent Citations
Application of NLP technology in data analysis
CN112836070A
Methods, apparatus and systems for annotation of text documents
US20200293712A1