Language storage method based on natural language

By cleaning, feature extraction and classification model training of natural language data, segmenting and judging its importance, the problem that the existing technology cannot accurately identify and store natural language privacy data separately, and safe storage of important data is achieved.

CN120030100APending Publication Date: 2025-05-23SHANDONG ENERGY GRP CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510051283.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-13
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The prior art cannot accurately identify the importance of private data in natural language and cannot store important data separately from non-important data, resulting in the inability to guarantee the security of private data.

Method used

By cleaning, extracting and training the stored natural language data, it is divided into important fragments and non-important fragments, and the storage medium and encryption method are selected based on the importance judgment.

Benefits of technology

It realizes the importance of natural language data, ensures the secure storage of important data, and improves the level of protection of privacy data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120030100A_ABST
    Figure CN120030100A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of data storage, and relates to a language storage method based on a natural language. The method specifically comprises the steps of performing data cleaning on to-be-stored natural language data, and extracting a set of all statement features in the cleaned to-be-stored natural language data; segmenting the cleaned natural language data to be stored into a plurality of segments according to the set of features of different statements, performing feature extraction on each segment, and extracting features for each segment including a primary feature and a secondary feature; dividing the importance of all fragments in the natural language data to be stored by training a classification model; according to the classification model, importance of fragments in the natural language data to be stored is judged, and corresponding storage media and encryption methods are selected to store the fragments. The storage mode is selected according to the importance judgment of the natural language data, and the security of important data is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the technical field of data storage and relates to a language storage method based on natural language. Background Art

[0002] In the process of saving natural language, some important data about privacy data will be involved, so the natural language involving important data needs to be further processed.

[0003] There is no method for judging the privacy data of natural language in the prior art. At the same time, the judgment of the privacy data of natural language is not accurate enough, and the importance of the privacy data cannot be accurately identified, and important data cannot be stored separately from non-important data. Summary of the invention

[0004] In order to solve the problems existing in the background technology, the present invention proposes a language storage method based on natural language.

[0005] In order to achieve the above object, the technical solution adopted by the present invention is as follows:

[0006] A language storage method based on natural language, characterized by comprising:

[0007] Perform data cleaning on the natural language data to be stored, extract different sentences based on the conjunctions, and extract a set of all sentence features in the cleaned natural language data to be stored;

[0008] The cleaned natural language data to be stored is divided into multiple segments according to the set of features of different sentences, and features are extracted for each segment. The features extracted for each segment include main features and secondary features;

[0009] The importance of all the segments in the natural language data to be stored is divided into important segments and unimportant segments by training a classification model;

[0010] Based on the classification model's judgment of the importance of a segment in the natural language data to be stored, a corresponding storage medium and encryption method are selected to store the segment.

[0011] Furthermore, the method for cleaning the natural language data to be stored is:

[0012] Delete special characters that cannot be compiled in the natural language data to be stored, and convert all the natural language data to be stored into a standardized format;

[0013] After formatting, the vocabulary in the natural language data to be stored is extracted, and the incomprehensible vocabulary except nouns in the natural language data to be stored is converted into standard vocabulary through context, and repeated Chinese characters and vocabulary are marked, and any original Chinese characters or vocabulary are retained, and repeated values ​​are deleted.

[0014] Furthermore, the specific method of dividing the cleaned natural language data to be stored into multiple segments according to the set of features of different sentences is:

[0015] Establish the feature value of each sentence in the cleaned natural language data to be stored by using the TF-IDF method;

[0016] Based on the different features of sentences in the natural language data to be stored, a feature value corresponding to each sentence in the natural language data to be stored and a topic corresponding to the feature value are added;

[0017] Use cosine similarity to calculate the similarity between each sentence. According to the similarity between all sentences, all similar sentences are aggregated and classified to obtain the natural language segmentation list to be stored, thus completing the segmentation operation of the natural language data to be stored.

[0018] Furthermore, the method for extracting features from each fragment is:

[0019] Segment each segment into multiple words, generate a list of natural language words to be stored, remove the connecting words in the list of natural language words to be stored, and obtain a set of natural language words to be stored;

[0020] Construct a word cloud map, treat each word in the natural language word set to be stored as a node in the word cloud map, and in the word cloud map, if any two words appear in the same sentence, establish a connection line between them, and use cosine similarity to determine the edge weight between the two nodes;

[0021] Set the initial weight of each node to 1, iterate each node, and update the weight of each node. The specific formula is:

[0022]

[0023] Where d is the damping coefficient, set its value to 0.85, M(v i ) points to v i All nodes of C(v j ) is the node v j The number of outgoing edges, PR(v i ) is the node v i The weight of each node is iterated multiple times until the weight of the node is less than the user-preset threshold;

[0024] Sort each node according to its weight, and select the first n nodes as key nodes, where n is a positive integer;

[0025] The n key nodes in each segment are aggregated, which are the keywords in the natural language data to be stored.

[0026] Furthermore, the main feature is that keywords are used as the characteristics of basic data analysis;

[0027] Secondary features are the words other than keywords that are used as features for basic data analysis.

[0028] Furthermore, the importance of all segments in the natural language data to be stored is divided by training a classification model, and the training method of the classification model is:

[0029] The user inputs any two different important fragments and submits them to the classification model. The classification model calculates the two different important data to obtain two benchmark scores. The two benchmark scores correspond to the two important data one by one.

[0030] The classification model analyzes the natural language data in the training set, obtains the calculated scores of the natural language data in the training set, compares and corrects the calculated scores with the benchmark scores, and obtains the standard scores of the important data;

[0031] The training set is divided according to the standard score, and the fragments are extracted. The classification results are judged by the mean square error formula. If the classification results do not meet expectations, the important fragments corresponding to the benchmark scores are replaced and the training set is reclassified until the classification results meet expectations.

[0032] Furthermore, the specific method of selecting the corresponding storage medium and encryption method to store the natural language is:

[0033] Asymmetrically encrypt important fragments and choose mechanical hard disk for storage;

[0034] For non-important fragments, no encryption is performed and any storage medium is selected for storage.

[0035] Compared with the prior art, the present invention has the following beneficial effects:

[0036] It is possible to judge important data and unimportant data of natural language data, and can further store the natural language data based on the judgment results, thereby ensuring the security of important natural language data. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 It is a flow chart of the method of the present invention. DETAILED DESCRIPTION

[0038] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0039] like Figure 1 As shown, the technical solution adopted by the present invention is as follows: a language storage method based on natural language, comprising the following steps:

[0040] The natural language data to be stored is cleaned, different sentences are formed according to the conjunctions, and a set of all sentence features of the cleaned natural language data to be stored is extracted.

[0041] The cleaned natural language data to be stored is divided into multiple segments according to a set of features of different sentences, and features are extracted for each segment. The features extracted for each segment include primary features and secondary features.

[0042] The importance of all segments in the natural language data to be stored is divided into important segments and unimportant segments by training a classification model.

[0043] Based on the classification model's judgment of the importance of a segment in the natural language data to be stored, a corresponding storage medium and encryption method are selected to store the segment.

[0044] When obtaining natural language data, it is inevitable that some dirty data will be difficult to analyze. Therefore, the first thing to do is to clean the stored natural language data.

[0045] Delete special characters that cannot be compiled in the natural language data to be stored, such as "¥, %, ^", etc. All natural language data to be stored need to be converted into a standardized format. Ensure that all data can be compiled and analyzed in the next step.

[0046] After formatting, the vocabulary in the natural language data to be stored is extracted, and the incomprehensible vocabulary in the natural language data to be stored, except for nouns, is converted into standard vocabulary through context. Some variant characters, non-standard vocabulary and dialects cannot be understood by computers and need to be converted into standard vocabulary. Repeated Chinese characters and vocabulary are marked, any original Chinese characters or vocabulary are retained, and duplicate values ​​are deleted. The subsequent analysis of repeated natural language data is reduced, the speed of natural language data analysis is improved, and resource waste is reduced.

[0047] After the original natural language data is cleaned, different sentences are extracted based on conjunctions. By extracting conjunctions such as "ma", "ne", "and", etc., the sentences are divided according to the conjunctions, avoiding the inability to complete the segmentation of the content after cleaning punctuation marks during data cleaning.

[0048] After different sentences are formed, a set of all sentence features in the cleaned natural language data to be stored is extracted.

[0049] The feature values of each sentence in the cleaned natural language data to be stored are established by the TF-IDF method.

[0050] Calculate the frequency of each word appearing in the natural language data. The formula is:

[0051]

[0052] where f(t, d) is the number of times the word t appears in the document d, ∑ k∈d f(k, d) is the total number of times all words appear in the document, and TF(t, d) is the frequency of the word t appearing in the document d. The higher the frequency, the more the current word can represent the characteristics of the current natural language data. On the contrary, the lower the frequency, the greater the deviation of the word from the characteristics of the current natural language data.

[0053] Count the number of times each word appears in all documents to calculate the inverse document frequency IDF. The specific formula is:

[0054]

[0055] where N is the total number of natural language data, |d ∈ D: t ∈ d| is the number of natural language data containing the word t, and the inverse document frequency IDF(t, D) is the rarity of the word t appearing in all documents.

[0056] Calculate the TF-IDF value for each word in each natural language data to form a feature vector. The specific formula is:

[0057] TF-IDF(t, d, D) = TF(t, d) × IDF(t, D)

[0058] Extract the feature vector of each word in any natural language data to obtain the feature vector of the sentence in the natural language database.

[0059] Based on the different features of the sentences in the natural language data to be stored, add the corresponding feature values and the topics corresponding to the feature values to each sentence in the natural language data to be stored. Extract a set of all sentence features in the natural language data to be stored.

[0060] After extracting the sets of features of different sentences, the cleaned natural language data to be stored is divided into multiple segments according to the sets of features of different sentences.

[0061] Use cosine similarity to calculate the similarity between each sentence. According to the similarity between all sentences, aggregate and classify all similar sentences to obtain the natural language segmentation list to be stored, and complete the segmentation operation of the natural language to be stored. Classify sentences with similar sentence features to complete the segmentation of the natural language data to be stored.

[0062] After completing the segmentation of the natural language data to be stored, it is necessary to extract features from each segment.

[0063] Each segment is divided into multiple words to generate a list of natural language words to be stored. The word segmentation can be completed through the "jieba" library in Python. The connecting words in the list of natural language words to be stored are removed to obtain a set of natural language words to be stored.

[0064] Construct a word cloud graph, and regard each word in the natural language vocabulary set to be stored as a node in the word cloud graph. In the word cloud graph, if any two words appear in the same sentence, a connecting line is established between them, and the cosine similarity is used to determine the edge weight between the two nodes.

[0065] Set the initial weight of each node to 1, iterate each node, and update the weight of each node. The specific formula is:

[0066]

[0067] Where d is the damping coefficient, set its value to 0.85, M(v i ) points to v i All nodes of C(v j ) is the node v j The number of outgoing edges, PR(v i ) is the node v i The weight of each node is calculated and multiple iterations are performed on each node until the weight of the node is less than the user-preset threshold.

[0068] Sort each node according to its weight, and select the first n nodes as key nodes, where n is a positive integer.

[0069] The n key nodes in each segment are aggregated, which are the keywords in the natural language data to be stored. The features extracted for each segment include main features and secondary features, and the main feature content is extracted here.

[0070] The main feature is the keyword as the feature of basic data analysis.

[0071] Secondary features are the words other than keywords that are used as features for basic data analysis.

[0072] After extracting features from the segments, a classification model is trained to classify the importance of all segments in the natural language data to be stored into important segments and unimportant segments.

[0073] The user inputs any two different important fragments and submits them to the classification model. The classification model calculates the two different important data and obtains two benchmark scores. The two benchmark scores correspond to the two important data one by one. The two parameters are used to complete the selection of the classification model benchmark score and initially build the classification model. The classification model can use a deep learning model.

[0074] The classification model analyzes the natural language data in the training set, obtains the calculated score of the natural language data in the training set, compares and corrects the calculated score with the benchmark score, and obtains the standard score of the important data. The training set data is used to analyze the preliminary classification model.

[0075] At this time, the classification model will generate a standard score based on the classification results. The standard score is the criterion for judging whether the natural language data is important data or unimportant data.

[0076] The training set is divided according to the standard score, and the segments are extracted. The classification results are judged by the mean square error formula, and the mean square error is used to design the loss function for the classification model. If the classification result does not meet expectations, the parameters of the classification model are updated, the important segments corresponding to the benchmark score are replaced, and the training set is reclassified until the classification result meets expectations.

[0077] Based on the classification model's judgment of the importance of a segment in the natural language data to be stored, a corresponding storage medium and encryption method are selected to store the segment.

[0078] The important fragments are processed by asymmetric encryption and mechanical hard disk is selected for storage. The data in the mechanical hard disk can also be restored when the hardware is damaged. Using mechanical hard disk for storage can greatly improve the security of data.

[0079] For non-important fragments, no encryption is performed and any storage medium is selected for storage.

[0080] Although the present invention has been described in detail with reference to the aforementioned embodiments, it is still possible for those skilled in the art to modify the technical solutions described in the aforementioned embodiments, or to make equivalent substitutions for some of the technical features therein. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the protection scope of the present invention.

Claims

1. A language storage method based on natural language, characterized in that: Included are: Perform data cleaning on the natural language data to be stored, extract different sentences based on the conjunctions, and extract a set of all sentence features in the cleaned natural language data to be stored; The cleaned natural language data to be stored is divided into multiple segments according to the set of features of different sentences, and features are extracted for each segment. The features extracted for each segment include main features and secondary features; The importance of all the segments in the natural language data to be stored is divided into important segments and unimportant segments by training a classification model; Based on the classification model's judgment of the importance of a segment in the natural language data to be stored, a corresponding storage medium and encryption method are selected to store the segment.

2. The method for storing language based on natural language according to claim 1, characterized in that: The method for data cleaning of the natural language data to be stored is: Delete special characters that cannot be compiled in the natural language data to be stored, and convert all the natural language data to be stored into a standardized format; After formatting, the vocabulary in the natural language data to be stored is extracted, and the incomprehensible vocabulary except nouns in the natural language data to be stored is converted into standard vocabulary through context, and repeated Chinese characters and vocabulary are marked, and any original Chinese characters or vocabulary are retained, and repeated values ​​are deleted.

3. The method for storing language based on natural language according to claim 2, characterized in that: The specific method of dividing the cleaned natural language data to be stored into multiple segments according to the set of features of different sentences is: Establish the feature value of each sentence in the cleaned natural language data to be stored by using the TF-IDF method; Based on the different features of sentences in the natural language data to be stored, a feature value corresponding to each sentence in the natural language data to be stored and a topic corresponding to the feature value are added; Use cosine similarity to calculate the similarity between each sentence. According to the similarity between all sentences, all similar sentences are aggregated and classified to obtain the natural language segmentation list to be stored, thus completing the segmentation operation of the natural language data to be stored.

4. The method for storing language based on natural language according to claim 3, characterized in that: The method for extracting features from each fragment is as follows: Segment each segment into multiple words, generate a list of natural language words to be stored, remove the connecting words in the list of natural language words to be stored, and obtain a set of natural language words to be stored; Construct a word cloud graph, treat each word in the natural language word set to be stored as a node in the word cloud graph, and in the word cloud graph, if any two words appear in the same sentence, establish a connection line between them, and use cosine similarity to determine the edge weight between the two nodes; Set the initial weight of each node to 1, iterate each node, and update the weight of each node. The specific formula is: Where d is the damping coefficient, set its value to 0.85, M(v i ) points to v i All nodes of C(v j ) is the node v j The number of outgoing edges, PR(v i ) is the node v i The weight of each node is iterated multiple times until the weight of the node is less than the user-preset threshold; Sort each node according to its weight, and select the first n nodes as key nodes, where n is a positive integer; The n key nodes in each segment are aggregated, which are the keywords in the natural language data to be stored.

5. The method for storing language based on natural language according to claim 4, characterized in that: The main feature is that keywords are used as the characteristics of basic data analysis; Secondary features are the words other than keywords that are used as features for basic data analysis.

6. The method for storing language based on natural language according to claim 1, characterized in that: The importance of all segments in the natural language data to be stored is divided by training a classification model, and the training method of the classification model is: The user inputs any two different important fragments and submits them to the classification model. The classification model calculates the two different important data to obtain two benchmark scores. The two benchmark scores correspond to the two important data one by one. The classification model analyzes the natural language data in the training set, obtains the calculated scores of the natural language data in the training set, compares and corrects the calculated scores with the benchmark scores, and obtains the standard scores of the important data; The training set is divided according to the standard score, and the fragments are extracted. The classification results are judged by the mean square error formula. If the classification results do not meet expectations, the important fragments corresponding to the benchmark scores are replaced and the training set is reclassified until the classification results meet expectations.

7. The method for storing language based on natural language according to claim 1, characterized in that: The specific method of selecting the corresponding storage medium and encryption method to store the natural language is: Asymmetrically encrypt important fragments and choose mechanical hard disk for storage; For non-important fragments, no encryption is performed and any storage medium is selected for storage.