A text training enhancement method and system based on deep learning

By proposing a text training enhancement method based on deep learning, a text training enhancement method based on deep learning is proposed. By combining sentence-sentence conversion and keyword-combining new data, the problem of sparse data is solved, and the fault tolerance of the model and the accuracy of information screening is improved.

CN113887724BActive Publication Date: 2025-06-17XIAMEN ANSCEN NETWORK TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111233752.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-22
Publication Date
2025-06-17
Estimated Expiration
2041-10-22

AI Technical Summary

Technical Problem

The prior art faces the problem of single data samples and small number in deep learning text training, resulting in low efficiency and poor effectiveness of model training.

Method used

A text training enhancement method based on deep learning is proposed. By conducting preliminary analysis and labeling of training text, it is divided into training set, verification set and test set, and deep learning training is carried out. If the test effect does not meet the requirements, the sentence-sentence conversion and keyword combination are carried out for the error data, new data are generated for strengthening, and re-training is carried out.

Benefits of technology

When the data samples are single and small, the data augmentation method improves the fault tolerance and overall performance of text training, and improves the accuracy of information screening.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113887724B_ABST
    Figure CN113887724B_ABST
Patent Text Reader

Abstract

The present invention provides a method and system for enhancing text training based on deep learning, including first obtaining corresponding text samples for specific requirements and performing preliminary processing on the text samples; then dividing the preprocessed data into a training set, a validation set, and a test set; finally converting it into machine language in a specific format for deep learning training to obtain a deep learning model, using the test set to verify the model results, strengthening the data with problems found after verification and adding it to the original text samples, and retraining to obtain a new model. Strengthening the data includes sentence pattern conversion, combination between different words, and creating sentence patterns applicable in different contexts for words with a high word frequency, ultimately strengthening the original text samples. This enables text training to still be carried out in the case of single data samples and a small number of data samples, thereby improving the accuracy of information discrimination.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of natural language processing, and in particular to a text training enhancement method and system based on deep learning. Background Art

[0002] In recent years, artificial intelligence centered on deep learning technology has received extensive attention. Both the academic community and the industrial community have regarded deep learning as the focus of research and application. The rapid development of deep learning technology is inseparable from the accumulation of massive data. There are endless applications and achievements of deep learning technology in the text field, each with its own advantages.

[0003] Natural Language Processing (NLP) is a kind of artificial intelligence that specializes in analyzing human language. The working principle of the whole process is as follows: First, collect a large amount of natural language and perform preprocessing on the natural language; Second, translate the natural language through deep learning algorithms; Finally, analyze the natural language and output the results. Then, in the whole process, collecting a large amount of natural language and preprocessing the language require a lot of manpower and material resources. And the result of natural language preprocessing also directly determines the effect of the final model.

[0004] With the development of artificial intelligence, the trend of deep learning is intensifying, and data is essential for deep learning research. At present, the data is sparse and extensive, resulting in the required data being too single, and the data volume cannot meet the requirements of deep learning training; at the same time, the process of sorting out data is time-consuming and laborious. In view of the situation of single data, low efficiency and poor results, this paper proposes a method for enhancing text training data based on deep learning, which can not only enrich the data quantity, but also expand the use of word contexts, improve the fault tolerance of data model training, and improve the overall performance of text data training.

[0005] This paper presents a text enhancement method based on deep learning training. First, corresponding text samples are obtained according to specific requirements, and the data is preliminarily processed based on manual retrieval. Secondly, the preprocessed data is divided into a training set, a validation set, and a test set. It is converted into machine language through a specific format and undergoes deep learning training to obtain a deep learning model. The test set is used to verify the model results and check its accuracy. Normally, the less work is done on the preliminary data, the greater the deviation of the model results. To make the model more powerful, the data needs to be strengthened. After obtaining the model results, the result data is sorted out, the correct data and problem data are checked, and the data is strengthened according to the problem data. The data strengthening methods include sentence pattern conversion and combination between different words. In addition, non-error data is also strengthened uniformly. High-frequency words (keywords) are obtained through manual retrieval, and the usage of keywords in different contexts is created to enhance the sample data and strengthen data interference. It solves the situation where text training, intelligence information acquisition, and important information discrimination can still be carried out in the case of single data samples and small numbers of data samples.

[0006] Currently, the original technologies for deep learning text training are as follows: First, the text to be trained is obtained, classified, and the central vectors of each type of text are obtained. Then, the text is further distinguished based on the central vectors. Deep convolutional neural network training is performed on the extracted text data of each category, and the data is continuously adjusted and the deviation is corrected. Finally, a model that can extract the required information entropy is obtained. The effect of this model is more based on a large amount of data. Only more representative text data and a larger amount of data can make the model more valuable. Summary of the Invention

[0007] The present invention proposes a text training enhancement method and system based on deep learning to solve the defects of the existing technologies mentioned above.

[0008] In one aspect, the present invention proposes a text training enhancement method based on deep learning, and the method includes the following steps:

[0009] S1: Conduct preliminary data analysis on the text to be trained to divide the text to be trained within a certain range, then retrieve the text to be trained within the certain range, obtain the location of the text to be trained, and then obtain the word frequency of each word therein. The words with a word frequency exceeding a certain number are used as keywords;

[0010] S2: After tagging the text to be trained, divide it into a training set, a validation set, and a test set. Preprocess the languages used in the texts in the training set, the validation set, and the test set to make them the languages that the machine can use. Then, use the preprocessed training set, validation set, and test set to perform deep learning training to obtain a trained result model, and use the test set to verify the test effect of the trained result model;

[0011] S3: If the test effect does not meet the required requirements, take out the data that goes wrong in the test effect and record it as problem data. After strengthening the problem data, jump to S2; if the test effect meets the required requirements, output the trained result model;

[0012] Strengthening the problem data includes: after converting the sentence pattern of the problem data, generating new data corresponding to the problem data and adding it to the text to be trained; after combining the keywords contained in the problem data in different degrees and orders, generating new data corresponding to the problem data and adding it to the text to be trained.

[0013] The above method first obtains the corresponding text samples for specific requirements and performs preliminary processing on the text samples; then divides the preprocessed data into a training set, a validation set, and a test set; finally, converts it into a machine language through a specific format, performs deep learning training to obtain a deep learning model, uses the test set to verify the model results, strengthens the data found to have problems after verification and adds it to the original text samples, and retrains to obtain a new model. Strengthening the data includes sentence pattern conversion, combination between different words, and creating sentence patterns applicable in different contexts for words with a high word frequency, ultimately strengthening the original text samples. This enables text training even when the data samples are single and the number of data samples is small, thereby improving the accuracy of information discrimination.

[0014] In a specific embodiment, the method further includes performing S4 after S1: setting sentence patterns that can be used in a variety of different contexts, using the sentence patterns as templates for creating samples, adding the words in the text to be trained into the sentence patterns respectively to obtain new samples, and using the new samples to strengthen the text to be trained.

[0015] In a specific embodiment, S4 specifically includes performing the following steps for each word in the text to be trained:

[0016] Set sentence patterns that can be used in a variety of different contexts. The sentence patterns contain fixed positions where the text is uncertain and any word can be filled in, and the text except for the fixed positions in the sentence patterns is certain information;

[0017] Fill each word in the text to be trained into the fixed positions respectively, and create positive samples containing each word according to the different usages of each word in different contexts; at the same time, create negative samples containing each word according to the opposite meanings of each word in different contexts;

[0018] Finally, use the positive samples to enhance the positive data of the text to be trained, and use the negative samples to enhance the interference data in the text to be trained.

[0019] In a specific embodiment, after converting the sentence pattern of the problem data, generate new data corresponding to the problem data and add it to the text to be trained, specifically including:

[0020] Express the sentences in the problem data using a variety of different methods through sentence pattern conversion, so as to generate multiple new sentences with the same meaning from one sentence, and add the new sentences to the text to be trained. In this method, the sentence pattern conversion realizes the strengthening of position information for the absolute position information used in the bert algorithm.

[0021] In a specific embodiment, after combining the keywords included in the problem data in different degrees and different orders, generate new data corresponding to the problem data and add it to the text to be trained, specifically including:

[0022] Split the keywords in the problem data to different extents, and then randomly combine the multiple words obtained after splitting to get multiple new words corresponding to the keywords. According to the multiple new words, change the sentences where the keywords are located into multiple new sentences and add them to the text to be trained.

[0023] In a specific embodiment, the method of expressing the sentences in the problem data using a variety of different methods through sentence pattern conversion specifically includes:

[0024] Conventional general method: Add preface information that does not affect the original meaning of the sentence before the sentence;

[0025] Sentence type conversion method: Change the affirmative sentence in the sentence into a double negative sentence; change the "ba" sentence in the sentence into a "bei" sentence; expand the sentence by adding several general adjectives; change the general sentence pattern in the sentence into a question sentence / exclamatory sentence.

[0026] In a specific embodiment, conduct preliminary data analysis on the text to be trained so as to divide the text to be trained within a certain range, specifically including:

[0027] Analyze the data in the text to be trained to obtain relevant information including keywords and subject matter in the data, and classify the text to be trained within a certain range according to the relevant information.

[0028] According to a second aspect of the present invention, there is provided a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a computer processor, the above method is implemented.

[0029] According to a third aspect of the present invention, there is provided a text training enhancement system based on deep learning, the system comprising:

[0030] A data analysis module: configured to perform preliminary data analysis on the text to be trained to divide the text to be trained within a certain range, then retrieve the text to be trained within the certain range, obtain the location of the text to be trained and then obtain the word frequency of each word therein, and use the words with a word frequency exceeding a certain number as keywords;

[0031] A model training module: configured to label the text to be trained and then divide it into a training set, a validation set, and a test set, preprocess the language used in the text in the training set, the validation set, and the test set into the language used by the machine, and then perform deep learning training on the preprocessed training set, validation set, and test set to obtain a training result model, and then use the test set to verify the test effect of the training result model;

[0032] A problem data enhancement module: configured to, if the test effect does not meet the required requirements, take out the data that goes wrong in the test effect and record it as problem data, enhance the problem data and then jump to the model training module; if the test effect meets the required requirements, output the training result model;

[0033] Enhancing the problem data includes: after converting the sentence pattern of the problem data, generating new data corresponding to the problem data and adding it to the text to be trained; after combining the keywords included in the problem data in different degrees and different orders, generating new data corresponding to the problem data and adding it to the text to be trained.

[0034] In a specific embodiment, the system further comprises:

[0035] A full-text data enhancement module: configured to set sentence patterns that can be used in a variety of different contexts, use the sentence patterns as templates for creating samples, add the words in the text to be trained into the sentence patterns respectively to obtain new samples, and use the new samples to enhance the text to be trained.

[0036] The present invention first obtains corresponding text samples according to specific requirements and performs preliminary processing on the text samples; then divides the preprocessed data into a training set, a validation set, and a test set; finally, converts it into machine language in a specific format and performs deep learning training to obtain a deep learning model. The test set is used to verify the model results, and the data with problems found after verification is strengthened and added to the original text samples, and training is performed again to obtain a new model. Strengthening the data includes converting sentence patterns, combining different words, and creating sentence patterns applicable in different contexts for words with a high word frequency, ultimately strengthening the original text samples. This enables text training to be carried out even when the data samples are single and the number of data samples is small, thereby improving the accuracy of information discrimination. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] The accompanying drawings are included to provide a further understanding of the embodiments and are incorporated into and constitute a part of this specification. The drawings illustrate the embodiments and, together with the description, are used to explain the principles of the present invention. Other embodiments and many of the intended advantages of the embodiments will be readily apparent as they become better understood by reference to the following detailed description. By reading the detailed description of the non-limiting embodiments with reference to the accompanying drawings below, other features, objects, and advantages of the present application will become more apparent:

[0038] Figure 1 is an exemplary system architecture diagram to which the present application can be applied;

[0039] Figure 2 is a flowchart of a method for enhancing text training based on deep learning according to an embodiment of the present invention;

[0040] Figure 3 is a framework diagram of a system for enhancing text training based on deep learning according to an embodiment of the present invention;

[0041] Figure 4 is a schematic structural diagram of a computer system of an electronic device suitable for implementing the embodiments of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0042] The present application will be further described in detail below with reference to the accompanying drawings and embodiments. It can be understood that the specific embodiments described herein are only for explaining the related invention and not for limiting the invention. It should also be noted that, for the sake of description, only parts related to the relevant invention are shown in the drawings.

[0043] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other. The present application will be described in detail below with reference to the drawings and embodiments.

[0044] Figure 1 An exemplary system architecture 100 for a text training enhancement method based on deep learning to which embodiments of the present application can be applied is shown.

[0045] As Figure 1 shown, the system architecture 100 may include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used to provide a medium for communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.

[0046] Users can use the terminal devices 101, 102, 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various applications may be installed on the terminal devices 101, 102, 103, such as data processing applications, data visualization applications, web browser applications, etc.

[0047] The terminal devices 101, 102, 103 may be hardware or software. When the terminal devices 101, 102, 103 are hardware, they may be various electronic devices, including but not limited to smartphones, tablets, laptop portable computers, and desktop computers, etc. When the terminal devices 101, 102, 103 are software, they may be installed in the above-listed electronic devices. It may be implemented as multiple software or software modules (such as software or software modules for providing distributed services), or may be implemented as a single software or software module. No specific limitation is made here.

[0048] The server 105 may be a server that provides various services, such as a background information processing server that provides support for the text to be trained displayed on the terminal devices 101, 102, 103. The background information processing server may process the obtained keywords and generate a processing result (such as new data corresponding to the problem data).

[0049] It should be noted that the method provided by the embodiments of the present application may be executed by the server 105, or may be executed by the terminal devices 101, 102, 103. The corresponding device is generally set in the server 105, or may also be set in the terminal devices 101, 102, 103.

[0050] It should be noted that the server may be hardware or software. When the server is hardware, it may be implemented as a distributed server cluster composed of multiple servers, or may be implemented as a single server. When the server is software, it may be implemented as multiple software or software modules (such as software or software modules for providing distributed services), or may be implemented as a single software or software module. No specific limitation is made here.

[0051] It should be understood that Figure 1 the numbers of the terminal devices, networks, and servers in [[ ]] are merely illustrative. According to the implementation requirements, there can be any number of terminal devices, networks, and servers.

[0052] A method for enhancing text training based on deep learning according to an embodiment of the present invention Figure 2 shows a flowchart of a method for enhancing text training based on deep learning according to an embodiment of the present invention. As Figure 2 shown, the method includes the following steps:

[0053] S1: Conduct preliminary data analysis on the text to be trained to divide the text to be trained within a certain range, then retrieve the text to be trained within the certain range, obtain the location of the text to be trained, and then obtain the word frequency of each word therein. The words with a word frequency exceeding a certain number are used as keywords.

[0054] In a specific embodiment, the conducting preliminary data analysis on the text to be trained to divide the text to be trained within a certain range specifically includes:

[0055] Analyze the data in the text to be trained to obtain relevant information including keywords and topic content in the data, and classify the text to be trained within a certain range according to the relevant information.

[0056] In a specific embodiment, first conduct simple data analysis on the text to be trained to divide the text to be trained within a certain range, then conduct a manual search on the text to be trained to more accurately grasp the location of the text. Tag the text to be trained, and divide the tagged text into a training set, a validation set, and a test set.

[0057] S2: After tagging the text to be trained, divide it into a training set, a validation set, and a test set. Preprocess the language used in the text in the training set, the validation set, and the test set into the language used by the machine, then use the preprocessed training set, validation set, and test set for deep learning training to obtain a training result model, and then use the test set to verify the test effect of the training result model.

[0058] In a specific embodiment, preprocess the divided text to be trained into the language required by the machine for deep learning training, obtain the model result after training, verify the model result, use the test set to verify the effect of the model, and finally, the actual recognition effect can be viewed through real data.

[0059] S3: If the test effect does not meet the required requirements, take out the data with errors in the test effect as problem data, strengthen the problem data, and then jump to S2; if the test effect meets the required requirements, output the trained result model.

[0060] Strengthening the problem data includes: after converting the sentence pattern of the problem data, generating new data corresponding to the problem data and adding it to the text to be trained; after combining the keywords included in the problem data in different degrees and different orders, generating new data corresponding to the problem data and adding it to the text to be trained.

[0061] In a specific embodiment, after converting the sentence pattern of the problem data, generating new data corresponding to the problem data and adding it to the text to be trained specifically includes:

[0062] Express the sentences in the problem data using a variety of different methods through sentence pattern conversion, so as to generate multiple new sentences with the same meaning from one sentence, and add the new sentences to the text to be trained.

[0063] In a specific embodiment, expressing the sentences in the problem data using a variety of different methods through sentence pattern conversion specifically includes:

[0064] Conventional general method: Add preface information that does not affect the original meaning of the sentence before the sentence.

[0065] Sentence type conversion method: Change the affirmative sentence in the sentence to a double negative sentence; change the "ba" sentence in the sentence to a "bei" sentence; expand the sentence by adding several general adjectives; change the general sentence pattern in the sentence to an interrogative sentence / exclamatory sentence.

[0066] In a specific embodiment, after combining the keywords included in the problem data in different degrees and different orders, generating new data corresponding to the problem data and adding it to the text to be trained specifically includes:

[0067] Split the keywords in the problem data to different degrees, and then randomly combine the multiple words obtained after splitting to get multiple new words corresponding to the keywords. According to the multiple new words, change the sentence where the keyword is located into multiple new sentences and add them to the text to be trained.

[0068] In a specific embodiment, according to the model effect data obtained in S2, check the problem data with errors, and strengthen the data according to the content of the problem data. On the one hand, convert the sentence pattern of the problem data, and on the other hand, combine different words in the problem data in different degrees and different orders.

[0069] The following uses actual sentences to illustrate the method described in S3:

[0070] (1) First, for the conversion of sentence patterns, the positional information is strengthened for the absolute positional information used in the bert algorithm. The following is an example: Normally, the following two sentences have the same meaning:

[0071] Example sentence 1: Have a cup of milk tea.

[0072] Example sentence 2: Hello, have a cup of milk tea.

[0073] The above two sentences have the same meaning, but currently the bert algorithm embeds the use of absolute positional information. Then, these two sentences will become different. Therefore, through basic sentence pattern conversion, sentences with the same meaning are transformed into different paraphrasing methods to reduce this deficiency and also increase the number of samples.

[0074] The following are examples of sentence pattern conversion:

[0075] Conventional general method: Add "Someone said", "He said", "Do you know?" etc. in front of all problem sentences. After adding these, it will not affect the original meaning of the sentence, but for positional information, the use of its absolute position will be reduced, which is also indirectly converting absolute positional information into relative positional information. However, after such changes, these "prefixes" will have too much weight and ignore the core content of the entire sample. Therefore, it is necessary to add the same prefix to other category samples proportionally.

[0076] Sentence type conversion method: Change affirmative sentences to double negative sentences, change ba sentences to bei sentences, expand the sentences, add some common adjectives, change general sentence patterns to interrogative sentences and exclamatory sentences.

[0077] (2) Secondly, for words, different degrees and different types of combinations are carried out. This approach is to avoid the over-occurrence of words, forming a fixed mindset for the machine and giving too much weight, resulting in the immediate recognition of the word as an inherent type once it appears. Therefore, word splitting and word combination are adopted. The following are examples:

[0078] Original sample sentence: He has been a good student since primary school. (In this sample, the main central scope is primary school, so whenever "primary school" appears, it is recognized as this type)

[0079] Error sentence: Xiaoming learned martial arts in primary school. (Here, "primary school" does not mean the original primary school)

[0080] Error sentence 2: Xiaoming bullied primary school students. (Here, "primary school students" and "primary school" do not mean the same thing)

[0081] In this case, for a machine, while combining the context, it will also assign higher weights to words with higher word frequencies, which is likely to generate problematic sentences. Therefore, the problematic data is sorted out, and the words with high word frequencies in the problematic data are split or combined to varying degrees to reduce the error probability.

[0082] In a specific embodiment, the method further includes performing S4 after S1: setting sentence patterns that can be used in a variety of different contexts, using the sentence patterns as templates for creating samples, adding the words in the text to be trained into the sentence patterns respectively to obtain new samples, and using the new samples to strengthen the text to be trained.

[0083] In a specific embodiment, S4 specifically includes performing the following steps for each word in the text to be trained:

[0084] Set sentence patterns that can be used in a variety of different contexts. The sentence patterns contain fixed positions where the text is uncertain and any word can be filled in, and the text in the sentence patterns except for the fixed positions is certain information;

[0085] Fill each word in the text to be trained into the fixed positions respectively, and create positive samples containing each word according to the different usages of each word in different situations; at the same time, create negative samples containing each word according to the opposite meanings of each word in different situations;

[0086] Finally, use the positive samples to enhance the positive data of the text to be trained, and use the negative samples to enhance the interference data in the text to be trained.

[0087] The following uses actual sentences to illustrate the method described in S4:

[0088] For the text to be trained obtained in S1, word frequency data is obtained through manual retrieval and algorithm tools, the usages of the same word frequency in different situations are created to increase positive samples, and at the same time, negative samples are increased according to the same word frequency to enhance the sample data and strengthen data interference. Finally, a more robust text model is obtained.

[0089] The same word has different meanings in different contexts; for example:

[0090] Example sentence 1: Pride goes before a fall;

[0091] Example sentence 2: As a Chinese, I am extremely proud that the women's volleyball team won the Olympic championship;

[0092] Example sentence 3: Having a telephone installed at home is no longer something new;

[0093] Example sentence 4: Fresh vegetables must be soaked in cold water before consumption.

[0094] Among them, the word "proud" in Example 1 and Example 2 is the same word, but the meanings are opposite. The same is true for Example 3 and Example 4. By strengthening the words with high word frequencies, the robustness of the sample data is also enhanced, and at the same time, the sample size is increased.

[0095] Figure 3 A framework diagram of a text training enhancement system based on deep learning according to an embodiment of the present invention is shown. The system includes a data analysis module 301, a model training module 302, a problem data enhancement module 303, and a full text data enhancement module 304.

[0096] In a specific embodiment, the data analysis module 301 is configured to perform preliminary data analysis on the text to be trained, thereby dividing the text to be trained within a certain range, then retrieving the text to be trained within the certain range, obtaining the location of the text to be trained, and then obtaining the word frequency of each word therein, and taking the words with a word frequency exceeding a certain number as keywords;

[0097] The model training module 302 is configured to label the text to be trained and then divide it into a training set, a validation set, and a test set. The language used in the text in the training set, the validation set, and the test set is preprocessed into a language used by a machine, and then the preprocessed training set, validation set, and test set are used for deep learning training to obtain a training result model, and then the test set is used to verify the test effect of the training result model;

[0098] The problem data enhancement module 303 is configured to, if the test effect does not meet the required requirements, take out the data that goes wrong in the test effect and record it as problem data, strengthen the problem data and then jump to the model training module; if the test effect meets the required requirements, output the training result model; strengthening the problem data includes: after converting the sentence pattern of the problem data, generating new data corresponding to the problem data and adding it to the text to be trained; after combining the keywords included in the problem data in different degrees and different orders, generating new data corresponding to the problem data and adding it to the text to be trained.

[0099] The full text data enhancement module 304 is configured to set sentence patterns that can be used in a variety of different contexts, and use the sentence patterns as templates for creating samples to add the words in the text to be trained into the sentence patterns respectively to obtain new samples, and use the new samples to strengthen the text to be trained.

[0100] This system first obtains corresponding text samples according to specific requirements and performs preliminary processing on the text samples; then divides the preprocessed data into a training set, a validation set, and a test set; finally, converts it into machine language in a specific format for deep learning training to obtain a deep learning model, uses the test set to verify the model results, strengthens the data with problems found after verification, adds it to the original text samples, and retrains to obtain a new model. Strengthening the data includes sentence pattern conversion, combination between different words, and creating sentence patterns applicable in different contexts for words with high word frequencies, ultimately strengthening the original text samples. This enables text training to still be carried out even when the data samples are single and the number of data samples is small, thereby improving the accuracy of information discrimination.

[0101] The following refers to Figure 4 , which shows a schematic structural diagram of a computer system 400 of an electronic device suitable for implementing the embodiments of the present application. Figure 4 The shown electronic device is merely an example and should not impose any limitation on the functions and usage scope of the embodiments of the present application.

[0102] As Figure 4 shown, the computer system 400 includes a central processing unit (CPU) 401, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 402 or the program loaded from the storage section 408 into the random access memory (RAM) 403. In the RAM 403, various programs and data required for the operation of the system 400 are also stored. The CPU 401, ROM 402, and RAM 403 are connected to each other via a bus 404. The input / output (I / O) interface 405 is also connected to the bus 404.

[0103] The following components are connected to the I / O interface 405: an input section 406 including a keyboard, a mouse, etc.; an output section 407 including such as a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 408 including a hard disk, etc.; and a communication section 409 including a network interface card such as a LAN card, a modem, etc. The communication section 409 performs communication processing via a network such as the Internet. A drive 410 is also connected to the I / O interface 405 as needed. A removable medium 411, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 410 as needed so that the computer program read from it can be installed into the storage section 408 as needed.

[0104] In particular, according to an embodiment of the present disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, an embodiment of the present disclosure includes a computer program product that includes a computer program carried on a computer-readable storage medium, and the computer program includes program code for performing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication section 409 and / or installed from the removable medium 411. When the computer program is executed by the central processing unit (CPU) 401, the above-described functions defined in the method of the present application are performed. It should be noted that the computer-readable storage medium described in the present application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, the computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. And in the present application, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries the computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable storage medium other than the computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable storage medium can be transmitted by any suitable medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.

[0105] Computer program code for performing the operations of this application can be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any kind of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (for example, by connecting through the Internet using an Internet service provider).

[0106] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks can occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0107] The modules described in the embodiments of this application can be implemented in software or in hardware. The described units can also be provided in a processor, and the names of these units do not, in some cases, constitute a limitation on the unit itself.

[0108] The embodiments of the present invention also relate to a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a computer processor, the method described above is implemented. The computer program contains program code for performing the method shown in the flowchart. It should be noted that the computer-readable medium of this application can be a computer-readable signal medium or a computer-readable medium or any combination of the two.

[0109] The present invention first obtains corresponding text samples according to specific requirements and performs preliminary processing on the text samples; then divides the preprocessed data into a training set, a validation set, and a test set; finally, converts it into machine language in a specific format for deep learning training to obtain a deep learning model, uses the test set to verify the model results, strengthens the data with problems found after verification and adds it to the original text samples, and retrains to obtain a new model. Strengthening the data includes sentence pattern conversion, combination between different words, and creating sentence patterns applicable in different contexts for words with a high word frequency, ultimately strengthening the original text samples. This enables text training to still be carried out even when the data samples are single and the number of data samples is small, thereby improving the accuracy of information discrimination.

[0110] The above description is only a preferred embodiment of the present application and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of the invention involved in the present application is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above inventive concept. For example, the technical solutions formed by mutually replacing the above features with the (but not limited to) technical features with similar functions disclosed in the present application.

Claims

1. A text training enhancement method based on deep learning, characterized in that, It includes the following steps: S1: Conduct preliminary data analysis on the text to be trained, thereby dividing the text to be trained within a certain range. Then, retrieve the text to be trained within the certain range, obtain the location of the text to be trained, and then obtain the word frequency of each word therein. Words with a word frequency exceeding a certain number are used as keywords; Set sentence patterns that can be used in a variety of different contexts. Using the sentence patterns as templates for creating samples, add the words in the text to be trained into the sentence patterns respectively to obtain new samples, and use the new samples to strengthen the text to be trained. Specifically, it further includes: performing the following steps for each word in the text to be trained: Set sentence patterns that can be used in a variety of different contexts. The sentence patterns contain fixed positions where the text is uncertain and any word can be filled in, and the text except the fixed positions in the sentence patterns is certain information; Fill each word in the text to be trained into the fixed positions respectively, and create positive samples containing each word according to the different usages of each word in different contexts; at the same time, create negative samples containing each word according to the opposite meanings of each word in different contexts; Finally, use the positive samples to enhance the positive data of the text to be trained, and use the negative samples to enhance the interference data in the text to be trained; S2: After labeling the text to be trained, divide it into a training set, a validation set, and a test set. Preprocess the language used in the text in the training set, the validation set, and the test set to become the language used by the machine, and then use the preprocessed training set, validation set, and test set for deep learning training to obtain a training result model, and then use the test set to verify the test effect of the training result model; S3: If the test effect does not meet the required requirements, take out the incorrect data in the test effect and record it as problem data. After strengthening the problem data, jump to S2; if the test effect meets the required requirements, output the training result model; Strengthening the problem data includes: after converting the sentence pattern of the problem data, generate new data corresponding to the problem data and add it to the text to be trained; after combining the keywords contained in the problem data in different degrees and different orders, generate new data corresponding to the problem data and add it to the text to be trained. Specifically, it includes: splitting the keywords in the problem data to different degrees, and then randomly combining the multiple words obtained after splitting to obtain multiple new words corresponding to the keywords. According to the multiple new words, change the sentence where the keyword is located into multiple new sentences and add them to the text to be trained.

2. The method according to claim 1, characterized in that, After converting the sentence pattern of the problem data, generate new data corresponding to the problem data and add it to the text to be trained. Specifically, it includes: Express the sentences in the problem data using a variety of different methods through sentence pattern conversion, so as to generate multiple new sentences with the same meaning from one sentence, and add the new sentences to the text to be trained.

3. The method according to claim 2, characterized in that, The statements in the problem data are expressed in multiple different ways through the conversion of statement patterns, specifically including: Conventional general method: adding preface information that does not affect the original meaning of the statement before the statement; Statement type conversion method: changing the affirmative sentence in the statement to a double negative sentence; changing the "ba" sentence in the statement to a "bei" sentence; expanding the statement by adding several general adjectives; changing the general sentence pattern in the statement to an interrogative sentence / exclamatory sentence.

4. The method according to claim 1, characterized in that, The preliminary data analysis is performed on the text to be trained, thereby classifying the text to be trained within a certain range, specifically including: Analyzing the data in the text to be trained to obtain relevant information including keywords and topic content in the data, and classifying the text to be trained within a certain range according to the relevant information.

5. A computer-readable storage medium, on which a computer program is stored, characterized in that, When the computer program is executed by a computer processor, it implements the method described in any one of claims 1 to 4.

6. A text training enhancement system based on deep learning, characterized in that, Including: Data analysis module: configured to perform preliminary data analysis on the text to be trained, thereby classifying the text to be trained within a certain range, then retrieving the text to be trained within the certain range, obtaining the location of the text to be trained and then obtaining the word frequency of each word therein, and taking the words with a word frequency exceeding a certain number as keywords; Setting sentence patterns that can be used in multiple different contexts, using the sentence patterns as templates for creating samples, adding the words in the text to be trained into the sentence patterns respectively to obtain new samples, and strengthening the text to be trained by using the new samples. Specifically, it further includes: performing the following steps for each word in the text to be trained: Setting sentence patterns that can be used in multiple different contexts, where the sentence patterns contain fixed positions with uncertain text and can be filled with any word, and the text in the sentence patterns except the fixed positions is certain information; Filling each word in the text to be trained into the fixed positions respectively, and creating positive samples containing each word according to the different usages of each word in different contexts; at the same time, creating negative samples containing each word according to the opposite meanings of each word in different contexts; Finally, strengthening the positive data of the text to be trained by using the positive samples, and strengthening the interference data in the text to be trained by using the negative samples; Model training module: configured to label the text to be trained and then divide it into a training set, a validation set, and a test set, preprocess the language used in the text in the training set, the validation set, and the test set to become the language used by the machine, and then perform deep learning training on the preprocessed training set, validation set, and test set to obtain a training result model, and then use the test set to verify the test effect of the training result model; Problem data strengthening module: configured to, if the test effect does not meet the required requirements, take out the data that goes wrong in the test effect and record it as problem data, strengthen the problem data and then jump to the model training module; if the test effect meets the required requirements, output the training result model; Enhancing the problem data includes: after converting the sentence pattern of the problem data, generating new data corresponding to the problem data and adding it to the text to be trained; after combining the keywords included in the problem data with different degrees and in different orders, generating new data corresponding to the problem data and adding it to the text to be trained. Specifically, it includes: splitting the keywords in the problem data to different degrees, then randomly combining the multiple words obtained after splitting to get multiple new words corresponding to the keywords, and changing the sentence where the keyword is located into multiple new sentences according to the multiple new words and adding them to the text to be trained.

Citation Information

Patent Citations

  • Training method and system of semantic comprehension system

    CN111680129A

  • Text enhancement method, text classification method and related devices

    CN112906392A