Corpus Expansion Method, Device, and Storage Medium
Through word segmentation processing and candidate replacement vocabulary scoring mechanisms, automatic generation and extended corpus solve the problem of time-consuming and labor-consuming manual entry, improve the quality and quantity of corpus, and is suitable for training of natural language processing models.
Patent Information
- Application Number
- CN202510282003.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-03-11
AI Technical Summary
The establishment of existing corpus requires manual entry, which consumes a lot of manpower and time, resulting in low corpus quality, especially in specific fields, which is difficult to obtain sufficient quantity and quality corpus data.
By performing word segmentation on the original corpus, candidate replacement vocabulary is obtained and replacement recommendation scores are calculated based on the number of word meanings and semantic similarity, candidate replacement vocabulary that meets the conditions is selected for replacement, and an extended corpus is generated.
It improves the accuracy and efficiency of corpus expansion, increases the number of corpus while maintaining semantic quality, reduces semantic changes, and is suitable for training of natural language processing models.
Smart Images

Figure CN119783672B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data processing, and particularly to a corpus expansion method, device, and storage medium. Background Art
[0002] With the continuous development of artificial intelligence and the increasing amount of data, natural language processing technology has been widely applied in various industries. The training effect of natural language processing models depends on corpus data, especially high-quality corpus data.
[0003] However, when establishing the existing corpus, it is necessary to manually input each piece of corpus, which not only wastes a lot of manpower, but also takes a long time, and it is easy to miss corpus or the corpus is not rich enough, resulting in low corpus quality. Especially for some characteristic fields (such as the medical field, the financial field, etc.), it is difficult to obtain sufficient quantity and quality of corpus data. Summary of the Invention
[0004] To solve the above technical problems, this application provides at least one corpus expansion method, device, and storage medium.
[0005] In the first aspect of this application, a corpus expansion method is provided. The method includes: performing word segmentation processing on the original corpus to obtain the original vocabulary corresponding to the original corpus; and obtaining candidate replacement vocabulary corresponding to the original vocabulary; where the original vocabulary and the candidate replacement vocabulary are marked with the number of semantic meanings, and the number of semantic meanings refers to the number of meanings contained in the vocabulary; determining a scoring influencing factor that matches the candidate replacement vocabulary based on the number of semantic meanings corresponding to the candidate replacement vocabulary and / or the original vocabulary; where the scoring influencing factor refers to a factor that affects the actual semantics of the corpus when replacing the candidate replacement vocabulary with the original vocabulary in the corpus; calculating a replacement recommendation score corresponding to the candidate replacement vocabulary using the scoring influencing factor, and selecting the candidate replacement vocabulary whose replacement recommendation score meets the preset conditions as the target replacement vocabulary corresponding to the original vocabulary; replacing the original vocabulary in the original corpus with the target replacement vocabulary corresponding to the original vocabulary to obtain an expanded corpus.
[0006] In an embodiment, the scoring influencing factor includes the lexical semantic similarity between the original vocabulary and the candidate replacement vocabulary and / or the sentence semantic similarity between the original corpus and the corpus after replacing the original vocabulary in the original corpus with the candidate replacement vocabulary; determining a scoring influencing factor that matches the candidate replacement vocabulary based on the number of semantic meanings corresponding to the candidate replacement vocabulary and / or the original vocabulary, including: if the number of semantic meanings corresponding to the candidate replacement vocabulary and / or the original vocabulary is less than or equal to the preset number threshold, then taking the lexical semantic similarity as the scoring influencing factor that matches the candidate replacement vocabulary; if the number of semantic meanings corresponding to the candidate replacement vocabulary and / or the original vocabulary is greater than the preset number threshold, then taking the lexical semantic similarity and the sentence semantic similarity as the scoring influencing factor that matches the candidate replacement vocabulary.
[0007] In one embodiment, the scoring influencing factors for candidate replacement word matching include lexical semantic similarity and sentence semantic similarity; calculating the replacement recommendation score corresponding to the candidate replacement word using the scoring influencing factors includes: mapping the lexical semantic similarity to obtain a first score, and mapping the sentence semantic similarity to obtain a second score; performing weighted summation on the first score and the second score, and using the weighted summation result as the replacement recommendation score corresponding to the candidate replacement word.
[0008] In one embodiment, selecting candidate replacement words whose replacement recommendation scores meet the preset conditions as the target replacement words corresponding to the original words includes: calculating the importance levels corresponding to each original word in the original corpus; setting a scoring threshold for each original word based on the importance level corresponding to each original word; where the importance level is positively correlated with the scoring threshold; for each original word, respectively selecting candidate replacement words whose replacement recommendation scores are greater than the scoring threshold corresponding to each original word, to obtain the target replacement words corresponding to each original word.
[0009] In one embodiment, the method further includes: extracting the semantics corresponding to the original corpus and / or the extended corpus to obtain standard semantic features; generating a corpus based on the standard semantic features to obtain a sentence conversion corpus different from the original corpus and / or the extended corpus; using the sentence conversion corpus as the new original corpus and / or the new extended corpus.
[0010] In one embodiment, before performing word segmentation on the original corpus to obtain the original words corresponding to the original corpus, the method further includes: obtaining the original structured data; extracting entities and / or attributes from the original structured data to obtain a full-element set; generating the original corpus based on the full-element set.
[0011] In one embodiment, generating the original corpus based on the full-element set includes: adding the entities or attributes in the full-element set to a preset corpus template to obtain an initial corpus; inputting the initial corpus into a pre-trained natural language model for natural language conversion to obtain the original corpus output by the natural language model.
[0012] In one embodiment, the method further includes: using the original corpus and the extended corpus as model training corpora; performing named entity recognition on the model training corpora to obtain key entities; and performing dependency syntactic analysis on the model training corpora to obtain the syntactic dependency relationships between each word in the model training corpora; identifying the entity association relationships between each key entity based on the syntactic dependency relationships between each word in the model training corpora; generating triples based on each key entity and the entity association relationships between each key entity to obtain a triple set; using the triple set to train a pre-trained language model to obtain a trained language model.
[0013] In the second aspect of the present application, a corpus expansion device is provided. The device includes: a word segmentation module for performing word segmentation on the original corpus to obtain the original vocabulary corresponding to the original corpus; and obtaining candidate replacement words corresponding to the original vocabulary; wherein, the original vocabulary and the candidate replacement words are correspondingly marked with the number of semantic meanings, and the number of semantic meanings refers to the number of meanings contained in the vocabulary; a factor matching module for determining the scoring influencing factors matching the candidate replacement words based on the number of semantic meanings corresponding to the candidate replacement words and / or the original vocabulary; wherein, the scoring influencing factors refer to the factors that affect the actual semantics of the corpus when the candidate replacement words are used to replace the original vocabulary in the corpus; a vocabulary selection module for calculating the replacement recommendation score corresponding to the candidate replacement words by using the scoring influencing factors, and selecting the candidate replacement words whose replacement recommendation scores meet the preset conditions as the target replacement words corresponding to the original vocabulary; a vocabulary replacement module for replacing the original vocabulary in the original corpus with the target replacement words corresponding to the original vocabulary to obtain an expanded corpus.
[0014] In the third aspect of the present application, an electronic device is provided, which includes a memory and a processor. The processor is used to execute the program instructions stored in the memory to implement the above-mentioned corpus expansion method.
[0015] In the fourth aspect of the present application, a computer-readable storage medium is provided, on which program instructions are stored. When the program instructions are executed by a processor, the above-mentioned corpus expansion method is implemented.
[0016] In the above solution, by performing word segmentation on the original corpus to obtain the original vocabulary corresponding to the original corpus; and obtaining candidate replacement words corresponding to the original vocabulary; wherein, the original vocabulary and the candidate replacement words are correspondingly marked with the number of semantic meanings; determining the scoring influencing factors matching the candidate replacement words based on the number of semantic meanings corresponding to the candidate replacement words and / or the original vocabulary; calculating the replacement recommendation score corresponding to the candidate replacement words by using the scoring influencing factors, and selecting the candidate replacement words whose replacement recommendation scores meet the preset conditions as the target replacement words corresponding to the original vocabulary; replacing the original vocabulary in the original corpus with the target replacement words corresponding to the original vocabulary to obtain an expanded corpus, the scoring influencing factors matching different candidate replacement words can be accurately obtained according to the number of semantic meanings corresponding to the candidate replacement words and / or the original vocabulary, the calculation accuracy of the replacement recommendation score can be improved, and further the accuracy of selecting the target replacement words for each original vocabulary can be improved. On the premise of increasing the number of corpus, the situation that the expanded corpus has semantic changes relative to the original corpus can be reduced, and the quality of the expanded corpus can be guaranteed.
[0017] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and do not limit the present application. Description of the Drawings
[0018] The accompanying drawings here are incorporated into the description and form a part of this description. These drawings show embodiments consistent with the present application and, together with the description, are used to explain the technical solutions of the present application.
[0019] Figure 1 It is a schematic diagram of the solution implementation environment shown in an exemplary embodiment of the present application;
[0020] Figure 2 It is a flowchart of the corpus expansion method shown in an exemplary embodiment of the present application;
[0021] Figure 3 It is a flowchart of training a language model using the original corpus and the expanded corpus shown in an exemplary embodiment of the present application;
[0022] Figure 4 It is a block diagram of the corpus expansion device shown in an exemplary embodiment of the present application;
[0023] Figure 5 It is a schematic structural diagram of an electronic device shown in an exemplary embodiment of the present application;
[0024] Figure 6 It is a schematic structural diagram of a computer-readable storage medium shown in an exemplary embodiment of the present application. Detailed implementation manners
[0025] The following describes the solutions of the embodiments of the present application in detail with reference to the accompanying drawings of the description.
[0026] In the following description, specific details such as specific system architectures, interfaces, and technologies are presented for the purpose of illustration rather than limitation, so as to understand the present application thoroughly.
[0027] The term "and / or" in this document is merely an association information describing associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this document generally represents an "or" relationship between the associated objects before and after. In addition, "multiple" in this document means two or more than two. In addition, the term "at least one" in this document means any one of multiple or any combination of at least two of multiple. For example, including at least one of A, B, and C can represent including any one or more elements selected from the set composed of A, B, and C.
[0028] The following describes the corpus expansion method provided by the embodiments of the present application.
[0029] Please refer to Figure 1 , Figure 1It is a schematic diagram of the solution implementation environment shown in an exemplary embodiment of the present application. The solution implementation environment may include a terminal 110 and a server 120, and the terminal 110 and the server 120 are communicatively connected to each other.
[0030] The number of terminals 110 may be one or more. The terminal 110 may be a smart phone, a tablet computer, a laptop computer, a desktop computer, etc., but is not limited thereto.
[0031] The server 120 may be an independent physical server, or a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms.
[0032] In one example, the server 120 may perform an expansion process on the original corpus obtained from the terminal 110 to obtain an expanded corpus. Of course, the server 120 may store the expanded corpus locally, send it back to the terminal 110, or transmit it to other terminals.
[0033] In one example, a client of a target application is installed and run in the terminal 110. For example, the target application may be an application providing a corpus expansion function, and the original corpus is expanded by using the target application to obtain an expanded corpus. The server 120 may be the background server of the target application for providing background services for the client of the target application.
[0034] For the corpus expansion method provided in the embodiments of the present application, the execution entity of each step may be the terminal 110, such as the client of the target application installed and run in the terminal 110, or the server 120, or the terminal 110 and the server 120 cooperate with each other to execute, that is, a part of the steps of the method are executed by the terminal 110 and the other part of the steps are executed by the server 120.
[0035] Please refer to Figure 2 , Figure 2 It is a flowchart of the corpus expansion method shown in an exemplary embodiment of the present application. The corpus expansion method may be applied to the Figure 1 shown implementation environment and is specifically executed by the server in the implementation environment. It should be understood that this method may also be applicable to other exemplary implementation environments and is specifically executed by devices in other implementation environments. The embodiments of the present application do not limit the implementation environment applicable to this method.
[0036] Such as Figure 2As shown, the corpus expansion method at least includes steps S210 to S240, which are introduced in detail as follows:
[0037] Step S210: Perform word segmentation on the original corpus to obtain the original words corresponding to the original corpus; and, obtain the candidate replacement words corresponding to the original words; wherein, the original words and the candidate replacement words are correspondingly marked with the number of word meanings, and the number of word meanings refers to the number of meanings contained in the word.
[0038] The original corpus refers to the text that needs to be expanded.
[0039] Perform word segmentation on the original corpus, and use the words obtained by word segmentation as the original words.
[0040] Algorithms such as the N-Gram model, the maximum matching algorithm, the Viterbi algorithm, etc. can be used to perform word segmentation on the original corpus, and the present application does not limit the specific manner of word segmentation.
[0041] For example, if the original corpus is "Product A high-performance smartphone", the original words obtained through word segmentation include "Product A", "high-performance", and "smartphone".
[0042] In addition, obtain the candidate replacement word set corresponding to the original word, and the candidate replacement word set contains one or more candidate replacement words.
[0043] Among them, the number of word meanings refers to the number of meanings contained in the word. For example, the meanings that "ink" can contain include: Meaning 1: A liquid containing pigments or dyes, used for writing or painting; Meaning 2: A person's culture and knowledge. Then the number of word meanings corresponding to "ink" is 2.
[0044] Specifically, obtain the candidate replacement word set corresponding to the original word, and the candidate replacement word set contains one or more candidate replacement words.
[0045] It should be noted that generally, there are multiple original words corresponding to the original corpus. It can be that each original word corresponds to the same candidate replacement word set, or each original word corresponds to a different candidate replacement word set.
[0046] For example, a vocabulary library is pre-constructed, and the vocabulary library contains multiple words.
[0047] The vocabulary library can be directly used as the candidate replacement word set corresponding to each original word.
[0048] It can also be to calculate the lexical feature similarity between the current original word and each word in the vocabulary, select the words in the vocabulary whose lexical feature similarity is greater than a preset threshold, or select the top N words in the vocabulary after sorting the lexical feature similarities in descending order, and use the selected words as the candidate replacement words corresponding to the current original word, so as to obtain the set of candidate replacement words corresponding to the current original word. By selecting candidate replacement words for each original word in the above manner, the set of candidate replacement words corresponding to each original word can be obtained respectively.
[0049] Step S220: Determine the scoring influencing factors that match the candidate replacement words based on the number of semantic meanings corresponding to the candidate replacement words and / or the original words; where the scoring influencing factors refer to the factors that affect the actual semantics of the corpus when replacing the original words in the original corpus with the candidate replacement words.
[0050] For example, some scoring influencing factors are given as examples:
[0051] Factor 1: The lexical semantic similarity between the original word and the candidate replacement word;
[0052] Factor 2: The sentence semantic similarity between the original corpus and the corpus after replacing the original words in the original corpus with the candidate replacement words;
[0053] Factor 3: The difference between the degree of meaning expressed by the original word and the degree of meaning expressed by the candidate replacement word;
[0054] Factor 4: The difference between the emotional color expressed by the original word and the emotional color expressed by the candidate replacement word.
[0055] Among them, the lexical semantic similarity can be obtained by calculating the vector distance between the feature vectors corresponding to the two words; the sentence semantic similarity can be obtained by calculating the vector distance between the feature vectors corresponding to the two corpora before and after replacement; the degree of meaning refers to the degree of semantics expressed by the word, such as "look down upon" and "despise", and "despise" expresses a heavier semantics than "look down upon"; the emotional color refers to the emotional tendency expressed by the word. For example, "achievement" is a commendatory term, "result" is a neutral term, and "consequence" is a derogatory term.
[0056] Of course, in addition to the above-mentioned scoring influencing factors given as examples, other scoring influencing factors can also be selected, and this application does not limit this.
[0057] Determine the scoring influencing factors that match the candidate replacement words based on the number of semantic meanings corresponding to the candidate replacement words and / or the original words.
[0058] It can be to determine the scoring influencing factors matching the candidate replacement word according to the number of semantic meanings corresponding to the candidate replacement word; it can also be to determine the scoring influencing factors matching the candidate replacement word according to the number of semantic meanings corresponding to the original word; it can also be to determine the scoring influencing factors matching the candidate replacement word according to the number of semantic meanings corresponding to the candidate replacement word and the original word. At this time, it can be to select the maximum value between the number of semantic meanings of the candidate replacement word and the number of semantic meanings of the original word as the number of semantic meanings referred to for finally matching the scoring influencing factors.
[0059] For example, according to the number of semantic meanings corresponding to the candidate replacement word and / or the original word, determine the number of scoring influencing factors to be considered, and select the corresponding number of scoring influencing factors from the above set of scoring influencing factors to obtain the scoring influencing factors matching the candidate replacement word. Among them, if the number of semantic meanings corresponding to the candidate replacement word and / or the original word is more, it indicates that the probability of the semantic meaning of the corpus changing after the candidate replacement word replaces the original word is higher. Therefore, it can be that the more the number of semantic meanings corresponding to the candidate replacement word and / or the original word, the more the number of corresponding scoring influencing factors.
[0060] Another example is that scoring influencing factors to be considered are pre-set for different semantic meaning number intervals. According to the semantic meaning number interval where the number of semantic meanings corresponding to the candidate replacement word and / or the original word is located, the scoring influencing factors matching the candidate replacement word are obtained. For example, it is set that the scoring influencing factors corresponding to the first number interval [0, 1] include factor one; the scoring influencing factors corresponding to the second number interval (1, 3] include factor two; the scoring influencing factors corresponding to the third number interval (3, 5] include factor one, factor two, and factor three.
[0061] Step S230: Calculate the replacement recommendation score corresponding to the candidate replacement word using the scoring influencing factors, and select the candidate replacement words whose replacement recommendation scores meet the preset conditions as the target replacement words corresponding to the original word.
[0062] Among them, the replacement recommendation score represents the recommendation degree that this candidate replacement word can replace the original word.
[0063] The higher the replacement recommendation score, the lower the possibility that the corpus obtained after replacing the candidate replacement word with the original word in the original corpus has a semantic change. On the contrary, the lower the replacement recommendation score, the higher the possibility that the corpus obtained after replacing the candidate replacement word with the original word in the original corpus has a semantic change.
[0064] Calculate the influence values corresponding to the selected scoring influencing factors, and convert the influence values into replacement recommendation scores.
[0065] For example, calculate the lexical semantic similarity between the original word and the candidate replacement word, and convert the lexical semantic similarity into a replacement recommendation score. Specifically, the higher the lexical semantic similarity, the higher the replacement recommendation score; the lower the lexical semantic similarity, the lower the replacement recommendation score.
[0066] Also for example, calculate the sentence semantic similarity between the original corpus and the corpus obtained by replacing the original words in the original corpus with candidate replacement words, and convert the sentence semantic similarity into a replacement recommendation score. Specifically, the higher the sentence semantic similarity, the higher the replacement recommendation score; the lower the sentence semantic similarity, the lower the replacement recommendation score.
[0067] Still for example, calculate the difference in the degree of meaning expressed by the original word and the candidate replacement word, and convert the difference in the degree of meaning into a replacement recommendation score. Specifically, the smaller the difference in the degree of meaning, the higher the replacement recommendation score; the larger the difference in the degree of meaning, the lower the replacement recommendation score.
[0068] Once again for example, calculate the difference in the emotional color expressed by the original word and the candidate replacement word, and convert the difference in the emotional color into a replacement recommendation score. Specifically, the smaller the difference in the emotional color, the higher the replacement recommendation score; the larger the difference in the emotional color, the lower the replacement recommendation score.
[0069] For the candidate replacement words corresponding to the current original word, use the scoring influencing factors matched by each candidate replacement word to calculate the replacement recommendation score corresponding to each candidate replacement word, and select the candidate replacement words whose replacement recommendation scores meet the preset conditions as the target replacement words corresponding to the original word.
[0070] For example, select the candidate replacement words whose replacement recommendation scores are greater than the preset threshold as the target replacement words corresponding to the original word.
[0071] Also for example, select the top M candidate replacement words after sorting the replacement recommendation scores in descending order as the target replacement words corresponding to the original word.
[0072] The preset conditions for judging the replacement recommendation score can be flexibly set according to the actual situation, and the present application does not limit this.
[0073] Step S240: Replace the original words in the original corpus with the target replacement words corresponding to the original words to obtain an extended corpus.
[0074] After obtaining the target replacement words corresponding to the original words, replace the original words in the original corpus and their corresponding target replacement words to obtain an extended corpus.
[0075] It should be noted that if an original word corresponds to multiple target replacement words, the original corpus is subjected to multiple word replacements in sequence to obtain multiple extended corpora.
[0076] Through the above method, one or more original words in the original corpus are replaced to obtain an extended corpus corresponding to the original corpus. Among them, if multiple original words in the original corpus are replaced, it can be that only one original word is replaced each time, or multiple original words can be replaced simultaneously each time. This application does not limit this.
[0077] Next, some embodiments of the present application will be described in detail.
[0078] In some embodiments, the original corpus used in this application can be a manually written corpus, and / or a corpus generated by a corpus generation neural network model, and / or a corpus automatically generated based on a corpus template.
[0079] Exemplarily, the steps for generating the original corpus may include: obtaining original structured data; extracting entities and / or attributes from the original structured data to obtain a complete element set; generating the original corpus based on the complete element set.
[0080] Among them, the original structured data can be any structured text data related to the corpus to be generated.
[0081] Entities or attributes are extracted from the original structured data to obtain a complete element set. It should be noted that when extracting entities and / or attributes from the original structured data, the relationships between the entities and / or attributes are also correspondingly extracted to facilitate subsequent corpus generation.
[0082] For example, all attributes, entities, and their relationships are extracted from the original structured data Draw to form a complete element set where represents the i-th entity or attribute, and the value of i ranges from 1 to n, where n is the number of entities and / or attributes extracted.
[0083] The original corpus is generated according to the complete element set E.
[0084] For example, the entities or attributes in the complete element set are added to a preset corpus template to obtain an initial corpus; the initial corpus is input into a pre-trained natural language model for natural language conversion to obtain the original corpus output by the natural language model.
[0085] To improve the standardization and consistency of corpus generation, a set of standardized preset corpus template sets T are constructed. Any preset corpus template in the preset corpus template set T predefines a sentence structure, and combines the entities and attributes in the complete element set E to generate an initial corpus .
[0086] Pre-set corpus template can be expressed in the form of a natural language structure containing placeholders, such as , by substituting the information in the full element set E into the pre-set corpus template , for the specific process, see formula 1 below:
[0087]
[0088] The generated initial corpus is used as the input of the natural language model.
[0089] For example, for a product entity in a database, the corpus generated to describe the product includes: The product is a smartphone with high performance.
[0090] Among them, the natural language model can be implemented based on network architectures such as the Text-To-Text Transfer Transformer (T5), or the Recurrent Neural Network (RNN), or the Transformer model, etc.
[0091] The training steps of the natural language model include:
[0092] In the design of the natural language model, the input data consists of natural language descriptions generated based on the full element set E (the input during application is the initial corpus), denoted as , where each natural language description is the text description of a specific entity or attribute. The expected output data is the natural language target description corresponding to the input data X, denoted as . Ensure that the natural language target description can accurately express the relevant entities, attributes, and their relationships, and is semantically complete and syntactically correct, so that the natural language model can learn high-quality language patterns.
[0093] During the training process of the natural language model, the goal is to learn a mapping function , whose role is to map the input data X to the expected output data Y. To achieve this goal, the natural language model can use the cross-entropy loss function to measure the difference between the actual output data and the expected output data of the model. Specifically, the cross-entropy loss function is defined as formula 2 below:
[0094]
[0095] Among them represents the model parameters, is the t-th word in the i-th expected output text sequence of the expected output data, is the t-th word in the actual output text sequence predicted by the natural language model, and it depends on the previously generated words and the input , where N is the total number of expected output text sequences in the expected output data, T is the total number of words in the text sequence, and L is the calculated model training loss.
[0096] Of course, in addition to calculating the model training loss using the cross-entropy loss function, other loss functions can also be used to calculate the model training loss, which is not limited in this application.
[0097] By minimizing this model training loss, the natural language model can gradually learn how to generate natural language sequences close to the target description.
[0098] The training of the natural language model uses standard supervised learning methods and is trained through the training set (X, Y). During the training process, the model will generate sentences based on the given input and update the parameters through backpropagation , so as to gradually optimize the text quality generated by the natural language model.
[0099] Optionally, to ensure the text generation ability of the natural language model, it is not only necessary to detect the performance of the natural language model on the training set, but also necessary to use the validation set to evaluate the performance of the natural language model to detect the text generation effect of the natural language model on unprocessed text data.
[0100] Among them, the performance evaluation of the validation set is obtained by calculating the difference between the actual output text and the expected output text, and combining indicators such as the language fluency, semantic accuracy, and text consistency of the actual output text. The performance evaluation results obtained from the validation set can be used to optimize the hyperparameters (such as learning rate, batch size) of the natural language model, and can also help determine whether it is necessary to further adjust the preprocessing process of the input data, etc., to optimize the natural language model and obtain a trained natural language model.
[0101] Then, the initial corpus is input into the pre-trained natural language model for natural language conversion to obtain the original corpus output by the natural language model.
[0102] The trained natural language model can generate accurate natural language descriptions, ensuring good adaptability and usability in applications in specific vertical fields.
[0103] Of course, in addition to the original corpus generation method shown in the above embodiments, other methods can also be used to generate the original corpus. For example, entities or attributes in the full element set E can be filled into a preset corpus template, and the filling result is directly used as the original corpus. For another example, entities or attributes in the full element set E can be flexibly selected and input into a natural language generation neural network model, and the natural language generation result is used as the original corpus. The present application does not limit the specific acquisition method of the original corpus.
[0104] After obtaining the original corpus, since the original corpus lacks the diversity of expression, it is necessary to expand the corpus.
[0105] The corpus is expanded by determining the target replacement words corresponding to the original words in the original corpus and replacing the original words and the target replacement words.
[0106] Specifically, based on the number of semantic meanings corresponding to the candidate replacement words and / or the original words, the scoring influencing factors matching the candidate replacement words are determined, so as to calculate the replacement recommendation score corresponding to the candidate replacement words by using the scoring influencing factors, and the candidate replacement words whose replacement recommendation score meets the preset conditions are selected as the target replacement words corresponding to the original words.
[0107] A specific embodiment of the matching scoring influencing factors is illustrated by way of example:
[0108] The scoring influencing factors include the lexical semantic similarity between the original word and the candidate replacement word and / or the sentence semantic similarity between the original corpus and the corpus after replacing the original word in the original corpus with the candidate replacement word. Determining the scoring influencing factors matching the candidate replacement words based on the number of semantic meanings corresponding to the candidate replacement words and / or the original words in step S220 includes: if the number of semantic meanings corresponding to the candidate replacement words and / or the original words is less than or equal to a preset quantity threshold, the lexical semantic similarity is used as the scoring influencing factor for the candidate replacement word; if the number of semantic meanings corresponding to the candidate replacement words and / or the original words is greater than the preset quantity threshold, the lexical semantic similarity and the sentence semantic similarity are used as the scoring influencing factors for the candidate replacement word.
[0109] For example, assuming that the preset quantity threshold is 1, if the number of semantic meanings corresponding to the candidate replacement word or the original word is 1, the lexical semantic similarity between the candidate replacement word and the original word is calculated; if the number of semantic meanings corresponding to the candidate replacement word or the original word is greater than 1, the sentence semantic similarity between the original corpus and the corpus after replacing the original word in the original corpus with the candidate replacement word is calculated, and the lexical semantic similarity between the candidate replacement word and the original word is calculated.
[0110] Of course, the above embodiments are only examples of a specific application scenario. The selection method of scoring influencing factors can be flexibly determined according to the actual application scenario, and the present application does not limit this.
[0111] Based on the above embodiments, if the scoring influencing factors matching the candidate replacement words include lexical semantic similarity and sentence semantic similarity, then calculating the replacement recommendation score corresponding to the candidate replacement words by using the scoring influencing factors includes: mapping the lexical semantic similarity to obtain a first score, and mapping the sentence semantic similarity to obtain a second score; performing weighted summation on the first score and the second score, and using the weighted summation result as the replacement recommendation score corresponding to the candidate replacement words.
[0112] Among them, the lexical semantic similarity is positively correlated with the first score, and the sentence semantic similarity is positively correlated with the second score.
[0113] The weight parameters used for performing weighted summation on the first score and the second score can be preset according to experience, or can be flexibly calculated according to specific circumstances.
[0114] For example, obtaining the number of word senses corresponding to the candidate replacement word, or the number of word senses corresponding to the original word, or the sum of the number of word senses corresponding to the candidate replacement word and the number of word senses corresponding to the original word, to obtain the target number of word senses, and setting the weight parameters used for performing weighted summation on the first score and the second score according to the target number of word senses. For example, the larger the target number of word senses, the greater the probability that the original corpus has a semantic change after word replacement. Therefore, the weight parameter corresponding to the second score is larger to reduce the probability of actual semantic change before and after word replacement in the original corpus.
[0115] An example of a specific embodiment for selecting the target replacement word is given:
[0116] In step S230, selecting the candidate replacement words whose replacement recommendation scores meet the preset conditions as the target replacement words corresponding to the original words includes:
[0117] Step S231: Calculating the importance levels corresponding to the respective original words in the original corpus.
[0118] For example, by analyzing the roles or functions played by the respective original words in the original corpus, determining the importance levels corresponding to the respective original words. For example, classifying the original words as object, subject, and predicate, and different classification results corresponding to different importance levels.
[0119] For another example, by counting the frequencies of the respective original words in all original corpora, obtaining the importance levels corresponding to the respective original words.
[0120] For another example, the importance of each original word in the original corpus is identified by using a pre-trained neural network model for identifying the importance of words.
[0121] This application does not limit the specific calculation method of the word importance.
[0122] Step S232: Based on the importance of each original word, a scoring threshold is set for each original word; wherein the importance and the scoring threshold are positively correlated.
[0123] That is, the higher the importance of the original words, the higher the scoring threshold, and the lower the importance of the original words, the lower the scoring threshold.
[0124] Step S233: for each original word, select candidate replacement words whose replacement recommendation scores are greater than the score thresholds corresponding to the original words, and obtain target replacement words corresponding to the original words.
[0125] According to the importance of each original word in the original corpus, the scoring threshold corresponding to each original word is flexibly set, and then the target replacement vocabulary is selected for each original corpus. By raising the scoring threshold of important original words, the semantic accuracy of the expanded corpus obtained by the final expansion is guaranteed, and by reasonably setting the scoring threshold of unimportant original words, it is ensured that there are enough target replacement words, thereby increasing the number of expanded corpora.
[0126] Of course, in addition to the above embodiments, other methods may be used to select target replacement words, such as selecting candidate replacement words that rank high after sorting the replacement recommendation scores in descending order as target replacement words, which is not limited in the present application.
[0127] In some embodiments, in addition to corpus expansion through vocabulary replacement, the following can also be done: extracting the semantics corresponding to the original corpus and / or the expanded corpus to obtain standard semantic features; generating corpus based on the standard semantic features to obtain sentence conversion corpus that is different from the original corpus and / or the expanded corpus; and using the sentence conversion corpus as new original corpus and / or new expanded corpus.
[0128] Exemplarily, a sentence conversion model is pre-trained, and the original corpus and / or extended corpus are input into the sentence conversion model to obtain a sentence conversion corpus output by the sentence conversion model that has the same semantics as the original corpus and / or extended corpus but a different expression form, and the sentence conversion corpus is used as the new original corpus and / or the new extended corpus.
[0129] Among them, the sentence conversion model can be implemented based on network architectures such as the BART (Bidirectional and Auto-Regressive Transformers) model, or the Recurrent Neural Network (RNN), or the Transformer model, etc. This application does not limit this.
[0130] Taking the BART model as an example, the BART model encodes the input corpus through a bidirectional encoder to obtain standard semantic features, and then decodes the standard semantic features through an auto-regressive decoder to generate new original corpus and / or new extended corpus with different expressions. Since the BART model can capture the global and local semantic features of sentences, the new corpus generated by it not only has significant differences in syntactic structure, but also can generate more diverse natural language expressions while maintaining the core semantics unchanged.
[0131] Through the above embodiments, the number of the finally extended corpus is further increased.
[0132] An example of the specific application scenario of the corpus is given:
[0133] Exemplarily, please refer to Figure 3 , Figure 3 which is a flowchart of training a language model using the original corpus and the extended corpus shown in an exemplary embodiment of this application. As Figure 3 shown, the specific steps include:
[0134] Step S310: Use the original corpus and the extended corpus as the model training corpus.
[0135] Step S320: Perform named entity recognition on the model training corpus to obtain key entities; and perform dependency syntactic analysis on the model training corpus to obtain the syntactic dependency relationships between the various words in the model training corpus.
[0136] The goal of named entity recognition is to identify words or phrases with practical meanings (such as person names, place names, organization names, etc.) from the model training corpus and associate them with other elements in the model training corpus.
[0137] By identifying the key entities, it can provide the necessary basic information for subsequent relation extraction to ensure that the key elements in the corpus can be captured.
[0138] Furthermore, perform dependency syntactic analysis on the model training corpus to obtain the syntactic dependency relationships between the various words in the model training corpus. Dependency syntactic analysis helps to capture the syntactic dependency structure between the words in the text and can help understand the syntactic relationships between entities.
[0139] Among them, a dependency parsing tool (such as SpaCy or Stanford Parser) can be used to perform dependency parsing on the model training corpus, and this application does not limit this.
[0140] Step S330: Based on the syntactic dependency relationships between various words in the model training corpus, identify the entity association relationships between various key entities.
[0141] After identifying the key entities, relation extraction needs to be carried out next. Relation extraction is the process of identifying the mutual relationships between entities. In this application, a dependency parsing tool is used to perform dependency parsing on the model training corpus, and a dependency tree is constructed based on the syntactic dependency relationships between various words in the analyzed model training corpus.
[0142] After obtaining the dependency tree of the model training corpus, a series of predefined rules are used to identify the entity association relationships between key entities.
[0143] For example, some relationships may be represented by specific verbs or prepositional phrases. The dependent path such as the predicate verb, object, or subject in the dependency relationship can be used as the basis for the rules, and these rules can cover common relationship patterns such as "belong to", "consist of", etc. Based on the dependency tree to match these dependency patterns and syntactic structures, the relationships between key entities can be effectively identified.
[0144] Illustrate with an example. If the model training corpus contains "Company A is a subsidiary of Company B", extract the two key entities "Company A" and "Company B", and match the entity association relationship of "subsidiary" according to the syntactic dependency relationships between various words.
[0145] Step S340: Generate triples based on each key entity and the entity association relationships between each key entity, and obtain a triple set.
[0146] The triple can be expressed as (entity, entity association relationship, entity).
[0147] Step S350: Use the triple set to train the pre-trained language model to obtain the trained language model.
[0148] The triples provide clear entity and relationship information for the model, which can enhance its performance in tasks such as knowledge reasoning and question answering systems.
[0149] Of course, in addition to the above use of the corpus for training the language model, it can also be applied in other application scenarios, and this application does not limit this.
[0150] The corpus expansion method provided by this application obtains the original words corresponding to the original corpus by performing word segmentation on the original corpus; and obtains the candidate replacement words corresponding to the original words, where the original words and the candidate replacement words are correspondingly marked with the number of word meanings. Based on the number of word meanings corresponding to the candidate replacement words and / or the original words, determine the scoring influencing factors that match the candidate replacement words. Calculate the replacement recommendation score corresponding to the candidate replacement words using the scoring influencing factors, and select the candidate replacement words whose replacement recommendation scores meet the preset conditions as the target replacement words corresponding to the original words. Replace the original words in the original corpus with the target replacement words corresponding to the original words to obtain the expanded corpus. The scoring influencing factors that match different candidate replacement words can be accurately obtained according to the number of word meanings corresponding to the candidate replacement words and / or the original words, improving the calculation accuracy of the replacement recommendation score, and further improving the accuracy of selecting the target replacement words for each original word. On the premise of increasing the corpus quantity, the situation of semantic changes in the expanded corpus relative to the original corpus is reduced, ensuring the quality of the expanded corpus.
[0151] Figure 4 It is a block diagram of a corpus expansion device shown in an exemplary embodiment of this application. As Figure 4 shown, the exemplary corpus expansion device 400 includes:
[0152] A word segmentation module 410, configured to perform word segmentation on the original corpus to obtain the original words corresponding to the original corpus; and obtain the candidate replacement words corresponding to the original words, where the original words and the candidate replacement words are correspondingly marked with the number of word meanings, and the number of word meanings refers to the number of meanings contained in the word.
[0153] A factor matching module 420, configured to determine the scoring influencing factors that match the candidate replacement words based on the number of word meanings corresponding to the candidate replacement words and / or the original words, where the scoring influencing factors refer to the factors that affect the actual semantics of the corpus when replacing the candidate replacement words with the original words in the corpus.
[0154] A word selection module 430, configured to calculate the replacement recommendation score corresponding to the candidate replacement words using the scoring influencing factors, and select the candidate replacement words whose replacement recommendation scores meet the preset conditions as the target replacement words corresponding to the original words.
[0155] A word replacement module 440, configured to replace the original words in the original corpus with the target replacement words corresponding to the original words to obtain the expanded corpus.
[0156] It should be noted that the corpus expansion device provided in the above embodiments and the corpus expansion method provided in the above embodiments belong to the same concept. The specific ways in which each module and unit perform operations have been described in detail in the method embodiments, and will not be repeated here. In practical applications, the corpus expansion device provided in the above embodiments can, according to needs, allocate the above functions to different functional modules, that is, divide the internal structure of the device into different functional modules to complete all or part of the functions described above. This is not limited herein.
[0157] Please refer to Figure 5 , Figure 5 which is a schematic structural diagram of an embodiment of an electronic device of the present application. The electronic device 500 includes a memory 501 and a processor 502. The processor 502 is used to execute program instructions stored in the memory 501 to implement the steps in any of the above method embodiments of the corpus expansion method. In a specific implementation scenario, the electronic device 500 may include, but is not limited to: a microcomputer, a server. In addition, the electronic device 500 may also include mobile devices such as a laptop computer, a tablet computer, etc., which are not limited herein.
[0158] Specifically, the processor 502 is used to control itself and the memory 501 to implement the steps in any of the above method embodiments of the corpus expansion method. The processor 502 may also be referred to as a Central Processing Unit (CPU). The processor 502 may be an integrated circuit chip with signal processing capabilities. The processor 502 may also be a general-purpose processor, a Digital Signal Processor (DSP), an Application-Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. Additionally, the processor 502 may be implemented jointly by integrated circuit chips.
[0159] Please refer to Figure 6 , Figure 6 which is a schematic structural diagram of an embodiment of a computer-readable storage medium of the present application. The computer-readable storage medium 600 stores program instructions 610 that can be run by a processor. The program instructions 610 are used to implement the steps in any of the above method embodiments of the corpus expansion method.
[0160] In some embodiments, the functions or modules included in the apparatus provided by the embodiments of the present disclosure can be used to execute the methods described in the foregoing method embodiments. The specific implementation can refer to the description of the foregoing method embodiments. For the sake of brevity, it will not be repeated here.
[0161] The foregoing descriptions of the various embodiments tend to emphasize the differences between the various embodiments. Their similarities or resemblances can be referred to each other. For the sake of brevity, they will not be repeated herein.
[0162] In several embodiments provided in the present application, it should be understood that the disclosed methods and apparatuses can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical or other forms.
[0163] In addition, the functional units in the various embodiments of the present application can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units. If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods in the various embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.
Claims
1. A corpus expansion method, characterized in that, The method includes: Performing word segmentation on the original corpus to obtain the original words corresponding to the original corpus; and obtaining candidate replacement words corresponding to the original words; wherein, the original words and the candidate replacement words are marked with the number of semantic meanings, and the number of semantic meanings refers to the number of meanings contained in the word; Determining the number of scoring influencing factors based on the number of semantic meanings corresponding to the candidate replacement words and / or the original words, and selecting the corresponding number of scoring influencing factors matching the candidate replacement words from the set of scoring influencing factors; alternatively, pre-setting scoring influencing factors to be considered for different ranges of the number of semantic meanings, and determining the scoring influencing factors matching the candidate replacement words based on the range of the number of semantic meanings where the candidate replacement words and / or the original words are located; wherein, the scoring influencing factors refer to the factors that affect the actual semantics of the corpus when the candidate replacement words replace the original words in the corpus; Calculating the replacement recommendation score corresponding to the candidate replacement words by using the matching scoring influencing factors, and selecting the candidate replacement words whose replacement recommendation scores meet the preset conditions as the target replacement words corresponding to the original words; Replacing the original words in the original corpus with the target replacement words corresponding to the original words to obtain an extended corpus.
2. The method according to claim 1, characterized in that The scoring influencing factors include the lexical semantic similarity between the original word and the candidate replacement word and / or the sentence semantic similarity between the original corpus and the corpus after replacing the original word in the original corpus with the candidate replacement word; If the number of semantic meanings corresponding to the candidate replacement word and / or the original word is less than or equal to the preset quantity threshold, then taking the lexical semantic similarity as the scoring influencing factor matching the candidate replacement word; If the number of semantic meanings corresponding to the candidate replacement word and / or the original word is greater than the preset quantity threshold, then taking the lexical semantic similarity and the sentence semantic similarity as the scoring influencing factors matching the candidate replacement word.
3. The method according to claim 2, wherein The scoring influencing factors matching the candidate replacement words include the lexical semantic similarity and the sentence semantic similarity; the calculating the replacement recommendation score corresponding to the candidate replacement words by using the matching scoring influencing factors includes: Mapping the lexical semantic similarity to obtain a first score, and mapping the sentence semantic similarity to obtain a second score; Performing weighted summation on the first score and the second score, and taking the weighted summation result as the replacement recommendation score corresponding to the candidate replacement word.
4. The method according to claim 1, wherein The selecting the candidate replacement words whose replacement recommendation scores meet the preset conditions as the target replacement words corresponding to the original words includes: Calculating the importance degrees corresponding to the respective original words in the original corpus; Based on the importance degrees corresponding to the respective original words, setting scoring thresholds for the respective original words; wherein, the importance degree is positively correlated with the scoring threshold; For each of the original words, candidate replacement words with a replacement recommendation score greater than the corresponding score threshold of each original word are respectively selected to obtain the target replacement words corresponding to each of the original words.
5. The method according to claim 1, wherein The method further includes: Extract the semantics corresponding to the original corpus and / or the extended corpus to obtain standard semantic features; Generate a corpus based on the standard semantic features to obtain a sentence conversion corpus different from the original corpus and / or the extended corpus; Use the sentence conversion corpus as the new original corpus and / or the new extended corpus.
6. The method according to claim 1, wherein Before performing word segmentation on the original corpus to obtain the original words corresponding to the original corpus, the method further includes: Obtain original structured data; Extract entities and / or attributes from the original structured data to obtain a full-element set; Generate an original corpus based on the full-element set.
7. The method according to claim 6, characterized in that, The generating the original corpus based on the full-element set includes: Add the entities or attributes in the full-element set to a preset corpus template to obtain an initial corpus; Input the initial corpus into a pre-trained natural language model for natural language conversion to obtain the original corpus output by the natural language model.
8. The method according to claim 1, wherein The method further includes: Use the original corpus and the extended corpus as model training corpora; Perform named entity recognition on the model training corpora to obtain key entities; and perform dependency syntactic analysis on the model training corpora to obtain the syntactic dependency relationships between the words in the model training corpora; Based on the syntactic dependency relationships between the words in the model training corpora, identify the entity association relationships between the key entities; Generate triples based on the key entities and the entity association relationships between the key entities to obtain a triple set; Use the triple set to train a pre-trained language model to obtain a trained language model.
9. An electronic device, characterized in that, The electronic device includes a memory and a processor, and the processor is configured to execute program instructions stored in the memory to implement the steps in the method according to any one of claims 1-8.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores program instructions, and the program instructions can be executed by a processor to implement the steps in the method according to any one of claims 1-8.
Citation Information
Patent Citations
Training corpus set construction method and device and text processing method and device
CN115033753A
Semantic analysis method and device and computer readable storage medium
CN118114681A