Sample processing method and device

By constructing the initial text meaning group and scene-oriented word list space of the question-and-answer model, the problems of resource waste and data redundancy in the data preparation process in the prior art are solved, and the prediction accuracy and processing capabilities of the model are improved.

CN113987147BActive Publication Date: 2025-06-10BEIJING KINGSOFT DIGITAL ENTERTAINMENT CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111256825.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-06-16
Publication Date
2025-06-10
Estimated Expiration
2041-06-16

AI Technical Summary

Technical Problem

In the prior art, in the data preparation process before training of the Q&A model, a large amount of manpower is required to participate in data processing and annotation, resulting in waste of resources and data redundancy, affecting the prediction accuracy of the model.

Method used

By constructing the initial text meaning group of the sample corpus and adding context labels to it, extracting the initial phrases, establishing the correspondence between the context labels and text meaning groups, and constructing a scene-oriented word list space to reduce redundant data and improve the adequate level of data preparation.

Benefits of technology

It effectively reduces the negative impact of redundant data on model training, saves storage resources, and improves the prediction accuracy and processing capabilities of the Q&A model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113987147B_ABST
    Figure CN113987147B_ABST
Patent Text Reader

Abstract

The present application provides a sample processing method and apparatus. The sample processing method includes: obtaining a sample corpus and constructing an initial text semantic group corresponding to the sample corpus; adding context tags to the sample corpus and extracting initial phrases corresponding to the initial text semantic group; establishing a correspondence between the context tags and the initial text semantic group; and constructing a scene-oriented vocabulary space corresponding to the sample corpus according to the correspondence and the initial phrases.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and particularly to a sample processing method and apparatus. Background Art

[0002] With the development of the artificial intelligence industry, the proportion of question-and-answer models in practical applications has gradually increased, and users have higher and higher requirements for the response accuracy and efficiency of question-and-answer models. In practical applications, the prediction accuracy of question-and-answer models depends on the quality and quantity of samples in the training stage. In the prior art, in the data preparation stage before training a question-and-answer model, manual participation is usually adopted to process and label data. This process not only consumes a large amount of human resources, but also generates a large amount of redundant data in the data processing stage due to the complex components contained in the sample corpus, resulting in excessive occupation of storage resources and having a certain impact on the accuracy of the question-and-answer model to be trained. Therefore, an effective solution is urgently needed to solve the above problems. Summary of the Invention

[0003] In view of this, embodiments of the present application provide a sample processing method to solve the technical defects existing in the prior art. Embodiments of the present application also provide a sample processing apparatus, a training method for a question-and-answer model, a training apparatus for a question-and-answer model, a computing device, and a computer-readable storage medium.

[0004] According to a first aspect of embodiments of the present application, there is provided a sample processing method, including:

[0005] Obtain a sample corpus and construct an initial text semantic group corresponding to the sample corpus;

[0006] Add a context label to the sample corpus and extract an initial phrase corresponding to the initial text semantic group;

[0007] Establish a correspondence between the context label and the initial text semantic group;

[0008] Construct a scene-oriented vocabulary space corresponding to the sample corpus according to the correspondence and the initial phrase.

[0009] Optionally, after the step of constructing a scene-oriented vocabulary space corresponding to the sample corpus according to the correspondence and the initial phrase, the method further includes:

[0010] Obtain a training sample and determine a sample phrase corresponding to the training sample;

[0011] Query the scene-oriented vocabulary space based on the sample phrase, and determine a target text semantic group corresponding to the training sample according to the query result;

[0012] Train the initial Q&A model using the target text semantic group and the training samples until a target Q&A model that meets the training stop condition is obtained.

[0013] Optionally, adding context labels to the sample corpus includes:

[0014] Extract multiple initial features of the sample corpus and preprocess the multiple initial features to obtain multiple target features;

[0015] Calculate the context similarity between each target feature and the sample corpus, and select at least one target feature as the context label according to the context similarity calculation result and add it to the sample corpus.

[0016] Optionally, querying the scenario-oriented thesaurus space based on the sample phrase and determining the target text semantic group corresponding to the training sample includes:

[0017] Map the sample phrase to the scenario-oriented thesaurus space and calculate the phrase similarity between the sample phrase and the context label;

[0018] Determine the target context label according to the phrase similarity calculation result, and use the initial text semantic group corresponding to the target context label as the target text semantic group.

[0019] Optionally, obtaining the training samples includes:

[0020] Obtain the training samples that have an association relationship with the sample corpus;

[0021] Among them, using the target text semantic group and the training samples to train the initial Q&A model until a target Q&A model that meets the training stop condition is obtained includes:

[0022] Train the initial Q&A model using the training samples that have an association relationship with the sample corpus and the target text semantic group until a target Q&A model that meets the training stop condition is obtained.

[0023] Optionally, determining the sample phrase corresponding to the training sample includes:

[0024] Parse the training sample to obtain the sample question text in the training sample;

[0025] Extract the first word unit and the second word unit from the sample question text, and construct the sample phrase based on the first word unit and the second word unit.

[0026] Optionally, training the initial Q&A model with the target text sense group and the training samples until a target Q&A model that meets the training stop condition is obtained includes:

[0027] Inputting the target text sense group and the sample question text in the training samples into the initial Q&A model for processing to obtain a predicted answer text;

[0028] Optimizing the initial Q&A model based on the predicted answer text and the sample answer text in the training samples until the target Q&A model that meets the training stop condition is obtained.

[0029] Optionally, inputting the target text sense group and the sample question text in the training samples into the initial Q&A model for processing to obtain a predicted answer text includes:

[0030] Generating a word unit vector and a scene label vector based on the sample question text, and generating a sense group vector based on the target text sense group;

[0031] Integrating the word unit vector and the scene label vector to obtain a sample question vector corresponding to the sample question text;

[0032] Inputting the sample question vector and the sense group vector into the initial Q&A model for processing to obtain the predicted answer text.

[0033] Optionally, inputting the sample question vector and the sense group vector into the initial Q&A model for processing to obtain the predicted answer text includes:

[0034] Inputting the sample question vector and the sense group vector into the initial Q&A model, and processing the sample question vector and the sense group vector through a fusion module in the initial Q&A model to obtain a fusion vector;

[0035] Inputting the fusion vector into an identification module in the initial Q&A model for processing to obtain an associated entity central word and a context scene distribution;

[0036] Processing the associated entity central word and the context scene distribution through an output layer in the initial Q&A model to obtain the predicted answer text.

[0037] Optionally, preprocessing the multiple initial features to obtain multiple target features includes:

[0038] Cleaning the multiple initial features, and determining the multiple target features according to the cleaning result;

[0039] Among them, selecting at least one target feature as the context label according to the context similarity calculation result includes:

[0040] Comparing the context similarity with a preset context similarity threshold, and selecting a target feature greater than or equal to the context similarity threshold as the context label; or

[0041] Selecting the target feature with the largest similarity as the context label according to the context similarity calculation result.

[0042] According to the second aspect of the embodiments of the present application, a sample processing device is provided, including:

[0043] An acquisition module, configured to acquire a sample corpus and construct an initial text sense group corresponding to the sample corpus;

[0044] An addition module, configured to add a context label to the sample corpus and extract an initial phrase corresponding to the initial text sense group;

[0045] An establishment module, configured to establish a correspondence between the context label and the initial text sense group;

[0046] A construction module, configured to construct a scene-oriented vocabulary space corresponding to the sample corpus according to the correspondence and the initial phrase.

[0047] According to the third aspect of the embodiments of the present application, a method for training a question-and-answer model is provided, including:

[0048] Acquiring a training sample and determining a sample phrase corresponding to the training sample;

[0049] Querying the scene-oriented vocabulary space in the above method based on the sample phrase, and determining a target text sense group corresponding to the training sample according to the query result;

[0050] Training an initial question-and-answer model by using the target text sense group and the training sample until a target question-and-answer model that meets the training stop condition is obtained.

[0051] According to the fourth aspect of the embodiments of the present application, a device for training a question-and-answer model is provided, including:

[0052] A sample acquisition module, configured to acquire a training sample and determine a sample phrase corresponding to the training sample;

[0053] A sense group determination module, configured to query the scene-oriented vocabulary space in the above method based on the sample phrase, and determine a target text sense group corresponding to the training sample according to the query result;

[0054] A training model module, configured to train an initial question-answering model by using the target text sense groups and the training samples until a target question-answering model that meets the training stop condition is obtained.

[0055] According to a fifth aspect of the embodiments of the present application, a computing device is provided, including:

[0056] A memory and a processor;

[0057] The memory is used to store computer-executable instructions, and when the processor executes the computer-executable instructions, the steps of the sample processing method or the training method of the question-answering model are implemented.

[0058] According to a sixth aspect of the embodiments of the present application, a computer-readable storage medium is provided, which stores computer-executable instructions, and when the instructions are executed by a processor, the steps of the sample processing method or the training method of the question-answering model are implemented.

[0059] For the sample processing method provided by the present application, in order to reduce the impact of redundant data on the model training stage and save storage resources; after obtaining the sample corpus, an initial text sense group corresponding to the sample corpus can be constructed, and then context labels are added to the sample corpus, and at the same time, the initial phrases corresponding to the initial text sense groups are extracted. Secondly, the correspondence between the context labels and the initial text sense groups is established, and finally, a scene-oriented vocabulary space corresponding to the sample corpus is constructed according to the correspondence and the initial phrases. By constructing the scene-oriented vocabulary space starting from the initial text sense groups of the sample corpus during the data preparation stage before model training, not only can the redundant data in the sample corpus occupying too many resources and the negative impacts be reduced, but also the richness of the scene-oriented vocabulary space can be ensured, thereby effectively improving the sufficiency of the data preparation stage and contributing to improving the prediction accuracy of the model in the model training stage. Description of the Drawings

[0060] Figure 1 is a flowchart of a training method for a question-answering model provided by an embodiment of the present application;

[0061] Figure 2 is a schematic structural diagram of a question-answering model in a training method for a question-answering model provided by an embodiment of the present application;

[0062] Figure 3 is a schematic structural diagram of a training device for a question-answering model provided by an embodiment of the present application;

[0063] Figure 4 is a flowchart of a text processing method provided by an embodiment of the present application;

[0064] Figure 5It is a processing flow chart provided by an embodiment of the present application for the ancient poetry Q&A scenario;

[0065] Figure 6 It is a schematic structural diagram of a text processing device provided by an embodiment of the present application;

[0066] Figure 7 It is a structural block diagram of a computing device provided by an embodiment of the present application. Detailed implementation manners

[0067] In the following description, many specific details are set forth in order to provide a thorough understanding of the present application. However, the present application can be implemented in many other ways different from those described herein, and those skilled in the art can make similar generalizations without departing from the connotation of the present application. Therefore, the present application is not limited by the specific implementations disclosed below.

[0068] The terms used in one or more embodiments of the present application are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of the present application. The singular forms "a", "the", and "said" used in one or more embodiments of the present application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of the present application refers to and includes any or all possible combinations of one or more of the associated listed items.

[0069] It should be understood that although the terms first, second, etc. may be used in one or more embodiments of the present application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of one or more embodiments of the present application, the first may also be referred to as the second, and similarly, the second may also be referred to as the first.

[0070] First, the noun terms related to one or more embodiments of the present invention are explained.

[0071] ERNIE: By modeling words, entities, and entity relationships in a large amount of data, it learns semantic knowledge in the real world; compared with BERT that learns local language co-occurrence semantic representations, ERNIE directly models semantic knowledge, enhancing the model's semantic representation ability.

[0072] LSTM: (Long Short-Term Memory) is a time-recurrent neural network designed specifically to solve the long-term dependence problem existing in general RNNs. All RNNs have a chain form of repeating neural network modules. In a standard RNN, this repeating structural module has a very simple structure, such as a tanh layer.

[0073] RNN: (Recurrent Neural Network) is a type of recurrent neural network that takes sequence data as input, performs recursion in the evolution direction of the sequence, and all nodes (recurrent units) are connected in a chain.

[0074] BiLSTM: (Bi-directional Long Short-Term Memory) is composed of a forward LSTM and a backward LSTM. It is applied to model context information in natural language processing tasks.

[0075] Phrase similarity: The similarity between two phrases can be calculated by the dot product between the word vectors of the two phrases.

[0076] Corpus: The basic element used in translation or language research scenarios, which is the basic unit that makes up a corpus.

[0077] LDA: (Latent Dirichlet Allocation) is a document topic generation model, also known as a three-layer Bayesian probability model, which includes three layers: words, topics, and documents. The so-called generation model means that it is considered that each word in an article is obtained through a process of "selecting a certain topic with a certain probability and then selecting a certain word from this topic with a certain probability". The distribution of documents to topics follows a multinomial distribution, and the distribution of topics to words follows a multinomial distribution.

[0078] Semantic Dependency Parsing (SDP): Analyzes the semantic associations between various language units in a sentence and presents the semantic associations in a dependency structure. Using semantic dependencies to depict sentence semantics does not require abstracting the vocabulary itself, but rather describes the vocabulary through the semantic framework it bears. The number of arguments is always much smaller than the number of vocabulary. The goal of semantic dependency parsing is to break through the constraints of the surface syntactic structure of the sentence and directly obtain the deep semantic information.

[0079] Sense group: Refers to each component divided according to meaning and structure in a sentence. Each component is called a sense group; the words within the same sense group are closely related and cannot be split randomly, otherwise it will cause misunderstanding.

[0080] Context: Refers to the environment in which language is used; the internal context refers to the relationship between a certain speech segment and a certain context, and the external context refers to the social environment of the language outside the speech segment.

[0081] In the present application, a method for training a question-and-answer model is provided. The present application also relates to a device for training a question-and-answer model, a method for text processing, a device for text processing, a computing device, and a computer-readable storage medium, which will be described in detail one by one in the following embodiments.

[0082] In practical applications, question-and-answer systems are applied in various fields. Due to different field characteristics, the difficulty of questions processed by question-and-answer systems in different fields varies. For example, in a classical poetry question-and-answer system, in order to accurately answer questions related to classical poetry, during the model preparation stage, classical poetry entities are usually extracted for training the question-and-answer model in the classical poetry question-and-answer system to make accurate answers. In the prior art, classical poetry entity extraction is usually divided into two categories, namely rule extraction and model extraction. The rule extraction method extracts entities by judging important words and sentences in the original poetry text and according to grammar rules and a poetry knowledge base. The model extraction method applies algorithms of natural language processing and generates a more concise and refined set of poetry entities through techniques such as self-attention, pre-training, and semantic group adaptation. Compared with the rule extraction method, the model extraction method is closer to the process of entity discovery and extraction by users. With the rise and research of deep neural networks, the entity extraction model algorithm based on the neural network Transformer has developed rapidly, achieved good results, and demonstrated strong model generalization ability.

[0083] However, due to the complexity of classical poetry itself, in the stage of classical poetry entity extraction, the existing extraction algorithms have relatively fuzzy processing granularity for poetry text entities, resulting in a lack of context recognition for the entity words of the output poetry corpus, and there are many repetitions of the output entity words, with serious data redundancy, which greatly affects the quality of entity extraction. Therefore, an effective solution is urgently needed to solve the above problems.

[0084] In view of this, the present application provides a method for training a question-and-answer model. After constructing the initial text semantic groups corresponding to the sample corpus, a scene-oriented vocabulary space corresponding to the sample corpus is generated based on the initial text semantic groups. At this time, sufficient corpus is prepared for model training. Then, training samples are obtained, and the sample phrases corresponding to the training samples are determined. The scene-oriented vocabulary space is queried using the sample phrases to determine the target text semantic groups corresponding to the training samples according to the query results. Finally, the initial question-and-answer model is trained based on the target text semantic groups and the training samples until the training stop condition is met, and the target question-and-answer model can be obtained. It realizes capturing the association between questions and corpus at the semantic level, effectively ensures the prediction accuracy of the trained question-and-answer model, and effectively improves the processing ability of the question-and-answer model by constructing a scene-oriented vocabulary space using rich sample corpus during the preparation stage, so as to accurately and efficiently complete the question-and-answer processing task.

[0085] Figure 1 The flowchart of a method for training a question-answering model provided according to an embodiment of the present application is shown, which specifically includes the following steps:

[0086] Step S102: Construct an initial text semantic group corresponding to the sample corpus, and generate a scenario-oriented word list space corresponding to the sample corpus based on the initial text semantic group.

[0087] Specifically, the sample corpus specifically refers to the text corpus that provides all sample information when training the question-answering model, and the sample corpora corresponding to different fields are also different. For example, in the field of classical poetry question answering, the sample corpus can be a text corpus composed of the main body of the poem, the title, the author, the author's information, the explanation of the main body of the poem, and the analysis of the poem; or in the field of sports knowledge question answering, the sample corpus can be a text corpus composed of sports event interpretations, sports stars, sports star information, and sports event locations; or in the field of person relationship question answering, the sample corpus can be a text corpus composed of persons, person information, family information, employment information, biographies of persons, and deeds of persons. In practical applications, the sample corpora used in different question-answering fields can be obtained and constructed according to actual needs, and this embodiment does not make any limitations here.

[0088] Furthermore, in order to train a question-answering model that meets the usage requirements, it is necessary to continuously provide samples for multiple rounds of iterative training, and at the same time, it is necessary to optimize the model in combination with the loss function. Therefore, in the data preparation stage, a large number of sample corpora need to be prepared to avoid the problems of model overfitting or incomplete training by increasing the sample richness.

[0089] Even further, the initial text semantic group specifically refers to the text semantic group constructed based on each sample corpus, that is, by dividing the sentences in each sample corpus into multiple components according to meaning and structure, and the multiple components can form the initial text semantic group corresponding to the corresponding sample corpus, which is used to construct the scenario-oriented word list space subsequently, and to assist the question-answering model in learning the fine-grained semantic association relationship between the text and the semantic group, so as to ensure the prediction ability of the model. It should be noted that since the data volume of the sample corpus is large, the sample corpus can be stored in the corpus corresponding to the field to facilitate the management and use of the sample corpus.

[0090] Correspondingly, the scenario orientation thesaurus space specifically refers to the expression relationship of the entity word association relationship constructed based on the initial text semantic groups corresponding to each sample corpus. By integrating the initial text semantic groups corresponding to each sample corpus, the initial text semantic group combination corresponding to the full-scale sample corpus can be formed to construct the scenario orientation thesaurus space. The scenario orientation thesaurus space contains the scenario information and context information of each sample corpus, which is used to locate the target text semantic group corresponding to the text subsequently, so as to give a precise answer to the question.

[0091] Furthermore, in the process of creating the scenario orientation thesaurus space based on the initial text semantic groups, since the components included in the initial text semantic groups corresponding to the sample corpus are relatively complex, if all the initial text semantic groups are directly used in full scale to form the scenario orientation thesaurus space, a large amount of redundant data will be generated, occupying too much storage resources. Therefore, in order to reduce the interference caused by the redundant data, the initial phrases will be preferentially extracted to complete the creation of the scenario orientation thesaurus space. In this embodiment, the specific implementation method is as follows in steps S1022 to S1024:

[0092] Step S1022: Add context labels to the sample corpus and extract the initial phrases in the initial text semantic group.

[0093] Specifically, the context label specifically refers to the label added to the sample corpus according to the language environment of the sample corpus. It should be noted that since the language environments corresponding to different sample corpora are relatively complex, multiple different context labels can be added to the sample corpus. For example, for the ancient poem "Thoughts in the Silent Night", it expresses the author's homesickness and the meaning of beautiful scenery. Therefore, when adding context labels to the sample corpus of the ancient poem "Thoughts in the Silent Night", homesickness context labels and scenery context labels can be added; or for the ancient poem "Farewell to Wang Lun", it expresses the author's parting feelings and the friendship between the author and "Wang Lun". Therefore, when adding context labels to the sample corpus of the ancient poem "Farewell to Wang Lun", parting context labels and friendship context labels can be added.

[0094] Based on this, the initial phrase specifically refers to the phrase composed of the nominal original word and the verb phrase stem extracted from the initial text semantic group, which is used to represent the core idea and key content of the sample corpus corresponding to the initial text semantic group, and at the same time can lay a foundation for the subsequent generation of the scenario orientation thesaurus space, ensuring that the answer can be accurately located from the semantics during the question answering stage and ensuring the correctness of the answer.

[0095] In practical applications, in the process of adding context labels to the sample corpus, considering that the context labels corresponding to the sample corpus may not be unique, the LDA feature engineering model can be used to process each sample corpus to accurately determine the context labels for each sample corpus, effectively improving the efficiency of adding context labels to the sample corpus.

[0096] At the same time, when extracting the initial phrases in the initial text meaning group, a semantic dependency analysis tool can also be used, that is, by analyzing the semantic associations between the sentence language units in the sample corpus, and presenting the semantic associations in a dependency structure, using language dependencies to characterize the sentence semantics, and realizing that in the case of the present application where the word units in the sample corpus do not need to be abstracted, the word units are described by the semantic framework borne by the word units, so as to avoid the constraints of the surface syntactic structure of the sentence and directly obtain the deep semantic information. For example, "甲吃苹果" is analyzed and processed by the semantic dependency analysis tool, and it is determined that the word unit "甲" has an Agt relationship with the word unit "吃", the word unit "吃" has an mTime relationship with the word unit "了", and the word unit "吃" has a Pat relationship with the word unit "苹果", wherein the Agt relationship represents the agent relationship, the mTime relationship represents the time mark, and the Pat relationship represents the patient relationship. Similarly, after extracting the initial phrases of each initial text meaning group, the internal relationship between the phrases can be stored for subsequent creation of a scene-oriented word table space.

[0097] Furthermore, in the process of adding context labels to the sample corpus, in order to ensure that the added context labels are more closely matched to the sample corpus, the context labels may be screened by calculating the context similarity. In this embodiment, the specific implementation is as follows:

[0098] Extracting a plurality of initial features of the sample corpus, and preprocessing the plurality of initial features to obtain a plurality of target features;

[0099] The context similarity between each target feature and the sample corpus is calculated, and at least one target feature is selected as the context label according to the context similarity calculation result, and is added to the sample corpus.

[0100] Specifically, the initial feature refers to the meaning expressed by the sample corpus in different dimensions, and the target feature refers to the feature expression obtained after preprocessing each initial feature. Preprocessing specifically refers to cleaning multiple initial features, that is, deleting repeated or redundant features to determine multiple target features from multiple initial features. Correspondingly, contextual similarity specifically refers to calculating the degree of similarity between each target feature and the corresponding sample corpus in the language environment dimension.

[0101] Based on this, after obtaining the sample corpus, in order to construct a richer scenario-oriented vocabulary space, multiple initial features of the sample corpus can be extracted at this time. Then, in order to avoid excessive computational pressure caused by data redundancy, the multiple initial features can be preprocessed to obtain multiple target features, where the number of multiple target features is less than or equal to the number of multiple initial features. Secondly, by calculating the context similarity between each target feature in the multiple target features and the sample corpus from which the initial features are extracted, at least one target feature can be selected as the context label according to the calculation result of the context similarity and added to the sample corpus.

[0102] In practical applications, in the process of selecting at least one target feature as the context label according to the calculation result of the context similarity, considering that the meanings expressed by the sample corpus in different dimensions are not the same, the context similarity can be compared with a preset context similarity threshold, and the target feature greater than or equal to the context similarity threshold can be selected as the context label of the sample corpus; in addition, after calculating the context similarity, the target feature with the largest context similarity can also be selected as the context label to ensure that each sample corpus has a unique context label. Specifically, the method for determining the context label can be selected according to the actual application scenario, and this embodiment does not make any limitation here.

[0103] In summary, by screening the context label starting from the features of the sample corpus, not only can the fit between the context label and the sample corpus be guaranteed, but also the richness of the subsequent constructed oriented scenario vocabulary space can be guaranteed, thereby improving the prediction ability of the question answering model.

[0104] Step S1024, establish the corresponding relationship between the context label and the initial text semantic group, and construct the scenario-oriented vocabulary space corresponding to the sample corpus according to the corresponding relationship and the initial phrase.

[0105] Specifically, on the basis of the above-mentioned completion of adding context labels to the sample corpus and extracting the initial phrases in the initial text semantic group, further, in order to ensure the richness of the constructed scenario-oriented vocabulary space, the corresponding relationship between the context label and the initial text semantic group can be established. That is, the context label is the label added to the sample corpus, and the initial text semantic group is constructed based on the sample corpus. Therefore, by determining each sample corpus, the corresponding relationship between the context label and the initial text semantic group can be determined. Then, according to this corresponding relationship and the initial phrases extracted from the initial text semantic group, the scenario-oriented vocabulary space corresponding to the sample corpus can be constructed for subsequent training of the question answering model.

[0106] This embodiment takes the training of the question answering model in the field of classical poetry as an example for illustration. The training methods of question answering models in other fields can refer to the corresponding description content of this embodiment, and will not be elaborated here.

[0107] For example, a classical poem corpus stores a large amount of corpus corresponding to classical poems, and the corpus corresponding to each classical poem contains the poem text, title, author, author information, poem text explanation, and poem analysis. Taking the classical poem corpus containing ten thousand classical poems and their corresponding content as an example, after determining the classical poem corpus contained in the classical poem corpus, at this time, a coarse-grained initial text semantic group of each classical poem can be established through a task heuristic topic classification algorithm, and at the same time, context labels can be added to each classical poem corpus through an LDA feature engineering model.

[0108] In the process of adding context labels, multiple initial features {homesickness; parting; frontier fortress; friendship; scenery; sadness... ambition} corresponding to each classical poem corpus can be extracted, and then the multiple initial features corresponding to each classical poem corpus are preprocessed to obtain the target features corresponding to each classical poem corpus. By calculating the context similarity between the classical poem corpus and its corresponding target features, the target features greater than the preset context similarity threshold are selected as the context labels of the classical poem corpus and added to its corresponding classical poem corpus. That is, the context labels corresponding to the classical poem corpus of "On Mission to the Frontier" include {frontier fortress; scenery; sad emotion}... The context labels corresponding to the classical poem corpus of "Thoughts in a Quiet Night" include {emotion; thought; scenery}.

[0109] Furthermore, after obtaining the context labels corresponding to each classical poem corpus, the corresponding relationship between the initial text semantic group corresponding to each classical poem context and the context labels can be established. At the same time, a semantic dependency analysis tool is used to extract the nominal original words and verb phrase stems of each initial text semantic group to form the initial word groups corresponding to each initial text semantic group. Combining this corresponding relationship and the initial word groups, a scene-oriented entity word list space corresponding to the classical poem corpus is constructed, which can be used to assist in the training of the classical poem question-answering model and the training of the classical poem question-answering system in the future.

[0110] In summary, in the process of establishing the scene-oriented word list space, by combining the context labels and initial word groups corresponding to the sample corpus, the semantic information of each sample corpus can be effectively guaranteed to be included in this space, which is convenient for the subsequent use in the training of the question-answering model.

[0111] Step S104, obtain training samples and determine the sample word groups corresponding to the training samples.

[0112] Specifically, based on the above-mentioned construction of the scenario-oriented vocabulary space from the sample corpus, further, in order to train a Q&A model that meets the requirements for the field corresponding to the sample corpus, after obtaining the training samples for this field, the sample phrases corresponding to the training samples can be determined to lay a foundation for subsequent determination of the target text semantic groups, so as to establish the relationship between semantic groups and different types of questions for the model, enabling the model to answer questions based on semantics.

[0113] Based on this, the training samples specifically refer to the samples used for subsequent training of the Q&A model, including sample questions and sample answers. Correspondingly, the sample phrases specifically refer to the entity phrases constructed based on the training samples, which are used to map to the scenario-oriented vocabulary space to find the text semantic groups corresponding to the samples in the vocabulary space for subsequent training of the Q&A model.

[0114] Further, in the process of determining the sample phrases corresponding to the training samples, since the scenario-oriented vocabulary space is constructed based on the initial text semantic groups, in order to ensure that the sample phrases can be successfully mapped to the scenario-oriented vocabulary space for subsequent determination of the target text semantic groups, the same architecture method will be selected to determine the sample phrases at this time, so as to ensure that the sample phrases can be successfully mapped to the scenario-oriented vocabulary space. In this embodiment, the specific implementation method is as follows:

[0115] Parse the training samples to obtain the sample question text in the training samples;

[0116] Extract the first word unit and the second word unit from the sample question text, and construct the sample phrase based on the first word unit and the second word unit.

[0117] Specifically, the sample question text specifically refers to the questions prepared in advance related to the field corresponding to the sample corpus. And in order to ensure the subsequent successful training of the Q&A model, the obtained training samples need to have a certain correlation with the sample corpus. This correlation specifically means that the sample question text included in the training samples is proposed based on the sample corpus, and the answer corresponding to this sample question text can also be determined from the sample corpus. Correspondingly, the first word unit specifically refers to the original noun extracted from the sample question text, and the second word unit specifically refers to the verb phrase stem extracted from the sample question text.

[0118] Based on this, after obtaining the training samples, the training samples can be parsed to obtain the sample question text in the training samples. Then, in order to be able to determine the target text semantic group associated with the sample question text from the scenario-oriented thesaurus space in the subsequent process and assist in the training of the complete question-answering model, the nominal original words and verb phrase stems can be extracted from the sample question text. Then, the two are integrated to construct the sample phrase group corresponding to the training sample, ensuring that the structure of the sample phrase group is the same as the structure of the scenario-oriented thesaurus space, so as to achieve the rapid determination of the target text semantic group and accelerate the training efficiency of the question-answering model.

[0119] Continuing with the above example, after constructing the scenario-oriented entity thesaurus space based on ten thousand classical poems and their corresponding contents, at this time, in order to be able to train a classical poem question-answering model for answering questions about classical poems, training samples containing sample question text and sample answer text can be obtained. Moreover, the sample question text included in the training samples is related to classical poems. The sample question text can include {What is the central idea of "Thoughts in the Silent Night"?}, {Please provide an ancient poem describing the scenery of the frontier fortress}, or {Who is the author of "Yellow Crane Tower"?}, etc. And in order to be able to determine the classical poem corpus associated with each sample question text, at this time, the nominal original words and verb phrase stems in each sample question text can be extracted respectively to generate the to-be-oriented associated entity phrase groups corresponding to each sample question text, which is convenient for subsequent mapping to the scenario-oriented entity thesaurus space and used to determine the target text semantic group corresponding to each sample question text.

[0120] Step S106, query the scenario-oriented thesaurus space based on the sample phrase group, and determine the target text semantic group corresponding to the training sample according to the query result.

[0121] Specifically, based on the above determination of the sample phrase group corresponding to the training sample, in order to ensure that the subsequent trained question-answering model can learn the semantic association relationship between the question text and the text semantic group, at this time, the scenario-oriented thesaurus space constructed in advance can be queried based on the sample phrase group corresponding to the training sample, so as to accurately determine the target text semantic group corresponding to the training sample according to the query result and be used for the subsequent training of the question-answering model.

[0122] Based on this, the target text semantic group specifically refers to the initial text semantic group with a relatively high degree of association with the training sample screened out from the scenario-oriented thesaurus space, and the answer corresponding to the sample question text in the training sample can be located from the sample corpus corresponding to the initial text semantic group.

[0123] Furthermore, in the process of determining the target text sense group, since the initial text sense groups corresponding to the sample corpora included in the scenario orientation vocabulary space are relatively complex, in order to accurately determine the target text sense group corresponding to the training sample, the target text sense group can be determined starting from the context label. In this embodiment, the specific implementation method is as follows:

[0124] Map the sample phrase to the scenario orientation vocabulary space, and calculate the phrase similarity between the sample phrase and the context label;

[0125] Determine the target context label according to the calculation result of the phrase similarity, and use the initial text sense group corresponding to the target context label as the target text sense group.

[0126] Specifically, the phrase similarity specifically refers to calculating the similarity between the sample phrase and the context labels included in the scenario orientation vocabulary space, and the target context label specifically refers to the context label with the highest phrase similarity to the sample phrase.

[0127] Based on this, after determining the sample phrase corresponding to the training sample, the sample phrase can be mapped to the scenario orientation vocabulary space, then calculate the phrase similarity between the sample phrase and the context labels included in the scenario orientation vocabulary space, then select the context label with the highest phrase similarity as the target context label, and finally determine the initial text sense group corresponding to the target context label in the scenario orientation vocabulary space as the target text sense group for subsequent training of the question answering model. Specifically, in the process of determining the target context label according to the phrase similarity, one or more target context labels can be determined. Correspondingly, the target text sense groups determined from the scenario orientation vocabulary space can also be one or more.

[0128] Continuing with the above example, after determining the to-be-oriented associated entity phrases corresponding to each sample question text, each to-be-oriented associated entity phrase can be mapped to the scenario-oriented entity vocabulary space. Then, calculate the phrase similarity between each to-be-oriented associated entity phrase and the phrase containing the context label in the scenario-oriented entity vocabulary space. According to the calculation results, determine the target context label associated with each sample question text. It is determined that the target context label associated with the sample question text "{What is the central idea of 'Thoughts in the Silent Night'?}" is "thought"; the target context label associated with the sample question text "{Please provide an ancient poem describing the scenery of the border area.}" is "border area";... The target context label associated with the sample question text "{Who is the author of 'Yellow Crane Tower'?}" is "scenery". At this time, the target text semantic groups corresponding to each sample question text can be determined by combining the sample phrases corresponding to the sample question text and the context labels. Among them, the target text semantic group corresponding to the sample question text "{What is the central idea of 'Thoughts in the Silent Night'?}" is the initial text semantic group corresponding to the classical poem 'Thoughts in the Silent Night', the target text semantic group corresponding to "{Please provide an ancient poem describing the scenery of the border area.}" is the initial text semantic group corresponding to the classical poem 'An envoy to the frontier',... The target text semantic group corresponding to "{Who is the author of 'Yellow Crane Tower'?}" is the initial text semantic group corresponding to the classical poem 'Yellow Crane Tower', which is used for subsequent training of the classical poem Q&A model.

[0129] In summary, by calculating the phrase similarity to determine the target text semantic group corresponding to the sample question text from the scenario-oriented vocabulary space, the accuracy of determining the target text semantic group can be effectively improved, and at the same time, it is ensured that the subsequent Q&A model can accurately learn the semantic association relationship, so as to train a Q&A model that meets the requirements.

[0130] Step S108, use the target text semantic group and the training samples to train the initial Q&A model until a target Q&A model that meets the training stop condition is obtained.

[0131] Specifically, based on the above determination of the target text semantic group corresponding to the training samples, further, at this time, the initial Q&A model can be trained by combining the target text semantic group and the training samples, so that the initial Q&A model can learn the fine-grained semantic association relationship between the sample question text in the training samples and the target text semantic group. Then, through continuous iteration and optimization, a target Q&A model that meets the training stop condition can be obtained. The training stop condition can be the number of training iterations or the comparison of loss values. In practical applications, the training stop condition can be set according to requirements, and this embodiment does not make any limitations here.

[0132] In practical applications, since the question-and-answer systems in different fields have different architectures, in order to train a question-and-answer model with better prediction accuracy for the question-and-answer system in that field, multiple modular components with different functions can be integrated into the question-and-answer system. For example, in the field of classical poetry question and answer, the involved classical poetry question-and-answer system can introduce a context recognition attention module. By setting context discrimination labels, a Chinese word segmentation module based on deep semantic units is adopted. At the same time, BiLSTM is used to establish the word-level hidden layer state layer distribution representation of the label statement and the poetry text statement, and calculate the attention vector matrix that fuses the label and text semantic information, which can quickly realize entity discovery and classification extraction. At the same time, a memory unit is configured to improve the accuracy of entity extraction and avoid repeated extraction of entity words related to the poetry question-and-answer context, thus effectively ensuring that accurate answers can be given to questions related to classical poetry.

[0133] In specific implementation, using BiLSTM actually analyzes the task from three dimensions: question type, high-frequency entity, and associated entity, establishes three corresponding learnable weight matrices, and through multiple rounds of iterative learning, establishes a semantic mapping matrix between the poetry corpus and different types of questions. According to the learning effect, a loss function is designed, different weights are assigned to the three matrices, and finally they are concatenated into a weight matrix for the real-time question-and-answer system.

[0134] Furthermore, in the process of training the initial question-and-answer model using the target text semantic group and training samples, since the initial question-and-answer model is a supervised model, it needs to be continuously optimized and parameter-tuned to complete the training of the model. In this embodiment, the training process is as follows in steps S1082 to S1084:

[0135] Step S1082, input the target text semantic group and the sample question text in the training samples into the initial question-and-answer model for processing to obtain a predicted answer text.

[0136] Specifically, the predicted answer text specifically refers to the text corresponding to the answer queried from the target text semantic group after the initial question-and-answer model performs prediction processing on the sample question text in the training samples.

[0137] Furthermore, during the training of the initial question-and-answer model, it is actually learning the semantic association relationship between the target text semantic group and the sample question text, and continuously optimizing the semantic association relationship by adjusting parameters, so as to obtain a target question-and-answer model that meets the training stop condition. In this embodiment, the specific implementation method is as follows:

[0138] Generate a word unit vector and a scene label vector based on the sample question text, and generate a semantic group vector based on the target text semantic group;

[0139] Integrate the word unit vector and the scenario label vector to obtain the sample question vector corresponding to the sample question text;

[0140] Input the sample question vector and the sense group vector into the initial Q&A model for processing to obtain the predicted answer text.

[0141] Specifically, the word unit vector specifically refers to the lexical and syntactic unit word vector constructed based on the sample question text. Correspondingly, the scenario label vector specifically refers to the label vector that can express the scenario of the sample question text constructed based on the sample question text. The sense group vector specifically refers to the vector representation constructed based on the sense groups of the target text and is used as the input of the model to facilitate model processing.

[0142] Based on this, after determining the sense groups of the target text and the sample question texts in the training samples, word unit vectors and scenario label vectors can be generated based on the sample question texts, and at the same time, sense group vectors can be generated based on the sense groups of the target text. Then, the word unit vectors and scenario label vectors are integrated to obtain the sample question vector corresponding to the sample question text. Finally, the content converted into vector representation (sample question vector and sense group vector) is input into the initial Q&A model for processing at the same time, and the predicted answer text corresponding to the sample question text can be predicted through the model to facilitate subsequent optimization of the model in combination with the sample answer texts in the training samples.

[0143] Furthermore, the process of the initial Q&A model predicting the answer from the vector representation is as follows:

[0144] Input the sample question vector and the sense group vector into the initial Q&A model, and process the sample question vector and the sense group vector through the fusion module in the initial Q&A model to obtain a fusion vector;

[0145] Input the fusion vector into the recognition module in the initial Q&A model for processing to obtain the associated entity central word and the context scenario distribution;

[0146] Process the associated entity central word and the context scenario distribution through the output layer in the initial Q&A model to obtain the predicted answer text.

[0147] Specifically, the fusion module specifically refers to the module that fuses the information of the sample question text and the sense groups of the target text inside the model. Correspondingly, the fusion vector is the vector representation obtained after the sample question vector and the sense group vector are fused. Correspondingly, the recognition module specifically refers to the module that recognizes the associated entity central word and the context scenario distribution in the vector after fusion inside the model. Correspondingly, the associated entity central word and the context scenario distribution specifically refer to the information used to locate the predicted answer text from the semantic dimension and the scenario dimension.

[0148] Based on this, after obtaining the sample question vector and the intent vector, the sample question vector and the intent vector can be input into the initial Q&A model. The fusion module in the initial Q&A model processes the sample question vector and the intent vector to obtain a fused vector. Then, the fused vector is input into the recognition module in the initial Q&A model for processing to obtain the associated entity central word and the context scenario distribution. Finally, the output layer in the initial Q&A model processes the associated entity central word and the context scenario distribution to obtain the predicted answer text.

[0149] Step S1084, optimize the initial Q&A model based on the predicted answer text and the sample answer text in the training sample until the target Q&A model that meets the training stop condition is obtained.

[0150] Specifically, after obtaining the predicted answer text, at this time, the loss function can be determined by combining the sample answer text in the training sample with the predicted answer text. Then, based on the loss function, the initial Q&A model is tuned / optimized, and the above model training process is continuously repeated to obtain the target Q&A model that meets the training stop condition.

[0151] See Figure 2 As shown, when the sample question text is "{What kind of feelings does 'Thoughts in the Silent Night' express by the author}", at this time, the lexical and syntactic unit word vectors and the scene label vectors corresponding to the sample question text can be extracted and fused to obtain the sample question vector. Then, the sample question vector is processed by the semantic dependency analysis tool in the classical poetry Q&A system, and the processed result is subjected to self-attention calculation by the text self-attention calculation unit to obtain the to-be-oriented associated entity phrase group. During this process, the Q&A model will train a task-oriented poetry corpus pointer weight matrix through the BiLSTM weight matrix training module, so that when performing Q&A processing, it can combine the text semantic groups and this matrix to locate the fine-grained poetry entity phrase group, that is, the predicted answer text can be mapped through the fine-grained poetry entity phrase group, and the model is optimized according to the predicted answer text until the classical poetry Q&A system that meets the training stop condition is obtained, so as to accurately give a reply when performing classical poetry Q&A processing.

[0152] The present application provides a method for training a question-and-answer model. After constructing the initial text semantic groups corresponding to the sample corpus, a scene-oriented vocabulary space corresponding to the sample corpus is generated based on the initial text semantic groups. At this time, sufficient corpus is prepared for model training. Then, training samples are obtained, and the sample phrases corresponding to the training samples are determined. The scene-oriented vocabulary space is queried using the sample phrases to determine the target text semantic groups corresponding to the training samples according to the query results. Finally, the initial question-and-answer model is trained based on the target text semantic groups and the training samples until the training stop condition is met, and the target question-and-answer model can be obtained. The association between questions and corpus is captured at the semantic level, effectively ensuring the prediction accuracy of the trained question-and-answer model. Moreover, a rich sample corpus is used to construct the scene-oriented vocabulary space in the preparation stage, effectively improving the processing ability of the question-and-answer model, thereby achieving accurate and efficient completion of the question-and-answer processing task.

[0153] Corresponding to the above method embodiment, the present application also provides an embodiment of a training device for a question-and-answer model. Figure 3 The structure diagram of a training device for a question-and-answer model provided by an embodiment of the present application is shown. As Figure 3 shown, the device includes:

[0154] A construction module 302, configured to construct the initial text semantic groups corresponding to the sample corpus, and generate the scene-oriented vocabulary space corresponding to the sample corpus based on the initial text semantic groups;

[0155] An acquisition module 304, configured to acquire training samples and determine the sample phrases corresponding to the training samples;

[0156] A determination module 306, configured to query the scene-oriented vocabulary space based on the sample phrases and determine the target text semantic groups corresponding to the training samples according to the query results;

[0157] A training module 308, configured to train the initial question-and-answer model using the target text semantic groups and the training samples until a target question-and-answer model that meets the training stop condition is obtained.

[0158] In an optional embodiment, the construction module 302 is further configured to:

[0159] Add context labels to the sample corpus and extract the initial phrases in the initial text semantic groups; establish the correspondence between the context labels and the initial text semantic groups, and construct the scene-oriented vocabulary space corresponding to the sample corpus according to the correspondence and the initial phrases.

[0160] In an optional embodiment, the construction module 302 is further configured to:

[0161] Extract multiple initial features of the sample corpus, and preprocess the multiple initial features to obtain multiple target features; calculate the context similarity between each target feature and the sample corpus, and select at least one target feature as the context label according to the calculation result of the context similarity, and add it to the sample corpus.

[0162] In an optional embodiment, the determining module 306 is further configured to:

[0163] Map the sample phrase to the scene-oriented vocabulary space, and calculate the phrase similarity between the sample phrase and the context label; determine the target context label according to the calculation result of the phrase similarity, and use the initial text semantic group corresponding to the target context label as the target text semantic group.

[0164] In an optional embodiment, the obtaining module 304 is further configured to:

[0165] Parse the training sample to obtain the sample problem text in the training sample; extract the first word unit and the second word unit in the sample problem text, and construct the sample phrase based on the first word unit and the second word unit.

[0166] In an optional embodiment, the training module 308 is further configured to:

[0167] Input the target text semantic group and the sample problem text in the training sample into the initial question-answering model for processing to obtain a predicted answer text; optimize the initial question-answering model based on the predicted answer text and the sample answer text in the training sample until the target question-answering model that meets the training stop condition is obtained.

[0168] In an optional embodiment, the training module 308 is further configured to:

[0169] Generate a word unit vector and a scene label vector based on the sample problem text, and generate a semantic group vector based on the target text semantic group; integrate the word unit vector and the scene label vector to obtain a sample problem vector corresponding to the sample problem text; input the sample problem vector and the semantic group vector into the initial question-answering model for processing to obtain the predicted answer text.

[0170] In an optional embodiment, the training module 308 is further configured to:

[0171] Input the sample question vector and the semantic group vector into the initial Q&A model, process the sample question vector and the semantic group vector through the fusion module in the initial Q&A model to obtain a fusion vector; input the fusion vector into the recognition module in the initial Q&A model for processing to obtain the associated entity central word and the context scene distribution; process the associated entity central word and the context scene distribution through the output layer in the initial Q&A model to obtain the predicted answer text.

[0172] The training device of the Q&A model provided in this application, after constructing the initial text semantic groups corresponding to the sample corpus, will generate a scene-oriented vocabulary space corresponding to the sample corpus based on the initial text semantic groups. At this time, sufficient corpus is prepared for model training; then obtain training samples, and determine the sample phrases corresponding to the training samples, query the scene-oriented vocabulary space using the sample phrases, so as to determine the target text semantic groups corresponding to the training samples according to the query results, and finally train the initial Q&A model according to the target text semantic groups and the training samples until the training stop condition is met, then the target Q&A model can be obtained; it realizes capturing the association between questions and corpus at the semantic level, effectively ensuring the prediction accuracy of the trained Q&A model, and using rich sample corpus to construct the scene-oriented vocabulary space in the preparation stage, effectively improving the processing ability of the Q&A model, so as to achieve accurate and efficient completion of the Q&A processing task.

[0173] The above is a schematic solution of a training device of a Q&A model in this embodiment. It should be noted that the technical solution of the training device of the Q&A model and the technical solution of the above Q&A model training method belong to the same concept. For the details not described in detail in the technical solution of the training device of the Q&A model, reference can be made to the description of the technical solution of the above Q&A model training method.

[0174] In addition, each component in the device embodiment should be understood as a functional module that must be established to implement each step of the program flow or each step of the method. Each functional module is not an actual functional division or separation limitation. The device claim defined by such a set of functional modules should be understood as mainly implementing the functional module architecture of the solution through the computer program recorded in the specification, rather than mainly implementing the physical device of the solution through hardware means.

[0175] This embodiment also provides a text processing method. Figure 4 The flowchart of a text processing method provided in an embodiment of the present application is shown, which specifically includes the following steps:

[0176] Step S402, obtain the question text uploaded by the user.

[0177] Step S404: Input the problem text into the target Q&A model in the training method of the Q&A model to obtain an answer text.

[0178] Step S406: Update the reply interface based on the answer text and display the updated reply interface to the user.

[0179] For example, when the problem text input by the user is {Please provide a frontier poem}, the problem text can be input into the classical poetry Q&A model for processing, and the answer text "On Mission to the Frontier" can be obtained according to the prediction result. At this time, in order to display the classical poem "On Mission to the Frontier" to the user, the reply interface displayed to the user can be updated based on the poem text of "On Mission to the Frontier", and the reply interface containing the poem text of "On Mission to the Frontier" can be displayed to the user according to the update result.

[0180] In summary, by using the target Q&A model obtained by the above training method to process the problem text, the reply accuracy can be effectively improved, and the response speed is faster, thereby improving the user experience.

[0181] The following combines the attached Figure 5 , taking the application of answering ancient poetry Q&A by the method provided in this application as an example, to further illustrate the method. Among them, Figure 5 Fig. shows a processing flow chart provided by an embodiment of the present application applied to the ancient poetry Q&A scenario, which specifically includes the following steps:

[0182] Step S502: Construct an initial text semantic group corresponding to the classical poetry corpus.

[0183] In the process of constructing the initial text semantic group corresponding to the classical poetry corpus, multiple components are divided from the sentences of each classical poetry corpus according to meaning and structure, and the multiple components can form the initial text semantic group corresponding to the corresponding sample corpus.

[0184] Step S504: Generate a scene-oriented entity vocabulary space based on the initial text semantic group.

[0185] After obtaining the initial text semantic group corresponding to the classical poetry corpus, at this time, context labels can be added to each classical poetry corpus, and at the same time, the initial phrases in each initial text semantic group are extracted. That is, the added context labels include {homesickness; parting; frontier; friendship; scenery; sadness... ambition}, and the initial phrases of the initial text semantic group are composed of nominal original words and verb phrase stems.

[0186] Furthermore, after determining the context labels and initial phrases, the corresponding relationship between the context labels and the initial question semantic group can be established, and then the scene-oriented entity vocabulary space corresponding to the classical poetry corpus can be constructed by using the corresponding relationship and the initial phrases.

[0187] Step S506: Obtain training samples and determine the to-be-oriented associated entity phrases corresponding to the sample question texts in the training samples.

[0188] After obtaining the training samples, the sample question texts can be extracted from the training samples, and the nominal original words and verb phrase stems in the sample question texts can be extracted. Then, the two are integrated to generate the to-be-oriented associated entity phrases for use in subsequent training of the classical poetry Q&A model.

[0189] Step S508: Map the to-be-oriented associated entity phrases to the scene-oriented entity vocabulary space and calculate the phrase similarity between the to-be-oriented associated entity phrases and the phrase groups of the context labels included in the space.

[0190] Step S510: Determine the target context label according to the calculation result of the phrase similarity, and use the initial text group corresponding to the target context label as the target text group corresponding to the sample question text.

[0191] After mapping the to-be-oriented associated entity phrases to the scene-oriented entity vocabulary space, the phrase similarity between the to-be-oriented associated entity phrases and the phrase groups of the context labels included in the scene-oriented entity vocabulary space can be calculated. Then, select the context label with the highest phrase similarity as the target context label. Then, determine the initial text group corresponding to this context label from the scene-oriented entity vocabulary space and select this initial text group as the target text group for subsequent model training.

[0192] Step S512: Use the target text group, the sample question text, and the sample answer text included in the training samples to train the classical poetry Q&A model until a target classical poetry Q&A model that meets the training stop condition is obtained.

[0193] At this time, the classical poetry Q&A model can be trained by combining the target text group, the sample question text, and the sample answer text included in the training samples, so that the classical poetry Q&A model can learn the fine-grained semantic association relationship between different question types and the target text group. Then, through continuous iteration and optimization, a target classical poetry Q&A model that meets the training stop condition can be obtained.

[0194] Step S514: Receive the question text to be answered input by the user.

[0195] When the classical poetry Q&A model training is completed, it can be reused. At this time, the question text to be answered input by the user is {Please provide an ancient poem describing homesickness}.

[0196] Step S516: Input the question text to be answered into the target classical poetry Q&A model to obtain the target answer text.

[0197] Step S518: Update the reply interface based on the target answer text and display the updated reply interface to the user.

[0198] Process the question text {Please provide an ancient poem describing homesickness} through the target classical poetry Q&A model, and obtain the answer text as "Thoughts in the Silent Night". At this time, update the reply interface based on the ancient poem "Thoughts in the Silent Night" and its corresponding text, and then the reply interface with the text of "Thoughts in the Silent Night" can be displayed to the user.

[0199] In summary, by training the Q&A model in the above manner, the prediction accuracy of the trained Q&A model is effectively guaranteed. And in the preparation stage, a scene-oriented vocabulary space is constructed using rich sample corpora, which effectively improves the processing ability of the Q&A model, thereby achieving accurate and efficient completion of the Q&A processing task. Corresponding to the above method embodiments, the present application also provides embodiments of a text processing device. FIG. 6 shows a schematic structural diagram of a text processing device provided by an embodiment of the present application. As Figure 6 shown, the device includes:

[0200] A text acquisition module 602, configured to acquire the question text uploaded by the user;

[0201] A text processing module 604, configured to input the question text into the target Q&A model in the training method of the Q&A model for processing to obtain an answer text;

[0202] An interface display module 606, configured to update the reply interface based on the answer text and display the updated reply interface to the user.

[0203] In summary, by using the target Q&A model obtained by the above training method to process the question text, the reply accuracy can be effectively improved, and the response speed is faster, thereby improving the user experience.

[0204] The above is a schematic solution of a text processing device in this embodiment. It should be noted that the technical solution of this text processing device and the technical solution of the above text processing method belong to the same concept. For the details not described in the technical solution of the text processing device, reference can be made to the description of the technical solution of the above text processing method. In addition, each component in the device embodiment should be understood as a functional module that must be established to implement each step of the program flow or each step of the method. The device claims defined by such a set of functional modules should be understood as a functional module framework that mainly realizes the solution through the computer program recorded in the specification, rather than an entity device that mainly realizes the solution through hardware means.

[0205] Figure 7 FIG. 700 shows a structural block diagram of a computing device 700 provided according to an embodiment of the present application. The components of the computing device 700 include but are not limited to a memory 710 and a processor 720. The processor 720 is connected to the memory 710 through a bus 730, and a database 750 is used to store data.

[0206] The computing device 700 further includes an access device 740, which enables the computing device 700 to communicate via one or more networks 760. Examples of these networks include the Public Switched Telephone Network (PSTN), Local Area Network (LAN), Wide Area Network (WAN), Personal Area Network (PAN), or a combination of communication networks such as the Internet. The access device 740 may include one or more of any type of wired or wireless network interfaces (e.g., Network Interface Card (NIC)), such as IEEE802.11 Wireless Local Area Network (WLAN) wireless interface, Worldwide Interoperability for Microwave Access (Wi-MAX) interface, Ethernet interface, Universal Serial Bus (USB) interface, cellular network interface, Bluetooth interface, Near Field Communication (NFC) interface, and so on.

[0207] In an embodiment of the present application, the above components of the computing device 700 and Figure 7 other components not shown in FIG. may also be connected to each other, for example, through a bus. It should be understood that Figure 7 the shown structural block diagram of the computing device is for illustrative purposes only and is not a limitation on the scope of the present application. Those skilled in the art can add or replace other components as needed.

[0208] The computing device 700 can be any type of stationary or mobile computing device, including mobile computers or mobile computing devices (e.g., tablet computers, personal digital assistants, laptop computers, notebook computers, netbooks, etc.), mobile phones (e.g., smartphones), wearable computing devices (e.g., smartwatches, smart glasses, etc.) or other types of mobile devices, or stationary computing devices such as desktop computers or PCs. The computing device 700 can also be a mobile or stationary server.

[0209] Among them, the processor 720 is used to execute computer-executable instructions of the training method of the question-and-answer model or the text processing method.

[0210] The above is a schematic solution of a computing device in this embodiment. It should be noted that the technical solution of this computing device and the technical solutions of the above question-and-answer model training method or text processing method belong to the same concept. For the details not described in the technical solution of the computing device, reference can be made to the descriptions of the technical solutions of the above question-and-answer model training method or text processing method.

[0211] An embodiment of the present application also provides a computer-readable storage medium storing computer instructions, which are executed by a processor for a training method of a question-and-answer model or a text processing method.

[0212] The above is a schematic solution of a computer-readable storage medium in this embodiment. It should be noted that the technical solution of this storage medium and the technical solutions of the above-mentioned training method of the question-and-answer model or the text processing method belong to the same concept. For the details not described in detail in the technical solution of the storage medium, reference can be made to the descriptions of the technical solutions of the above-mentioned training method of the question-and-answer model or the text processing method.

[0213] The above describes specific embodiments of the present application. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in a different order than in the embodiments and still achieve the desired results. Additionally, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0214] The computer instructions include computer program code, which can be in the form of source code, object code, executable files, or some intermediate forms, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.

[0215] It should be noted that for the foregoing method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present application is not limited by the described action sequence, because according to the present application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present application.

[0216] In the above embodiments, the descriptions of the various embodiments have their own focuses. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0217] The preferred embodiments of the present application disclosed above are only used to help illustrate the present application. The alternative embodiments do not describe all the details in detail, nor do they limit the invention to the specific embodiments described. Obviously, many modifications and variations can be made according to the content of the present application. These embodiments are selected and specifically described in the present application to better explain the principle and practical application of the present application, so that those skilled in the art can well understand and utilize the present application. The present application is only limited by the claims and their full scope and equivalents.

Claims

1. A sample processing method, characterized in that, comprising: obtaining a sample corpus and constructing an initial text sense group corresponding to the sample corpus; adding a context label to the sample corpus and extracting an initial phrase corresponding to the initial text sense group; establishing a correspondence between the context label and the initial text sense group; constructing a scene-oriented vocabulary space corresponding to the sample corpus according to the correspondence and the initial phrase, wherein the scene-oriented vocabulary space contains scene information and context information corresponding to the sample corpus.

2. The sample processing method according to claim 1, characterized in that, after the step of constructing the scene-oriented vocabulary space corresponding to the sample corpus according to the correspondence and the initial phrase, further comprising: obtaining a training sample and determining a sample phrase corresponding to the training sample; querying the scene-oriented vocabulary space based on the sample phrase, and determining a target text sense group corresponding to the training sample according to the query result; training an initial question-answering model by using the target text sense group and the training sample until a target question-answering model that meets the training stop condition is obtained.

3. The sample processing method according to claim 1, characterized in that, the adding a context label to the sample corpus includes: extracting a plurality of initial features of the sample corpus and preprocessing the plurality of initial features to obtain a plurality of target features; calculating the context similarity between each target feature and the sample corpus, and selecting at least one target feature as the context label according to the context similarity calculation result and adding it to the sample corpus.

4. The sample processing method according to claim 2, characterized in that, the querying the scene-oriented vocabulary space based on the sample phrase and determining a target text sense group corresponding to the training sample according to the query result includes: mapping the sample phrase to the scene-oriented vocabulary space and calculating the phrase similarity between the sample phrase and the context label; determining a target context label according to the phrase similarity calculation result, and using the initial text sense group corresponding to the target context label as the target text sense group.

5. The sample processing method according to claim 2, characterized in that, the obtaining a training sample includes: obtaining the training sample having an association relationship with the sample corpus; wherein, the training the initial question-answering model by using the target text sense group and the training sample until a target question-answering model that meets the training stop condition is obtained includes: training the initial question-answering model by using the training sample having an association relationship with the sample corpus and the target text sense group until a target question-answering model that meets the training stop condition is obtained.

6. The sample processing method according to claim 2, characterized in that, the determining a sample phrase corresponding to the training sample includes: parsing the training sample to obtain a sample question text in the training sample; extracting a first word unit and a second word unit from the sample question text, and constructing the sample phrase based on the first word unit and the second word unit.

7. The sample processing method according to claim 6, wherein, the training of the initial Q&A model using the target text sense group and the training samples until a target Q&A model that meets the training stop condition is obtained includes: inputting the target text sense group and the sample question text in the training samples into the initial Q&A model for processing to obtain a predicted answer text; optimizing the initial Q&A model based on the predicted answer text and the sample answer text in the training samples until the target Q&A model that meets the training stop condition is obtained.

8. The sample processing method according to claim 7, wherein, the inputting the target text sense group and the sample question text in the training samples into the initial Q&A model for processing to obtain a predicted answer text includes: generating a word unit vector and a scene label vector based on the sample question text, and generating a sense group vector based on the target text sense group; integrating the word unit vector and the scene label vector to obtain a sample question vector corresponding to the sample question text; inputting the sample question vector and the sense group vector into the initial Q&A model for processing to obtain the predicted answer text.

9. The sample processing method according to claim 8, wherein, the inputting the sample question vector and the sense group vector into the initial Q&A model for processing to obtain the predicted answer text includes: inputting the sample question vector and the sense group vector into the initial Q&A model, and processing the sample question vector and the sense group vector through a fusion module in the initial Q&A model to obtain a fusion vector; inputting the fusion vector into an identification module in the initial Q&A model for processing to obtain an associated entity central word and a context scene distribution; processing the associated entity central word and the context scene distribution through an output layer in the initial Q&A model to obtain the predicted answer text.

10. The sample processing method according to claim 3, wherein, the preprocessing of the multiple initial features to obtain multiple target features includes: cleaning the multiple initial features and determining the multiple target features according to the cleaning result; wherein, the selecting at least one target feature as the context label according to the context similarity calculation result includes: comparing the context similarity with a preset context similarity threshold, and selecting a target feature greater than or equal to the context similarity threshold as the context label; or selecting the target feature with the largest similarity as the context label according to the context similarity calculation result.

11. A sample processing device, wherein, it includes: an acquisition module configured to acquire a sample corpus and construct an initial text sense group corresponding to the sample corpus; an addition module configured to add a context label to the sample corpus and extract an initial phrase corresponding to the initial text sense group; a establishment module configured to establish a correspondence between the context label and the initial text sense group; A construction module, configured to construct a scenario-oriented vocabulary space corresponding to the sample corpus according to the corresponding relationship and the initial phrase, wherein the scenario-oriented vocabulary space includes scenario information and context information corresponding to the sample corpus.

12. A method for training a question-and-answer model, characterized in that, it includes: Obtaining a training sample and determining a sample phrase corresponding to the training sample; Querying the scenario-oriented vocabulary space in the method according to any one of claims 1-10 based on the sample phrase, and determining the target text semantic group corresponding to the training sample according to the query result; Using the target text semantic group and the training sample to train an initial question-and-answer model until a target question-and-answer model that meets the training stop condition is obtained.

13. A training device for a question-and-answer model, characterized in that, it includes: A sample acquisition module, configured to obtain a training sample and determine a sample phrase corresponding to the training sample; A semantic group determination module, configured to query the scenario-oriented vocabulary space in the method according to any one of claims 1-10 based on the sample phrase, and determine the target text semantic group corresponding to the training sample according to the query result; A model training module, configured to use the target text semantic group and the training sample to train an initial question-and-answer model until a target question-and-answer model that meets the training stop condition is obtained.

14. A computing device, characterized in that, it includes: A memory and a processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the steps of the method according to any one of claims 1 to 10 or 12.

15. A computer-readable storage medium storing computer instructions, characterized in that, when the instructions are executed by a processor, the steps of the method according to any one of claims 1 to 10 or 12 are implemented.

Citation Information

Patent Citations

  • Training methods and devices for question-answering models

    CN113127624B