Long text training data generation method, related device and computer program product

By obtaining long text source data and using large models to generate and verify answers, the problem of long text training data relying on manual labeling is solved, efficient and reliable training data generation is achieved, and the performance of large models in long text processing tasks is improved.

CN120804714APending Publication Date: 2025-10-17IFLYTEK CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510987255.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-17
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

In existing technologies, long text training data relies on manual labeling, resulting in high labor costs, low labeling efficiency, difficulty in expanding the scale of training data, and insufficient generalization and context understanding capabilities of large models in long text processing tasks.

Method used

By obtaining long text source data, using a large model to generate relevant questions and answers, and performing self-consistency verification based on the text similarity between the answers, the most credible answer is determined and long text training data is generated.

Benefits of technology

It significantly reduces the cost of data collection and annotation, improves the efficiency and quality of training data configuration for long text processing tasks, provides a reliable and efficient training data solution for large models in long text processing tasks, and improves model performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120804714A_ABST
    Figure CN120804714A_ABST
Patent Text Reader

Abstract

The invention discloses a long text training data generation method, a related device and a computer program product, and relates to the field of artificial intelligence, and the method comprises the steps: firstly obtaining long text source data, and then generating related questions and corresponding answers of the long text source data through the generation capability of a large language model, the method comprises the following steps: generating long text source data, performing answer self-consistency verification on the basis of the similarity between the generated answers, determining the answer with the highest credibility as the final answer, and generating long text training data by utilizing the long text source data, related questions and the corresponding final answer, thereby realizing a long text training data generation task. The training data configuration efficiency and quality suitable for the long text processing task are improved, and a basis is provided for optimizing the model performance of a large model on the long text processing task.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of artificial intelligence, and more particularly, to a long text training data generation method, related device and computer program product. BACKGROUND

[0002] With the rapid development of deep learning, artificial intelligence and other technologies, large models have made significant progress in natural language processing and are widely used in machine translation, automatic abstract generation, text classification and semantic search tasks.

[0003] However, due to the size and quality of the training data, the generalization ability and context understanding ability of large models in long text processing tasks are still insufficient. To optimize the performance of the model, a large amount of long text training data needs to be configured.

[0004] Current long text training data usually depends on manual annotation, which has the problems of high labor cost, low annotation efficiency and difficulty in expanding the scale of training data. SUMMARY

[0005] In view of the above problems, the present application is proposed to provide a long text training data generation method, related device and computer program product to realize the automatic generation task of long text training data. The specific scheme is as follows:

[0006] In a first aspect, a long text training data generation method is provided, comprising:

[0007] obtaining long text source data, the text length of the long text source data being greater than a preset length value;

[0008] calling a large model to instruct the large model to generate a related question of the long text source data;

[0009] calling the large model to instruct the large model to generate an answer corresponding to the related question, the number of generated answers being not less than 3;

[0010] For each generated answer, the confidence of the current answer is calculated based on the text similarity between the current answer and each remaining answer. The answer with the highest confidence is determined as the final answer corresponding to the related question.

[0011] Using the long text source data, the related question and the corresponding answer, long text training data is generated.

[0012] In a second aspect, a long text training data generation device is provided, comprising:

[0013] a source data acquisition unit configured to obtain long text source data, the text length of the long text source data being greater than a preset length value;

[0014] a question generation unit configured to invoke the large model to instruct the large model to generate a relevant question of the long text source data;

[0015] an answer generation unit configured to invoke the large model to instruct the large model to generate an answer corresponding to the relevant question, and the number of generated answers is not less than 3;

[0016] an answer screening unit configured to, for each of the generated answers, calculate a credibility of a current answer based on a text similarity between the current answer and each of the remaining answers, and determine an answer with the highest credibility as a final answer corresponding to the relevant question;

[0017] a training data generation unit configured to generate long text training data by using the long text source data, the relevant question, and the final answer.

[0018] In a third aspect, an electronic device is provided, including a memory and a processor.

[0019] The memory is configured to store a program.

[0020] The processor is configured to execute the program to implement each step of the long text training data generation method described in the first aspect.

[0021] In a fourth aspect, a readable storage medium is provided, which stores a computer program. When the computer program is executed by a processor, each step of the long text training data generation method described in the first aspect is implemented.

[0022] In a fifth aspect, a computer program product is provided, which includes a computer program. When the computer program is executed by a processor, each step of the long text training data generation method described in the first aspect is implemented.

[0023] By the above technical solution, the long text source data is first obtained, and then the large language model is used to generate the relevant question and the corresponding answer of the long text source data. Based on the similarity between each of the generated answers, the answer self-consistency is checked, and the answer with the highest credibility is determined as the final answer. Then, the long text training data is generated by using the long text source data, the relevant question, and the final answer. The long text training data generation task is implemented, the efficiency and quality of the training data configuration suitable for the long text processing task are improved, and the basis for optimizing the model performance of the large model in the long text processing task is provided. BRIEF DESCRIPTION OF DRAWINGS

[0024] Various other advantages and benefits will become apparent to those of ordinary skill in the art, upon reading the following detailed description of the preferred embodiments. The accompanying drawings are included to provide a description of preferred embodiments, and are not meant to limit the present application. Furthermore, the same reference numerals in different drawings are intended to represent the same components throughout the several drawings. In the drawings:

[0025] Figure 1 An embodiment system architecture schematic diagram of the long text training data generation method provided by the embodiments of the present application is shown in the following figure:

[0026] Figure 2 A flowchart of the long text training data generation method provided by the embodiments of the present application is shown in the following figure:

[0027] Figure 3 A structure schematic diagram of the long text training data generation device provided by the embodiments of the present application is shown in the following figure:

[0028] Figure 4 A structure schematic diagram of the electronic device provided by the embodiments of the present application is shown in the following figure. DETAILED DESCRIPTION

[0029] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor fall within the scope of protection of the present application.

[0030] The training data generated by the long text training data generation scheme provided by the present application can be used as the training data of a generative artificial intelligence model (such as a large model), especially the training data of an artificial intelligence model for implementing a long text processing task.

[0031] The present application provides a long text training data generation method, which can be applied to a system architecture as shown in the following figure. Figure 1 The system can include a terminal 100 and a server 200. The server 200 can include one or more servers (as an example, one server is described in the following figure). Figure 1

[0032] The terminal 100 or the server 200 can be used alone to execute the long text training data generation method provided by the embodiments of the present application. In addition, the terminal 100 and the server 200 can also be used cooperatively to execute the long text training data generation method provided by the embodiments of the present application. The terminal 100 in the embodiments of the present application can be a mobile phone, a computer, etc., and the embodiments of the present application do not make any limitation on this.

[0033] ​The embodiment of the application provides a long text training data generation method. The long text training data generation method is applied to a computer device, and the computer device can be a terminal 100 in Figure 1 or a system composed of the terminal 100 and a server 200. Referring to Figure 2 , the long text training data generation method specifically includes the following steps:

[0034] Step S101, acquiring long text source data.

[0035] The text length of the long text source data is greater than a preset length value. Optionally, the preset length value can be 64000, that is, the long text source data includes at least 64k tokens (a basic unit in the field of natural language processing, which can be referred to as a word element). It should be noted that the long text source data in the long text training data can be equivalent to the input text in the actual application scenario of the large model except the question raised by the user, and is the basis for the large model to accurately answer the question.

[0036] Step S102, calling a large model to instruct the large model to generate a related question of the long text source data.

[0037] When the large model is called to generate the related question, the prompt word prompt can be configured based on the text data. On this basis, the prompt word prompt can also be configured in combination with at least one of the following prompt methods: role setting, task description, few-shot prompt, and thinking chain. For example, the prompt word used when the large model is called to generate the related question can include: "You are a long text content understanding expert, and now you are given a text content. Please extract the key content and ask questions. The given text is as follows: {context1}; please think step by step, understand the text content, and ask meaningful questions. For example: According to the article content, what trend will the economic form present this year?". Wherein {context1} represents a placeholder for embedding text data, and in some possible implementations, context1 can refer to the aforementioned long text source data.

[0038] The application obtains the question part in the long text training data by asking questions based on the long text source data through the generation capability of the large model.

[0039] Step S103, calling a large model to instruct the large model to generate an answer corresponding to the related question.

[0040] After the relevant question is generated, the text content and the question can be embedded into an answer generation prompt for the large language model to generate an answer. For example, the prompt used when calling the large model to generate an answer can include: "You are an expert in long text content understanding and a question and answer master. Now give you a piece of text content and a question. Please read and understand the text, and answer the question according to the text information. The following is the given text: {context2}; the following is the given question: {question}; please think step by step, understand the intention of the question, and output the answer to the question in clear logic and layout." Where {context2} represents a placeholder for embedding text content, and {question} represents a placeholder for embedding a question (i.e., the relevant question described above). In some possible implementations, context2 can refer to the long text source data described above.

[0041] The number of generated answers is not less than 3. It should be noted that the subsequent application determines the credibility of the current answer based on the text similarity of the answer pair containing the current answer and one of the remaining answers. Therefore, in order to distinguish the credibility of different answers, at least 3 answers are required. In addition, the above step S103 can be implemented in a way that more than one answer is generated each time the large model is called, or in a way that only one answer is generated each time the large model is called, and the present application does not limit it.

[0042] Step S104, for each generated answer, based on the text similarity between the current answer and each of the remaining answers, calculate the credibility of the current answer; determine the answer with the highest credibility as the final answer corresponding to the relevant question.

[0043] It should be noted that for calculation and other objective reasoning tasks, the answer corresponding to the question is relatively objective and unique, so the consistency of multiple answers can be determined by comparing the conclusion part of the extracted answer. Even if multiple answers are inconsistent, the final answer can also be determined by counting votes. However, for long text processing tasks such as question and answer and text generation, it is difficult to extract a part of the information from the complete answer as the final conclusion, and different answers may contain partially consistent content, that is, the answer consistency checking task is difficult to implement by counting votes.

[0044] To obtain a high-reliability answer suitable for a long text processing task, the present application first performs multiple thoughts based on the text content and related questions by the large model capability, obtains multiple answers to the same related question, and then performs self-consistency checking and screening among the multiple answers based on the principle of minority obeying majority to determine the final answer to the question. Specifically, the aforementioned principle means that if all the generated answers are consistent, the answer generated by the model is reliable; if there is a difference between the generated answers, the higher the similarity of an answer to the remaining answers and the more the remaining answers agree with the answer, the higher the reliability of the answer. This principle solves the problem that the traditional statistical voting method cannot achieve the answer consistency checking task of long text data by using fuzzy logic.

[0045] Step S105, generating long text training data by using long text source data, related questions and corresponding final answers.

[0046] By the above scheme, first, the long text source data is obtained, then the related questions and corresponding answers of the long text source data are generated by using the generation capability of the language model, and the self-consistency checking of the answers is performed based on the similarity between the contents of the generated answers to determine the answer with the highest credibility as the final answer. Then, the long text training data is generated by using the long text source data, related questions and corresponding final answers, the long text training data self-determination generation task is realized, the data collection and labeling cost is significantly reduced, the training data configuration efficiency and quality suitable for long text processing tasks are improved, a reliable and efficient training data solution for large model application is provided, which helps to improve the model training effect, thereby promoting the application of large-scale language model in natural language processing tasks.

[0047] In one or more embodiments provided by the present application, the step S101 of obtaining long text source data can include:

[0048] Step S201, obtaining original text data.

[0049] The original text data can be text data obtained directly or indirectly through a network or other means, and after subsequent processing and screening, the long text source data in the long text training data is finally obtained.

[0050] In a possible implementation, the step S201 of obtaining original text data can include:

[0051] Step S301, obtaining text data or non-text data.

[0052] The aforementioned text data and non-text data can be obtained by collecting pure text, book pictures, long video audio and other modal data, and on this basis, the obtained data is preprocessed according to the corresponding type of data source, such as cross-modal conversion, to obtain original text data.

[0053] In step S302, the obtained data is preprocessed to obtain original text data.

[0054] Next, the preprocessing methods for different modal data are described respectively.

[0055] When the obtained data is text data, the preprocessing includes invalid character and non-standard character removal. It should be noted that current network applications, etc. usually adopt different encoding methods when encoding text content to meet different special control needs. For example, for spaces, ASCII code uses '\u00A0' to represent non-breaking space, '\u200B' to represent zero-width space, '\u202F' to represent narrow non-breaking space, '\u2060' to represent word space, and '\u3000' to represent full-width space. That is, the text data, i.e. the aforementioned pure text data, often contains invalid or non-standard content. The same text content represented by different encodings causes obstacles to the semantic understanding of large language models. In order to ensure the quality of training data, the embodiment pre-processes the obtained text data to replace or remove invalid and non-standard characters in the pure text data.

[0056] When the obtained non-text data is speech data, the preprocessing includes speech recognition processing and spoken expression rewriting processing. It should be noted that the text content converted from long audio data often contains spoken content, and long text training data that is not written in a formal way is difficult to optimize the general performance of the model. Based on this, the embodiment removes and rewrites the spoken content of the long audio data. Optionally, the aforementioned spoken expression rewriting processing can be realized by calling a large model to instruct the large model to rewrite the spoken expression in the speech recognition content.

[0057] When the obtained non-text data is text image, the preprocessing includes text recognition processing.

[0058] Based on this, the embodiment pre-processes the obtained text and non-text data of different types to obtain original text data of different data sources, enriches the source of long text training data, and to some extent solves the problem of scarcity of high-quality long text data resources, providing a basis for generating a large amount of high-quality long text training data.

[0059] It should be noted that the long text has a high possibility of containing repeated content. If the long text data containing a large amount of repeated content is used as training data, the large language model will learn the repeated language patterns, which will adversely affect the performance of the model. Therefore, it is necessary to remove the repetitive paragraphs in the original text data.

[0060] Step S202, segmenting the original text data, calculating the difference degree between the text segments, and removing the repetitive text data according to the difference degree.

[0061] In one possible implementation, the Hamming distance can be used to represent the difference degree between the text segments. Specifically, the foregoing step S202 can include:

[0062] First, use the SimHash algorithm to encode each text segment obtained by segmentation to obtain the SimHash encoding of each text segment;

[0063] Second, calculate the Hamming distance between the SimHash encodings of each two text segments to represent the difference degree between the text segments;

[0064] Third, determine the text segment at the later position as the repetitive segment when the Hamming distance between the two text segments is zero, and remove the repetitive segment in the original text data.

[0065] For example, taking text segment a and text segment b as an example, the SimHash encoding of text segment a is represented as A, and the SimHash encoding of text segment b is represented as B. The Hamming distance between the two text segments can be represented as: In the formula, count is a counting operation, I represents the i-th bit of the SimHash encoding, and I represents the length of the encoding. That is, the number of inconsistent bits in the SimHash encodings of the two text segments is counted, and the counting result is divided by the total length of the encoding, which is the Hamming distance between the two text segments. The greater the Hamming distance, the greater the inconsistency between the two SimHash encodings, and the greater the difference between the two text segments. Conversely, the smaller the Hamming distance, the smaller the difference between the two text segments. Specifically, if the Hamming distance between the SimHash encodings of the two text segments is zero, it means that the SimHash encodings of the two text segments are consistent, and the two text segments are repetitive, which needs to be removed to ensure data quality.

[0066] Step S203, when the text length of the de-duplicated text data is greater than the preset length value, determining the de-duplicated text data as long text source data.

[0067] Based on the above, in the embodiment, when obtaining long text source data, the original text data is de-duplicated to remove repetitive fragments in a single text. On this basis, the text length of the de-duplicated text data is screened, and finally high-quality long text source data is obtained, which provides a basis for ensuring the quality of training data generation.

[0068] Optionally, to optimize the performance of the model for processing long text data in a specified language, the text character length of the de-duplicated text data is judged, and the de-duplicated text data whose text length is greater than the preset length value and whose language is the specified language is determined as long text source data to filter non-long text data and non-specified language data. For example, a pre-trained fast text classification model (such as a FastText model) can be called to detect the language of the de-duplicated text data.

[0069] In one or more embodiments provided in the present application, the step S102 described above, calling a large model to instruct the large model to generate long text source data, can include:

[0070] Step S401, determining the keywords of the long text source data.

[0071] The keywords described in the present application can refer to the core discussion content of the long text source data, which has practical significance and usually has a high frequency of occurrence.

[0072] Step S402, slicing the long text source data, and retaining the text slice containing the keywords as short text data.

[0073] It should be noted that the slicing in step S402 is different from the segmentation described above, and the text slice obtained by slicing has a smaller text granularity. For example, the slicing can be performed according to special characters such as line breaks, periods, question marks, etc.

[0074] Step S403, calling a large model to instruct the large model to generate a question based on the short text data, and taking the generated question as a related question of the long text source data.

[0075] That is, in some possible implementations, context1 in the prompt example described above can refer to long text source data that has been processed based on keyword-based text reduction, i.e., short text data described above.

[0076] Compared with the scheme of directly sending the long text source data into the large model to generate related problems, the embodiment reduces the input text length when the large model implements the problem generation task, and to some extent solves the problem of low quality of questions caused by the limited length of text that the large language model can accurately understand and jointly remember. In addition, by retaining the text slice containing the keyword and eliminating the text slice not containing the keyword, the embodiment reduces the length of the text while ensuring the topic coherence and information integrity of the text to some extent, which helps the large model to better handle the problem generation task and thus improves the generation quality of related problems.

[0077] In a possible implementation, the step S401 of determining the keyword of the long text source data can include:

[0078] The step S501 is to perform Term Frequency-Inverse Document Frequency (TF-IDF) statistics on each word in the long text source data, and extract a theme word from the long text source data according to the statistical result.

[0079] The theme word extraction method based on TF-IDF refers to determining the theme word by comprehensively considering the occurrence of the word in the long text source data and the occurrence of the word in other documents. For example, the TF-IDF value of a certain word T can be expressed as: In the formula, D represents the long text source data, freq(T, D) represents the number of occurrences of the word T in the long text source data D, size(D) represents the total number of words in the long text source data D, df(T) represents the total number of documents containing the word T in the corpus, and N represents the total number of texts in the corpus. The corpus refers to the corpus matching the language of the long text source data, and the specific type of the corpus is not limited in the present application.

[0080] Based on the above, the greater the number of occurrences of a certain word in the long text source data, the greater the corresponding TF-IDF value; the more documents containing a certain word in the corpus, the smaller the corresponding TF-IDF value; that is, the word with a larger TF-IDF value has a higher frequency of occurrence in the current document (i.e., the long text source data) and a higher uniqueness in the entire corpus. Based on this, the present application extracts the word with high frequency and uniqueness in the long text source data as a theme word by comparing the current document and the corpus document, which eliminates the possibility of using common words such as "is" and "of" as keywords, and realizes the representative keyword extraction task that conforms to the theme of the current document.

[0081] The step S502 is to call the large model to instruct the large model to perform word association based on the theme word, and obtain a related theme word.

[0082] Exemplarily, the associated topic words derived by association can include 3-5 words.

[0083] In step S503, the topic words and the associated topic words are determined as the keywords of the long text source data.

[0084] The above scheme jointly forms a keyword set from the topic words extracted from the long text source data and the associated topic words derived by association, enriches the basis for subsequent processing, reduces the possibility of losing key information when subsequently performing text reduction, and provides a basis for improving the quality of question generation.

[0085] In one or more embodiments provided in the present application, the step S103 of calling the large model to instruct the large model to generate answers corresponding to the associated questions can include:

[0086] The large model is called at least three times to instruct the large model to answer the associated questions and generate corresponding answers.

[0087] In some possible implementations, the aforementioned calling of the large model at least three times to instruct the large model to answer the associated questions can include: calling the large model at least three times to instruct the large model to generate answers according to the short text data and the associated questions. That is, the context2 in the aforementioned answer generation prompt words can refer to the short text data.

[0088] The above scheme improves the accuracy of the large language model in understanding the text content of the long text source data by using the short text data, and provides a basis for improving the quality of answer generation.

[0089] When calling the large model to generate questions and answers, the present application fully considers the capability boundary of the large model, reduces the length of the text that needs to be understood and processed by the large model when generating associated questions and corresponding answers through the text reduction scheme based on keywords, solves the problems of forgetting the text content and incomplete answers when the large model processes long text input data, and ensures the information integrity of the reduced long text source data (i.e., short text data) through multiple keyword acquisition methods.

[0090] In the process of calling the large model to generate answers multiple times, the sampling temperature of the called large model is set to a preset high-temperature sampling parameter value.

[0091] Exemplarily, the sampling probability based on the temperature parameter can be represented as: ; in the formula, P( ) represents the sampling probability, y t represents the output of the model at the t-th step of the decoding, i represents the i-th word in the inference vocabulary, z t,idenotes the original sampling probability of the i-th word in the vocabulary generated by the model at the t-th step, T denotes a temperature parameter, and j denotes the total number of words in the vocabulary. As shown in the above formula, when T = 1, the original sampling probability is unchanged; when T < 1, the temperature sampling can amplify the difference between high-probability words and low-probability words, so that the probability distribution is more sharp, and accordingly a generation result with higher certainty can be obtained; when T > 1, the temperature sampling will reduce the probability difference between words, so that the probability distribution is more smooth, and accordingly the diversity of the generation result can be increased. Based on this, the high-temperature sampling parameter value mentioned above can be greater than 1, so as to improve the output diversity of the large model in the multiple answer generation process.

[0092] By using a sampling temperature with a large value, the multiple answer generation process mentioned above can be equivalent to that the large model repeatedly thinks in different ways multiple times to obtain multiple versions of answers; on the basis of the diversified answer generation results, the answers are compared, and the self-consistency between the answers is verified and the answers are selected, if the consistency of an answer with each of the remaining answers is relatively high, the answer is determined as the final answer, so that the answer generation task with reliable quality is realized, and the content reliability of the answer label in the training data is improved.

[0093] In one or more embodiments provided in the present application, the step S104 of calculating the reliability of the current answer based on the text similarity between the current answer and each of the remaining answers can include:

[0094] In step S601, the encoding vectors of the generated answers are obtained.

[0095] For example, the encoding vector mentioned above can be obtained by calling the pre-trained language model BERT (Bidirectional Encoder Representations from Transformers, BERT) to process the answer text of the generated answer.

[0096] In step S602, the cosine similarity between the current answer and each of the remaining answers is calculated according to the respective encoding vectors.

[0097] The cosine similarity can refer to the cosine value of the included angle θ between the encoding vectors of two answer texts. For example, the cosine similarity of two answers can be expressed as ; in the formula, i denotes the i-th bit in the encoding vector, and n denotes the total length of the encoding vector. As shown in the above formula, the smaller the included angle between the encoding vectors of two answers, the greater the corresponding cosine similarity, and the more similar the encoding vectors of two answers. Similarly, the cosine distance of two answers can be expressed as 1-cosθ, and the greater the cosine distance of two answers, the less similar the encoding vectors of two answers.

[0098] Based on the foregoing, the cosine similarity can be used to represent the similarity between the encoding vectors of the two answer texts, which belongs to a kind of vector similarity. In other possible implementations, the similarity between the encoding vectors of the two answer texts can also be represented based on the distance between the two encoding vectors and other vector similarities.

[0099] In step S603, the text similarity between the current answer and each of the other answers is determined based on the respective cosine similarities.

[0100] Specifically, the cosine similarity between the two answers can be used as the text similarity, or the cosine similarity between the two answers can be calculated by numerical mapping to obtain the text similarity. For example, the calculated cosine similarity can be mapped to a non-negative interval by numerical mapping to obtain the text similarity between the current answer and each of the other answers. At this time, any text similarity is positively correlated with the corresponding cosine similarity, and any text similarity is non-negative.

[0101] In one possible implementation, determining the text similarity between the current answer and each of the other answers based on the respective cosine similarities can include:

[0102] Numerically mapping each of the cosine similarities to obtain the text similarity between the current answer and each of the other answers, wherein any text similarity is positively correlated with the corresponding cosine similarity.

[0103] That is, as the cosine similarity increases, the corresponding text similarity increases at an accelerating rate. It should be noted that what is important for the final answer is whether the current answer is sufficiently similar to the other answers, rather than whether the current answer is sufficiently dissimilar to the other answers. Based on this, the numerical mapping calculation described above can increase the contribution of items with higher cosine similarity in determining the final answer to optimize the self-consistency verification quality of the answer.

[0104] In one possible implementation, the cosine similarity between the encoding vectors x1 and x2 of the two answers can be calculated by numerical mapping as follows: where e is a natural constant, and the other parameters can be referred to the foregoing. Similarly, the cosine similarity between the encoding vectors x1 and x2 of the two answers can also be calculated by numerical mapping as follows: where the exponential part is the negative of the cosine distance. The greater the cosine distance, the lower the corresponding text similarity, and the smaller the cosine distance, the higher the corresponding text similarity.

[0105] In other possible implementations, a corresponding correction coefficient can also be configured for the cosine similarity between the answer encodings. The greater the cosine similarity, the greater the correction coefficient, to achieve the numerical mapping purpose described above.

[0106] Step S604, calculate the statistical value of the text similarity between the current answer and each remaining answer as the credibility of the current answer.

[0107] For example, the statistical value described above can be an average value, a median value, or other statistical values, and the present application is not limited thereto.

[0108] The above scheme compares each generated answer content, calculates the text similarity between each current answer and each remaining answer, determines the recognition degree of the remaining answers to the current answer, and then determines the credibility of the current answer by synthesizing the text similarity of the current answer and each remaining answer. The final answer most recognized by each remaining answer is determined, which has high interpretability and reliability. In addition, in the face of special cases, such as an answer containing the content of each remaining answer and the content of each remaining answer being mutually exclusive, the answer can still be determined as the final answer based on the comprehensive text similarity between the answers, so as to determine the final answer with high content integrity.

[0109] In summary, on the basis of the long text source data acquisition step, the question generation step and the answer generation step, the complete long text source data is combined with the related questions and the corresponding answers determined after screening to complete the configuration task of a set of long text training data.

[0110] The long text training data generation device provided by the embodiment of the present application is described below. The long text training data generation device described below can be referred to in conjunction with the long text training data generation method described above.

[0111] Referring to Figure 3 , Figure 3 A long text training data generation device structure diagram disclosed by the embodiment of the present application.

[0112] As Figure 3 shown, the device can include:

[0113] The source data acquisition unit 11 is configured to acquire long text source data, wherein the text length of the long text source data is greater than a preset length value;

[0114] The question generation unit 12 is configured to call a large model to instruct the large model to generate a question related to the long text source data;

[0115] The answer generation unit 13 is configured to call a large model to instruct the large model to generate an answer corresponding to the related question, and the number of generated answers is not less than 3;

[0116] The answer screening unit 14 is configured to calculate a credibility of each generated answer based on a text similarity between the current answer and each of the remaining answers, and determine the answer with the highest credibility as the final answer corresponding to the related question.

[0117] The training data generation unit 15 is configured to generate long text training data by using the long text source data, the related question and the final answer.

[0118] In one or more embodiments provided in the present application, the process of calculating the credibility of the current answer based on the text similarity between the current answer and each of the remaining answers by the answer screening unit 14 can include:

[0119] obtaining an encoding vector of each generated answer;

[0120] calculating a cosine similarity between the current answer and each of the remaining answers according to each of the encoding vectors, and determining a text similarity between the current answer and each of the remaining answers based on each of the cosine similarities;

[0121] calculating a statistical value of the text similarity between the current answer and each of the remaining answers as the credibility of the current answer.

[0122] In one or more embodiments provided in the present application, the process of determining the text similarity between the current answer and each of the remaining answers based on each of the cosine similarities by the answer screening unit 14 can include:

[0123] performing numerical mapping calculation on each of the cosine similarities to obtain the text similarity between the current answer and each of the remaining answers, wherein any text similarity is positively correlated with the corresponding cosine similarity.

[0124] In one or more embodiments provided in the present application, the process of calling a large model by the question generation unit 12 to instruct the large model to generate the related question of the long text source data can include:

[0125] determining a keyword of the long text source data;

[0126] slicing the long text source data, and retaining a text slice containing the keyword as short text data;

[0127] calling the large model to instruct the large model to generate a question according to the short text data, and taking the generated question as the related question of the long text source data.

[0128] In one or more embodiments provided in the present application, the process of determining the keyword of the long text source data by the question generation unit 12 can include:

[0129] counting word frequency-inverse text frequency indexes of each word item in the long text source data, and extracting a theme word from the long text source data according to a counting result;

[0130] calling a large model to instruct the large model to perform word association according to the theme word, to obtain a related theme word;

[0131] determining the theme word and the related theme word as keywords of the long text source data.

[0132] In one or more embodiments provided in the present application, the process of calling a large model to instruct the large model to generate an answer corresponding to the related question by the answer generation unit 13 can include:

[0133] calling the large model to instruct the large model to answer the related question at least three times to generate a corresponding answer;

[0134] wherein a sampling temperature of the called large model is set to a preset high-temperature sampling parameter value.

[0135] In one or more embodiments provided in the present application, the process of obtaining long text source data by the source data obtaining unit 11 can include:

[0136] obtaining original text data;

[0137] segmenting the original text data and calculating a difference degree between text segments;

[0138] performing a deduplication process on the original text data according to the difference degree;

[0139] when a text length of the deduplicated text data is greater than the preset length value, determining the deduplicated text data as long text source data.

[0140] In one or more embodiments provided in the present application, the process of obtaining original text data by the source data obtaining unit 11 can include:

[0141] obtaining text data or non-text data;

[0142] performing preprocessing on the obtained data to obtain original text data; wherein when the obtained data is text data, the preprocessing includes invalid character and non-standard character removal; when the obtained non-text data is voice data, the preprocessing includes voice recognition processing and spoken expression rewriting processing; when the obtained non-text data is a text image, the preprocessing includes text recognition processing.

[0143] Each unit in the long text training data generation device can be implemented in whole or in part through software, hardware, or a combination thereof. Each of the above units can be embedded in or independent of a processor in a computer device in the form of hardware, or can be stored in a memory in the computer device in the form of software, so that the processor can call and execute the corresponding operations of each of the above units.

[0144] An electronic device is also provided in an embodiment of the present application. Figure 4 As shown, it shows a schematic diagram of the structure of an electronic device suitable for implementing the embodiment of the present application. The electronic device in the embodiment of the present application may include but is not limited to fixed terminals such as mobile phones, tablet computers, learning machines, teaching large screens, wearable devices, etc. Figure 4 The electronic device shown is merely an example and should not limit the functions and scope of use of the embodiments of the present application.

[0145] like Figure 4 As shown, the electronic device may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 1, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 2 or the program loaded from the storage device 8 to the random access memory (RAM) 3 to implement the long text training data generation method of the aforementioned embodiment of the present application. When the electronic device is powered on, the RAM 3 also stores various programs and data required for the operation of the electronic device. The processing device 1, ROM 2 and RAM 3 are connected to each other via a bus 4. The input / output (I / O) interface 5 is also connected to the bus 4.

[0146] Typically, the following devices may be connected to the I / O interface 5: an input device 6 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 7 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 8 including, for example, a memory card, a hard disk, etc.; and a communication device 9. The communication device 9 may allow the electronic device to communicate with other devices wirelessly or by wire to exchange data. Figure 4 The electronic device is shown with various devices, but it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed instead.

[0147] An embodiment of the present application also provides a computer program product including computer-readable instructions. When the computer-readable instructions are executed on an electronic device, the electronic device implements any of the long text training data generation methods provided in the embodiments of the present application.

[0148] The embodiment of the present application further provides a computer readable storage medium, the storage medium carries one or more computer programs, when the one or more computer programs are executed by an electronic device, the electronic device can realize any long text training data generation method provided by the embodiment of the present application.

[0149] In addition, it should be noted that the above-described apparatus embodiments are merely illustrative, wherein the units described as separate components can or can not be physically separated, and the components displayed as units can or can not be physical units, i.e., can be located in one place, or can be distributed on multiple network units. Part or all of the modules can be selected according to actual needs to achieve the purpose of the embodiment. In addition, the connection relationship between the modules in the apparatus embodiment provided by the present application indicates that there is a communication connection between them, which can be implemented as one or more communication buses or signal lines.

[0150] Through the description of the above embodiments, those skilled in the art can clearly understand that the present application can be realized by means of software and necessary general hardware, and of course, it can also be realized by special hardware including special integrated circuits, special CPUs, special memories, special components, etc. Generally, functions completed by computer programs can be easily realized by corresponding hardware, and the specific hardware structure for realizing the same function can also be various, such as analog circuit, digital circuit or special circuit, etc. However, for the present application, software program implementation is a better embodiment. Based on this understanding, the technical solutions of the present application can be embodied in the form of a software product, which is stored in a readable storage medium, such as a computer floppy disk, U disk, mobile hard disk, ROM, RAM, magnetic disk or optical disk, etc., including a plurality of instructions to make a computer device (which can be a personal computer, training device, or network device, etc.) execute the methods described in various embodiments of the present application.

[0151] In the above embodiments, all or part can be realized by software, hardware, firmware or any combination thereof. When realized by software, it can be realized in the form of a computer program product in whole or in part.

[0152] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the flow or function described in the embodiments of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium, for example, the computer instructions can be transmitted from one website, computer, training device or data center to another website, computer, training device or data center through wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that a computer can store or a data storage device such as a training device, a data center, etc. integrated with one or more available media sets. The available media can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid state disk (SSD)), etc.

[0153] The various embodiments in the specification are described in a progressive manner, each embodiment focuses on the difference from other embodiments, the various embodiments can be combined as needed, and the same or similar parts refer to each other.

Claims

1. A method for generating long text training data, characterized in that: include: Acquire long text source data, where the text length of the long text source data is greater than a preset length value; Calling the big model to instruct the big model to generate relevant questions of the long text source data; Calling the big model to instruct the big model to generate answers corresponding to the relevant questions, and the number of generated answers is not less than 3; For each generated answer, the credibility of the current answer is calculated based on the text similarity between the current answer and each of the remaining answers; the answer with the highest credibility is determined as the final answer corresponding to the relevant question; Long text training data is generated using the long text source data, the relevant questions and the corresponding final answers.

2. The method for generating long text training data according to claim 1, wherein: The step of calculating the credibility of the current answer based on the text similarity between the current answer and the remaining answers includes: Obtain the encoding vector of each generated answer; Calculating the cosine similarity between the current answer and each of the remaining answers based on each of the encoding vectors; and determining the text similarity between the current answer and each of the remaining answers based on each of the cosine similarities; Calculate the statistical value of the text similarity between the current answer and each of the remaining answers as the credibility of the current answer.

3. The method for generating long text training data according to claim 2, wherein: Determining text similarities between the current answer and the remaining answers based on the cosine similarities includes: Numerical mapping calculation is performed on each of the cosine similarities to obtain text similarities between the current answer and each of the remaining answers; wherein any text similarity is positively correlated with the corresponding cosine similarity.

4. The method for generating long text training data according to any one of claims 1 to 3, characterized in that: The calling of the large model to instruct the large model to generate the long text source data includes: determining keywords of the long text source data; Slicing the long text source data, and retaining the text slices containing the keywords as short text data; The large model is called to instruct the large model to generate questions according to the short text data, and the generated questions are used as relevant questions of the long text source data.

5. The method for generating long text training data according to claim 4, wherein: Determine the keywords of the long text source data, including: Performing word frequency-inverse text frequency index statistics on each word in the long text source data, and extracting subject words from the long text source data based on the statistical results; Calling the big model to instruct the big model to perform word association based on the subject words to obtain relevant subject words; The subject word and the related subject words are determined as keywords of the long text source data.

6. The method for generating long text training data according to any one of claims 1 to 3, characterized in that: The calling of the large model to instruct the large model to generate answers corresponding to the relevant questions includes: Calling the big model at least three times to instruct the big model to answer the relevant questions and generate corresponding answers; Among them, the sampling temperature of the called large model is set to the preset high-temperature sampling parameter value.

7. The method for generating long text training data according to any one of claims 1 to 3, characterized in that: The obtaining of long text source data includes: Get the original text data; Segmenting the original text data and calculating the degree of difference between the text segments; Performing deduplication processing on the original text data according to the degree of difference; When the text length of the deduplicated text data is greater than the preset length value, the deduplicated text data is determined as long text source data.

8. The method for generating long text training data according to claim 7, wherein: Get raw text data, including: Obtain text data or non-text data; Preprocess the acquired data to obtain original text data; Among them, when the acquired data is text data, the preprocessing includes removing invalid characters and non-standard characters; when the acquired non-text data is voice data, the preprocessing includes voice recognition processing and spoken expression rewriting processing; when the acquired non-text data is a text image, the preprocessing includes text recognition processing.

9. A long text training data generation device, characterized in that: include: A source data acquisition unit, configured to acquire long text source data, wherein the length of the long text source data is greater than a preset length value; a question generating unit, configured to call the large model to instruct the large model to generate questions related to the long text source data; An answer generation unit, configured to call the large model to instruct the large model to generate answers corresponding to the relevant questions, with the number of generated answers being no less than 3; an answer screening unit, configured to calculate, for each generated answer, the credibility of the current answer based on the text similarity between the current answer and each of the remaining answers; The answer with the highest credibility is determined as the answer corresponding to the relevant question. final Answer; The training data generating unit is used to generate long text training data using the long text source data, the relevant questions and the corresponding final answers.

10. An electronic device, characterized in that: include: memory and processor; The memory is used to store programs; The processor is configured to execute the program to implement each step of the method for generating long text training data according to any one of claims 1 to 8.

11. A readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, each step of the method for generating long text training data according to any one of claims 1 to 8 is implemented.

12. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, each step of the method for generating long text training data according to any one of claims 1 to 8 is implemented.

Citation Information

Cited By

  • Data generation method and related product

    CN121478919A