A screening-based acquisition and editing method, system and electronic device based on product documents

By slicing and screening the product documents, the ChatGLM2 model is used to generate Q&A pairs and screen unqualified pairs, which solves the problems of insufficient generation quality and high cost in the existing technology, and realizes efficient Q&A pair generation and key information extraction.

CN117851561BActive Publication Date: 2025-07-11HUBEI PUBLIC INFORMATION IND CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311724872.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-12-15
Publication Date
2025-07-11
Estimated Expiration
2043-12-15

AI Technical Summary

Technical Problem

The existing Q&A fine-tune the generation model through a large amount of manual annotation data, resulting in insufficient generation quality and high cost, which cannot be used for most professional fields.

Method used

The screening and editing method based on product documents is adopted, and the document data is obtained for slice processing, and the ChatGLM2 model is used to extract key information and generate Q&A pairs. Combined with the screening model, the quality of Q&A pairs with unqualified quality is improved.

Benefits of technology

It effectively solves the problems of high cost and insufficient quality of manual labeling in the generation of Q&A pairs, improves the quality of Q&A pairs, provides efficient key information extraction and screening, and is suitable for customer solutions for front-line customer service personnel.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117851561B_ABST
    Figure CN117851561B_ABST
Patent Text Reader

Abstract

The present invention discloses a screening-based editing and compilation method, system and electronic device based on product documents. The method includes: obtaining document data of a product, performing slicing processing on the document data to obtain corresponding sliced statement data, and storing it in a local knowledge base; according to each sliced statement data in the local knowledge base, giving a first prompt statement predefined by the ChatGLM2 model, extracting corresponding product key information as the answer of the Q&A pair, and generating a product key information list; giving a second prompt statement predefined by the large language model ChatGLM2, and according to the product key information list, making the ChatGLM2 model generate questions corresponding to each product key information, and organizing to obtain the basic Q&A pairs of the product key information list; screening the basic Q&A pairs of the document based on a screening model to obtain the final edited and compiled file. The present invention improves the quality of the Q&A pairs generated by the model, and the finally obtained edited and compiled file can be used by customer service personnel.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of natural language processing, and in particular, to a screening-based editing method, system, and electronic device for product documents. Background Art

[0002] With the rapid development of natural language processing technology, using a dialogue generation model to generate question-and-answer pairs based on product documents can effectively help front-line customer service personnel summarize key information in long and complex product documents. A large language model is a language model composed of an artificial neural network with many parameters (usually billions of weights or more), which is trained on a large amount of unlabeled text using self-supervised learning or semi-supervised learning. With the development of technology, large language models have shown performance beyond previous dialogue models in various fields, especially in low-resource scenarios lacking labeled data.

[0003] Most existing question-and-answer pair generation models are fine-tuned on a generation model using a large amount of manually labeled data, which often requires a large amount of labor costs for data annotation. However, training a question-and-answer pair generation model by annotating questions and answers will result in insufficient generation quality, and the manual annotation cost is extremely expensive and cannot be applied to most professional fields. Summary of the Invention

[0004] To overcome the deficiencies of related products in the prior art, the present invention proposes a screening-based editing method, system, and electronic device for product documents.

[0005] In a first aspect, the present invention provides a screening-based editing method for product documents, including: S100. Obtain the document data of the product, perform slicing processing on the document data to obtain corresponding sliced statement data, and store it in the local knowledge base;

[0006] S200. According to each sliced statement data in the local knowledge base, give a first predefined prompt statement of the ChatGLM2 model, extract the corresponding product key information as the answer of the question-and-answer pair, and generate a list of product key information;

[0007] S300. Give a second predefined prompt statement of the large language model ChatGLM2, and according to the list of product key information, make the ChatGLM2 model generate questions corresponding to each product key information, and organize to obtain the basic document question-and-answer pairs corresponding to the list of product key information;

[0008] S400. Screen the basic document question-and-answer pairs based on a screening model, screen out the question-and-answer pairs with unqualified quality, and obtain the final editing file.

[0009] Optionally, step S100 specifically includes:

[0010] S101. Obtain the document data of the product, where the document data only contains text statements;

[0011] S102. Split the statements in the document data according to punctuation marks and length to obtain sliced statement data. First, split according to full stops. When the length of a statement exceeds two hundred, split using commas. If there is no comma, perform forced splitting to obtain sliced statement data with a length of two hundred;

[0012] S103. Store the obtained sliced statement data in the local knowledge base.

[0013] Optionally, step S200 specifically includes:

[0014] S201. Pre-define a first prompt statement as the guiding rule for the ChatGLM2 model to extract corresponding product key information;

[0015] S202. Combine the first prompt statement with at least one sliced statement data to form a sequence;

[0016] S203. Input the obtained sequence into the ChatGLM2 model to obtain the corresponding product key information as the answer to the Q&A pair;

[0017] S204. Extract the product key information from each sliced statement data in the local knowledge base respectively to generate a product key information list.

[0018] Optionally, step S300 specifically includes:

[0019] S301: Pre-define a second prompt statement as the guiding rule for the ChatGLM2 model to extract corresponding questions based on the corresponding product key information;

[0020] S302: Combine at least one product key information in the product key information list with the pre-defined second prompt statement and input it into the ChatGLM2 model to obtain the question corresponding to this product key information;

[0021] S303: Repeat step S302 to organize and obtain the basic Q&A pairs corresponding to the product key information list.

[0022] Optionally, step S400 specifically includes:

[0023] S401. Select a sliced statement data and splice it with one Q&A pair in the basic Q&A pairs of the document to obtain three sequences S1, S2, and S3. Sequence S1 is used to determine whether the question corresponds to this sliced statement data, sequence S2 is used to determine whether the answer is fabricated, and sequence S3 is used to determine whether the question and the answer match;

[0024] S402. Input the three obtained sequences S1, S2, and S3 into the encoding model to obtain three corresponding encodings. After splicing them and passing through a fully connected layer and Sigmoid, a score is obtained.

[0025] S403. Perform the screening steps of S401 and S402 on all the question-and-answer pairs obtained from the slice statement data.

[0026] S404. Execute step S403 on all the slice statement data obtained from the document data to obtain the final compiled file.

[0027] Optionally, the method further includes:

[0028] S500. Organize the obtained basic question-and-answer pairs of the document and train the screening model.

[0029] Optionally, step S500 specifically includes:

[0030] S510. Obtain the basic document question-and-answer pairs through steps S100 - S300.

[0031] S520. Screen all the question-and-answer pairs in the basic document question-and-answer pairs through preset screening rules to obtain qualified question-and-answer pairs and unqualified question-and-answer pairs respectively. Among them, the qualified question-and-answer pairs are used as positive examples, and the unqualified question-and-answer pairs are used as negative examples. Perform binary classification training on the screening model according to the positive and negative examples.

[0032] Optionally, step S520 specifically includes:

[0033] S521. Use the obtained qualified question-and-answer pairs as positive examples and all unqualified question-and-answer pairs as negative examples.

[0034] S522. Divide the obtained positive and negative examples into a training set, a validation set, and a test set according to 8:1:1 respectively.

[0035] S523. Use the training set, the validation set, and the test set to train the screening model and evaluate the trained screening model.

[0036] In a second aspect, the present invention further provides a screening-based compilation system for product documents, which is applied to the screening-based compilation method for product documents described in any one of the above, and includes:

[0037] A storage unit, configured to obtain the document data of the product, perform slicing processing on the document data to obtain corresponding slice statement data, and store it in the local knowledge base;

[0038] An answer extraction unit, configured to extract corresponding product key information as the answer to the Q&A pair and generate a list of product key information according to each slice statement data in the local knowledge base and a first prompt statement predefined by the ChatGLM2 model;

[0039] A question extraction unit, configured to generate questions corresponding to each piece of product key information according to the list of product key information by using a second prompt statement predefined by the large language model ChatGLM2, and organize all Q&A pairs corresponding to the list of product key information;

[0040] An editing and screening unit, configured to screen all Q&A pairs based on a screening model, and filter out unqualified Q&A pairs among them to obtain a final editing file.

[0041] In a third aspect, the present invention further provides an electronic device, including: a memory and a processor, which are communicatively connected to each other, wherein the memory stores computer instructions, and the processor executes the computer instructions.

[0042] Compared with the prior art, the present invention has the following advantages:

[0043] In the embodiment of the present invention, the screening-based editing method based on product documents trains a screening model through the obtained Q&A pairs. At the same time, based on this screening model, by splicing slice statement data with questions, slice statement data with answers, and questions with answers, the local relationship features are captured, effectively solving problems such as whether there is fact fabrication, irrelevance to the question, and mismatch between questions and answers. And use this screening model to judge the generated Q&A pairs, improving the quality of the Q&A pairs generated by the model. The final editing file obtained can be used by customer service personnel. Description of the Drawings

[0044] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required to be used in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0045] Figure 1 It is a schematic flowchart of the screening-based editing method based on product documents of the present invention;

[0046] Figure 2 It is a schematic flowchart of the screening of the screening-based editing method based on product documents of the present invention;

[0047] Figure 3 It is a schematic diagram of the training and evaluation of the screening model;

[0048] Figure 4 This is a schematic diagram of the principle structure of the screening-based acquisition and editing system based on product documents according to the present invention. Specific embodiments

[0049] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. The accompanying drawings show preferred embodiments of the present invention. The present invention can be implemented in many different forms and is not limited to the embodiments described herein. On the contrary, the purpose of providing these embodiments is to make the understanding of the disclosure content of the present invention more thorough and comprehensive.

[0050] Refer to Figure 1 as shown Figure 1 This is a schematic flowchart of the screening-based acquisition and editing method based on product documents according to the present invention. The screening-based acquisition and editing method based on product documents includes the following steps:

[0051] S100. Obtain the document data of the product, perform slicing processing on the document data to obtain corresponding sliced statement data, and store it in the local knowledge base;

[0052] S200. According to each sliced statement data in the local knowledge base, give the first predefined prompt statement of the ChatGLM2 model, extract the corresponding key product information as the answer to the question and answer pair, and generate a list of key product information;

[0053] S300. Give the second predefined prompt statement of the large language model ChatGLM2, and according to the list of key product information, make the ChatGLM2 model generate questions corresponding to each key product information, and organize to obtain the basic document question and answer pair corresponding to the list of key product information;

[0054] S400. Based on the screening model, screen the basic document question and answer pair, screen out the question and answer pairs with unqualified quality, and obtain the final acquisition and editing file.

[0055] Combined with Figure 2 as shown, in the embodiment of the present invention, step S100 specifically includes:

[0056] S101. Obtain the document data of the product, and the document data only contains text statements; by obtaining the document data of one or more products, these document data are intended to introduce the same product, and the format of the document data includes PDF, DOC, TXT, DOCX, and the document data only contains text statements.

[0057] S102. Split the sentences in the document data according to punctuation marks and length to obtain sliced sentence data. First, split according to full stops. When the length of a sentence exceeds two hundred, split it using commas. If there are no commas, perform a forced split to obtain sliced sentence data with a length of two hundred.

[0058] For example, as shown in Formula 1-1 below, the complete document D is divided into x1, x2,..., x n and so on, n sliced sentence data,

[0059] D=(x1, x2,..., x n )#(1-1)

[0060] x i =(w1, w2..., w l )#(1-2)

[0061] And as shown in Formula 1-2, the sliced sentence data x i consists of l words (w1, w2..., w l ).

[0062] S103. Store the sliced sentence data obtained after splitting in the local knowledge base.

[0063] In the embodiment of the present invention, step S200 specifically includes:

[0064] S201. Pre-define a first prompt statement as a guiding rule for the ChatGLM2 model to extract corresponding product key information; in the embodiment of the present invention, by pre-defining the first prompt statement, it is used to guide ChatGLM2 to summarize product key information from the sentences obtained in step S102, and requires generating as many and complete key information extracts as possible. For example:

[0065] P1 = "It is required to extract as many and complete key information as possible from this background information"

[0066] P1 is the first prompt statement defined for this step.

[0067] S202. Combine the first prompt statement with at least one sliced sentence data to form a sequence; fill a sliced sentence data obtained in step S102 and the original pre-defined first prompt statement into the prompt template as a complete prompt. For example: P2 = I, P1, Example

[0068] I = "Given the background information [x i "

[0069] Example = "Input: XXXXX\n\n, Output: 1.XXXXX\N2.XXXXX\n"

[0070] P2 is the obtained complete prompt, I is the prompt of the given information, and x i is the previous statement obtained in step S102. Example is the output sample of the given document, which is used to standardize the output format of the model.

[0071] S203. Input the obtained sequence into the ChatGLM2 model to obtain the corresponding product key information as the answer to the Q&A pair; input the complete prompt obtained in S202 into the ChatGLM2 model to obtain several pieces of product key information. For example, A1, A2,..., A z = ChatGLM2(P2), where A i is the product key information.

[0072] S204. Extract the product key information from each slice statement data in the local knowledge base respectively to generate a list of product key information.

[0073] In the embodiment of the present invention, step S300 specifically includes:

[0074] S301: Pre-define a second prompt statement as the guiding rule for the ChatGLM2 model to extract the corresponding questions according to the corresponding product key information; in the embodiment of this method, by pre-defining the second prompt statement, it is used to guide ChatGLM2 to summarize the questions corresponding to the key information according to the key information and the knowledge base. For example:

[0075] p3 = taking A i as the answer to the question, then what is this question

[0076] P4 = P3, A i

[0077] where p3 is the pre-defined second prompt statement, and then it is combined with one piece of product key information A i obtained in step S203, that is, P4.

[0078] S302: Combine at least one piece of product key information in the list of product key information with the pre-defined second prompt statement, and input it into the ChatGLM2 model to obtain the question corresponding to this piece of product key information; for example: Q i = ChatGLM2(P4).

[0079] S303: Repeat step S302 to organize and obtain the basic Q&A pairs of the document corresponding to the list of product key information; for example,

[0080] Output = {(Q1, A1), (Q2, A2),...}#(1-7)

[0081] Where Q is the question and A is the answer. In this way, customer service staff can quickly obtain key information from long product documents and answer customers' questions. The following gives a sample of output question-answer pairs, as shown below:

[0082] Q1: What are the charging standards for the Star Card Flow Edition package?

[0083] A1: The Star Card Flow Edition package includes three charging standards, with monthly basic fees of (19 yuan) / (29 yuan) / (39 yuan) packages.

[0084] Combined with Figure 3 shown, in the embodiment of the present invention, step S400 specifically includes:

[0085] S401. Select a slice statement data and splice it with one of the question-answer pairs in the document basic question-answer pairs to obtain three sequences S1, S2, and S3. Sequence S1 is used to determine whether the question corresponds to the slice statement data, sequence S2 is used to determine whether the answer is fabricated, and sequence S3 is used to determine whether the question and the answer match; for example: S1 = [CLS]text[SEP]Question, S2 = [CLS]text[SEP]Answer, S3 = [CLS]Answer[SEP]Question, where text is the slice statement data, Question and Answer are its corresponding question-answer pair, and [CLS] and [SEP] are delimiters.

[0086] S402. Input the obtained three sequences S1, S2, and S3 into the encoding model, and correspondingly obtain three encodings H1, H2, and H3. After splicing them, pass through a fully connected layer and Sigmoid to obtain a score; the three encodings H1, H2, and H3 are respectively:

[0087] H1 = Roberta(S1)

[0088] H2 = Roberta(S2)

[0089] H2 = Roberta(S3)

[0090] The obtained score score is:

[0091] In the embodiment of the present invention, by using the representations H 1cls , H 2cls , H 3cls of the special start character [CLS] of the three sequences, after splicing, pass through a fully connected layer FFN and Sigmoid to obtain a score. Through three forms of encoding, the model can effectively capture local features and determine whether there are problems such as fabricating facts, being irrelevant to the question, and the question and answer not matching.

[0092] S403. Screen all the question-and-answer pairs obtained from the slice statement data through the steps S401 and S402;

[0093] S404. Execute step S403 for all the slice statement data obtained from the document data to obtain the final editing file.

[0094] In the embodiment of the present invention, the method further includes:

[0095] S500. Sort out the obtained basic question-and-answer pairs of the document and train a screening model.

[0096] Optionally, step S500 specifically includes:

[0097] S510. Obtain basic document question-and-answer pairs through steps S100-S300;

[0098] S520. Screen all the question-and-answer pairs in the basic document question-and-answer pairs through a preset screening rule to respectively obtain qualified question-and-answer pairs and unqualified question-and-answer pairs, where the qualified question-and-answer pairs are used as positive examples and the unqualified question-and-answer pairs are used as negative examples (including using some question-and-answer pairs not generated according to this sentence as negative examples and using some answers of qualified question-and-answer pairs after deleting some content as negative examples), use Roberta trained by Harbin Institute of Technology as the base model and perform binary classification training on the screening model according to the positive examples and negative examples.

[0099] In the embodiment of the present invention, step S520 specifically includes:

[0100] S521. Use the obtained qualified question-and-answer pairs as positive examples and all unqualified question-and-answer pairs as negative examples;

[0101] S522. Divide the obtained positive examples and negative examples into a training set, a validation set, and a test set according to 8:1:1 respectively;

[0102] S523. Use the training set, the validation set, and the test set to train the screening model and evaluate the trained screening model.

[0103] In the embodiment of the present invention, the screening model is trained by using the training set, the validation set, and the test set, and the training objective is: Loss = -(y * log(y')+(1 - y) * log(1 - y'))

[0104] where y' is the probability that the question-and-answer pair is a positive example, that is, score, and y is the label of the sample.

[0105] And evaluate the trained model, and the evaluation results are shown in the following table:

[0106]

[0107]

[0108] In the embodiment of the present invention, the screening-based editing and compilation method based on product documents trains a screening model through the obtained question-and-answer pairs. At the same time, based on this screening model, by splicing the sliced statement data with the question, the sliced statement data with the answer, and the question with the answer, the local relationship features are captured, effectively solving problems such as whether there is fact fabrication, irrelevance to the question, and question-answer mismatch. And use this screening model to judge the generated question-and-answer pairs, improving the quality of the question-and-answer pairs generated by the model. The finally obtained editing and compilation file can be used by customer service personnel.

[0109] Based on the above embodiment, the embodiment of the present invention also provides a screening-based editing and compilation system based on product documents. Refer to Figure 2 As shown, the screening-based editing and compilation system based on product documents described in the embodiment of the present invention includes:

[0110] A storage unit 10, configured to obtain the document data of the product, perform slicing processing on the document data to obtain corresponding sliced statement data, and store it in the local knowledge base;

[0111] An answer extraction unit 20, configured to, according to each sliced statement data in the local knowledge base, give a first prompt statement predefined by the ChatGLM2 model, extract the corresponding product key information as the answer of the question-and-answer pair, and generate a list of product key information;

[0112] A question extraction unit 30, configured to give a second prompt statement predefined by the large language model ChatGLM2, and according to the list of product key information, make the ChatGLM2 model generate questions corresponding to each product key information, and organize all the question-and-answer pairs corresponding to the list of product key information;

[0113] An editing and compilation screening unit 40, configured to screen all the question-and-answer pairs based on the screening model, screen out the question-and-answer pairs with unqualified quality, and obtain the final editing and compilation file.

[0114] The screening-based editing and compilation system based on product documents described in the embodiment of the present invention can execute the screening-based editing and compilation method based on product documents provided in the above embodiment. The screening-based editing and compilation system based on product documents has the corresponding functional steps and beneficial effects of the material level monitoring method described in the above embodiment. For details, please refer to the embodiment of the screening-based editing and compilation method based on product documents. The embodiment of the present invention will not be elaborated here.

[0115] An embodiment of the present invention further provides an electronic device, which may include a processor and a memory, and the processor and the memory may be connected through a bus or other means. The processor may be a central processing unit (CPU). The processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., or a combination of the above types of chips. The memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs, non-transitory computer-executable programs, and modules, such as program instructions / modules corresponding to the material level monitoring method in the embodiment of the present invention. The processor executes various functional applications and data processing of the processor by running the non-transitory software programs, instructions, and modules stored in the memory, that is, implements the screening-based customer service intelligent editing method of the large language model in the above method embodiment.

[0116] The memory may include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created by the processor, etc. In addition, the memory may include high-speed random access memory and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. The one or more modules are stored in the memory and, when executed by the processor, execute the screening-based customer service intelligent editing method of the large language model in the embodiment shown in Figure 1 The above electronic device details can be referred to specifically Figure 1For the corresponding relevant descriptions and effects in the illustrated embodiments, they are understood here and will not be elaborated further. Those skilled in the art can understand that to implement all or part of the processes in the methods of the above embodiments, it can be completed by instructing relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, the storage medium can be a Read-Only Memory (ROM), a Random Access Memory (RAM), a Flash Memory, a Hard Disk Drive (abbreviation: HDD), or a Solid-State Drive (SSD), etc.; the storage medium can also include a combination of the above types of memories.

[0117] The content not described in detail in this specification belongs to the prior art well known to those skilled in the art. The above are only embodiments of the present invention, but do not limit the patent scope of the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, for those skilled in the art, they can still modify the technical solutions described in the foregoing specific embodiments, or perform equivalent replacements for some of the technical features. Any equivalent structure made by using the content of the specification and drawings of the present invention, directly or indirectly applied in other related technical fields, is similarly within the scope of the patent protection of the present invention.

Claims

1. A screening-based editing and collecting method based on product documents, characterized in that, Including: S100. Obtain the document data of the product, perform slicing processing on the document data to obtain corresponding sliced statement data, and store it in the local knowledge base; S200. According to each sliced statement data in the local knowledge base, given the first predefined prompt statement of the ChatGLM2 model, extract the corresponding key product information as the answer of the Q&A pair, and generate a list of key product information; S300. Given the second predefined prompt statement of the large language model ChatGLM2, make the ChatGLM2 model generate questions corresponding to each key product information according to the list of key product information, and organize to obtain the basic document Q&A pairs corresponding to the list of key product information; S400. Screen the basic document Q&A pairs based on the screening model, screen out the Q&A pairs with unqualified quality, and obtain the final editing file; Step S400 specifically includes: S401. Select a sliced statement data and splice it with a Q&A pair in the basic document Q&A pairs to obtain three sequences S1, S2, and S3. Sequence S1 is used to judge whether the question corresponds to the sliced statement data, sequence S2 is used to judge whether the answer is fabricated, and sequence S3 is used to judge whether the question and the answer match; S402. Input the obtained three sequences S1, S2, and S3 into the encoding model, obtain three corresponding encodings, splice them, and then obtain a score after passing through a fully connected layer and Sigmoid; S403. Perform the screening of steps S401 and S402 on all Q&A pairs obtained from the sliced statement data; S404. Execute step S403 on all sliced statement data obtained from the document data to obtain the final editing file.

2. The screening-based acquisition and editing method based on product documents according to claim 1, wherein Step S100 specifically includes: S101. Obtain the document data of the product, and the document data only contains text statements; S102. Split the statements in the document data according to punctuation marks and length to obtain sliced statement data. First, split according to the period. When the length of the statement exceeds two hundred, split it with a comma. If there is no comma, select forced splitting to split out sliced statement data with a length of two hundred; S103. Store the sliced statement data obtained after splitting into the local knowledge base.

3. The screening-based editing and collection method based on product documents according to claim 1, wherein Step S200 specifically includes: S201. Predefine the first prompt statement as the guiding rule for the ChatGLM2 model to extract the corresponding key product information; S202. Combine the first prompt statement with at least one sliced statement data to form a sequence; S203. Input the obtained sequence into the ChatGLM2 model to obtain the corresponding key product information as the answer of the Q&A pair; S204. Extract the key product information from each sliced statement data in the local knowledge base respectively to generate a list of key product information.

4. The screening-based editing and collecting method based on product documents according to claim 1, wherein Step S300 specifically includes: S301: Predefine the second prompt statement as the guiding rule for the ChatGLM2 model to extract the corresponding question according to the corresponding key product information; S302: Combine at least one piece of product key information in the product key information list with a predefined second prompt statement, and input it into the ChatGLM2 model to obtain the question corresponding to this piece of product key information. S303: Repeat step S302 to organize and obtain the basic Q&A pairs corresponding to the product key information list.

5. The screening-based acquisition and editing method based on product documents according to claim 1, wherein The method further includes: S500. Organize the obtained basic document Q&A pairs and train a screening model.

6. The screening-based acquisition and editing method based on product documentation according to claim 5, characterized in that, Step S500 specifically includes: S510. Obtain basic document Q&A pairs through steps S100 - S300. S520. Screen all the Q&A pairs in the basic document Q&A pairs through a preset screening rule to obtain qualified Q&A pairs and unqualified Q&A pairs respectively. Among them, the qualified Q&A pairs are used as positive examples, and the unqualified Q&A pairs are used as negative examples. Perform binary classification training on the screening model according to the positive examples and negative examples.

7. The screening-based editing and collection method based on product documentation according to claim 6, wherein Step S520 specifically includes: S521. Use the obtained qualified Q&A pairs as positive examples and all unqualified Q&A pairs as negative examples. S522. Divide the obtained positive examples and negative examples into a training set, a validation set, and a test set according to 8:1:1 respectively. S523. Use the training set, the validation set, and the test set to train the screening model and evaluate the trained screening model.

8. A screening-based editing and collection system based on product documents, which is applied to the screening-based editing and collection method based on product documents according to any one of claims 1-7, and is characterized in that, It includes: A storage unit for obtaining the document data of the product, slicing the document data to obtain corresponding sliced statement data, and storing it in the local knowledge base. An answer extraction unit for, according to each sliced statement data in the local knowledge base, giving a predefined first prompt statement to the ChatGLM2 model, extracting the corresponding product key information as the answer of the Q&A pair, and generating a product key information list. A question extraction unit for giving a predefined second prompt statement to the large language model ChatGLM2, and according to the product key information list, making the ChatGLM2 model generate the question corresponding to each piece of product key information, and organizing and obtaining all the Q&A pairs corresponding to the product key information list. An editing and screening unit for screening all the Q&A pairs based on the screening model, screening out the Q&A pairs with unqualified quality, and obtaining the final editing file.

9. An electronic device, characterized in that, It includes: A memory and a processor, which are communicatively connected to each other. The memory stores computer instructions, and the processor executes the computer instructions to execute the method according to any one of claims 1 - 7.

Citation Information

Patent Citations

  • Method, device and system for constructing question and answer library by using large language model and medium

    CN117216205A

  • Question and answer method and device based on long document, storage medium and equipment

    CN117216208A