Question and answer pair data generation method and device, electronic equipment and storage medium

By analyzing and segmenting multimodal data, generating structured information and using large models, the problem of low efficiency in updating the enterprise knowledge base is solved, and efficient question-and-answer data generation and automatic update of the knowledge base is achieved.

CN120296038APending Publication Date: 2025-07-11GUANGZHOU DULING TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510331103.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-07-11

Smart Images

  • Figure CN120296038A_ABST
    Figure CN120296038A_ABST
Patent Text Reader

Abstract

The invention discloses a question and answer pair data generation method and device, electronic equipment and a storage medium, and relates to the field of computers, in particular to the technical field of artificial intelligence such as large models, natural language processing and computer vision. According to the specific implementation scheme, firstly, obtained multi-modal data are analyzed, key information in the multi-modal data is determined, then structured information corresponding to the multi-modal data is generated according to the key information, then the segmentation granularity of current information is determined, the structured information is segmented based on the segmentation granularity of the current information, and the segmented information is obtained. And finally, inputting the information sets into the large model to obtain question and answer pairs generated by the large model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technologies, and in particular to artificial intelligence technology fields such as large models, natural language processing, and computer vision. Specifically, it relates to a method, apparatus, electronic device, and storage medium for generating question-and-answer pair data. Background Art

[0002] Currently, enterprise knowledge bases are usually constructed and updated based on question-and-answer pair data regularly annotated manually. However, the efficiency of generating question-and-answer pair data in this way is low, resulting in a long update cycle of the knowledge base, poor timeliness, and inability to meet real-time business requirements. Summary of the Invention

[0003] The present disclosure provides a method, apparatus, electronic device, and storage medium for generating question-and-answer pair data. The specific solutions are as follows:

[0004] According to one aspect of the present disclosure, there is provided a method for generating question-and-answer pair data, characterized by comprising:

[0005] Parsing the obtained multimodal data to determine the key information in the multimodal data;

[0006] Generating structured information corresponding to the multimodal data according to the key information;

[0007] Determining the segmentation granularity of the current information;

[0008] Based on the segmentation granularity of the current information, segmenting the structured information to obtain a plurality of information sets;

[0009] Inputting the information sets into a large model to obtain question-and-answer pairs generated by the large model.

[0010] According to another aspect of the present disclosure, there is provided a device for generating question-and-answer pair data, characterized by comprising:

[0011] A first determination module for parsing the obtained multimodal data to determine the key information in the multimodal data;

[0012] A generation module for generating structured information corresponding to the multimodal data according to the key information;

[0013] A second determination module for determining the segmentation granularity of the current information;

[0014] A segmentation module for segmenting the structured information based on the segmentation granularity of the current information to obtain a plurality of information sets;

[0015] A processing module for inputting the information sets into a large model to obtain question-and-answer pairs generated by the large model.

[0016] According to another aspect of the present disclosure, there is provided an electronic device, including:

[0017] at least one processor; and

[0018] a memory communicatively connected to the at least one processor; wherein,

[0019] the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method described in the above embodiments.

[0020] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the method described in the above embodiments.

[0021] According to another aspect of the present disclosure, there is provided a computer program product including a computer program, and the computer program implements the steps of the method described in the above embodiments when executed by a processor.

[0022] The method, device, electronic device, and storage medium for generating question-and-answer pair data provided by the present disclosure have the following beneficial effects: First, the obtained multimodal data is parsed to determine the key information in the multimodal data, then structured information corresponding to the multimodal data is generated according to the key information, then the segmentation granularity of the current information is determined, and based on the segmentation granularity of the current information, the structured information is segmented to obtain a plurality of information sets, and finally the information sets are input into a large model to obtain question-and-answer pairs generated by the large model. Thus, by parsing the multimodal data, determining the corresponding key information, generating structured information according to the key information, segmenting the structured information based on the current segmentation granularity, and inputting the obtained information sets after segmentation into the large model to obtain question-and-answer pairs, the quality and efficiency of generating question-and-answer pair data are improved, providing conditions for realizing the automatic update of the knowledge base and improving the timeliness of the knowledge base.

[0023] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:

[0025] Figure 1 is a flowchart of a method for generating question-and-answer pair data provided by an embodiment of the present disclosure;

[0026] Figure 2 Schematic flowchart of the question-answer pair data generation method provided in another embodiment of the present disclosure;

[0027] Figure 3 Schematic flowchart of the question-answer pair data generation method provided in another embodiment of the present disclosure;

[0028] Figure 4 Schematic flowchart of the question-answer pair data generation method provided in another embodiment of the present disclosure;

[0029] Figure 5 Schematic flowchart of the question-answer pair data generation method provided in another embodiment of the present disclosure;

[0030] Figure 6 Schematic flowchart of the question-answer pair data generation method provided in another embodiment of the present disclosure;

[0031] Figure 7 Schematic flowchart of generating and updating question-answer pairs in the question-answer pair data generation method proposed in an embodiment of the present disclosure;

[0032] Figure 8 Schematic architecture diagram of the question-answer pair data generation method proposed in an embodiment of the present disclosure;

[0033] Figure 9 Schematic structure diagram of the question-answer pair data generation device provided in an embodiment of the present disclosure;

[0034] Figure 10 Block diagram of an electronic device for implementing the question-answer pair data generation method of the embodiments of the present disclosure. Detailed implementation manners

[0035] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted below.

[0036] The embodiments of the present disclosure relate to the fields of artificial intelligence technologies such as large models, natural language processing, and computer vision.

[0037] Artificial Intelligence, abbreviated as AI in English, is a new technical science that studies, develops theories, methods, technologies, and application systems for simulating, extending, and expanding human intelligence.

[0038] The large model can also be called the Foundation Model. The model extracts knowledge from hundreds of millions of corpus or images, learns, and then produces a large model with hundreds of millions of parameters.

[0039] Natural Language Processing (NLP) is an interdisciplinary field in computer science, artificial intelligence, and linguistics. It mainly studies how to enable computers to understand, process, generate, and simulate human language capabilities, so as to achieve the ability to have natural conversations with humans.

[0040] Computer vision refers to using cameras and computers to replace human eyes for tasks such as object recognition, tracking, and measurement in machine vision, and further performing graphics processing to make the images processed by the computer more suitable for human eyes to observe or be transmitted to instruments for detection.

[0041] It should be noted that in the technical solution of the present disclosure, the acquisition, storage, use, processing, etc. of data all comply with the relevant regulations of national laws and regulations and do not violate public order and good customs.

[0042] Next, the method, device, electronic device, and storage medium for generating question-and-answer pair data according to the embodiments of the present disclosure will be described with reference to the accompanying drawings.

[0043] Figure 1 It is a schematic flowchart of the method for generating question-and-answer pair data provided by an embodiment of the present disclosure.

[0044] As Figure 1 shown, the method for generating question-and-answer pair data includes:

[0045] Step 101, parse the acquired multimodal data to determine the key information in the multimodal data.

[0046] Among them, the multimodal data can be real-time instant message data, which can include various modal data such as text, pictures, voices, and files. The present disclosure does not limit this.

[0047] Among them, instant messages can also be called "Instant Messaging (IM) messages", "IM messages", etc. The present disclosure does not limit this.

[0048] In the present disclosure, when parsing the acquired multimodal messages to determine the key information in the multimodal data, different methods can be used to parse different modal data to determine the key information in the multimodal data.

[0049] For example, taking the multimodal data including text data, picture data, voice data, and file data as an example for illustration.

[0050] For text data, based on natural language processing techniques, entities, relationships, semantic tags, etc. in the text can be extracted, and then the text data can be parsed to obtain key information in the text data;

[0051] For image data, based on computer vision models, text, tables, key objects, etc. in the image data can be recognized, and the key information in the image data can be parsed through image recognition techniques;

[0052] For speech data, based on automatic speech recognition techniques, after converting the speech data into text data, combined with context semantic analysis, the key information in the speech data can be extracted;

[0053] For file data, the document structure is parsed, and key information is extracted to support the parsing and processing of multiple file formats, etc. The present disclosure does not limit this.

[0054] Thus, by parsing multi-modal data, a data foundation is provided for constructing a knowledge base question-answering system containing multi-modal data, thereby improving the question-answering accuracy.

[0055] Step 102, generate structured information corresponding to the multi-modal data according to the key information.

[0056] In the present disclosure, after determining the key information in the multi-modal data, the key information can be mapped to unified knowledge graph nodes, integrating information of different modalities and sources into a unified knowledge graph, associating knowledge graph nodes such as time stamps, senders, and conversation contexts, and then generating structured information corresponding to the multi-modal data based on the knowledge graph nodes. For example, taking the scenario: User A sent a picture B at the current moment t, the obtained multi-modal data is picture B, parsing picture B, extracting the key information M of picture B, using the key information M of picture B as a node, associating with the sender node User A and the t moment node, and then generating the corresponding structured information based on these nodes: User A sent a picture B at the t moment, and the key information of picture B is M. The present disclosure does not limit this.

[0057] It should be noted that when generating structured information corresponding to the multi-modal data according to the key information, in the case where the multi-modal data is group chat data, since the group chat data has no specific recipient, the generated structured information can include a time stamp node, a sender node, and a key information node. In the case where the multi-modal data is private chat data, since the private chat data has a corresponding recipient, the generated structured information can include a time stamp node, a sender node, a recipient node, and a key information node. That is to say, the nodes in the structured information corresponding to the multi-modal data in different scenarios may be different. The present disclosure does not limit this.

[0058] Step 103, determine the segmentation granularity of the current information.

[0059] Among them, the segmentation granularity can be the granularity for segmenting structured information, which can be pre-set or can also be determined according to needs, and can include multi-dimensional segmentation granularities. For example, it can include time segmentation granularity, semantic segmentation granularity, quantity segmentation granularity, etc. Among them, the quantity segmentation granularity is the granularity for segmenting the entity quantity or data volume in the structured information, and the present disclosure does not limit this.

[0060] In the present disclosure, after generating the structured information corresponding to the multi-modal data, in order to ensure that when segmenting the current information, each segmented information segment can be more complete, thereby improving the quality of generating question-and-answer pair data, the segmentation granularity of the current information can be determined first.

[0061] Step 104, based on the segmentation granularity of the current information, segment the structured information to obtain multiple information sets.

[0062] Among them, the information set can be the information segment after segmenting the structured information.

[0063] In the present disclosure, after determining the segmentation granularity of the current information, the structured information can be segmented based on the segmentation granularity to obtain multiple information sets corresponding to the structured information, thereby providing a data basis for generating question-and-answer pair data. For example, after the segmentation granularity includes time segmentation granularity, semantic segmentation granularity, and quantity segmentation granularity for entities, the structured information can be segmented separately based on each segmentation granularity to obtain the information set corresponding to the time dimension, the information set corresponding to the semantic dimension, and the information set corresponding to the quantity dimension. The present disclosure does not limit this.

[0064] Step 105, input the information set into the large model to obtain the question-and-answer pairs generated by the large model.

[0065] It should be noted that the specific type and structure of the large model used to generate question-and-answer pairs in the present disclosure can be pre-set or can also be determined according to actual needs. For example, the large model can be a generative question-and-answer large model, or can also be a retrieval-augmented generation model, etc. The present disclosure does not limit this.

[0066] In the present disclosure, after obtaining multiple information sets, the information sets can be input into the large model to obtain the question-and-answer pairs generated by the large model, improve the efficiency of generating question-and-answer pair data, and after generating the question-and-answer pairs, the question-and-answer pairs are automatically added to the knowledge base to realize the automatic update of the knowledge base and improve the timeliness of the knowledge base.

[0067] Among them, the knowledge base can be pre-set and can be used to store question-and-answer pairs, providing a data basis for providing users with accurate, fast, and high-quality answers.

[0068] In the embodiments of the present disclosure, first, the obtained multi-modal data is parsed to determine the key information in the multi-modal data. Then, based on the key information, structured information corresponding to the multi-modal data is generated. After that, the segmentation granularity of the current information is determined, and based on the segmentation granularity of the current information, the structured information is segmented to obtain multiple information sets. Finally, the information sets are input into the large model to obtain the question-and-answer pairs generated by the large model. Thus, by parsing the multi-modal data, determining its corresponding key information, generating structured information based on the key information, segmenting the structured information based on the current segmentation granularity, and inputting the obtained information sets after segmentation into the large model to obtain question-and-answer pairs, the quality and efficiency of the generation of question-and-answer pair data are improved, providing conditions for realizing the automatic update of the knowledge base and improving the timeliness of the knowledge base.

[0069] Figure 2 It is a schematic flowchart of a method for generating question-and-answer pair data provided by another embodiment of the present disclosure.

[0070] As Figure 2 shown, the method for generating question-and-answer pair data includes:

[0071] Step 201: Parse the obtained multi-modal data to determine the key information in the multi-modal data.

[0072] Step 202: Generate structured information corresponding to the multi-modal data according to the key information.

[0073] Among them, the specific implementation forms of steps 201 to 202 can refer to the detailed descriptions in other embodiments of the present disclosure and will not be specifically elaborated here.

[0074] Step 203: Determine the time segmentation granularity of the current information according to the time period to which the moment when the multi-modal data is generated belongs.

[0075] It should be noted that the specific division of the time period to which the moment of generating multimodal data belongs can be made according to the interaction frequency between the knowledge base question-answering system and the user. For example, the time period during which the knowledge base question-answering system has a high-frequency interaction with the user can be divided into one time period, and the time period during which the interaction frequency between the knowledge base question-answering system and the user is relatively low can be divided into one time period. Taking the knowledge base question-answering system of an enterprise as an example, usually, it can be divided according to the work and rest rules of the enterprise. For example, during the morning shift period of the enterprise (such as from 8 am to 4 pm), the interaction frequency between the knowledge base question-answering system and the user is relatively high, and during the night shift period (such as from 4 pm to 0 am), the interaction frequency between the knowledge base question-answering system and the user is relatively low, etc., so that the determined time segmentation granularity can adapt to the business needs and user interaction patterns of different time periods of the enterprise, improving the accuracy and practicality of the segmentation of the information set. The present disclosure does not limit this.

[0076] In the present disclosure, after generating the structured information corresponding to the multimodal data, in order to improve the accuracy of segmenting the structured information from the time dimension, the time segmentation granularity of the current information can be determined first according to the time period to which the moment of generating the multimodal data belongs. For example, in the case where the time period to which the moment of generating the multimodal data belongs is the high-frequency interaction period between the knowledge base question-answering system and the user, in order to capture the user's behavior characteristics and more accurately segment the structured information, the determined time segmentation granularity of the current information can be finer than the time segmentation granularity corresponding to the time period during which the interaction frequency between the knowledge base question-answering system and the user is relatively low. The present disclosure does not limit this.

[0077] Step 204, perform semantic clustering on the structured information to determine the topic distribution information corresponding to the structured information.

[0078] Among them, the topic distribution information may include topics included in the structured information, information corresponding to each topic, and the relevance between each topic, etc. The present disclosure does not limit this.

[0079] In the present disclosure, after generating the structured information corresponding to the multimodal data, in order to better understand the user's intentions and needs, the structured information can be semantically clustered through the LDA model to determine the topic distribution information corresponding to the structured information.

[0080] Among them, LDA is the abbreviation of the Latent Dirichlet Allocation model, which is a document topic generation model.

[0081] Step 205, determine the semantic segmentation granularity of the current information based on the topic distribution information.

[0082] In the present disclosure, after determining the topic distribution information corresponding to the structured information, in order to classify similar information under the same topic, the semantic segmentation granularity of the current information can be determined based on the topic distribution information, so that the user's intention and requirements can be understood more accurately and reliably.

[0083] Step 206: Determine the quantity segmentation granularity of the current information according to the number of entities and / or the amount of data included in the structured information.

[0084] Among them, the quantity segmentation granularity can be the segmentation granularity of the entities included in the structured information, or it can also be the segmentation granularity of the data volume of the structured information. The specific segmentation granularity can be determined according to needs. For example, the quantity segmentation granularity can be that the information set after segmentation contains messages sent by 3 to 5 senders, or it can also be that the information set after segmentation contains 3 messages, etc. The present disclosure does not limit this.

[0085] In the present disclosure, after generating the structured information corresponding to the multimodal data, in order to control the information density of the information set obtained after segmentation to ensure that the information set obtained after segmentation is neither too long nor too short, the quantity segmentation granularity of the current information can be determined according to the number of entities and / or the amount of data included in the structured information. For example, the structured information can be segmented according to each information set containing 3 to 5 entities, so that each information set obtained after segmentation contains 3 to 5 entities, etc. The present disclosure does not limit this.

[0086] In some possible implementation forms, the segmentation granularity of the current information can also be determined according to the source of the multimodal data. For example, the segmentation granularity of the current information can be determined according to whether the multimodal data is from the same source, so as to provide conditions for improving the quality of generating question-and-answer pair data. The present disclosure does not limit this.

[0087] Among them, the source of the multimodal data can be from a knowledge base question-and-answer system, or it can also be from product documents, or it can also be from a customer service system, etc. The present disclosure does not limit this.

[0088] It should be noted that different sources of multimodal data may have different segmentation granularities for time, semantics, etc. The present disclosure does not limit this.

[0089] Among them, step 203, step 204 to step 205, and step 206 are parallel steps and can be executed in parallel, or can also be executed sequentially, etc. The specific execution order can be determined according to pre-setting or actual needs. For example, step 203, step 204 to step 205, and step 206 can be executed in parallel, that is, the time segmentation granularity, semantic segmentation granularity, and quantity segmentation granularity of the current information can be determined simultaneously, or step 203 can be executed first, then step 204 to step 205, and finally step 206, that is, the time segmentation granularity, semantic segmentation granularity, and quantity segmentation granularity of the current information can be determined sequentially, etc. The present disclosure does not limit this.

[0090] Step 207, based on the segmentation granularity of the current information, segment the structured information to obtain multiple information sets.

[0091] In the present disclosure, after determining the segmentation granularity of the current information, the structured information can be segmented separately based on the time segmentation granularity to capture the user's behavior characteristics in different time periods, can be segmented separately based on the semantic segmentation granularity to classify the information belonging to the same theme, and can also be segmented separately based on the quantity segmentation granularity, so that the obtained information sets after segmentation are neither too long nor too short. Thus, by segmenting the structured information in multiple dimensions, the multi-modal data can be analyzed and understood more accurately and reliably, and then the user's intention and needs can be better understood, and the quality of the generated question-and-answer pairs can be improved.

[0092] It should be noted that when segmenting the structured information based on the time segmentation granularity, semantic segmentation granularity, and quantity segmentation granularity, it can be performed in parallel, or can also be performed sequentially, etc. That is, the time segmentation, semantic segmentation, and quantity segmentation of the structured information can be performed separately at the same time, or the time segmentation of the structured information can be performed separately first, then the semantic segmentation of the structured information can be performed separately, and finally the quantity segmentation of the structured information can be performed separately, etc. The specific execution order can be pre-set, or can also be determined according to actual needs. The present disclosure does not limit this.

[0093] In some possible implementation forms, after obtaining multiple information sets, in order to ensure that the time segmentation granularity can adapt to the business requirements and user interaction modes (such as whether the interaction frequency is high or not) in different time periods, and then improve the accuracy and practicality of the segmentation of the structured information in the time dimension, the description information corresponding to each time period can be determined first according to these multiple information sets, where the description information includes at least one of the following: the total number of information sets, the amount of data included in each information set, the number of entities included in each information set, and then the time segmentation granularity corresponding to each time period can be updated according to the description information corresponding to each time period.

[0094] For example, in the case where the amount of data or the number of entities included in each information set is large, it can be determined that the time period to which the multimodal data belongs is the time period with a high interaction frequency with the user. Since the amount of data and the number of entities are large, in order to ensure that the amount of data or the number of entities included in each information set is neither too large nor too small, and to improve the efficiency and quality of generating question-and-answer pairs, at this time, the time segmentation granularity of the time period to which the multimodal data belongs can be set to a finer granularity, and the structured information can be segmented more precisely in the time dimension, so as to realize the dynamic adjustment of the time segmentation granularity, so that the time segmentation granularity corresponding to each time period can better adapt to the business requirements and user interaction patterns of that time period, and improve the accuracy of segmenting the structured information in the time dimension.

[0095] Step 208: Input the information set into the large model to obtain the question-and-answer pairs generated by the large model.

[0096] Among them, for the specific implementation form of step 208, reference can be made to the detailed descriptions in other embodiments of the present disclosure, and details will not be described here.

[0097] In the embodiments of the present disclosure, first, the obtained multimodal data is parsed to determine the key information in the multimodal data, and based on the key information, the structured information corresponding to the multimodal data is generated. Then, according to the time period to which the moment when the multimodal data is generated belongs, the time segmentation granularity of the current information is determined, and the structured information is semantically clustered to determine the topic distribution information corresponding to the structured information. Based on the topic distribution information, the semantic segmentation granularity of the current information is determined. After that, according to the number of entities and / or the amount of data included in the structured information, the quantity segmentation granularity of the current information is determined, and based on the segmentation granularity of the current information, the structured information is segmented to obtain multiple information sets. Finally, the information sets are input into the large model to obtain the question-and-answer pairs generated by the large model. Thus, after generating the structured information corresponding to the multimodal data, the time segmentation granularity, semantic segmentation granularity, and quantity segmentation granularity of the current information are determined, and based on the multi-dimensional segmentation granularity, the structured information is segmented, improving the accuracy and reliability of analyzing and understanding the multimodal data, and providing conditions for improving the quality and accuracy of the generated question-and-answer pairs.

[0098] Figure 3 It is a schematic flowchart of a method for generating question-and-answer pair data provided by another embodiment of the present disclosure.

[0099] As Figure 3 shown, the method for generating question-and-answer pair data includes:

[0100] Step 301: Parse the obtained multimodal data to determine the key information in the multimodal data.

[0101] Step 302: Generate structured information corresponding to the multimodal data according to the key information.

[0102] Step 303: Determine the segmentation granularity of the current information.

[0103] Among them, the specific implementation forms of steps 301 to 303 can refer to the detailed descriptions in other embodiments of the present disclosure, and will not be specifically elaborated here.

[0104] Step 304: Segment the structured information respectively based on the segmentation granularity corresponding to each dimension to obtain multiple information sets.

[0105] In the present disclosure, after determining the segmentation granularity of the current information, the structured information can be segmented separately based on the segmentation granularity corresponding to the time dimension, semantic dimension, and quantity dimension respectively to obtain the information sets corresponding to each dimension.

[0106] Step 305: When the content included in at least two of the multiple information sets is the same, retain any one of the at least two information sets.

[0107] In the present disclosure, after obtaining the multiple information sets, when the content included in at least two of the multiple information sets is the same, in order to avoid data redundancy and improve the efficiency of generating question-and-answer pair data, any one of the at least two information sets can be retained.

[0108] In some possible implementation forms, before segmenting the structured information respectively based on the segmentation granularity corresponding to each dimension, it is also possible to determine whether the information sets obtained for each dimension are the same based on the segmentation granularity corresponding to each dimension. If they are the same, the structured information can be segmented only based on the segmentation granularity corresponding to any one dimension. The present disclosure does not limit this.

[0109] Step 306: Input the information set into the large model to obtain the question-and-answer pairs generated by the large model.

[0110] Among them, the specific implementation form of step 306 can refer to the detailed descriptions in other embodiments of the present disclosure, and will not be specifically elaborated here.

[0111] In the embodiments of the present disclosure, first, the obtained multimodal data is parsed to determine the key information in the multimodal data, and based on the key information, structured information corresponding to the multimodal data is generated. Then, the segmentation granularity of the current information is determined, and the structured information is segmented based on the segmentation granularity corresponding to each dimension to obtain a plurality of information sets. After that, when the content included in at least two of the plurality of information sets is the same, any one of the at least two information sets is retained. Finally, the information sets are input into a large model to obtain the question-and-answer pairs generated by the large model. Thus, after generating the structured information corresponding to the multimodal data, based on the segmentation granularity of each dimension corresponding to the current information, the structured information is segmented, and when there are information sets with the same content among the plurality of information sets obtained after segmentation, only one of the information sets with the same content is retained, thereby effectively avoiding data redundancy and providing conditions for improving the efficiency of generating question-and-answer pairs.

[0112] Figure 4 It is a schematic flowchart of a method for generating question-and-answer pair data provided in another embodiment of the present disclosure.

[0113] As Figure 4 shown, the method for generating question-and-answer pair data includes:

[0114] Step 401, parse the obtained multimodal data to determine the key information in the multimodal data.

[0115] Step 402, generate structured information corresponding to the multimodal data based on the key information.

[0116] Step 403, determine the segmentation granularity of the current information.

[0117] Step 404, segment the structured information based on the segmentation granularity of the current information to obtain a plurality of information sets.

[0118] Among them, the specific implementation forms of steps 401 to 404 can refer to the detailed descriptions in other embodiments of the present disclosure, and will not be specifically elaborated here.

[0119] Step 405, determine the prompt template according to the source of the multimodal data.

[0120] Among them, the prompt template can be used for the large model to generate question-and-answer pairs. Its format and content can be preset, or can also be determined according to actual needs. For example, the prompt template can be: {Q: [question], A: [answer], Origion: [original conversation], Questioner: [question asker], Resolver: [question solver], Participant: [participants in the discussion], IsResolved: [whether it is resolved]}, where the content in "[]" is the content to be generated by the large model, and the present disclosure does not limit this.

[0121] In the present disclosure, after splitting the structured information to obtain multiple information sets, the information sets can be input into the large model to obtain the question-and-answer pairs generated by the large model. In order to improve the quality and standardization of the question-and-answer pairs generated by the large model, before inputting the information sets into the large model, the prompt template can be determined first based on the source of the multimodal data.

[0122] Step 406, based on the prompt template, input the information sets into the large model to obtain the question-and-answer pairs generated by the large model.

[0123] In the present disclosure, after determining the prompt template, the information sets can be input into the large model based on the prompt template, thereby obtaining the question-and-answer pairs generated by the large model, improving the quality and standardization of the generated question-and-answer pairs.

[0124] In the embodiments of the present disclosure, first, the obtained multimodal data is parsed to determine the key information in the multimodal data, and the structured information corresponding to the multimodal data is generated according to the key information. Then, the segmentation granularity of the current information is determined, and based on the segmentation granularity of the current information, the structured information is segmented to obtain multiple information sets. Finally, the prompt template is determined according to the source of the multimodal data, and based on the prompt template, the information sets are input into the large model to obtain the question-and-answer pairs generated by the large model. Thus, after generating the structured information corresponding to the multimodal data, the structured information is segmented according to the segmentation granularity of the current information, and the prompt template is determined according to the source of the multimodal data. Based on the prompt template, the obtained information sets after segmentation are output to the large model to obtain the question-and-answer pairs generated by the large model, thereby improving the performance of the large model, improving the quality of the generated question-and-answer pairs, and providing conditions for enhancing the user experience.

[0125] Figure 5 It is a schematic flowchart of a method for generating question-and-answer pair data provided by another embodiment of the present disclosure.

[0126] As Figure 5 shown, the method for generating question-and-answer pair data includes:

[0127] Step 501, parse the obtained multimodal data to determine the key information in the multimodal data.

[0128] Step 502: Generate structured information corresponding to the multimodal data according to the key information.

[0129] Step 503: Determine the segmentation granularity of the current information.

[0130] Step 504: Based on the segmentation granularity of the current information, segment the structured information to obtain multiple information sets.

[0131] Step 505: Input the information sets into the large model to obtain the Q&A pairs generated by the large model.

[0132] Among them, for the specific implementation forms of steps 501 to 505, reference can be made to the detailed descriptions in other embodiments of the present disclosure, and details will not be elaborated here.

[0133] Step 506: When the confidence level of the Q&A pair is less than the confidence threshold, determine the conflict reason corresponding to the Q&A pair.

[0134] Among them, the confidence threshold can be the confidence critical value for judging the credibility of the Q&A pair. It can be preset or determined according to actual needs. The present disclosure does not make any limitations thereto.

[0135] Among them, the conflict reason can include whether the Q&A pair conflicts with the reference knowledge in the knowledge base, whether the Q&A pair answers off-topic, whether the Q&A pair is semantically incoherent, etc. The present disclosure does not make any limitations thereto.

[0136] It should be noted that when generating Q&A pairs, the large model can also evaluate the confidence level of the generated Q&A pairs from multiple dimensions such as the source authority and relevance of the multimodal data and whether the Q&A pair conflicts with the content in the knowledge base, so as to improve the accuracy and reliability of the confidence level of the Q&A pairs.

[0137] In the present disclosure, after obtaining the Q&A pairs generated by the large model, when the confidence level of the Q&A pair is less than the confidence threshold, the large model can be used to detect conflicts in the Q&A pair and determine the conflict reason corresponding to the Q&A pair, thereby providing conditions for updating and correcting the Q&A pair.

[0138] Step 507: Update the Q&A pair based on the conflict reason to obtain the updated Q&A pair.

[0139] In the present disclosure, after determining the conflict reason corresponding to the Q&A pair, the Q&A pair can be updated based on the conflict reason to obtain the updated Q&A pair, thereby improving the quality of the Q&A pair.

[0140] In some possible implementation forms, when updating the Q&A pair based on the conflict reason to obtain the updated Q&A pair, the conflict reason and the Q&A pair can be input into the large model to obtain the Q&A pair updated by the large model. For example, when the conflict reason is that the Q&A pair misses the point or the semantics of the Q&A pair is not smooth, the conflict reason and the Q&A pair can be input into the large model, and the large model can be used to correct and update the Q&A pair, etc., so as to automatically correct the Q&A pair by the large model, effectively reducing the workload of manual review and the operation and maintenance cost.

[0141] In some possible implementation forms, when updating the Q&A pair based on the conflict reason to obtain the updated Q&A pair, the Q&A pair can also be sent to the associated business personnel based on the conflict reason to obtain the updated Q&A pair returned by the business personnel. For example, when the conflict reason is that the Q&A pair conflicts with the reference knowledge in the knowledge base, at this time, the Q&A pair can be sent to the associated business personnel for manual review to determine whether the Q&A pair is accurate, so as to effectively avoid the situation of incorrect judgment by the large model and improve the accuracy and efficiency of Q&A pair data generation.

[0142] In some possible implementation forms, when updating the Q&A pair based on the conflict reason to obtain the updated Q&A pair, when the conflict reason indicates that the Q&A pair conflicts with at least one reference knowledge, the Q&A pair can also be updated based on the source of the at least one reference knowledge and the source and quantity of the reference knowledge that does not conflict with the Q&A pair to obtain the updated Q&A pair. For example, when the source of the reference knowledge that does not conflict with the Q&A pair is more authoritative and the quantity is larger, the Q&A pair can remain unchanged; when the source of the reference knowledge that conflicts with the Q&A pair is more authoritative and the quantity is larger, the Q&A pair can be updated based on the conflicting reference knowledge to obtain the updated Q&A pair, thereby improving the accuracy of the updated Q&A pair. The present disclosure does not limit this.

[0143] In the present disclosure, when updating the Q&A pair based on the conflict reason, it can be updated by the large model, or manually reviewed and updated, or updated according to the source and quantity of the reference knowledge in the knowledge base, so as to update and correct the Q&A pair by multiple methods, ensuring the accuracy and reliability of the updated Q&A pair and improving the flexibility of the Q&A pair data generation method.

[0144] In the embodiments of the present disclosure, first, the obtained multimodal data is parsed to determine the key information in the multimodal data, and based on the key information, structured information corresponding to the multimodal data is generated. Then, the segmentation granularity of the current information is determined, and based on the segmentation granularity of the current information, the structured information is segmented to obtain multiple information sets. After that, the information sets are input into a large model to obtain the Q&A pairs generated by the large model. When the confidence level of the Q&A pair is less than the confidence threshold, the conflict reason corresponding to the Q&A pair is determined. Finally, based on the conflict reason, the Q&A pair is updated to obtain the updated Q&A pair. Thus, after using the segmentation granularity of the current information to segment the structured information corresponding to the multimodal data to obtain information sets, based on the information sets, a large model is used to generate Q&A pairs, and the confidence level of the Q&A pairs is evaluated. When the confidence level of the Q&A pair is low, the large model is used to detect conflicts in the Q&A pair to obtain the conflict reason, and based on the conflict reason, the Q&A pair is updated, thereby improving the quality and reliability of the obtained Q&A pairs.

[0145] Figure 6 FIG. is a schematic flowchart of a method for generating Q&A pair data provided by another embodiment of the present disclosure.

[0146] As Figure 6 shown, the method for generating Q&A pair data includes:

[0147] Step 601, parse the obtained multimodal data to determine the key information in the multimodal data.

[0148] Step 602, generate structured information corresponding to the multimodal data based on the key information.

[0149] Step 603, determine the segmentation granularity of the current information.

[0150] Step 604, based on the segmentation granularity of the current information, segment the structured information to obtain multiple information sets.

[0151] Step 605, input the information sets into a large model to obtain the Q&A pairs generated by the large model.

[0152] Step 606, when the confidence level of the Q&A pair is less than the confidence threshold, determine the conflict reason corresponding to the Q&A pair.

[0153] Among them, for the specific implementation forms of steps 601 to 606, reference can be made to the detailed descriptions in other embodiments of the present disclosure, and no further details will be provided here.

[0154] Step 607, in the case that the conflict reason indication Q&A pair conflicts with at least one reference knowledge, determine the first weight of at least one reference knowledge and the second weight of the reference knowledge that does not conflict, based on the source of at least one reference knowledge, and the source and quantity of the reference knowledge that does not conflict with the Q&A pair.

[0155] In the present disclosure, after determining the conflict reason corresponding to the Q&A pair, in the case that the conflict reason indication Q&A pair conflicts with at least one reference knowledge, it can be determined that there are currently multiple versions of answers corresponding to the question in this Q&A pair. At this time, in order to determine a more accurate and reliable answer, the source of the reference knowledge that conflicts with the Q&A pair, and the source and quantity of the reference knowledge that does not conflict with the Q&A pair can be determined first, and then based on the authority of the source of the conflicting reference knowledge, determine the first weight of the conflicting reference knowledge, and based on the source and quantity of the non-conflicting reference knowledge, determine the second weight of the non-conflicting reference knowledge.

[0156] It should be noted that when determining the first weight of at least one reference knowledge and the second weight of the reference knowledge that does not conflict, based on the source of at least one reference knowledge, and the source and quantity of the reference knowledge that does not conflict with the Q&A pair, the weight can be determined based on the authority and quantity of the source. The higher the authority of the source of the reference knowledge and the larger the quantity, the higher the accuracy and reliability of the reference knowledge can be indicated, and the corresponding weight is also larger. The present disclosure does not limit this.

[0157] Step 608, in the case that the first weight is greater than the second weight, update the Q&A pair based on at least one reference knowledge to obtain an updated Q&A pair.

[0158] In the present disclosure, in the case that the first weight is greater than the second weight, it can be determined that the weight of the reference knowledge that conflicts with the Q&A pair is greater than the weight of the reference knowledge that does not conflict with the Q&A pair. The accuracy and quality of the answer in the Q&A pair generated by the large model may be low. At this time, the Q&A pair can be updated and corrected based on at least one reference knowledge that conflicts with the Q&A pair to improve the reliability and quality of the Q&A pair.

[0159] In some possible implementation forms, in the case that the second weight is greater than the first weight, it can be determined that the weight of the reference knowledge that does not conflict with the Q&A pair is greater than the weight of the reference knowledge that conflicts with the Q&A pair. The accuracy and quality of the answer in the Q&A pair are high. At this time, the Q&A pair can be kept unchanged.

[0160] It should be noted that when the first weight is equal to the second weight, the specific processing method for the question-and-answer pair can be preset or determined according to actual needs. For example, when the first weight is equal to the second weight, it can be submitted to manual review, or the question-and-answer pair can be kept unchanged, or the question-and-answer pair can be updated based on the reference knowledge conflicting with the question-and-answer pair, etc. The present disclosure does not limit this.

[0161] The following combines Figure 7 to illustrate by way of example the process of generating and updating question-and-answer pairs using a large model in the question-and-answer pair data generation method proposed in the embodiments of the present disclosure. Figure 7 It is a schematic flowchart of generating and updating question-and-answer pairs in the question-and-answer pair data generation method proposed in the embodiments of the present disclosure.

[0162] Figure 7 In

[0163] For example Figure 7 As shown, after obtaining the segmented information set, the information set can be stored in a list containing the historical information set, the list can be incrementally updated, and then the information set can be input into the large model to obtain the question-and-answer pairs generated by the large model, and the large model can be used to evaluate the confidence of the question-and-answer pairs.

[0164] When the confidence of the question-and-answer pair is greater than or equal to the confidence threshold, the question-and-answer pair is automatically added to the knowledge base. When the confidence of the question-and-answer pair is less than the confidence threshold, the large model can be used to detect conflicts of the question-and-answer pair. When there is a conflict between the question-and-answer pair and the reference knowledge in the knowledge base, the question-and-answer pair can be sent to the associated business personnel for manual review and update. Or when there is no conflict between the question-and-answer pair and the reference knowledge in the knowledge base, the question-and-answer pair can be updated using the large model based on the conflict reason (such as answering off-topic, etc.) and the question-and-answer pair.

[0165] After that, the updated question-and-answer pairs are added to the knowledge base, and the knowledge in the knowledge base is incrementally updated, thereby improving the generation quality of the question-and-answer pair data, realizing the automatic update of the knowledge base, and effectively ensuring the timeliness of the knowledge base.

[0166] In the embodiments of the present disclosure, first, the obtained multimodal data is parsed to determine the key information in the multimodal data, and based on the key information, structured information corresponding to the multimodal data is generated. Then, the segmentation granularity of the current information is determined, and based on the segmentation granularity of the current information, the structured information is segmented to obtain multiple information sets. After that, the information sets are input into a large model to obtain question-and-answer pairs generated by the large model. When the confidence level of the question-and-answer pair is less than the confidence threshold, the conflict reason corresponding to the question-and-answer pair is determined. When the conflict reason indicates that the question-and-answer pair conflicts with at least one reference knowledge, based on the source of the at least one reference knowledge and the source and quantity of the reference knowledge that does not conflict with the question-and-answer pair, the first weight of the at least one reference knowledge and the second weight of the non-conflicting reference knowledge are determined. Finally, when the first weight is greater than the second weight, the question-and-answer pair is updated based on the at least one reference knowledge to obtain an updated question-and-answer pair. Thus, after obtaining the segmented information sets corresponding to the multimodal data and inputting the information sets into the large model to generate question-and-answer pairs, when the confidence level of the question-and-answer pair is low, the large model is used to detect conflicts in the question-and-answer pair. When the question-and-answer pair conflicts with the reference knowledge in the knowledge base, based on the source of the conflicting reference knowledge, the weight of the conflicting reference knowledge is determined, and based on the source and quantity of the non-conflicting reference knowledge, the weight of the non-conflicting reference knowledge is determined. When the weight of the conflicting reference knowledge is greater than the weight of the non-conflicting reference knowledge, the question-and-answer pair is updated based on the conflicting reference knowledge, thereby improving the quality and accuracy of the finally obtained question-and-answer pair.

[0167] Next, in conjunction with Figure 8 , an example is given to illustrate the overall architecture of the question-and-answer pair data generation method proposed in the embodiments of the present disclosure. Figure 8 It is a schematic diagram of the architecture of the question-and-answer pair data generation method proposed in the embodiments of the present disclosure.

[0168] As Figure 8 shown, the architecture of the question-and-answer pair data generation method includes a knowledge extraction domain, a knowledge optimization domain, and a knowledge evolution domain.

[0169] In the present disclosure, after obtaining the multimodal data, first, through the knowledge extraction domain, the multimodal data is parsed to obtain the key information of the multimodal data. Based on the key information, a knowledge graph is constructed and stored, and then based on the knowledge graph, structured information corresponding to the multimodal data is generated.

[0170] Then, through the knowledge optimization domain, multi-dimensional (time, semantics, quantity) analysis is performed on the structured information, the segmentation granularity (such as the time segmentation granularity) is adjusted, the current time segmentation granularity, semantic segmentation granularity, and quantity segmentation granularity corresponding to the structured information are determined, and based on the segmentation granularity of each dimension, the structured information is segmented separately to obtain the segmented information sets of the structured information.

[0171] Finally, through the knowledge evolution domain, the information is input into the large model to generate question-answer pairs, and the large model is used to evaluate the confidence of the question-answer pairs. When the confidence of the question-answer pairs meets the standard (such as the confidence is greater than or equal to the confidence threshold), the question-answer pairs are automatically added to the knowledge base. When the confidence of the question-answer pairs does not meet the standard (such as the confidence is less than the confidence threshold), the large model is used to trigger conflict detection and process the conflicts of the question-answer pairs. Thus, the question-answer pairs can be corrected and updated based on the conflict reasons, improving the accuracy and reliability of the question-answer pair data generation method.

[0172] Therefore, by generating question-answer pairs through the architecture of the question-answer pair data generation method proposed in this disclosure, evaluating the confidence of the question-answer pairs, automatically storing the question-answer pairs with qualified confidence, and correcting, updating and then storing the question-answer pairs with unqualified confidence, not only can the workload of manual work be effectively reduced, the efficiency and quality of question-answer pair generation be improved, but also the automatic update of the knowledge base be realized, the timeliness of the knowledge base be improved, and thus the question-answer quality be ensured and the user experience be enhanced.

[0173] To implement the above embodiments, an embodiment of this disclosure further provides a question-answer pair data generation device.

[0174] Figure 9 The following is a schematic structural diagram of the question-answer pair data generation device provided by an embodiment of this disclosure.

[0175] As Figure 9 shown, the question-answer pair data generation device 900 includes: a first determination module 901, a generation module 902, a second determination module 903, a segmentation module 904, and a processing module 905.

[0176] The first determination module 901 is configured to parse the acquired multimodal data and determine the key information in the multimodal data;

[0177] The generation module 902 is configured to generate structured information corresponding to the multimodal data according to the key information;

[0178] The second determination module 903 is configured to determine the segmentation granularity of the current information;

[0179] The segmentation module 904 is configured to segment the structured information based on the segmentation granularity of the current information to obtain a plurality of information sets;

[0180] The processing module 905 is configured to input the information sets into the large model to obtain question-answer pairs generated by the large model.

[0181] Optionally, the second determination module 903 is specifically configured to:

[0182] Determine the time segmentation granularity of the current information according to the time period to which the moment when the multimodal data is generated belongs.

[0183] Optionally, the above-mentioned segmentation module 904 is further configured to:

[0184] Determine the description information corresponding to each time period according to multiple information sets, where the description information includes at least one of the following: the total number of information sets, the amount of data included in each information set, and the number of entities included in each information set;

[0185] Update the time segmentation granularity corresponding to each time period according to the description information corresponding to each time period.

[0186] Optionally, the above-mentioned second determination module 903 is specifically configured to:

[0187] Perform semantic clustering on the structured information to determine the topic distribution information corresponding to the structured information;

[0188] Determine the semantic segmentation granularity of the current information based on the topic distribution information.

[0189] Optionally, the above-mentioned second determination module 903 is specifically configured to:

[0190] Determine the quantity segmentation granularity of the current information according to the number of entities and / or the amount of data included in the structured information.

[0191] Optionally, the above-mentioned second determination module 903 is specifically configured to:

[0192] Determine the segmentation granularity of the current information according to the source of the multimodal data.

[0193] Optionally, the above-mentioned segmentation module 904 is specifically configured to:

[0194] Segment the structured information based on the segmentation granularity corresponding to each dimension respectively to obtain multiple information sets;

[0195] In the case where the content included in at least two of the multiple information sets is the same, retain any one of the at least two information sets.

[0196] Optionally, the above-mentioned processing module 905 is specifically configured to:

[0197] Determine the prompt template according to the source of the multimodal data;

[0198] Input the information set into the large model based on the prompt template to obtain the question-and-answer pairs generated by the large model.

[0199] Optionally, the above-mentioned processing module 905 is further configured to:

[0200] When the confidence of the question-and-answer pair is less than the confidence threshold, determine the conflict reason corresponding to the question-and-answer pair;

[0201] Based on the conflict reason, update the question-and-answer pair to obtain the updated question-and-answer pair.

[0202] Optionally, the above processing module 905 is further configured to perform any of the following:

[0203] Input the conflict reason and the question-and-answer pair into the large model to obtain the updated question-and-answer pair of the large model;

[0204] Based on the conflict reason, send the question-and-answer pair to the associated business personnel to obtain the updated question-and-answer pair returned by the business personnel;

[0205] When the conflict reason indicates that the question-and-answer pair conflicts with at least one reference knowledge, update the question-and-answer pair based on the source of at least one reference knowledge and the source and quantity of the reference knowledge that does not conflict with the question-and-answer pair to obtain the updated question-and-answer pair.

[0206] Optionally, the above processing module 905 is further configured to:

[0207] Based on the source of at least one reference knowledge and the source and quantity of the reference knowledge that does not conflict with the question-and-answer pair, determine the first weight of at least one reference knowledge and the second weight of the reference knowledge that does not conflict;

[0208] When the first weight is greater than the second weight, update the question-and-answer pair based on at least one reference knowledge to obtain the updated question-and-answer pair.

[0209] It should be noted that the explanatory description of the foregoing embodiments of the question-and-answer pair data generation method also applies to the question-and-answer pair data generation device of this embodiment, so it will not be repeated here.

[0210] In the embodiments of the present disclosure, first, the obtained multimodal data is parsed to determine the key information in the multimodal data, then according to the key information, the structured information corresponding to the multimodal data is generated, then the segmentation granularity of the current information is determined, and based on the segmentation granularity of the current information, the structured information is segmented to obtain a plurality of information sets, and finally the information sets are input into the large model to obtain the question-and-answer pairs generated by the large model. Thus, by parsing the multimodal data, determining the corresponding key information, generating structured information according to the key information, segmenting the structured information based on the current segmentation granularity, and inputting the obtained information sets after segmentation into the large model to obtain the question-and-answer pairs, the quality and efficiency of the question-and-answer pair data generation are improved, providing conditions for realizing the automatic update of the knowledge base and improving the timeliness of the knowledge base.

[0211] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0212] Figure 10 FIG. shows a schematic block diagram of an exemplary electronic device 1000 that can be used to implement embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, for example, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, for example, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0213] As Figure 10 shown, the device 1000 includes a computing unit 1001 that can perform various appropriate actions and processes according to a computer program stored in a ROM (Read-Only Memory) 1002 or a computer program loaded from a storage unit 1008 into a RAM (Random Access Memory) 1003. In the RAM 1003, various programs and data required for the operation of the device 1000 can also be stored. The computing unit 1001, the ROM 1002, and the RAM 1003 are connected to each other via a bus 1004. An I / O (Input / Output) interface 1005 is also connected to the bus 1004.

[0214] A plurality of components in the device 1000 are connected to the I / O interface 1005, including: an input unit 1006, such as a keyboard, a mouse, etc.; an output unit 1007, such as various types of displays, speakers, etc.; a storage unit 1008, such as a magnetic disk, an optical disk, etc.; and a communication unit 1009, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 1009 allows the device 1000 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0215] The computing unit 1001 can be various general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 1001 include, but are not limited to, CPU (Central Processing Unit), GPU (Graphic Processing Units), various dedicated AI (Artificial Intelligence) computing chips, various computing units running machine learning model algorithms, DSP (Digital Signal Processor), and any suitable processor, controller, microcontroller, etc. The computing unit 1001 executes the various methods and processes described above, such as the Q&A pair data generation method. For example, in some embodiments, the Q&A pair data generation method can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 1008. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 1000 via the ROM 1002 and / or the communication unit 1009. When the computer program is loaded into the RAM 1003 and executed by the computing unit 1001, one or more steps of the Q&A pair data generation method described above can be executed. Alternatively, in other embodiments, the computing unit 1001 can be configured to execute the Q&A pair data generation method by any other suitable means (e.g., by means of firmware).

[0216] Various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuitry, integrated circuit systems, FPGA (Field Programmable Gate Array), ASIC (Application-Specific Integrated Circuit), ASSP (Application Specific Standard Product), SOC (System On Chip), CPLD (Complex Programmable Logic Device), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special or general programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0217] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing devices, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program codes can be executed entirely on the machine, partially on the machine, executed partially on the machine as an independent software package and partially on a remote machine, or executed entirely on a remote machine or server.

[0218] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media would include electrical connections based on one or more wires, portable computer disks, hard disks, RAM, ROM, EPROM (Electrically Programmable Read-Only-Memory), or flash memory, optical fibers, CD-ROM (Compact Disc Read-Only Memory), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0219] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (Cathode-Ray Tube) or an LCD (Liquid Crystal Display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0220] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: LAN (Local Area Network), WAN (Wide Area Network), the Internet, and blockchain networks.

[0221] A computer system can include a client and a server. The client and the server are generally far from each other and usually interact through a communication network. The client-server relationship is created by computer programs running on respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services (Virtual Private Server). The server can also be a server of a distributed system, or a server combined with blockchain.

[0222] According to an embodiment of the present disclosure, the present disclosure also provides a computer program product, which, when executed by an instruction processor in the computer program product, executes the Q&A pair data generation method proposed in the above embodiments of the present disclosure.

[0223] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in the present disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in the present disclosure can be achieved. No limitation is imposed herein.

[0224] The above specific embodiments do not constitute a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present disclosure shall be included within the protection scope of the present disclosure.

Claims

1. A method for generating question-and-answer pair data, characterized in that, Including: Parsing the obtained multimodal data to determine the key information in the multimodal data; Generating structured information corresponding to the multimodal data according to the key information; Determining the segmentation granularity of the current information; Based on the segmentation granularity of the current information, segmenting the structured information to obtain multiple information sets; Inputting the information sets into a large model to obtain question-answer pairs generated by the large model.

2. The method according to claim 1, wherein The determining the segmentation granularity of the current information includes: Determining the time segmentation granularity of the current information according to the time period to which the moment when the multimodal data is generated belongs.

3. The method according to claim 2, characterized in that, After obtaining the multiple information sets, it further includes: Determining the description information corresponding to each time period according to the multiple information sets, where the description information includes at least one of the following: the total number of information sets, the amount of data contained in each information set, and the number of entities contained in each information set; Updating the time segmentation granularity corresponding to each time period according to the description information corresponding to each time period.

4. The method according to claim 1, characterized in that, The determining the segmentation granularity of the current information includes: Performing semantic clustering on the structured information to determine the topic distribution information corresponding to the structured information; Based on the topic distribution information, determining the semantic segmentation granularity of the current information.

5. The method according to claim 1, wherein The determining the segmentation granularity of the current information includes: Determining the quantity segmentation granularity of the current information according to the number of entities and / or the amount of data contained in the structured information.

6. The method according to claim 1, wherein, The determining the segmentation granularity of the current information includes: Determining the segmentation granularity of the current information according to the source of the multimodal data.

7. The method according to any one of claims 1-6, characterized in that, The segmenting the structured information based on the segmentation granularity of the current information to obtain multiple information sets includes: Respectively segmenting the structured information based on the segmentation granularity corresponding to each dimension to obtain multiple information sets; When the content contained in at least two information sets among the multiple information sets is the same, retaining any one of the at least two information sets.

8. The method according to any one of claims 1-6, characterized in that, The inputting the information sets into a large model to obtain question-answer pairs generated by the large model includes: Determining a prompt template according to the source of the multimodal data; Based on the prompt template, inputting the information sets into a large model to obtain question-answer pairs generated by the large model.

9. The method according to any one of claims 1-6, characterized in that, After the inputting the information sets into a large model to obtain question-answer pairs generated by the large model, it further includes: When the confidence level of the question-answer pair is less than the confidence threshold, determining the conflict reason corresponding to the question-answer pair; Updating the question-answer pair based on the conflict reason to obtain an updated question-answer pair.

10. The method according to claim 9, wherein The updating the question-answer pair based on the conflict reason to obtain an updated question-answer pair includes any one of the following: Inputting the conflict reason and the question-answer pair into the large model to obtain an updated question-answer pair generated by the large model; Based on the conflict reason, sending the question-answer pair to the associated business personnel to obtain an updated question-answer pair returned by the business personnel; In the case where the conflict reason indicates that the Q&A pair conflicts with at least one reference knowledge, update the Q&A pair based on the source of the at least one reference knowledge, and the source and quantity of the reference knowledge that does not conflict with the Q&A pair, so as to obtain an updated Q&A pair.

11. The method according to claim 10, wherein The updating the Q&A pair based on the source of the at least one reference knowledge, and the source and quantity of the reference knowledge that does not conflict with the Q&A pair, so as to obtain an updated Q&A pair, includes: Determine a first weight of the at least one reference knowledge and a second weight of the non-conflicting reference knowledge based on the source of the at least one reference knowledge, and the source and quantity of the reference knowledge that does not conflict with the Q&A pair; In the case where the first weight is greater than the second weight, update the Q&A pair based on the at least one reference knowledge, so as to obtain an updated Q&A pair.

12. A question-and-answer pair data generation device, characterized in that Includes: A first determination module, configured to parse the acquired multimodal data and determine key information in the multimodal data; A generation module, configured to generate structured information corresponding to the multimodal data according to the key information; A second determination module, configured to determine the segmentation granularity of the current information; A segmentation module, configured to segment the structured information based on the segmentation granularity of the current information to obtain a plurality of information sets; A processing module, configured to input the information sets into a large model to obtain Q&A pairs generated by the large model.

13. An electronic device, comprising: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 1-11.

14. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method according to any one of claims 1-11.

15. A computer program product, comprising a computer program, where the computer program, when executed by a processor, implements the steps of the method according to any one of claims 1-11.