Video conversation Q&A data generation method, device, electronic device and medium

By determining the target video description information and prompt words in the video conversation question and answer data generation, and automatically generating question and answer data using a large language model, the problem of manual construction is solved, and fast and accurate question and answer data generation is achieved.

CN116842216BActive Publication Date: 2025-07-22BEIJING YOUZHUJU NETWORK TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310813786.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-07-04
Publication Date
2025-07-22
Estimated Expiration
2043-07-04

AI Technical Summary

Technical Problem

In the prior art, the construction of video dialogue Q&A data requires manual and complete viewing of videos, which takes a long time and is difficult to ensure accuracy.

Method used

By determining the target video description information and the target prompt words, the Q&A model based on a large language model is automatically generated, and the Q&A model is guided to generate Q&A according to the video description information.

Benefits of technology

It realizes the rapid and accurate generation of video conversation Q&A data, reduces manual construction time, and improves data accuracy and diversity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116842216B_ABST
    Figure CN116842216B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure disclose a method, apparatus, electronic device, and medium for generating video dialogue Q&A data. The method includes: determining target video description information corresponding to a target video; determining target prompt words adopted by a target Q&A model, where the target Q&A model is pre-configured based on a large language model, and the target prompt words can guide the target Q&A model to approach the desired dialogue Q&A effect according to the target video description information when performing a dialogue Q&A generation task; and outputting dialogue Q&A data associated with the target video through the target Q&A model based on the target video description information and the target prompt words. This solution solves the problem that it takes a long time to construct dialogue Q&A data because it requires manual viewing of the entire video before manually constructing the dialogue Q&A data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to the technical field of video processing, and in particular, to a method, apparatus, electronic device, and medium for generating video dialogue question-and-answer data. Background Art

[0002] With the development of video content understanding, video dialogue question-and-answer has become a relatively important technology. Video dialogue question-and-answer refers to parsing the question answer according to the input video and the question about the video.

[0003] The basis of video dialogue question-and-answer is the annotation of video dialogue question-and-answer data. At present, the construction of video dialogue question-and-answer data is mainly achieved through manual construction based on the video. However, when manually constructing video dialogue question-and-answer data, it is necessary to watch the entire video before starting to write a description of the video. Since videos often range from a few minutes to several hours, it is very time-consuming. At the same time, it is more difficult to describe videos, and a certain understanding of the field of video content is required to accurately describe the video content, which may lead to poor quality of the constructed video dialogue question-and-answer data and affect subsequent video dialogue question-and-answer. Summary of the Invention

[0004] Embodiments of the present disclosure provide a method, apparatus, electronic device, and medium for generating video dialogue question-and-answer data, which solve the problem of being unable to quickly generate accurate video dialogue question-and-answer data.

[0005] In a first aspect, embodiments of the present disclosure provide a method for generating video dialogue question-and-answer data, the method including:

[0006] Determine the target video description information corresponding to the target video;

[0007] Determine the target prompt word adopted by the target question-and-answer model, where the target question-and-answer model is pre-configured based on a large language model, and the target prompt word can guide the target question-and-answer model to the desired dialogue question-and-answer effect according to the target video description information when performing the dialogue question-and-answer generation task;

[0008] Based on the target video description information and the target prompt word, output the dialogue question-and-answer data associated with the target video through the target question-and-answer model.

[0009] In a second aspect, embodiments of the present disclosure further provide a device for generating video dialogue question-and-answer data, the device including:

[0010] A description information determination module, configured to determine the target video description information corresponding to the target video;

[0011] A prompt information determination module, configured to determine a target prompt word adopted by a target question-and-answer model, where the target question-and-answer model is pre-configured based on a large language model, and the target prompt word can guide the target question-and-answer model to approach an expected dialogue question-and-answer effect according to the target video description information when performing a dialogue question-and-answer generation task;

[0012] A data generation module, configured to output dialogue question-and-answer data associated with the target video through the target question-and-answer model based on the target video description information and the target prompt word.

[0013] In a third aspect, an electronic device is further provided in the embodiments of the present disclosure. The electronic device includes:

[0014] At least one processor; and

[0015] A memory communicatively connected to the at least one processor; wherein,

[0016] The memory stores a computer program executable by the at least one processor, and when the computer program is executed by the at least one processor, the at least one processor is enabled to execute the video dialogue question-and-answer data generation method according to any one of the above embodiments.

[0017] In a fourth aspect, a computer-readable medium is further provided in the embodiments of the present disclosure. The computer-readable medium stores computer instructions, and when the computer instructions are executed by a processor, the video dialogue question-and-answer data generation method according to any one of the above embodiments is implemented.

[0018] In the technical solution of the embodiments of the present disclosure, when constructing video dialogue question-and-answer data, the target video description information and the target prompt word that can guide the target question-and-answer model to approach the expected dialogue question-and-answer effect according to the target video description information when performing the dialogue question-and-answer generation task are determined. In this way, the target question-and-answer model pre-configured based on the large language model can be used to automatically generate dialogue question-and-answer data under the dialogue question-and-answer guidance of the target prompt word according to the target video description information, solving the problem that it takes a long time to construct dialogue question-and-answer data manually after watching the video completely. The construction can be carried out in real time by analyzing the video in real time to obtain the video description information. At the same time, since the prompt word is used to guide the question-and-answer model to construct dialogue question-and-answer data during the construction process, the accuracy of the constructed dialogue question-and-answer data can be guaranteed to a certain extent.

[0019] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. Description of the Drawings

[0020] In conjunction with the accompanying drawings and with reference to the following specific embodiments, the above and other features, advantages, and aspects of the various embodiments of the present disclosure will become more apparent. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic and that the original and the elements are not necessarily drawn to scale.

[0021] Figure 1 It is a flowchart of a method for generating video dialogue question-and-answer data provided by an embodiment of the present disclosure;

[0022] Figure 2 It is a schematic diagram showing the display of video dialogue question-and-answer data on a display interface applicable to an embodiment of the present disclosure;

[0023] Figure 3 It is a block diagram of a structure of a device for generating video dialogue question-and-answer data provided by an embodiment of the present disclosure;

[0024] Figure 4 It is a block diagram of a structure of an electronic device for implementing the method for generating video dialogue question-and-answer data of an embodiment of the present disclosure. Specific Embodiments

[0025] The embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.

[0026] It should be understood that the various steps recited in the method embodiments of the present disclosure can be executed in a different order and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this regard.

[0027] The term "including" and its variations used herein are open-ended, that is, "including but not limited to". The term "based on" is "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description.

[0028] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules, or units, and are not used to limit the order of the functions performed by these devices, modules, or units or their interdependent relationships.

[0029] It should be noted that the modifiers "one" and "multiple" mentioned in this disclosure are illustrative rather than restrictive. Those skilled in the art should understand that, unless otherwise clearly specified in the context, it should be understood as "one or more".

[0030] The names of the messages or information exchanged between multiple devices in the embodiments of this disclosure are only for illustrative purposes and do not limit the scope of these messages or information.

[0031] It can be understood that before using the technical solutions disclosed in the embodiments of this disclosure, the types, usage scopes, usage scenarios, etc. of the personal information involved in this disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.

[0032] For example, when responding to receiving an active request from the user, a prompt message is sent to the user to clearly prompt the user that the operation requested by the user will require obtaining and using the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, application program, server, or storage medium that performs the operations of the technical solutions of this disclosure based on the prompt message.

[0033] As an optional but non-limiting implementation manner, the manner of sending a prompt message to the user in response to receiving an active request from the user can be, for example, in the form of a pop-up window, and the prompt message can be presented in text in the pop-up window. In addition, the pop-up window can also carry selection controls for the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0034] It can be understood that the above process of notifying and obtaining the user's authorization is only illustrative and does not limit the implementation manners of this disclosure. Other manners that meet relevant laws and regulations can also be applied to the implementation manners of this disclosure.

[0035] It can be understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of corresponding laws, regulations, and related provisions.

[0036] In the following embodiments, optional features and examples are provided in each embodiment. The features described in the embodiments can be combined to form multiple optional solutions. Each numbered embodiment should not be regarded as only one technical solution. In addition, without conflict, the embodiments in this disclosure and the features in the embodiments can be combined with each other.

[0037] Figure 1The flowchart of a method for generating video dialogue Q&A data provided by an embodiment of the present disclosure. The technical solution of this embodiment is applicable to the situation of constructing video dialogue Q&A data using videos. This method can be executed by a video dialogue Q&A data generation device, which can be implemented by software and / or hardware and is generally integrated in any electronic device with network communication functions, including but not limited to devices such as computers and personal digital assistants.

[0038] As Figure 1 shown, the method for generating video dialogue Q&A data in this embodiment may include the following processes S110 - S130:

[0039] S110. Determine the target video description information corresponding to the target video.

[0040] Video dialogue Q&A data may refer to data for conducting Q&A conversations about videos. As videos, especially short videos, become a lifestyle for more and more people, the understanding and derivation of video content are becoming increasingly important. For example, generating titles, abstracts, and descriptions for videos, discussing some themes or details in the video, including discussing the things, events, and related background knowledge in the video.

[0041] However, when constructing video dialogue Q&A data for a video, it is usually necessary to watch the video completely manually before starting to construct the video dialogue Q&A data for the video, resulting in a relatively long construction time for the video dialogue Q&A data. And when describing the video, it is necessary to have a certain understanding of the field of the video content to describe the video content more accurately, which further leads to a large construction difficulty for the video dialogue Q&A data and the accuracy of the constructed video dialogue Q&A data cannot be guaranteed.

[0042] Therefore, when obtaining the video description information of the target video, according to the preset video description information requirements, the video information that meets the requirements can be extracted from the target video through video analysis, and then the target video description information corresponding to the target video can be obtained. Among them, the target video description information includes video title description, the main body of things in a single-frame video picture and the local detail events between the main bodies of things, the overall detail events between the main bodies of things expressed before and after consecutive multi-frame video pictures, the position of the main body of things in the video picture, and the dialogue text content of the main body of things in the video picture.

[0043] As an optional but non-limiting implementation manner, determining the target video description information corresponding to the target video may include steps A1 - A3:

[0044] Step A1. Detect whether there are main bodies of things in at least two video pictures extracted from the target video, and determine the position of the main body of things in the video picture when it is detected that there are main bodies of things.

[0045] Step A2: When it is detected that the main subject appears in the video frame and text captions appear in the video frame, determine the text captions that appear as the dialogue text content of the main subject in the video frame.

[0046] Step A3: When it is detected that the main subject appears in the video frame and matching audio exists in the video frame, convert the matching audio into text and then determine it as the dialogue text content of the main subject in the video frame.

[0047] By adopting the above optional method, it is not necessary for a human to watch the entire video. Only by importing the target video and performing the video according to the preset video description information requirements, the video description information that meets the requirements can be automatically obtained, and the obtained video description information can be added in time for analysis. Moreover, as the video is continuously analyzed, new video description information can be continuously updated and added, which can greatly reduce the acquisition time of the video description information, thereby reducing the construction time of the video dialogue Q&A data. And particularly importantly, due to the use of automatic video analysis, useful content can be mined as much as possible to avoid omission to provide more basis for the generation of subsequent dialogue Q&A data.

[0048] S120: Determine the target prompt words used by the target Q&A model. The target Q&A model is pre-configured based on a large language model, and the target prompt words can guide the target Q&A model to achieve the desired dialogue Q&A effect according to the target video description information when performing the dialogue Q&A generation task.

[0049] As an optional but non-limiting implementation method, the target prompt words can be configured with a first prompt message, a second prompt message, a third prompt message, a fourth prompt message, and a fifth prompt message. The first prompt message is used to instruct the target Q&A model to simulate watching the target video to perform the dialogue Q&A generation task of creating questions and answering. The second prompt message is used to instruct the target Q&A model to use the details included in the target video description information when performing the dialogue Q&A generation task so that the answer conforms to the target video description information. The third prompt message is used to instruct the target Q&A model to give a clear answer when performing the dialogue Q&A generation task. The fourth prompt message is used to instruct the target Q&A model to ask questions of a preset type when performing the dialogue Q&A generation task. The fifth prompt message is used to instruct the target Q&A model to give an answer including a detailed reasoning process when performing the dialogue Q&A generation task.

[0050] Exemplarily, the first prompt information can specifically be a prompt word that expects the target question-and-answer model to simulate the target video viewing, so that the target question-and-answer model can perform a question-and-answer dialogue question-and-answer task just like the user normally watches the target video. The second prompt information can specifically be a prompt word that expects the target question-and-answer model to use the detailed content contained in the video data description information as much as possible in the question-and-answer process of performing the dialogue question-and-answer generation task to ensure that the answer can be more in line with the actual content of the video data. The third prompt information can specifically expect the target question-and-answer model to give a clear answer in combination with the target video details and the interaction between the characters when performing the dialogue question-and-answer generation task: it can be a prompt word of the answer directly derived or indirectly inferred from the target video description information, and it can be determined that the question does not exist at all or cannot be answered after the video. The fourth prompt information can specifically expect the target question-and-answer model to prompt a prompt word that asks questions involving time perception and reasoning, and asks complex questions related to the video content when performing the dialogue question-and-answer generation task. The fifth prompt information can specifically be a prompt for the expected target question-answering model when performing the dialogue question-answering generation task. Since the video description will be received when watching the video, please give priority to asking about the visual changes over time and the reasons or causes behind these changes, rather than a prompt word for the question that can be inferred from a single picture.

[0051] Further, a complete example of the target prompt is given below. As an AI visual assistant, you are watching a video, and the content of the video is described at the end of this text. Based on these descriptions, your task is to answer all questions as if you were watching the video directly. Create a conversation with at least three rounds of Q&A, and try to have as many rounds of Q&A as possible. The person asking the questions is referred to as the "questioner", and the person answering the questions is referred to as the "AI visual assistant". The conversation should make use of as much information from the video as possible. Ensure that the answers reflect the tone of a visual AI assistant that is actively watching the video and answering questions. The conversation includes various questions and corresponding answers. Combine questions related to the visual content of the video, such as details of the video content, interactions between people, etc. The questions must have clear answers: questions that can be directly observed or indirectly inferred from the video content, and questions that can be determined to be completely absent or unanswerable in the video. Next, include questions related to time perception and reasoning, such as asking what a person did before or after an event, or asking for the specific timestamp of certain events or actions. Also include complex questions related to the video content, such as asking about the background knowledge of objects or actions in the video, discussing the events that occurred in the video, delving into counterfactual topics (e.g., what would happen if a person lost their phone while playing with it in the video), predicting how the story or scenario of the video will develop. Since you will receive the video description while watching the video, give priority to asking questions about visual changes over time and the reasons or causes behind these changes, rather than questions that can be inferred from a single frame. Remember not to ask about uncertain details. When answering complex questions, provide detailed answers, combined with detailed examples or reasoning steps to make the content more persuasive and structured. Use multiple paragraphs if necessary. If a question cannot be answered based on the given description, use "The provided video does not present such information".

[0052] As an optional but non-limiting implementation manner, determining the target prompt adopted by the target Q&A model may include the following steps:

[0053] In response to the selection operation of the target video, determine the application scenario of the dialogue Q&A data corresponding to the target video; determine the target prompt that matches the application scenario of the dialogue Q&A data corresponding to the target video from the candidate prompts associated with the target Q&A model.

[0054] S130. Based on the target video description information and the target prompt, output the dialogue Q&A data associated with the target video through the target Q&A model.

[0055] As an optional but non-limiting implementation manner, based on the target video description information and the target prompt, outputting the video dialogue Q&A data associated with the target video through the target Q&A model includes the following steps:

[0056] Add the target video description information at the preset position indicated by the target prompt to obtain the target input information of the target question-and-answer model; based on the target input information, control the target question-and-answer model to execute the dialogue question-and-answer generation task, and output the video dialogue question-and-answer data of the target video according to the execution of the dialogue question-and-answer generation task.

[0057] For a target prompt similar to the one provided above, a video description information field can be configured at the end position of the target prompt. After the video description information field, the text content corresponding to the target video description information can be specifically written. Furthermore, the added target video description information and the target prompt are combined into a complete text content as the target input information of the target question-and-answer model.

[0058] Input the target input information containing the target video description information and the target prompt into the target question-and-answer model. The target question-and-answer model can understand the target video according to the target video description information, automatically execute the dialogue question-and-answer generation task, and create a conversation task with at least three rounds of dialogue question-and-answers, thereby forming the video dialogue question-and-answer data of question and answer (see Figure 2 the dialogue question-and-answer schematic diagram shown on the question-and-answer output interface), and enable the target question-and-answer model to perform the dialogue question-and-answer generation task under the guidance of the target prompt.

[0059] As an optional but non-limiting implementation manner, after outputting the dialogue question-and-answer data associated with the target video through the target question-and-answer model, the following steps are further included:

[0060] In response to the screening operation of the dialogue question-and-answer data associated with the target video, determine the target dialogue question-and-answer data from the dialogue question-and-answer data associated with the target video.

[0061] In response to the editing operation of the target dialogue question-and-answer data, adjust and replace the answer corresponding to the question in the target dialogue question-and-answer data so that the adjusted and replaced answer fits the target video description information.

[0062] By adopting the above method, user interaction is actually introduced, ensuring that the output dialogue question-and-answer data associated with the target video has both the efficient generation ability of the machine and the fine review ability of the human, so as to achieve the quality and diversity of the generated dialogue question-and-answer data associated with the target video. By introducing user interaction, allowing users to select or modify the text, not only improves the quality of the dialogue question-and-answer data associated with the target video, but also increases the diversity of the generated dialogue question-and-answer data associated with the target video.

[0063] In the technical solution of the embodiment of the present disclosure, when constructing video dialogue Q&A data, the target video description information and the target prompt words that can guide the target Q&A model to the expected dialogue Q&A effect according to the target video description information during the execution of the dialogue Q&A generation task are determined. In this way, the target Q&A model pre-configured based on the large language model can automatically generate dialogue Q&A data according to the target video description information under the guidance of the target prompt words for dialogue Q&A, solving the problem that it takes a long time to construct dialogue Q&A data manually after watching the video completely. The construction can be carried out in real time by analyzing the video in real time to obtain the video description information. At the same time, since the Q&A model is guided by prompt words to construct dialogue Q&A data during the construction process, the accuracy of the constructed dialogue Q&A data can be guaranteed to a certain extent.

[0064] Figure 3 The following is a structural block diagram of a video dialogue Q&A data generation device provided by an embodiment of the present disclosure. The technical solution of this embodiment is applicable to the situation of constructing video dialogue Q&A data using videos. The device can be implemented by software and / or hardware and is generally integrated in any electronic device with network communication functions, including but not limited to devices such as computers and personal digital assistants.

[0065] As Figure 3 shown, the video dialogue Q&A data generation device of this embodiment may include the following:

[0066] A description information determination module 310, configured to determine target video description information corresponding to the target video;

[0067] A prompt information determination module 320, configured to determine target prompt words adopted by the target Q&A model. The target Q&A model is pre-configured based on a large language model, and the target prompt words can guide the target Q&A model to the expected dialogue Q&A effect according to the target video description information during the execution of the dialogue Q&A generation task;

[0068] A data generation module 330, configured to output dialogue Q&A data associated with the target video through the target Q&A model based on the target video description information and the target prompt words.

[0069] Based on the above embodiment, optionally, the video description information includes video title description, the main body of things in a single-frame video picture and the local detail events between the main bodies of things, the overall detail events between the main bodies of things expressed before and after consecutive multi-frame video pictures, the position of the main body of things in the video picture, and the dialogue text content of the main body of things in the video picture.

[0070] Based on the above embodiment, optionally, determining the target video description information corresponding to the target video includes:

[0071] Detect whether a subject appears in at least two video frames extracted from the target video, and determine the position of the subject in the video frame when it is detected that the subject appears;

[0072] When it is detected that a subject appears in the video frame and text captions appear in the video frame, determine the dialogue text content of the subject in the video frame as the text captions that appear;

[0073] When it is detected that a subject appears in the video frame and there is matching audio in the video frame, convert the matching audio into text and then determine it as the dialogue text content of the subject in the video frame.

[0074] Based on the above embodiments, optionally, determining the target prompt words adopted by the target question-and-answer model includes:

[0075] In response to a selection operation on the target video, determine the application scenario of the dialogue question-and-answer data corresponding to the target video;

[0076] Determine the target prompt words that match the application scenario of the dialogue question-and-answer data corresponding to the target video from the candidate prompt words associated with the target question-and-answer model.

[0077] Based on the above embodiments, optionally, based on the target video description information and the target prompt words, output the video dialogue question-and-answer data associated with the target video through the target question-and-answer model, including:

[0078] Add the target video description information at the preset position indicated by the target prompt word to obtain the target input information of the target question-and-answer model;

[0079] Based on the target input information, control the target question-and-answer model to execute the dialogue question-and-answer generation task, and output the video dialogue question-and-answer data of the target video according to the execution of the dialogue question-and-answer generation task.

[0080] Based on the above embodiments, optionally, the target prompt includes a first prompt message, a second prompt message, a third prompt message, a fourth prompt message, and a fifth prompt message. The first prompt message is used to instruct the target question-and-answer model to simulate watching the target video to perform a dialogue question-and-answer generation task of creating questions and answers. The second prompt message is used to instruct the target question-and-answer model to use the details included in the target video description information to make the answers conform to the target video description information when performing the dialogue question-and-answer generation task. The third prompt message is used to instruct the target question-and-answer model to give clear answers when performing the dialogue question-and-answer generation task. The fourth prompt message is used to instruct the target question-and-answer model to ask questions of a preset type when performing the dialogue question-and-answer generation task. The fifth prompt message is used to instruct the target question-and-answer model to give answers including a detailed reasoning process when performing the dialogue question-and-answer generation task.

[0081] Based on the above embodiments, optionally, after outputting the dialogue question-and-answer data associated with the target video by the target question-and-answer model, the following steps are further included:

[0082] In response to a screening operation on the dialogue question-and-answer data associated with the target video, determine target dialogue question-and-answer data from the dialogue question-and-answer data associated with the target video;

[0083] In response to an editing operation on the target dialogue question-and-answer data, adjust and replace the answer corresponding to the question in the target dialogue question-and-answer data so that the adjusted and replaced answer conforms to the target video description.

[0084] In the technical solution of the embodiments of the present disclosure, when constructing video dialogue question-and-answer data, the target video description information and the target prompt that can guide the target question-and-answer model to conform to the target video description information in the dialogue question-and-answer generation task are determined. In this way, the target question-and-answer model pre-configured based on the large language model can automatically generate dialogue question-and-answer data according to the target video description information under the guidance of the target prompt in the dialogue question-and-answer, solving the problem that it takes a long time to construct dialogue question-and-answer data manually after watching the video completely. The construction can be carried out in real time by analyzing the video to obtain the video description information. At the same time, since the prompt is used to guide the question-and-answer model to construct dialogue question-and-answer data during the construction process, the accuracy of the constructed dialogue question-and-answer data can be guaranteed to a certain extent.

[0085] The video dialogue question-and-answer data generation device provided by the embodiments of the present disclosure can execute the video dialogue question-and-answer data generation method provided by any embodiment of the present disclosure, and has corresponding functional modules and beneficial effects for executing the video dialogue question-and-answer data generation method.

[0086] It should be noted that the various units and modules included in the above-mentioned device are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific names of the functional units are only for the convenience of mutual distinction and do not limit the protection scope of the embodiments of the present disclosure.

[0087] Reference is made below to Figure 4 , which shows a schematic structural diagram of an electronic device 800 suitable for implementing the embodiments of the present disclosure. The terminal device in the embodiments of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 4 The electronic device shown is only an example and should not impose any limitations on the functions and usage scope of the embodiments of the present disclosure.

[0088] As Figure 4 shown, the electronic device 800 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 801, which may perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 802 or a program loaded from a storage device 806 into a random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the electronic device 800 are also stored. The processing device 801, the ROM 802, and the RAM 803 are connected to each other through a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.

[0089] Generally, the following devices may be connected to the I / O interface 805: an input device 806 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 807 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 806 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 809. The communication device 809 may allow the electronic device 800 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 4 the electronic device 800 with various devices is shown, it should be understood that it is not required to implement or include all the shown devices. More or fewer devices may be alternatively implemented or included.

[0090] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only for illustrative purposes and do not limit the scope of these messages or information.

[0091] The electronic device provided by the embodiments of the present disclosure and the video dialogue question-answer data generation method provided by the above embodiments belong to the same inventive concept. For technical details not described in detail in this embodiment, reference may be made to the above embodiments, and this embodiment has the same beneficial effects as the above embodiments.

[0092] Specifically, according to the embodiments of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, the embodiments of the present disclosure include a computer program product that includes a computer program carried on a non-transitory computer-readable medium. The computer program includes program code for executing the video dialogue question-answer data generation method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 809, or installed from the storage device 806, or installed from the ROM 802. When the computer program is executed by the processing device 801, it executes the above functions defined in the video dialogue question-answer data generation method of the embodiments of the present disclosure.

[0093] The embodiments of the present disclosure provide a computer storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the video dialogue question-answer data generation method provided by the above embodiments.

[0094] It should be noted that the above-mentioned computer-readable medium in the present disclosure may be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium may be any tangible medium that contains or stores a program, which can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present disclosure, the computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium may also be any computer-readable medium other than the computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0095] In some embodiments, the client and the server can communicate using any currently known or future-developed network protocol such as HTTP (HyperText Transfer Protocol), and can be interconnected with digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include local area networks ("LANs"), wide area networks ("WANs"), the Internet (e.g., the Internet), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed networks.

[0096] The above-mentioned computer-readable medium may be included in the above-mentioned electronic device; or it may exist separately and not be assembled into the electronic device.

[0097] The above computer-readable medium stores one or more programs that, when executed by the electronic device, cause the electronic device to: determine target video description information corresponding to a target video; determine target prompt words used by a target question-and-answer model, where the target question-and-answer model is pre-configured based on a large language model, and the target prompt words can guide the target question-and-answer model to achieve an expected dialogue question-and-answer effect according to the target video description information when performing a dialogue question-and-answer generation task; and output dialogue question-and-answer data associated with the target video through the target question-and-answer model based on the target video description information and the target prompt words.

[0098] Alternatively, the above computer-readable medium stores one or more programs that, when executed by the electronic device, cause the electronic device to execute the video dialogue question-and-answer data generation method described in any one of the above embodiments.

[0099] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages or combinations thereof. The programming languages include, but are not limited to, object-oriented programming languages such as Java, Smalltalk, and C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may execute entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0100] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that, in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.

[0101] The units involved in the embodiments of the present disclosure can be implemented in software or in hardware. Among them, the name of a unit does not constitute a limitation on the unit itself in some cases. For example, the first acquisition unit can also be described as "the unit for acquiring at least two Internet protocol addresses".

[0102] The functions described above herein can be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include: Field Programmable Gate Arrays (FPGA), Application Specific Integrated Circuits (ASIC), Application Specific Standard Products (ASSP), System on a Chip (SOC), Complex Programmable Logic Devices (CPLD), and the like.

[0103] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a Random Access Memory (RAM), a Read-Only Memory (ROM), an Erasable Programmable Read-Only Memory (EPROM or Flash Memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0104] The above description is only a preferred embodiment of the present disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, the technical solutions formed by mutually replacing the above features with the (but not limited to) technical features having similar functions disclosed in the present disclosure.

[0105] Moreover, although the operations are depicted in a particular order, this should not be construed as requiring that the operations be performed in the particular order shown or in sequential order. In certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although several specific implementation details are included in the foregoing discussion, these should not be construed as limitations on the scope of the present disclosure. Certain features that are described in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment may also be implemented separately or in any suitable sub-combination in multiple embodiments.

[0106] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.

Claims

1. A method for generating video conversation Q&A data, characterized in that, The method includes: Determine the target video description information corresponding to the target video; Determine the target prompt words adopted by the target question-and-answer model, where the target question-and-answer model is pre-configured based on a large language model, and the target prompt words can guide the target question-and-answer model to the desired dialogue question-and-answer effect according to the target video description information when performing the dialogue question-and-answer generation task; the target prompt words are configured with the first prompt information, the second prompt information, the third prompt information, the fourth prompt information, and the fifth prompt information. The first prompt information is used to instruct the target question-and-answer model to simulate watching the target video to perform the dialogue question-and-answer generation task of creating questions and answers. The second prompt information is used to instruct the target question-and-answer model to use the details included in the target video description information when performing the dialogue question-and-answer generation task so that the answer conforms to the target video description information. The third prompt information is used to instruct the target question-and-answer model to give a clear answer when performing the dialogue question-and-answer generation task. The fourth prompt information is used to instruct the target question-and-answer model to ask questions of a preset type when performing the dialogue question-and-answer generation task. The fifth prompt information is used to instruct the target question-and-answer model to give an answer including a detailed reasoning process when performing the dialogue question-and-answer generation task; Based on the target video description information and the target prompt words, output the dialogue question-and-answer data associated with the target video through the target question-and-answer model.

2. The method according to claim 1, wherein The video description information includes video title description, the main body of things in a single-frame video picture and the local detail events between the main bodies of things, the overall detail events between the main bodies of things expressed before and after consecutive multi-frame video pictures, the position of the main body of things in the video picture, and the dialogue text content of the main body of things in the video picture.

3. The method according to claim 2, wherein Determine the target video description information corresponding to the target video, including: Detect whether the main body of things appears in at least two video pictures extracted from the target video, and determine the position of the main body of things in the video picture when detecting the appearance of the main body of things; When detecting that the main body of things appears in the video picture and detecting that there are text subtitles in the video picture, determine the text subtitles that appear as the dialogue text content of the main body of things in the video picture; When detecting that the main body of things appears in the video picture and detecting that there is matching audio in the video picture, convert the matching audio into text and then determine it as the dialogue text content of the main body of things in the video picture.

4. The method according to claim 1, wherein Determine the target prompt words adopted by the target question-and-answer model, including: In response to the selection operation of the target video, determine the application scenario of the dialogue question-and-answer data corresponding to the target video; Determine the target prompt words that match the application scenario of the dialogue question-and-answer data corresponding to the target video from the candidate prompt words associated with the target question-and-answer model.

5. The method according to claim 1, characterized in that, Based on the target video description information and the target prompt words, output the video dialogue question-and-answer data associated with the target video through the target question-and-answer model, including: Add the target video description information at the preset position indicated by the target prompt words to obtain the target input information of the target question-and-answer model; Based on the target input information, control the target question-and-answer model to execute a dialogue question-and-answer generation task, and output the video dialogue question-and-answer data of the target video according to the execution of the dialogue question-and-answer generation task.

6. The method according to claim 1, wherein After outputting the dialogue question-and-answer data associated with the target video through the target question-and-answer model, it further includes: In response to the screening operation of the dialogue question-and-answer data associated with the target video, determine the target dialogue question-and-answer data from the dialogue question-and-answer data associated with the target video; In response to the editing operation of the target dialogue question-and-answer data, adjust and replace the answer corresponding to the question in the target dialogue question-and-answer data so that the adjusted and replaced answer conforms to the target video description.

7. A video conversation Q&A data generation device, characterized in that, The device includes: A description information determination module for determining the target video description information corresponding to the target video; A prompt information determination module for determining the target prompt words used by the target question-and-answer model. The target question-and-answer model is pre-configured based on a large language model, and the target prompt words can guide the target question-and-answer model to the desired dialogue question-and-answer effect according to the target video description information when executing the dialogue question-and-answer generation task; the target prompt words are configured with the first prompt information, the second prompt information, the third prompt information, the fourth prompt information, and the fifth prompt information. The first prompt information is used to instruct the target question-and-answer model to simulate watching the target video to execute the dialogue question-and-answer generation task of creating questions and answering them. The second prompt information is used to instruct the target question-and-answer model to use the details included in the target video description information when executing the dialogue question-and-answer generation task so that the answer conforms to the target video description information. The third prompt information is used to instruct the target question-and-answer model to give a clear answer when executing the dialogue question-and-answer generation task. The fourth prompt information is used to instruct the target question-and-answer model to ask questions of a preset type when executing the dialogue question-and-answer generation task. The fifth prompt information is used to instruct the target question-and-answer model to give an answer including a detailed reasoning process when executing the dialogue question-and-answer generation task; A data generation module for outputting the dialogue question-and-answer data associated with the target video through the target question-and-answer model based on the target video description information and the target prompt words.

8. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor. When the computer program is executed by the at least one processor, the at least one processor can execute the video dialogue question-and-answer data generation method according to any one of claims 1-6.

9. A computer-readable medium, characterized in that, The computer-readable medium stores computer instructions for causing a processor to implement the video dialogue question-and-answer data generation method according to any one of claims 1-6 when executed.

Citation Information

Patent Citations

  • Visual question answering method and device based on regularization and dual learning

    CN116049371A

  • Dialogue generation method, deep learning model training method, device and equipment

    CN116303962A