Data generation method and device, electronic equipment and storage medium

By acquiring and transforming multimodal data to generate two-stage question-answer pairs, the problem of scarce text data in vertical domain databases is solved, improving data quality and the ability to intelligently apply domain knowledge.

CN120851168APending Publication Date: 2025-10-28GUANGZHOU HUYA INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511004713.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-21
Publication Date
2025-10-28

AI Technical Summary

Technical Problem

Databases in vertical sectors such as healthcare, finance, and gaming contain relatively little plain text data, which is of low quality. This results in insufficient domain knowledge, making it difficult to form high-quality question-and-answer pairs and hindering the development of intelligent technologies.

Method used

By acquiring raw multimodal data, converting it into plain text data, and using a large model to generate two-stage question-answer pairs, the core knowledge is first extracted, and then complex answers are deeply analyzed to generate high-quality question-answer pairs as data in the database.

Benefits of technology

It significantly improved the quantity and quality of text data in the database, enriched the knowledge expression forms in vertical fields, and provided strong support for intelligent applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120851168A_ABST
    Figure CN120851168A_ABST
Patent Text Reader

Abstract

The invention provides a data generation method and device, electronic equipment and a storage medium, and the method comprises the steps: obtaining original multi-modal data related to target content, and converting the original multi-modal data into plain text data; generating a first-stage question and answer pair according to the plain text data and the prompt information; generating a plurality of second-stage question and answer pairs by using the first reply text in the first-stage question and answer pairs and the new prompt information; wherein the question text in the second-stage question and answer pair is generated by disassembling the first answer text; and taking the combination result of the second reply text in each second-stage question and answer pair and the question text of the first-stage question and answer pair as in-library data of the target content. The problem of scarcity of text data in the vertical field in the database can be effectively relieved, and the data quality of the vertical field in the database is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of large model technology, and more specifically, to a data generation method, apparatus, electronic device, and storage medium. Background Art

[0002] With the rapid development of artificial intelligence technology, the demand for high-quality text data is growing in vertical industries such as healthcare, finance, and gaming.

[0003] However, in current databases of vertical industries such as healthcare, finance, and gaming, the amount of plain text data is far less than the total information contained in publicly available web pages, images, and videos, resulting in a limited amount of directly usable text data within the databases. Furthermore, the existing text data in these databases generally suffers from problems such as disorganized formatting, redundancy, noise interference, and logical gaps, making it difficult to directly map into high-quality knowledge representations like question-and-answer (QA) pairs. This leads to insufficient domain knowledge and hinders the intelligent development of these vertical industries.

[0004] Therefore, how to improve the amount and quality of usable text data related to domain knowledge in the database is a technical problem that needs to be solved. Summary of the Invention

[0005] In view of this, the purpose of this invention is to provide a data generation method, apparatus, electronic device, and storage medium capable of generating high-quality text data and improving the quantity and quality of text data in vertical domains within a database. To achieve the above objective, the technical solution adopted by this invention is as follows: In a first aspect, the present invention provides a data generation method, the method comprising: acquiring raw multimodal data related to target content and converting the raw multimodal data into plain text data; generating a first-stage question-and-answer pair based on the plain text data and prompt information; generating multiple second-stage question-and-answer pairs using a first response text in the first-stage question-and-answer pair and new prompt information; wherein the question text in the second-stage question-and-answer pair is generated by disassembling the first response text; and using the combination result of the second response text in each second-stage question-and-answer pair and the question text of the first-stage question-and-answer pair as in-database data of the target content.

[0006] Secondly, the present invention provides a data generation apparatus, comprising: a data acquisition and conversion module, configured to acquire raw multimodal data related to target content and convert the raw multimodal data into plain text data; a data generation module, configured to generate a first-stage question-and-answer pair based on the plain text data and prompt information; generate multiple second-stage question-and-answer pairs using the first response text in the first-stage question-and-answer pair and new prompt information; wherein the question text in the second-stage question-and-answer pair is generated by disassembling the first response text; and the combination result of the second response text in each second-stage question-and-answer pair and the question text of the first-stage question-and-answer pair are used as the database data of the target content.

[0007] Thirdly, the present invention provides an electronic device including a processor and a memory, the memory storing machine-executable instructions executable by the processor, the processor executing the machine-executable instructions to implement the data generation method described in any of the foregoing embodiments.

[0008] Fourthly, the present invention provides a storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the data generation method as described in any of the foregoing embodiments.

[0009] The data generation method, apparatus, electronic device, and storage medium provided by this invention first acquire raw multimodal data related to the target content and convert it into plain text data. Based on this, a first-stage question-and-answer pair is generated by combining pre-designed prompts. This first-stage question-and-answer pair extracts and summarizes core knowledge closely related to the target content, laying a solid foundation for subsequent processing. Furthermore, in the second stage, questions are posed again to the response text generated in the first stage, thereby achieving in-depth analysis and semantic mining of complex answers, significantly improving the accuracy and logical completeness of the responses. Finally, the refined response text generated in the second stage is combined with the questions posed in the first stage to form high-quality question-and-answer pairs, which serve as in-database data for the target content. This process effectively alleviates the problem of scarce vertical domain text data in databases, enriches the knowledge expression forms of vertical domains in databases, and provides strong support for the intelligent application of knowledge in this domain.

[0010] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description

[0011] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0012] Figure 1 This is a schematic diagram illustrating an application scenario provided by an embodiment of the present invention; Figure 2 A schematic flowchart illustrating the data generation method provided in an embodiment of the present invention; Figure 3 A functional block diagram of a data generation device provided in an embodiment of the present invention; Figure 4 This is a structural block diagram of an electronic device provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0013] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.

[0014] Therefore, the following detailed description of the embodiments of the invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the invention without inventive effort are within the scope of protection of the invention.

[0015] It should be noted that relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0016] Please see Figure 1 , Figure 1This is a schematic diagram of an application scenario provided by an embodiment of the present invention. The server 20 can be in the cloud. The server 20 can connect to one or more terminals 10 via a local area network (LAN), a wide area network (WAN), the Internet, or other types of data networks. The terminals 10 can be client devices, including but not limited to: smartphones, tablets, laptops, PDAs, personal computers, smart home devices, and in-vehicle devices. The server 20 and the terminals 10 can exchange information via a network 30.

[0017] Database 40 can maintain data from various vertical fields (such as healthcare, finance, and gaming), which can include various types such as images, videos, audio, and plain text. Database 40 can also store one or more services or software applications that can run and execute the data generation method provided in this embodiment of the invention. Server 20 can read data from database 40 and apply it to real-world scenarios, and can also generate new data and store it in database 40.

[0018] In this embodiment of the invention, Figure 1 The application scenarios shown can be applied in, but are not limited to, specific industries or professional fields such as healthcare, finance, and gaming. For example, in the gaming field, server 20 can be a game server, and database 40 can maintain game information such as hero characters, skills, and quest flows. Server 20 can read relevant game data from database 40 and use it in actual game operations. At the same time, server 20 can also generate data related to game knowledge and store it in database 40.

[0019] Given the limited quantity and low quality of available text data in vertical domains within databases, this invention proposes an efficient data generation method. The data generated by this method is highly targeted and practical, effectively compensating for the shortcomings of insufficient text data quantity and poor quality in specific domains, thereby significantly improving the overall quality and usability of the database.

[0020] See Figure 2 , Figure 2 This is a schematic flowchart illustrating the data generation method provided in an embodiment of the present invention. The execution subject of this method can be an electronic device, such as... Figure 1 Server 20. The method includes steps S201 to S204, as described below: S201: Obtain the raw multimodal data related to the target content and convert the raw multimodal data into plain text data; In this embodiment of the invention, the target content can be content within a specific field, which may include, but is not limited to, finance, healthcare, and gaming. For example, the target content can be game objects (such as heroes) in the gaming field, disease symptoms and medical equipment in the healthcare field, or investment strategies and financial products in the financial field.

[0021] S202: Generate the first-stage question-and-answer pair based on plain text data and prompts; S203: Generate multiple second-stage question-and-answer pairs using the first response text and new prompt information from the first-stage question-and-answer pair; In this process, the question text in the second-stage question-and-answer pair is generated by disassembling the first response text. S204: Combine the combined results of the second response texts in each second-stage question-and-answer pair with the question texts of the first-stage question-and-answer pair as the target content's in-database data.

[0022] In the aforementioned data generation method, the raw multimodal data related to the target content is first acquired and converted into plain text data. Based on this, pre-designed prompts are combined to generate a first-stage question-and-answer pair. This first-stage pair extracts and summarizes core knowledge closely related to the target content, laying a solid foundation for subsequent processing. Further, in the second stage, questions are posed again to the response text generated in the first stage, enabling in-depth analysis and semantic mining of complex answers, significantly improving the accuracy and logical completeness of the responses. Finally, the refined response text generated in the second stage is combined with the questions posed in the first stage to form high-quality question-and-answer pairs, which serve as in-database data for the target content. This process effectively alleviates the scarcity of vertical domain text data in databases, enriches the knowledge representation forms of vertical domains in databases, and provides strong support for the intelligent application of knowledge in this domain.

[0023] Next, the embodiments of the present invention will provide a clear and detailed explanation of the above data generation process.

[0024] In step S201, the target content can be content within a specific domain specified by relevant personnel. For example, the target content could be game objects (such as game heroes, weapons, skills, etc.) within the gaming domain. Raw multimodal data refers to unprocessed data in various forms related to the target content, and this raw multimodal data can originate from various websites. For instance, in the gaming domain, relevant personnel can use search engines, gaming communities, and other channels to select official game websites, large gaming forums, game strategy websites, etc., as data sources. For each data source, web scraping technology can be used to collect all raw data related to the target content, and this raw data typically exists in various forms such as text, images, and videos.

[0025] For example, in the gaming industry, raw data can be scraped from web pages such as official game websites, large game forums, and game strategy websites. Of course, raw data from other fields can also be collected using similar methods, and this is not a limitation here.

[0026] Optionally, for easier data management and use, the raw data can be stored in a database for unified maintenance. This raw data related to the target content can be maintained in a table format by creating dedicated database tables. For example, if the target content is a hero in a game, the table could contain fields such as hero name, webpage URL, raw data text, image links, and video links.

[0027] In this embodiment of the invention, the original multimodal data may include, but is not limited to, plain text, images, videos, and audio. For example, in the gaming field, the original multimodal data may include game logs, player chat logs (plain text), hero artwork, skill effect images, etc. (images), hero background animations, skill demonstration videos, etc. (videos), and player voice, NPC (Non-Player Character) dialogue, etc. (audios). As another example, in the financial field, plain text data may include financial news reports, transaction records, etc.; images may include financial statement charts, promotional posters of financial institutions, etc.; and video data may include financial product introduction videos, etc., and so on. These are not all listed here.

[0028] Considering that plain text data is suitable for large-scale models in various specific domains, we can first convert these original data formats into plain text data.

[0029] In this embodiment of the invention, for raw data in the text modality, since it is presented in text form, the main content and key information of the text can be extracted directly and quickly. For raw data in the non-text modality, different data parsing methods are required to convert the non-text data into text content. For example, for web page text data, the text can be extracted directly; for non-text modality data such as images and videos, relevant text information can be extracted specifically. In addition, for the converted plain text data, some preprocessing operations are required to improve the data quality. Therefore, as an example, step S201 can be implemented according to the following process: Step a1: Generate semantic description text corresponding to the non-text modal data; Step a2: Preprocess the semantic description text and text modal data to obtain plain text data.

[0030] In step a1, non-textual modal data can include images, videos, audio, etc. The raw data of non-textual modal data is relatively complex, requiring targeted data parsing to obtain the text. Taking images as an example, precise understanding and semantic parsing of the image content are needed. Videos, on the other hand, add a time dimension to the image data, requiring not only parsing the content of each frame but also parsing the text corresponding to the audio data within the video.

[0031] In this embodiment of the invention, the text generation process of the original data of image and video classes is described below.

[0032] For image-based data, such as visual resources like hero concept art, skill effect images, and skin illustrations in the gaming field, the corresponding text generation process is shown in steps 1 to 3: Step 1: Generate visual description text corresponding to the image data; In this embodiment of the invention, visual descriptive text is a textual expression of the image content. Optionally, a pre-trained image captioning model can be used to process the input image data to automatically generate visual descriptive text corresponding to its content. For example, in the gaming field, for a special effects image of a certain skill, the generated descriptive text could be "Golden energy converges into a giant sword, vertically striking the ground, causing real damage marked with a highlighted effect to enemy heroes within range."

[0033] Step 2: Perform character recognition on image data to extract recognizable text; Understandably, some images may also contain text related to the target content, such as skill images in the gaming field including information like "true damage" or "equipment name." Therefore, character recognition is also needed to extract this type of text from the images.

[0034] For example, as a case study, text in an image can be identified using open-source text recognition technology, and then the text can be extracted. It should be understood that character recognition is a mature, existing technology, and will not be elaborated upon here.

[0035] Alternatively, in order to improve text quality, text correction algorithms, such as spell check models, can be used to reduce noise in the character recognition results.

[0036] Step 3: Fuse and deduplicate the visual description text and the identifiable text to obtain the semantic description text.

[0037] In this embodiment of the invention, during the fusion of visual descriptive text and identifiable text extracted from an image, the visual descriptive text and the identifiable text extracted from the image can be spliced ​​together according to predefined splicing rules to achieve the fusion purpose. For example, as an example, they can be spliced ​​in the format of "[visual descriptive text] + [identifiable text]".

[0038] To ensure data quality, for fused text data, duplicate content (such as skill names appearing simultaneously) that appears in both visual description text and identifiable text can be removed by rule matching, and then a semantic description of pure text for image data can be output.

[0039] For video data, such as dynamic resources like hero background animations and skill demonstration videos in the gaming field, and dynamic resources like medical image demonstrations in the medical field, the corresponding text generation process is shown in steps 1 to 3: Step 1: Extract keyframes from video data and generate visual description text for the keyframes; In this embodiment of the invention, keyframes can be extracted from the video at a preset extraction frequency, such as one frame every five frames. The extraction frequency can be flexibly set by relevant personnel. However, during the extraction process, for some action-oriented videos, motion regions in the video frames can be detected first (e.g., optical flow detection), and then keyframes containing significant motion changes can be extracted to reduce data redundancy.

[0040] Optionally, but not limited to, keyframe extraction can be performed using a video processing library.

[0041] The extraction of keyframes is similar to the method of generating visual descriptive text from image data, and will not be repeated here.

[0042] Step 2: Perform speech recognition on the audio of the video data to obtain the audio text and the timestamp information of the audio text; In this embodiment of the invention, speech recognition can be used to transcribe video audio tracks sentence by sentence, generating narration text with timestamps. For example, in the gaming field, 00:03-00:08 corresponds to "Next, we will show the ultimate skill of a certain hero." Optionally, speech recognition can be performed using existing mature speech recognition models.

[0043] Step 3: Combine the visual description text of the keyframe and the audio text corresponding to the keyframe's timestamp into a semantic description text.

[0044] It can be understood that the visual description text of the key frames is obtained through step 1, and the audio text corresponding to the time stamps is obtained through step 2. Since each key frame also corresponds to a time stamp, the audio text corresponding to the key frames can be obtained simultaneously. The visual description text of the key frames and the audio text are combined to form a semantic description text in the form of "dynamic scene description + voice commentary".

[0045] In the embodiments of the present invention, in addition to the above-mentioned image and video types, targeted data parsing and text generation can be performed on data of other modalities. For example, for audio data, the semantic description text corresponding to the audio data can be obtained in a similar manner to the implementation of step 2 above, which will not be elaborated here.

[0046] Next, the embodiments of the present invention will introduce the process of preprocessing the text corresponding to the above various modality data.

[0047] In one implementation, during the text preprocessing process, regular expressions can be used to remove special characters (such as @, #, $, etc.), garbled characters, and whitespace characters (including line breaks, tab characters, etc.) from the text. For example, using the powerful regular expression function provided in programming tools, a regular expression rule can be written, such as "re.sub(r'[^\w\s]|[\u4e00-\u9fa5]', '', text)", to remove the content in the text except for Chinese characters, letters, numbers, and whitespace characters.

[0048] In addition, in order to further purify the text data, the stop word list of natural language processing tools can also be used to remove the stop words (such as "的", "了", "是", etc.) in the text. It should be noted that the meaning of the above regular expression rule is to match any non-word character, non-whitespace character, and non-Chinese character, and replace it with an empty string, so as to achieve the purpose of removing these characters.

[0049] In one implementation, during the text preprocessing process, formatting tags (such as HTML tags, CSS style tags, etc.) can also be removed first. For example, for the text containing Hero Skills: A certain skill name after the cleaning process, the resulting text is: "Hero Skill: A certain skill name".

[0050] Based on the preprocessed plain text data described above, this embodiment of the invention designs a two-stage knowledge injection mechanism. The first stage of data generation is completed through step S202. Considering that the input data in this stage is relatively messy and does not form complete natural sentences, the generated data is relatively general and the answer includes a lot of information, step S203 is designed to execute the second stage of data generation process. The final generated data is hierarchical, concise, and semantically accurate, which improves the quality of relevant data of the target content.

[0051] First, let's introduce step S202, which is the first stage of data generation process.

[0052] In step S202, a large model is first used to generate first-stage question-and-answer pairs (QA pairs) based on plain text data and prompts. Each QA pair includes a question text and a response text. To distinguish it from the response text generated in the second stage, the response text in the first-stage QA pair is referred to as the first response text, and the response text in the subsequent second-stage QA pair is referred to as the second response text. It should be understood that this notation is merely for differentiation and does not limit the content of the response text.

[0053] In this embodiment of the invention, the prompt message is a prompt template designed by relevant personnel based on plain text data. For example: "Please extract the question from the following description. The answer to the question should encompass as much content as possible in the description and be output in the format 'Question: [Question Content]\nAnswer: [Answer Content]'. Description: [Specific Text Content]". Based on this prompt message, the preprocessed plain text data is sequentially filled into the Prompt template, and then sent to the large model via the API interface of the large model. The large model understands and analyzes the plain text data according to the prompt message and generates QA pairs.

[0054] Considering that the questions in the QA pairs generated in the first stage are relatively general and the answers encompass a lot of information, the second stage of data generation process, i.e., step S203, is carried out based on the QA pairs generated in the first stage.

[0055] In step S203, this embodiment of the invention will use the first response text generated in the first stage to continue generating QA pairs. This is equivalent to conducting in-depth analysis and decomposition of the complex answers generated in the first stage, allowing the large model to generate semantically accurate response texts, which is more conducive to the model's learning and understanding of knowledge in the game domain. Specifically, step S203 can be implemented as follows: Step b1: Semantically decompose the first response text into multiple subtexts; In this embodiment of the invention, after extracting the QA pairs in the first stage, the processing of the Answer content (i.e., the first response text) is as follows: First, syntactic analysis and semantic understanding are performed using natural language processing tools to decompose the long answer into multiple semantically independent clauses, i.e., the subtext in this embodiment of the invention.

[0056] Step b2: Obtain the hint information for each subtext; In this embodiment of the invention, the prompt information is designed by relevant personnel based on sub-texts. Different types of prompt templates are designed for each sub-text. For example, for a clause describing a skill, the template "Based on '[clause content]', generate a question and its answer about the skill. The answer should be concise and clear, in the format 'Question: [Question content] Answer: [Answer content]'" is used; while for a clause describing a hero's background, the template "Please extract key information from '[clause content]' to form a question and answer, in the format 'Question: [Question content] Answer: [Answer content]'" is used.

[0057] Step b3: Based on the new prompt information, generate the question text and the answer text for the question text.

[0058] Similar to the first stage, the large model generates a question about each subtext and the corresponding answer text based on the prompts for that subtext.

[0059] Step b4: Combine the question text and the answer text to form a second-stage question-and-answer pair.

[0060] It is understandable that the questions in the second stage are generated by decomposing the responses in the first stage. This embodiment of the invention integrates the response texts from the second stage to obtain a semantically accurate response text for the questions in the first stage. See step S204. Its implementation may include steps c1 to c3, as explained below: Step c1: Determine the concatenation order of each second response text; Step c2: Concatenate the second response text according to the concatenation order to obtain the combined result; Step c3: Combine the combined result with the question text from the first stage question-answer pair to form a new question-answer pair, which will be used as data in the database.

[0061] In the above implementation, in each newly generated QA pair in the second stage, the order of the subtext corresponding to the QA pair in the second stage in the first response text can be used as the splicing order of the second response text. In other words, the subtext is spliced ​​sequentially according to the logical order in the original answer (i.e., the first response text in the first stage) to obtain the combination result. The combination result is then combined with the question text of the QA pair in the first stage to finally form a complete QA pair.

[0062] It is understood that in the second stage of this embodiment of the invention, questions are raised again based on the response text generated in the first stage, thereby achieving in-depth analysis and semantic mining of complex answers, significantly improving the accuracy and logical completeness of the responses. Finally, the refined response text generated in the second stage is combined with the questions raised in the first stage to form a high-quality question-answer pair.

[0063] In summary, the data generation method provided by this invention has the following advantages: First, by using a pre-trained image description generation model and video structured parsing, this invention transforms non-textual information such as visual effects and dynamic demonstrations of the target content into computable semantic text. This expands the knowledge source in the data acquisition stage from single webpage text to multimedia formats such as images and videos, solving the problem of incomplete coverage of specific domain knowledge by traditional text data. Second, by crawling webpage data and performing two-stage processing, this invention fully mines specific domain knowledge in existing network resources and generates a large amount of directly usable pure text data, effectively solving the problem of insufficient usable text data in vertical domains. Furthermore, the two-stage data generation method provided by this invention not only extracts key domain knowledge as a whole to form preliminary QA pairs, but also conducts in-depth analysis and decomposition of complex answers. It also uses a carefully designed Prompt template to guide a large model to generate hierarchical, concise, and semantically accurate response texts. These refined response texts are then combined with the questions raised in the first stage to form high-quality question-answer pairs, serving as in-library data for the target content. This process effectively alleviates the problem of scarce text data in vertical domains in databases, enriches the knowledge representation forms in vertical domains in databases, and provides strong support for the intelligent application of knowledge in this domain.

[0064] To perform the corresponding steps in the above embodiments and various possible methods, an implementation of the data generation device 30 is given below. Please refer to [link to relevant documentation]. Figure 3 , Figure 3This is a functional block diagram of a data generation device provided in an embodiment of the present invention. It should be noted that the basic principle and technical effects of the data generation device 30 provided in this embodiment are the same as those in the above embodiments. For the sake of brevity, any parts not mentioned in the embodiments of the present invention can be referred to the corresponding content in the above embodiments. The data generation device 30 includes: The data acquisition and conversion module 301 is used to acquire raw multimodal data related to the target content and convert the raw multimodal data into plain text data; Data generation 302 is used to generate a first-stage question-and-answer pair based on plain text data and prompt information; multiple second-stage question-and-answer pairs are generated using the first response text in the first-stage question-and-answer pair and new prompt information; wherein, the question text in the second-stage question-and-answer pair is generated by disassembling the first response text; the combination result of the second response text in each second-stage question-and-answer pair and the question text in the first-stage question-and-answer pair are used as the in-database data of the target content.

[0065] It is understandable that the data acquisition and transformation module 301 and the data generation module 302 can execute collaboratively. Figure 2 Each step in the process is used to achieve the corresponding technical effect.

[0066] It should be noted that the data generation device 30 provided in this embodiment of the invention can be specific hardware on the device or software or firmware installed on the device. The implementation principle and technical effects of the device provided in this embodiment of the invention are the same as those in the foregoing method embodiments. For the sake of brevity, any parts not mentioned in the device embodiments can be referred to the corresponding content in the foregoing method embodiments. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can all be referred to the corresponding processes in the above method embodiments, and will not be repeated here.

[0067] Optionally, the above modules can be stored in the form of software or firmware. Figure 4 The memory shown is in the operating system (OS) of the electronic device 40, and can be... Figure 4 The processor executes the commands. Meanwhile, the data and program code required to execute these modules can be stored in memory.

[0068] Please see Figure 4 , Figure 4 This is a structural block diagram of an electronic device provided in an embodiment of the present invention. Optionally, the electronic device 40 may be... Figure 1The server 20 includes a memory 401, a processor 402, and a communication interface 403. The memory 401, processor 402, and communication interface 403 are electrically connected to each other directly or indirectly to realize data transmission or interaction. For example, these components can be electrically connected to each other through one or more communication buses or signal lines.

[0069] Optionally, the bus can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of representation, Figure 4 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.

[0070] In this embodiment of the invention, the processor 402 may be a general-purpose processor, a digital signal processor, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components, capable of implementing or executing the methods, steps, and logic block diagrams disclosed in this embodiment of the invention. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the methods disclosed in this embodiment of the invention can be directly manifested as execution by the hardware processor, or execution by a combination of hardware and software modules within the processor. The software modules may reside in the memory 401, and the processor 402 reads the program instructions from the memory 401 and, in conjunction with its hardware, completes the steps of the aforementioned methods.

[0071] In this embodiment of the invention, the memory 401 can be a non-volatile memory, such as a hard disk drive (HDD) or a solid-state drive (SSD), or it can be volatile memory, such as RAM. The memory can also be any other medium capable of carrying or storing desired executable program code having an instruction or data structure form and accessible by a computer, but is not limited thereto. The memory in this embodiment of the invention can also be a circuit or any other device capable of implementing a storage function for storing instructions and / or data.

[0072] The memory 401 can be used to store software programs and modules, such as the instructions / modules of the data generation device 30 provided in this embodiment of the invention. These can be stored in the memory 401 in the form of software or firmware, or embedded in the operating system (OS) of the electronic device 40. The processor 402 executes various functional applications and data processing by executing the software programs and modules stored in the memory 401. The communication interface 403 can be used to communicate with other node devices for signaling or data.

[0073] I understand. Figure 4 The structure shown is for illustrative purposes only; the electronic device 4 may also include components that are more advanced than those shown. Figure 4 The more or fewer components shown, or having the same Figure 4 The different configurations shown. Figure 4 The components shown can be implemented using hardware, software, or a combination thereof.

[0074] Based on the above embodiments, the present invention also provides a readable storage medium storing a computer program. When the computer program is executed by a computer, it causes the computer to perform the data generation method provided in the above embodiments. For specific implementation details, please refer to the method embodiments, which will not be repeated here.

[0075] The present invention can also provide a computer program product for executing a data generation method, including a computer-readable storage medium storing program code. The instructions included in the program code can be used to execute the methods in the preceding method embodiments. For specific implementation, please refer to the method embodiments, which will not be repeated here.

[0076] In the embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. The apparatus embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and there may be other division methods in actual implementation. Furthermore, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Additionally, the coupling or direct coupling or communication connection shown or discussed may be through some communication interface; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0077] Furthermore, the units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the objectives of the embodiments of the present invention, depending on actual needs.

[0078] Furthermore, the functional modules in the various embodiments of this application can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.

[0079] It should be noted that if the function is implemented as a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, electronic device, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes: USB flash drive, mobile hard drive, read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk, and other media capable of storing program code.

Claims

1. A data generation method, characterized in that, The method includes: Acquire raw multimodal data related to the target content and convert the raw multimodal data into plain text data; A first-stage question-and-answer pair is generated based on the plain text data and prompts. Multiple second-stage question-answer pairs are generated using the first response text and new prompt information from the first-stage question-answer pair; wherein, the question text in the second-stage question-answer pair is generated by disassembling the first response text; The combination of the second response text in each second-stage question-and-answer pair and the question text in the first-stage question-and-answer pair are used as the in-library data of the target content.

2. The data generation method according to claim 1, characterized in that, Multiple second-stage question-answer pairs are generated using the first response text and new prompt information from the first-stage question-answer pair, including: The semantics of the first response text are decomposed into multiple subtexts; Obtain new prompt information for each of the subtexts; Based on the new prompt information, generate a question text about the subtext and an answer text to the question text; The question text and the answer text to the question text are combined to form the second-stage question-answer pair.

3. The data generation method according to claim 2, characterized in that, The combination of the second response texts from each second-stage question-and-answer pair and the question texts from the first-stage question-and-answer pair are used as the database data for the target content, including: Determine the concatenation order of each second response text; The second response text is concatenated according to the concatenation order to obtain the combined result; The combined result and the question text from the first stage question-answer pair are combined to form a new question-answer pair, which is used as data in the library.

4. The data generation method according to claim 3, characterized in that, Determining the concatenation order of each second response text includes: The order in which the subtexts corresponding to the second-stage question-and-answer pairs are placed in the first response text is used as the splicing order of the second response text.

5. The data generation method according to any one of claims 1-4, characterized in that, The original multimodal data is then converted into plain text data, including: Generate semantic description text corresponding to non-text modal data; The semantic description text and text modal data are preprocessed to obtain the plain text data.

6. The data generation method according to claim 5, characterized in that, Generate semantic description text corresponding to non-text modal data, including: Generate visual description text corresponding to image-type data; Perform character recognition on the image data to extract recognizable text; The visual description text and the identifiable text are fused and deduplicated to obtain the semantic description text.

7. The data generation method according to claim 5, characterized in that, Generating semantic description text corresponding to non-text modal data also includes: Extract keyframes from video data and generate visual description text for the keyframes; Perform speech recognition on the audio of the video data to obtain the audio text and the timestamp information of the audio text; The visual description text of the keyframe and the audio text corresponding to the timestamp of the keyframe are combined to form the semantic description text.

8. A data generation apparatus, characterized in that, include: The data acquisition and conversion module is used to acquire raw multimodal data related to the target content and convert the raw multimodal data into plain text data; The data generation module is used to generate a first-stage question-and-answer pair based on the plain text data and prompt information; and to generate multiple second-stage question-and-answer pairs using the first response text and new prompt information in the first-stage question-and-answer pair; wherein the question text in the second-stage question-and-answer pair is generated by disassembling the first response text. The combination of the second response text in each second-stage question-and-answer pair and the question text in the first-stage question-and-answer pair are used as the in-library data of the target content.

9. An electronic device, characterized in that, It includes a processor and a memory, the memory storing machine-executable instructions that can be executed by the processor to implement the data generation method according to any one of claims 1-7.

10. A storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the data generation method as described in any one of claims 1-7.