Data generation method and apparatus, electronic device, and medium

By evaluating and filtering the quality of web page content, high-quality <instructions, answers> alignment data is generated, which solves the problem of low data acquisition efficiency in training of large language models, and achieves efficient and professional data generation.

CN116992112BActive Publication Date: 2025-07-25BAIDU ONLINE NETWORK TECH (BEIJIBG) CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202310804597.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-30
Publication Date
2025-07-25
Estimated Expiration
2043-06-30

AI Technical Summary

Technical Problem

When training large language models, it is difficult to efficiently obtain high-quality alignment data. Manual labeling is inefficient and quality depends on labeling personnel. Automatic mining methods are difficult to cover a variety of fields and task types.

Method used

By obtaining the web page content corresponding to the target generation task, evaluating its content quality, timeliness and authority, filtering out high-quality web page content, and generating problem instructions based on sub-categorization, using training samples to improve data generation efficiency and professionalism.

Benefits of technology

Generate high-quality aligned data in a short time, improves data generation efficiency, ensures professionalism and literary quality, and is suitable for a variety of fields and task types.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116992112B_ABST
    Figure CN116992112B_ABST
Patent Text Reader

Abstract

The present disclosure provides a data generation method, apparatus, electronic device, computer-readable storage medium, and computer program product, which relate to the field of artificial intelligence, and particularly to the fields of deep learning and natural language processing technologies. The implementation solution is as follows: obtaining a plurality of web page contents corresponding to a first document type, where the first document type corresponds to a target generation task; obtaining the score of each web page content among the plurality of web page contents for evaluating at least one of the content quality, timeliness, and authority of the corresponding web page content; filtering the plurality of web page contents based on the scores to obtain at least one web page content whose score exceeds a preset threshold; for each of the at least one web page content: determining a second document type corresponding to the web page content, where the second document type is a subtype of the first document type; and generating a question instruction corresponding to the web page content based on the second document type, with the web page content serving as the answer information corresponding to the question instruction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence, and more particularly to the fields of deep learning and natural language processing technology. Specifically, it relates to a data generation method, apparatus, electronic device, computer-readable storage medium, and computer program product. Background Art

[0002] Artificial intelligence is a discipline that studies how to make a computer simulate certain human thinking processes and intelligent behaviors (such as learning, reasoning, thinking, planning, etc.). It has both hardware-level technologies and software-level technologies. Artificial intelligence hardware technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, and big data processing. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, speech recognition technology, natural language processing technology, and machine learning / deep learning, big data processing technology, and knowledge graph technology.

[0003] In recent years, large language models have refreshed the best results of multiple natural language processing tasks. The large language models have demonstrated powerful natural language understanding and generation capabilities and are widely used in fields such as question answering, search, and intelligent creation. Among them, the high-quality aligned data in the form of <instruction, answer> used in the training process is the key to stimulating the capabilities of the language model and meeting the requirements of downstream applications. Summary of the Invention

[0004] The present disclosure provides a data generation method, apparatus, electronic device, computer-readable storage medium, and computer program product.

[0005] According to one aspect of the present disclosure, there is provided a data generation method, including: obtaining a plurality of web page contents corresponding to a first document type, where the first document type corresponds to a target generation task; obtaining a score for each of the plurality of web page contents, where the score is used to evaluate at least one of the content quality, timeliness, and authority of the corresponding web page content; and filtering the plurality of web page contents based on the score to obtain at least one web page content whose score exceeds a preset threshold. For each of the at least one web page content, perform the following operations: determining a second document type corresponding to the web page content, where the second document type is a subtype of the first document type; and generating a question instruction corresponding to the web page content based on the second document type, where the web page content serves as the answer information corresponding to the question instruction.

[0006] According to another aspect of the present disclosure, there is provided a model training method, including: obtaining a sample to be trained, where the sample to be trained includes web page content and a question instruction corresponding to the web page content, and the question instruction is obtained according to the method of the present disclosure; and training a fourth model based on the sample to be trained.

[0007] According to another aspect of the present disclosure, there is provided a data generation device, including: a first obtaining unit configured to obtain a plurality of web page contents corresponding to a first document type, where the first document type corresponds to a target generation task; a second obtaining unit configured to obtain a score of each of the plurality of web page contents, and the score is used to evaluate at least one of the content quality, timeliness, and authority of the corresponding web page content; and a filtering unit configured to filter the plurality of web page contents based on the score to obtain at least one web page content whose score exceeds a preset threshold; a generating unit configured to, for each of the at least one web page contents, perform the following operations: determining a second document type corresponding to the web page content, where the second document type is a subtype of the first document type; and generating a question instruction corresponding to the web page content based on the second document type, where the web page content serves as answer information corresponding to the question instruction.

[0008] According to another aspect of the present disclosure, there is provided a model training device, including: a second obtaining unit configured to obtain a sample to be trained, where the sample to be trained includes web page content and a question instruction corresponding to the web page content, and the question instruction is obtained according to the method of the present disclosure; and a training unit configured to train a fourth model based on the sample to be trained.

[0009] According to another aspect of the present disclosure, there is provided an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the method of the present disclosure.

[0010] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions for causing a computer to execute the method of the present disclosure.

[0011] According to another aspect of the present disclosure, there is provided a computer program product including a computer program which, when executed by a processor, implements the method of the present disclosure.

[0012] According to one or more embodiments of the present disclosure, for web content that meets the task requirements, high-quality web content is filtered out therefrom, and candidate question instructions are determined by identifying its sub-categories or attributes. Thus, high-quality <instruction, answer> alignment data can be generated in a short time, improving the data generation efficiency and ensuring a certain degree of professionalism and literary quality.

[0013] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become readily apparent from the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] The drawings exemplarily illustrate embodiments and form a part of the specification, and are used together with the written description of the specification to explain the exemplary embodiments of the embodiments. The illustrated embodiments are for illustrative purposes only and do not limit the scope of the claims. In all the drawings, the same reference numerals refer to similar but not necessarily identical elements.

[0015] Figure 1 A schematic diagram of an exemplary system in which various methods described herein can be implemented according to an embodiment of the present disclosure is shown;

[0016] Figure 2 A flowchart of a data generation method according to an embodiment of the present disclosure is shown;

[0017] Figure 3 A flowchart of a model training method according to an embodiment of the present disclosure is shown;

[0018] Figure 4 A flowchart of a fourth model pre-training method according to an embodiment of the present disclosure is shown;

[0019] Figure 5 A block diagram of the structure of a data generation device according to an embodiment of the present disclosure is shown;

[0020] Figure 6 A block diagram of the structure of a model training device according to an embodiment of the present disclosure is shown; and

[0021] Figure 7 A block diagram of the structure of an exemplary electronic device capable of implementing the embodiments of the present disclosure is shown. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0022] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0023] In the present disclosure, unless otherwise specified, the terms "first", "second", etc. are used to describe various elements and are not intended to limit the positional relationship, temporal relationship, or importance relationship of these elements. Such terms are only used to distinguish one element from another. In some examples, the first element and the second element may refer to the same instance of the element, and in certain cases, based on the context description, they may also refer to different instances.

[0024] The terms used in the description of the various examples in the present disclosure are only for the purpose of describing specific examples and are not intended to be limiting. Unless the context clearly indicates otherwise, if the number of elements is not specifically limited, the element may be one or more. In addition, the term "and / or" used in the present disclosure covers any one of the listed items and all possible combinations.

[0025] Embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings.

[0026] Figure 1 A schematic diagram of an exemplary system 100 in which various methods and apparatuses described herein can be implemented according to an embodiment of the present disclosure is shown. Referring to Figure 1 , the system 100 includes one or more client devices 101, 102, 103, 104, 105, and 106, a server 120, and one or more communication networks 110 that couple the one or more client devices to the server 120. The client devices 101, 102, 103, 104, 105, and 106 can be configured to execute one or more applications.

[0027] In an embodiment of the present disclosure, the server 120 can run one or more services or software applications that enable the execution of data generation methods and model training methods.

[0028] In certain embodiments, the server 120 can also provide other services or software applications, and these services or software applications can include non-virtual environments and virtual environments. In certain embodiments, these services can be provided as web-based services or cloud services, for example, provided to users of the client devices 101, 102, 103, 104, 105, and / or 106 under a software as a service (SaaS) model.

[0029] In Figure 1 the configuration shown, server 120 may include one or more components that implement the functions performed by server 120. These components may include software components, hardware components, or a combination thereof that may be executed by one or more processors. Users operating client devices 101, 102, 103, 104, 105, and / or 106 may in turn utilize one or more client applications to interact with server 120 to utilize the services provided by these components. It should be understood that a variety of different system configurations are possible, which may differ from system 100. Thus, Figure 1 is an example of a system for implementing the various methods described herein and is not intended to be limiting.

[0030] Users may use client devices 101, 102, 103, 104, 105, and / or 106 to determine user inputs such as target generation tasks. The client device may provide an interface that enables the user of the client device to interact with the client device. The client device may also output information to the user via the interface. Although Figure 1 only six client devices are depicted, those skilled in the art will be able to understand that the present disclosure may support any number of client devices.

[0031] Client devices 101, 102, 103, 104, 105, and / or 106 may include various types of computing devices such as portable handheld devices, general-purpose computers (such as personal computers and laptop computers), workstation computers, wearable devices, smart screen devices, self-service terminal devices, service robots, gaming systems, thin clients, various messaging devices, sensors, or other sensing devices, etc. These computing devices may run various types and versions of software applications and operating systems such as MICROSOFT Windows, APPLE iOS, UNIX-like operating systems, Linux, or Linux-like operating systems (such as GOOGLE Chrome OS); or include various mobile operating systems such as MICROSOFT WindowsMobile OS, iOS, Windows Phone, Android. Portable handheld devices may include cellular phones, smart phones, tablets, personal digital assistants (PDAs), etc. Wearable devices may include head-mounted displays (such as smart glasses) and other devices. Gaming systems may include various handheld gaming devices, Internet-enabled gaming devices, etc. The client device is capable of executing various different applications such as various Internet-related applications, communication applications (such as email applications), short message service (SMS) applications, and may use various communication protocols.

[0032] Network 110 can be any type of network well-known to those skilled in the art, which can support data communication using any one of a variety of available protocols (including but not limited to TCP / IP, SNA, IPX, etc.). By way of example only, one or more networks 110 can be a local area network (LAN), an Ethernet-based network, Token Ring, a wide area network (WAN), the Internet, a virtual network, a virtual private network (VPN), an intranet, an extranet, a blockchain network, a public switched telephone network (PSTN), an infrared network, a wireless network (such as Bluetooth, WIFI), and / or any combination of these and / or other networks.

[0033] Server 120 can include one or more general-purpose computers, dedicated server computers (such as PC (personal computer) servers, UNIX servers, midrange servers), blade servers, mainframes, server clusters, or any other suitable arrangement and / or combination. Server 120 can include one or more virtual machines running a virtual operating system, or other computing architectures involving virtualization (such as one or more flexible pools of logical storage devices that can be virtualized to maintain virtual storage devices of the server). In various embodiments, server 120 can run one or more services or software applications that provide the functions described below.

[0034] The computing units in server 120 can run one or more operating systems including any of the above operating systems as well as any commercially available server operating systems. Server 120 can also run any one of a variety of additional server applications and / or middleware applications, including HTTP servers, FTP servers, CGI servers, JAVA servers, database servers, etc.

[0035] In some embodiments, server 120 can include one or more applications to analyze and merge data feeds and / or event updates received from users of client devices 101, 102, 103, 104, 105, and 106. Server 120 can also include one or more applications to display data feeds and / or real-time events via one or more display devices of client devices 101, 102, 103, 104, 105, and 106.

[0036] In some embodiments, the server 120 can be a server of a distributed system or a server combined with a blockchain. The server 120 can also be a cloud server or an intelligent cloud computing server or an intelligent cloud host with artificial intelligence technology. A cloud server is a host product in a cloud computing service system, which is used to solve the defects of high management difficulty and weak business scalability existing in traditional physical hosts and virtual private server (VPS) services.

[0037] The system 100 may further include one or more databases 130. In certain embodiments, these databases can be used to store data and other information. For example, one or more of the databases 130 can be used to store information such as training data and model parameters. The databases 130 can reside in various locations. For example, the database used by the server 120 can be local to the server 120, or can be remote from the server 120 and can communicate with the server 120 via a network-based or dedicated connection. The databases 130 can be of different types. In certain embodiments, the database used by the server 120 can be a relational database, for example. One or more of these databases can store, update, and retrieve data to and from the database in response to commands.

[0038] In certain embodiments, one or more of the databases 130 can also be used by an application to store application data. The database used by the application can be a different type of database, such as a key-value store, an object store, or a conventional store supported by a file system.

[0039] Figure 1 The system 100 can be configured and operated in various ways so that various methods and apparatuses described according to the present disclosure can be applied.

[0040] The alignment data formed as <instruction, answer> is applied to the training phase of the large language model. Among them, <instruction> can be, for example, a text describing a clear requirement and can also be regarded as a question; while <answer> is the text content that can meet the requirement or the answer to the question.

[0041] Currently, the acquisition of <instruction, answer> alignment data usually adopts the methods of manual annotation or automatic mining. Among them, the method of manual annotation first manually writes or collects instructions, and then manually writes the corresponding answer, thereby obtaining <instruction, answer> alignment data. The method of automatic mining is to directly extract the aligned <instruction, answer> data from Internet texts through rules or models.

[0042] Different from the data annotation requirements of traditional single natural language processing tasks, the training stage of large language models (such as the fine-tuning stage) usually requires aligned data of multiple domains and multiple tasks, posing greater challenges to the large-scale generation of high-quality aligned data.

[0043] In this context, both manual annotation and automatic mining methods face their respective limitations: The manual annotation method is inefficient, and the annotation quality is closely related to the ability of the annotators. The requirements for the ability of annotators for <instruction, answer> aligned data in different domains vary. For example, the annotation of <instruction, answer> aligned data in the legal field requires the annotator to have sufficient legal knowledge. For instructions in long text creation such as poetry creation, novel creation, and script creation, manual annotation usually cannot generate high-quality <instruction, answer> aligned data in a short time. Although the automatic mining method is efficient, the supported task types are relatively limited. Most unsupervised texts on the Internet do not exist in the form of question-answer pairs, that is, many high-quality contents that can be used as answers do not have corresponding instructions. Therefore, it is also difficult to obtain <instruction, answer> aligned data suitable for large language model fine-tuning through the automatic mining method.

[0044] Therefore, according to an embodiment of the present disclosure, a data annotation method is provided. Figure 2 The flowchart of the data annotation method according to an embodiment of the present disclosure is shown. As Figure 2 shown, method 200 includes: obtaining a plurality of web page contents corresponding to a first document type, where the first document type corresponds to a target generation task (step 210); obtaining a score for each of the plurality of web page contents, where the score is used to evaluate at least one of the content quality, timeliness, and authority of the corresponding web page content (step 220); and filtering the plurality of web page contents based on the score to obtain at least one web page content whose score exceeds a preset threshold (step 230); for each of the at least one web page content, perform the following operations (step 240): determining a second document type corresponding to the web page content, where the second document type is a subtype of the first document type (step 2401); and generating a question instruction corresponding to the web page content based on the second document type, where the web page content serves as the answer information corresponding to the question instruction (step 2402).

[0045] According to an embodiment of the present disclosure, for web page contents that meet the task requirements, high-quality web page contents are filtered out, and candidate question instructions are determined by determining their subcategories or attributes. Thus, high-quality <instruction, answer> aligned data can be generated in a short time, improving the data generation efficiency and ensuring a certain degree of professionalism and literary grace.

[0046] In some embodiments, the first quantity of web page content can be obtained by means of a web crawler. A crawler is an application or script for crawling web pages, mainly including traditional crawlers and focused crawlers. Further, search engines usually use crawlers to crawl web pages, and analyze, filter, and index the crawled web page content for web page query and retrieval.

[0047] In some embodiments, web page content can also be obtained through preset web page address information, which may include the address url of the web page, etc. The web page can be located according to the web page url. For example, the web page structure information can be obtained by using the scrapy-redis framework through the web page address information. As a distributed crawler framework, scrapy-redis is usually used for large-scale website data collection.

[0048] In the present disclosure, the first document type can correspond to a target generation task. For example, the target generation task can be a legal data generation task, a poetry data generation task, etc. Therefore, the first document type can be news, poetry, law, entertainment, military, etc.

[0049] In some examples, the html code corresponding to the web page block can be obtained, and all text content and / or image content, etc. can be extracted therefrom. In some examples, the obtained web page content can also be part of the web page, such as a text fragment. For example, the source code of the web page block, that is, the html code, can be obtained first, and all text nodes in the web page block are split into fragments in sequence. If the text contains a delimiter, the parts before and after the delimiter respectively form fragments, and thus text fragments are obtained. Further, this part of the content can be the core text information in the web page. For example, based on a preset extraction method, the core text information is obtained from the web page structure information. The web page structure information is equivalent to the web page source code, and is composed of elements such as web page html, head, body, and corresponding text, picture information, style css, and script javescript.

[0050] For example, the core component of the core text information is the body text information. Since in a web page, the text density and symbol density in the area where the body is located are the largest, in other areas, such as the top navigation bar, the side advertisement bar, and the bottom information bar area, the text is scarce and there are almost no text symbols. Therefore, the core text area can be determined by using the text density and symbol density as the judgment basis.

[0051] Thus, according to some embodiments, obtaining the score of each of the multiple web page contents includes: inputting the multiple web page contents into a trained first model respectively to obtain the score of each web page content. The score is used to evaluate at least one of the content quality, timeliness, and authority of the corresponding web page content.

[0052] Exemplarily, different aspects of the web page content can be scored separately. The different aspects can be, for example, the author, creation time, content quality, etc. For example, in the poem data generation task, in terms of the author, the score of ancient great poets is higher than that of other ordinary authors. By adding up the scores of each aspect, the final score of the web page content is obtained.

[0053] In some examples, corresponding weights can also be set for the different aspects, and the final score of the web page content is obtained by weighted summation.

[0054] By obtaining the score of the web page content, filtering of low-quality or irrelevant web page content based on the web page content or other auxiliary information is achieved, thereby ensuring that the obtained web page content contains high-quality documents with strong professionalism and good literary grace.

[0055] According to some embodiments, the second document type is category data obtained after deep semantic refinement of production clues. For example, when the first document type is poetry, the second document type can be genre, theme, image, etc. Exemplarily, the genre can be five-character quatrains, seven-character quatrains, seven-character regulated verses, etc., the theme can be reminiscing about the past, chanting things, landscape and pastoral, boudoir resentment, farewell, etc., and the image can be plum blossoms, the moon, snow, etc.

[0056] It can be understood that the second document type can be one or more. For example, it can be only the poetry genre, or it can be the poetry genre and images, etc., without limitation here.

[0057] According to some embodiments, determining the second document type corresponding to the web page content includes: inputting the web page content into a trained second model to obtain the second document type corresponding to the web page content.

[0058] Exemplarily, the second model can be a classification model to determine its corresponding second document type. And when the second document type is multiple, each second document type can be determined through the corresponding second model respectively. For example, input the web page content into the second model for determining the genre to obtain the corresponding genre; further input the web page content into the second model for determining the theme to obtain the corresponding theme.

[0059] According to some embodiments, generating a question instruction corresponding to the web page content based on the second document type includes: filling information into a preset template based on the second document type to generate a question instruction corresponding to the web page content.

[0060] In some examples, one or more question instruction templates may be preset to select a suitable or corresponding template based on the determined second document type for information filling to generate a question instruction corresponding to the web page content. For example, the question template may be: Please write a poem about XX (theme) containing XXX (image). Thus, the information of the determined second document type including the theme and the image can be filled into the template to obtain a question instruction such as "Please write a poem about'sending off' containing 'the moon'".

[0061] According to some embodiments, generating a question instruction corresponding to the web page content based on the second document type includes: inputting the second document type into a trained third model to generate a question instruction corresponding to the web page content.

[0062] In some embodiments, the third model may be a generative language model, such as GPT, UniLM, etc. By inputting one or more keywords of the determined second document type into the third model, a semantically coherent text including the one or more keywords is obtained as the question instruction.

[0063] According to some embodiments, the method according to the present disclosure may further include: scoring the matching degree between the generated question instruction and the web page content to obtain a corresponding score value; and fine-tuning at least one of the question instruction and the web page content based on the score value.

[0064] In some examples, the generated question instruction may be scored by a trained network model to obtain a corresponding score value. For example, the network model may be trained with <answer, instruction> alignment data that has been manually reviewed and annotated as samples so that it can evaluate the alignment degree between the instruction and the answer and obtain a corresponding score value. Fine-tuning aims to ensure the quality of each of the candidate instruction and the candidate answer and the alignment degree between them. In some examples, first, the overall quality of the candidate instruction and the candidate answer needs to be evaluated, and then after rewriting, it is ensured to meet the final output requirements, or the low-quality <answer, instruction> pair is directly discarded.

[0065] According to the embodiments of the present disclosure, first, the web page content that meets the task requirements is screened and filtered, possible candidate answers are extracted from it, and candidate instructions are determined by determining their sub-categories or attributes. Finally, after reviewing and rewriting the candidate instructions and candidate answers, high-quality <instruction, answer> data is obtained.

[0066] According to an embodiment of the present disclosure, as Figure 3 shown, a model training method 300 is further provided, including: obtaining a sample to be trained, where the sample to be trained includes web page content and a question instruction corresponding to the web page content, and the question instruction is obtained according to the method described in any one of the above embodiments (step 310); and training a fourth model based on the sample to be trained (step 320).

[0067] According to some embodiments, the fourth model is a pre-trained network model. As Figure 4 shown, the fourth model is pre-trained through the following steps: obtaining a sample text corpus (step 410); at least dividing the sample text corpus into two texts that are semantically coherent before and after (step 420); inputting the first text of the two texts into the fourth model to obtain the text predicted by the fourth model, where the predicted text corresponds to the second text of the two texts (step 430); and adjusting the parameters of the fourth model based on the second text and the predicted text (step 440).

[0068] According to an embodiment of the present disclosure, as Figure 5 shown, a data generation device 500 is further provided, including: a first acquisition unit 510 configured to acquire a plurality of web page contents corresponding to a first document type, where the first document type corresponds to a target generation task; a second acquisition unit 520 configured to acquire the score of each web page content in the plurality of web page contents, and the score is used to evaluate at least one of the content quality, timeliness, and authority of the corresponding web page content; and a filtering unit 530 configured to filter the plurality of web page contents based on the score to obtain at least one web page content whose score exceeds a preset threshold; a generation unit 540 configured to perform the following operations for each of the at least one web page content: determining a second document type corresponding to the web page content, where the second document type is a subtype of the first document type; and generating a question instruction corresponding to the web page content based on the second document type, where the web page content serves as the answer information corresponding to the question instruction.

[0069] Here, the operations of the above units 510-540 of the data generation device 500 are respectively similar to the operations of steps 210-240 described above, and will not be repeated here.

[0070] According to an embodiment of the present disclosure, as Figure 6As shown in the figure, a model training device 600 is further provided, including: a third acquisition unit 610 configured to acquire a sample to be trained, where the sample to be trained includes web content and a question instruction corresponding to the web content, and the question instruction is obtained according to the method described in any of the above embodiments; and a training unit 620 configured to train a fourth model based on the sample to be trained.

[0071] Here, the operations of the above units 610-620 of the model training device 600 are respectively similar to the operations of steps 310-320 described above, and will not be elaborated here.

[0072] In the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision, and disclosure of the user's personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0073] According to an embodiment of the present disclosure, an electronic device, a readable storage medium, and a computer program product are further provided.

[0074] Reference Figure 7 , the structural block diagram of an electronic device 700 that can be used as the server or client of the present disclosure will now be described. It is an example of a hardware device that can be applied to various aspects of the present disclosure. The electronic device is intended to represent various forms of digital electronic computer devices, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processors, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are only examples and are not intended to limit the implementation of the present disclosure described and / or required herein.

[0075] As Figure 7 shown, the electronic device 700 includes a computing unit 701, which can execute various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 702 or the computer program loaded from the storage unit 708 into the random access memory (RAM) 703. In the RAM 703, various programs and data required for the operation of the electronic device 700 can also be stored. The computing unit 701, the ROM 702, and the RAM 703 are connected to each other through a bus 704. The input / output (I / O) interface 705 is also connected to the bus 704.

[0076] A plurality of components in the electronic device 700 are connected to the I / O interface 705, including: an input unit 706, an output unit 707, a storage unit 708, and a communication unit 709. The input unit 706 can be any type of device capable of inputting information into the electronic device 700. The input unit 706 can receive input digital or character information, and generate key signal inputs related to the user settings and / or function controls of the electronic device, and can include, but are not limited to, a mouse, a keyboard, a touch screen, a trackpad, a trackball, a joystick, a microphone, and / or a remote control. The output unit 707 can be any type of device capable of presenting information, and can include, but are not limited to, a display, a speaker, a video / audio output terminal, a vibrator, and / or a printer. The storage unit 708 can include, but are not limited to, magnetic disks and optical discs. The communication unit 709 allows the electronic device 700 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks, and can include, but are not limited to, a modem, a network card, an infrared communication device, a wireless communication transceiver, and / or a chipset, such as a Bluetooth device, an 802.11 device, a WiFi device, a WiMax device, a cellular communication device, and / or the like.

[0077] The computing unit 701 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 701 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 701 executes the various methods and processes described above, such as method 200 or 300. For example, in some embodiments, method 200 or 300 can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 708. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 700 via the ROM 702 and / or the communication unit 709. When the computer program is loaded into the RAM 703 and executed by the computing unit 701, one or more steps of the method 200 or 300 described above can be executed. Alternatively, in other embodiments, the computing unit 701 can be configured to execute method 200 or 300 in any other suitable manner (e.g., by means of firmware).

[0078] The various embodiments of the systems and techniques described above in this specification can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems-on-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that receives data and instructions from, and transmits data and instructions to, a storage system, at least one input device, and at least one output device.

[0079] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus, such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The program code can be executed entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine or entirely on the remote machine or server.

[0080] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0081] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and a pointing device (e.g., a mouse or a trackball), by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).

[0082] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), the Internet, and blockchain network.

[0083] A computer system can include a client and a server. The client and the server are generally far from each other and usually interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, or a server of a distributed system, or a server incorporating blockchain.

[0084] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and no limitation is made herein.

[0085] Although embodiments or examples of the present disclosure have been described with reference to the accompanying drawings, it should be understood that the above methods, systems, and devices are merely exemplary embodiments or examples, and the scope of the present invention is not limited by these embodiments or examples, but is only defined by the authorized claims and their equivalent scope. Various elements in the embodiments or examples may be omitted or replaced by their equivalent elements. In addition, the steps may be executed in an order different from that described in the present disclosure. Further, the various elements in the embodiments or examples may be combined in various ways. Importantly, with the evolution of technology, many of the elements described herein may be replaced by equivalent elements that emerge after the present disclosure.

Claims

1. A data generation method, comprising: Obtaining a plurality of web page contents corresponding to a first document type, wherein the first document type corresponds to a target generation task; Obtaining a score for each of the plurality of web page contents, the score being used to evaluate at least one of the content quality, timeliness, and authority of the corresponding web page content; And Filtering the plurality of web page contents based on the score to obtain at least one web page content whose score exceeds a preset threshold; For each of the at least one web page content, perform the following operations: Determining a second document type corresponding to the web page content, wherein the second document type is a subtype of the first document type, and the second document type is category data obtained after deep semantic refinement of the web page content; And Generating a question instruction corresponding to the web page content based on the second document type, including: filling information into a preset template based on the second document type to generate a question instruction corresponding to the web page content, wherein the web page content serves as the answer information corresponding to the question instruction.

2. The method according to claim 1, wherein, The second document type is category data obtained after deep semantic refinement of the web page content.

3. The method according to claim 1 or 2, wherein Obtaining a score for each of the plurality of web page contents includes: inputting the plurality of web page contents into a trained first model respectively to obtain the score for each web page content.

4. The method according to claim 1, wherein, Determining the second document type corresponding to the web page content includes: inputting the web page content into a trained second model to obtain the second document type corresponding to the web page content.

5. The method according to claim 1, wherein, Generating a question instruction corresponding to the web page content based on the second document type includes: inputting the second document type into a trained third model to generate a question instruction corresponding to the web page content.

6. The method according to claim 1, further comprising: Scoring the matching degree between the generated question instruction and the web page content to obtain a corresponding score value; And Fine-tuning at least one of the question instruction and the web page content based on the score value.

7. A model training method, comprising: Obtaining a training sample to be trained, wherein the training sample to be trained includes a web page content and a question instruction corresponding to the web page content, and the question instruction is obtained according to the method as claimed in any one of claims 1-6; and Training a fourth model based on the training sample to be trained.

8. The method according to claim 7, wherein The fourth model is a pre-trained network model, wherein the fourth model is pre-trained through the following steps: Obtaining a sample text corpus; At least dividing the sample text corpus into two texts that are semantically coherent before and after; Inputting the first text of the two texts into the fourth model to obtain the text predicted by the fourth model, wherein the predicted text corresponds to the second text of the two texts; and Adjusting the parameters of the fourth model based on the second text and the predicted text.

9. A data generation device, comprising: A first acquisition unit, configured to acquire a plurality of web page contents corresponding to a first document type, where the first document type corresponds to a target generation task; A second acquisition unit, configured to acquire a score for each of the plurality of web page contents, where the score is used to evaluate at least one of the content quality, timeliness, and authority of the corresponding web page content; And A filtering unit, configured to filter the plurality of web page contents based on the score to obtain at least one web page content whose score exceeds a preset threshold; A generation unit, configured to perform the following operations for each of the at least one web page content: Determine a second document type corresponding to the web page content, where the second document type is a subtype of the first document type, and the second document type is category data obtained after deep semantic refinement of the web page content; And Based on the second document type, generate a question instruction corresponding to the web page content, including: filling information into a preset template based on the second document type to generate a question instruction corresponding to the web page content, where the web page content serves as the answer information corresponding to the question instruction.

10. The apparatus according to claim 9, wherein, The second document type is category data obtained after deep semantic refinement of the web page content.

11. The device according to claim 9 or 10, wherein, The second acquisition unit is configured to: input the plurality of web page contents into a trained first model respectively to obtain the score of each web page content.

12. The device according to claim 9, wherein, The operation of determining the second document type corresponding to the web page content includes: inputting the web page content into a trained second model to obtain the second document type corresponding to the web page content.

13. The device according to claim 9, wherein The operation of generating a question instruction corresponding to the web page content based on the second document type includes: inputting the second document type into a trained third model to generate a question instruction corresponding to the web page content.

14. The apparatus according to claim 9, further comprising: A scoring unit, configured to score the matching degree between the generated question instruction and the web page content to obtain a corresponding score value; And A fine-tuning unit, configured to fine-tune at least one of the question instruction and the web page content based on the score value.

15. A model training apparatus, comprising: A third acquisition unit, configured to acquire a training sample to be trained, where the training sample to be trained includes a web page content and a question instruction corresponding to the web page content, and the question instruction is obtained according to any one of claims 1-6; and A training unit, configured to train a fourth model based on the training sample to be trained.

16. The apparatus according to claim 15, wherein, The fourth model is a pre-trained network model, where the fourth model is pre-trained through the following steps: Acquire a sample text corpus; At least divide the sample text corpus into two texts that are semantically coherent before and after; Input the first text of the two texts into the fourth model to obtain the text predicted by the fourth model, where the predicted text corresponds to the second text of the two texts; and Based on the second text and the predicted text, adjust the parameters of the fourth model.

17. An electronic device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method according to any one of claims 1-8.

18. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method according to any one of claims 1-8.

19. A computer program product comprising a computer program, wherein, The computer program, when executed by a processor, implements the method according to any one of claims 1-8.

Citation Information

Patent Citations

  • Data high-speed processing conversion communication method and device

    CN108399205A

  • Problem generation method and device and storage medium

    CN109726274A

  • Problem generation model training method, problem generation method and related equipment thereof

    CN111639163A

  • Problem generation method and device based on judicial court trial and computing equipment

    CN113569040A

  • Question and answer page recommendation method and device, equipment and storage medium

    CN116303910A