Method, device, medium and electronic equipment for generating large model question and answer data
By obtaining the summary content of seed entities and the associated entity set, the target question is generated, which solves the problems of high cost and low quality in question-and-answer database generation in the FAQ system, and realizes high-quality and low-illusion question-and-answer data generation.
Patent Information
- Application Number
- CN202511447856.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-10
- Publication Date
- 2026-01-23
- Estimated Expiration
- 2045-10-10
AI Technical Summary
Existing FAQ systems suffer from high costs and low quality in generating question-and-answer databases, making it difficult to generate high-quality, low-illusion question-and-answer data.
By acquiring seed entities, searching to obtain summary content, determining the set of associated entities, and generating target questions and seed entities based on the set of associated entities, a large model question-answering data is formed.
High-quality large-scale model question-and-answer data was generated, avoiding answer illusion and improving the accuracy and consistency of the question-and-answer data.
Smart Images

Figure CN120910112B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of artificial intelligence, in particular, to a large model question and answer data generation method and device, medium and electronic equipment. BACKGROUND
[0002] With the rapid development of artificial intelligence, FAQ (Frequently Asked Questions) based on intelligent scenarios has gradually become intelligent, and more and more users can use FAQ systems for unmanned problem solving. However, the FAQ question and answer database generation has problems such as high cost and low generation quality. SUMMARY
[0003] This summary is provided to introduce a selection of concepts that are further described below in the detailed description. This summary is not intended to identify key or essential features of the claimed subject matter, nor is it used to limit the scope of the claimed subject matter's scope.
[0004] In a first aspect, the present disclosure provides a large model question and answer data generation method, comprising:
[0005] obtaining a seed entity, and searching the seed entity to obtain summary content, the summary content being summary content of a plurality of search results related to the seed entity;
[0006] determining a set of associated entities of the seed entity based on the summary content, the set of associated entities including a plurality of associated entities;
[0007] generating a target question according to the set of associated entities, and generating the large model question and answer data from the target question and the seed entity.
[0008] In a second aspect, the present disclosure provides a large model question and answer data generation device, comprising:
[0009] a search module configured to obtain a seed entity, and search the seed entity to obtain summary content, the summary content being summary content of a plurality of search results related to the seed entity;
[0010] a determination module configured to determine a set of associated entities of the seed entity based on the summary content, the set of associated entities including a plurality of associated entities;
[0011] a generation module configured to generate a target question according to the set of associated entities, and generate the large model question and answer data from the target question and the seed entity.
[0012] In a third aspect, the present disclosure provides a computer readable medium having stored thereon a computer program which, when executed by a processing device, implements the steps of the method of the first aspect.
[0013] In a fourth aspect, the present disclosure provides an electronic device comprising:
[0014] a storage device having stored thereon a computer program;
[0015] a processing device configured to execute the computer program in the storage device to implement the steps of the method of the first aspect.
[0016] In a fifth aspect, the present disclosure provides a computer program product comprising a computer program which, when executed by a processor, implements the steps of the method of the first aspect.
[0017] Through the above technical solution, in the case of extracting a seed entity, the associated entity set associated with the seed entity can be obtained by searching the seed entity, the associated entity set includes a plurality of associated entities, and on this basis, the target question is generated according to the associated entities. Since the answer to the generated target question is unique, that is, the answer to the target question is the seed entity, the illusion problem of the answer can be avoided to some extent, and since the associated entities are obtained by searching the summary content, high-quality large model question and answer data can be generated.
[0018] Other features and advantages of the present disclosure will be described in detail in the following detailed description. BRIEF DESCRIPTION OF DRAWINGS
[0019] The above and other features, advantages, and aspects of embodiments of the present disclosure will become more apparent by describing in detail exemplary embodiments thereof with reference to the attached drawings in which:
[0020] Figure 1 is a flowchart of a large model question and answer data generation method according to an embodiment of the present disclosure.
[0021] Figure 2 is an example diagram of searching for a page number in a large model question and answer data generation method according to an embodiment of the present disclosure.
[0022] Figure 3 is an example diagram of a pure text summary page obtained by preprocessing summary content in a large model question and answer data generation method according to an embodiment of the present disclosure.
[0023] Figure 4is an example diagram of an entity list in a large model question and answer data generation method according to an embodiment of the present disclosure.
[0024] Figure 5 is an example diagram of a triple relationship list in a large model question and answer data generation method according to an embodiment of the present disclosure.
[0025] Figure 6 is an example diagram of an attribute list in a large model question and answer data generation method according to an embodiment of the present disclosure.
[0026] Figure 7 is an example diagram of a knowledge graph constructed in a large model question and answer data generation method according to an embodiment of the present disclosure.
[0027] Figure 8 is an example diagram of relationship chain extension in a large model question and answer data generation method according to an embodiment of the present disclosure.
[0028] Figure 9 is an example diagram of input data in a large model question and answer data generation method according to an embodiment of the present disclosure.
[0029] Figure 10 is an example diagram of another input data in a large model question and answer data generation method according to an embodiment of the present disclosure.
[0030] Figure 11 is a specific flowchart of generating FAQ in a large model question and answer data generation method according to an embodiment of the present disclosure.
[0031] Figure 12 is a block diagram of a large model question and answer data generation device according to an embodiment of the present disclosure.
[0032] Figure 13 is a structural schematic diagram of an electronic device according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0033] Embodiments of the present disclosure will be described in more detail by referring to the drawings. Although certain embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments set forth herein, but rather the embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for exemplary purposes only and are not intended to limit the scope of protection of the present disclosure.
[0034] It should be understood that each step recited in the method embodiments of the present disclosure can be performed in different orders and / or in parallel. In addition, the method embodiments can include additional steps and / or omit performing the steps shown. The scope of the present disclosure is not limited in this respect.
[0035] The term "comprising" and variations thereof as used herein are open-ended, that is, "comprising but not limited to." The term "based on" means "based, at least in part, on." The term "one embodiment" means "at least one embodiment." The term "another embodiment" means "at least one additional embodiment." The term "some embodiments" means "at least some embodiments." Related terms have corresponding meanings.
[0036] It should be noted that the terms "first", "second", and the like in the present disclosure are merely used to distinguish different devices, modules or units, and do not imply the order or interdependence of the functions performed by these devices, modules or units.
[0037] It should be noted that the terms "one", "multiple" in the present disclosure are illustrative and not restrictive, and those skilled in the art should understand that unless otherwise explicitly stated in the context, it should be understood as "one or more".
[0038] The names of the messages or information exchanged between the devices in the embodiments of the present disclosure are only for illustrative purposes, and are not intended to limit the scope of the messages or information.
[0039] For example, in response to receiving an active request of a user, a prompt message is sent to the user to explicitly prompt the user that the operation requested by the user will require obtaining and using personal information of the user. Thus, the user can voluntarily choose whether to provide personal information to the software or hardware such as electronic devices, application programs, servers or storage media that perform the operation of the technical solutions of the present disclosure according to the prompt message.
[0040] As an optional but non-limiting implementation, in response to receiving an active request of a user, the way of sending a prompt message to the user may, for example, be a pop-up window, in which the prompt message can be presented in the form of text. In addition, the pop-up window can also carry selection controls for the user to select "agree" or "disagree" to provide personal information to the electronic device.
[0041] At the same time, it can be understood that the data involved in the technical solutions (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of relevant laws, regulations and provisions.
[0042] Large language models (LLMs) are widely used in customer service, education, search engines, and other scenarios. In related technologies, large language models mainly rely on pre-training corpus. When facing real-time and frequently updated information, there are usually problems such as knowledge aging, answer illusion, and retrieval generation link fragmentation. Among them, knowledge aging refers to the inability of large language models to grasp the latest information; answer illusion refers to the deviation of generated content from real information; and retrieval generation link fragmentation refers to the difficulty of large language models in combining external knowledge for dynamic updating during data generation.
[0043] Therefore, to solve the above problems, the embodiments of the present disclosure provide a large model question and answer data generation method, device, medium and electronic equipment to generate high-quality and low-illusion large model question and answer data.
[0044] The present disclosure will be further explained and described below in conjunction with the accompanying drawings.
[0045] Figure 1 FIG. 1 is a flowchart of a large model question and answer data generation method according to an embodiment of the present disclosure. The large model question and answer data generation method can be applied to an electronic device with processing capability, such as a terminal or a server. The large model question and answer data generation method can be executed by a large model question and answer data generation device, which can be implemented by software and / or hardware, and the software and / or hardware can be configured in the electronic device. Referring to FIG. 1, the large model question and answer data generation method can include the following steps. Figure 1 The large model question and answer data generation method can include the following steps.
[0046] In step S110, a seed entity is obtained, and the seed entity is searched to obtain summary content, which is the summary content of a plurality of search results related to the seed entity.
[0047] In the embodiments of the present disclosure, the seed entity can be obtained by entity extraction on the input question, for example, the user input search question is “how to take care of green plants”, and here, green plants can be used as a seed entity. Alternatively, the seed entity can also be an entity extracted from an entity library when constructing a large model question and answer data. The entity library can be a pre-constructed entity database, which can include a plurality of benchmark entities. In the process of generating a large model question and answer data, any entity can be extracted from these benchmark entities as a seed entity.
[0048] In the embodiments of the present disclosure, the entity library can be pre-constructed, and the entity library can also be updated in real time during the generation of the large model question and answer data, that is, the entity library can be updated based on the seed entity and the entity association set. The acquisition of the entity association set will be described in detail in the following embodiments, which will not be described here.
[0049] The entity library can include an entity set, a relationship set, and an attribute set. The entity set can include a plurality of reference entities, which can each be a seed entity, and each reference entity can be configured with an entity label. Here, the entity set can be used to store information related to the entity. For example, the entity set can include an entity name, the number of times the entity name is used as a subject (counts), the number of times the entity name is used as an object (counto), and whether to end the search, and the like. The entity set can be as shown in Table 1.
[0050] Table 1
[0051]
[0052] Based on Table 1, the entity set can include, in addition to the entity name of the reference entity, the number of times the reference entity is used as a subject, the number of times the reference entity is used as an object, and whether to end the search for the entity. When the value of searched is t, the search for the reference entity is ended. The value of searched can be determined based on the values of counts and counto. For example, when the values of counts and / or counto are greater than a preset value, the value of searched can be set to t. The information in the entity set described above can be used to construct subsequent large model question and answer data.
[0053] The relationship set is mainly used to record the relationships of the reference entities in the entity set. The relationship set can include SPO (Subject Predict Object) triples. The relationship set corresponding to each reference entity in Table 1 can be as shown in Table 2.
[0054] Table 2
[0055]
[0056] Based on Table 2, both person A and person D are characters involved in work A (a novel), and work A is created by author A. In addition, person B, person C, and person D all have relationships with person A (the protagonist).
[0057] The attribute set is used to store the attributes of each reference entity in the entity set. Based on the example described above, the attributes of each reference entity can be as shown in Table 3.
[0058] Table 3
[0059]
[0060] Based on Table 3, it can be seen that the attribute set can include attributes specific to each reference entity, such as the character traits of person D are "simple, kind, sensible, and diligent".
[0061] In the embodiments of the present disclosure, each reference entity can be configured with an entity label, wherein the entity label can include at least one of popularity information, a domain label, a task type label, and a rejection label of the entity. The popularity information can include a popularity category and a popularity score, the popularity category can include small and non-small (mass), and the higher the popularity score is, the more popular the entity is.
[0062] In the embodiments of the present disclosure, the score of the entity can be obtained by searching engine result score, encyclopedia coverage, social platform heat, and whether it frequently appears in mainstream media. When it is determined that the popularity score is greater than a preset score, the entity is determined to be a popular entity (non-small entity). Conversely, if it is determined that the popularity score is less than or equal to the preset score, the entity can be determined to be a small entity.
[0063] For example, the popularity score of the first reference entity is less than 45, and it is determined to be a small entity.
[0064] The domain label refers to the domain type to which the entity belongs, which can include a first-level domain and a second-level domain; the task type label refers to the scene type to which the entity belongs, which can include a first-level scene and a second-level scene; the rejection label can be a pre-rejection entity label, which is used to indicate whether the entity is clear and explicit. If the entity is ambiguous or the definition is ambiguous, the value of the rejection label can be set to True. Conversely, if it is determined that the entity is not ambiguous or ambiguous, the rejection label of the entity can be set to False.
[0065] In order to better understand the information in the entity label, the present disclosure provides Table 4 as shown below:
[0066] Table 4
[0067]
[0068] Based on Table 4, it can be seen that the first-level domain to which the entity belongs can include person, place, product, organization, time, concept, biology, and technology; the second-level domain can include scientific person, university, legal concept, and film and television work; the first-level scene can include general knowledge Q&A, enterprise service and internal work, business analysis and industry research, professional discipline research, news investigation and interpretation, and personal life entertainment; the second-level scene can include science and technology, social humanities, economic business, cultural art, life entertainment, comprehensive and other, and game strategy tutorial.
[0069] In the process of obtaining the label information of each entity, the embodiment of the present disclosure can determine whether the reference entity satisfies the elimination condition, and if the reference entity satisfies the elimination condition, the embodiment of the present disclosure can set the elimination label of the reference entity to a True value. Here, the elimination condition can include at least one of the following: the reference entity has ambiguity; the reference entity is a fuzzy category entity; and the reference entity is a broad category generalization.
[0070] Here, having ambiguity can mean that the ambiguity value of the reference entity is higher than an ambiguity threshold value, that is, when it is determined that the reference entity is named with multiple irrelevant semantics or entities, the elimination label thereof can be set to a True value; the fuzzy category entity refers to an entity that does not have a specific meaning, such as “a person”, “an organization”, and “important person” and the like are all fuzzy category entities; and the broad category generalization refers to a category rather than a specific object, such as the concept of “animal” and “literary work” is broad and cannot be specifically positioned.
[0071] That is, in the process of obtaining the elimination label of the reference entity, the embodiment of the present disclosure can determine whether the reference entity satisfies the at least one condition, and if so, the elimination label thereof can be set to a True value, and if not, the elimination label thereof can be set to a False value.
[0072] As an optional way, when it is determined that the value of the elimination label of the reference entity is a True value, the embodiment of the present disclosure can perform clarity processing on the reference entity, that is, clarity processing on the reference entity whose elimination label value is True. Here, the clarity processing can be modifying the name of the reference entity, so that it changes from an ambiguous entity to a non-ambiguous entity, or changes from a fuzzy entity to a clear entity. For example, the clarity processing can be adding a limit word to the entity name, such as processing the entity “Chang’an” to “Chang’an District”.
[0073] Alternatively, in the process of obtaining the seed entity, the embodiment of the present disclosure can determine whether the elimination label of the seed entity is a first specified value (False), and if it is determined that the value of the elimination label is the first specified value, the embodiment of the present disclosure can take it as a seed entity. That is, the embodiment of the present disclosure can select a seed entity from a plurality of reference entities based on the entity label, and the value of the elimination label corresponding to the seed entity is the first specified value. Conversely, when the value of the elimination label of the selected entity is a second specified value (True), the embodiment of the present disclosure does not take it as a seed entity, or can perform clarity processing on it and take the clarity-processed entity as a seed entity.
[0074] After obtaining the seed entity, the embodiment of the present disclosure can search the seed entity to obtain the summary content, which can be the summary content of multiple search results related to the seed entity. Here, the search of the seed entity can be implemented through a search engine, and the summary content is the summary content of multiple search results output by the search engine after inputting the seed entity.
[0075] That is, the embodiment of the present disclosure can retrieve the seed entity by calling a search engine to obtain the description of the seed entity existing in the Internet, that is, the relationship information of the seed entity with other entities existing in the Internet. In addition, the embodiment of the present disclosure can also set search parameters before searching the seed entity, which can include an offset and a retrieval result recall amount, so that a proper amount of retrieval results can be obtained through the search parameters.
[0076] Please refer to Figure 2 , the seed entity is "green vine", and the search page obtained by the search engine can include a first search result 201, a second search result 202 and a third search result 203. The summary content of these three search results can be used as the summary content of the seed entity "green vine".
[0077] It should be noted that after obtaining the summary content of multiple search results related to the seed entity, the embodiment of the present disclosure can preprocess these summary contents. Among them, the preprocessing can be splicing, cleaning and standardizing the multiple summary contents obtained by searching.
[0078] For example, the preprocessing at least includes at least one of the following: removing blank and redundant guides, standardizing multiple backslashes and guide symbols, merging multiple backslashes, cleaning escape guides and backslashes, removing illegal control characters, repairing format problems, removing trailing redundant commas, repairing missing values and key-value pairs, and determining whether a field is a valid field.
[0079] By splicing, cleaning and standardizing the summary content obtained by searching, a pure text summary page that is more convenient for parameter passing can be obtained. For example, by splicing and cleaning the summary content of multiple search results as shown in Figure 2 , a pure text summary page 301 as shown in Figure 3 can be obtained.
[0080] On this basis, the embodiment of the present disclosure can extract the relationship information of the seed entity with other entities existing in the Internet based on the preprocessed summary content, that is, the embodiment of the present disclosure can determine the associated entity set of the seed entity based on the summary content, that is, step S120 is entered.
[0081] In step S120, a set of associated entities of the seed entity is determined based on the summary content, and the set of associated entities includes a plurality of associated entities.
[0082] In the embodiments of the present disclosure, the set of associated entities can include a plurality of associated entities, which can be entities directly associated with the seed entity, or can be indirect entities associated with the seed entity. For example, the plurality of associated entities can include a first target associated entity and other associated entities, where the first target associated entity can be an entity directly associated with the seed entity, and the other associated entities can be other entities indirectly associated with the seed entity.
[0083] In the process of determining the set of associated entities of the seed entity based on the summary content, the embodiments of the present disclosure can extract a plurality of candidate entities related to the seed entity in the summary content, and determine the relationship information between each candidate entity and the seed entity. On this basis, the first target associated entity is determined from the plurality of candidate entities according to the relationship information, and the other associated entities are determined based on the first target associated entity, and the set of associated entities is obtained.
[0084] Here, the set of associated entities can include an entity list, a triple list and an attribute list, that is, when determining the relationship information between each candidate entity and the seed entity, the embodiments of the present disclosure can construct an entity list based on the plurality of candidate entities. On this basis, the relationship between each candidate entity and the seed entity is extracted, and a triple relationship list (SPO triple) is constructed based on the relationship. At the same time, the embodiments of the present disclosure can obtain the attribute information of each candidate entity, and based on the attribute information, an attribute list can be constructed.
[0085] As an example, the seed entity is entity A, after searching the entity A and obtaining a pure text summary page convenient for parameter passing, the embodiments of the present disclosure can construct a triple about the entity A according to the search information. Among them, four candidate entities related to the entity A are determined by extracting the summary content, and the entity list constructed based on the four candidate entities can be as shown in Figure 4 , it can be known that the candidate entities related to the entity A include person B, person C, work A and place A. Here, the candidate entities can be persons, organizations, works and places, etc. Figure 4
[0086] Among them, person B is the wife of person A, person C is the daughter of person A, work A is a movie starring person A, and place A is a key event done by person A in the place. Based on these relationships, the embodiments of the present disclosure can obtain an example diagram as shown in Figure 5 , based on Figure 5 It can be known that the triple relationship list can include multiple triple relationship lists related to the seed entity. In addition, if it is detected that the summary content includes a time, the triple relationship list can include a time field.
[0087] Optionally, the attribute list of the multiple candidate entities associated with the person A can be as shown in Figure 6 Figure 6 It can be known that the attributes in the attribute list are the attributes of each entity itself, such as the attribute of the person B itself is the wife of the person A, and the attribute of the work A is a police film shot by the person A.
[0088] It should be noted that in the process of extracting the candidate entities related to the seed entity, the maximum number of candidate entities can be set by the embodiments of the present disclosure to ensure the accuracy of subsequent large model question and answer data generation. For example, if it is detected that the number of candidate entities in the summary content exceeds the specified number, the embodiments of the present disclosure can filter the specified number of entities as candidate entities. Here, the filtering of candidate entities can be determined based on the closeness between the entities and the seed entity. The higher the closeness, the more likely the entity is a candidate entity of the seed entity. Conversely, if the closeness between the candidate entity and the seed entity is small, the embodiments of the present disclosure can not regard it as a candidate entity.
[0089] Optionally, if it is detected that the number of candidate entities of the summary content is less than the specified number, the embodiments of the present disclosure can regard each entity as a candidate entity, and obtain the relationship list and the attribute list between each candidate entity and the seed entity.
[0090] It should be noted that after retrieving N sets of triple relationships each time, the embodiments of the present disclosure can store them in the “relationship set” of the entity set, and store the attribute relationship in the attribute set. In addition, the “subject” and “object” in the triple can be stored in the “entity set”. After storage, the embodiments of the present disclosure can perform a deduplication operation on the “entity set” to avoid entity duplication.
[0091] In addition, the single entity is retrieved for N times in a loop, and when the target entity in the entity set appears as “subject” in the triple, the retrieval is stopped. In this way, it can be ensured that the entity chain list is long enough, that is, more rich entities are obtained, and repeated operations are avoided.
[0092] As another optional manner, after the plurality of candidate entities associated with the seed entity are acquired, the embodiment of the disclosure can screen the plurality of candidate entities to select the first target associated entity. In the screening process, the embodiment of the disclosure can adopt the strategy of edge priority. For example, the embodiment of the disclosure can construct a knowledge graph based on the triple list, and can determine the first target associated entity according to the number of edges of the candidate entity in the knowledge graph. In other words, the embodiment of the disclosure can extract knowledge graph information from the summary content, and can screen the first target associated entity based on the knowledge graph information.
[0093] In order to better illustrate the acquisition process of the first target associated entity (entity B), the embodiment of the disclosure gives an example diagram as shown in FIG. 3, when the movie A is the seed entity, the candidate entities associated with it are person A and person B, wherein person A is the director of the movie A, that is, he directed the movie A, and person B is the leading actor of the movie A. Through the search, the embodiment of the disclosure can acquire the relationship information between the seed entity and the candidate entities, and the relationship information between the seed entity and the candidate entities is as shown in Table 1. Figure 7 As shown in Table 1, the relationship information between the seed entity and the candidate entities is as shown in Table 1. Figure 7 As shown in Table 1, the relationship information between the seed entity and the candidate entities is as shown in Table 1.
[0094] It should be noted that in the process of determining the first target associated entity from the plurality of candidate entities according to the relationship information between each candidate entity and the seed entity, the embodiment of the disclosure can also acquire the computing power of the electronic device first, that is, determine whether the computing power of the electronic device matches the number of the plurality of candidate entities. If the computing power of the electronic device matches the number of the plurality of candidate entities, the embodiment of the disclosure can take each candidate entity as the first target associated entity. On the contrary, if the computing power of the electronic device does not match the number of the plurality of candidate entities, the embodiment of the disclosure can determine the number of entity expansions that match the computing power of the electronic device based on the computing power of the electronic device, and select a corresponding number of candidate entities from the plurality of candidate entities as the first target associated entity based on the number of entity expansions.
[0095] As an example, five candidate entities associated with the seed entity are acquired through the search, and after it is determined that the computing power of the electronic device can match the subsequent expansion of the five candidate entities, the five candidate entities can be taken as the first target associated entity, that is, the five candidate entities can be searched respectively, and the expansion entities associated with each candidate entity can be determined respectively.
[0096] As another example, if ten candidate entities associated with the seed entity are obtained through the search, and it is determined through detection that the computing power of the electronic device cannot match the subsequent expansion of the ten candidate entities, the ten candidate entities can be screened, and five candidate entities can be selected as the first target associated entities. As can be seen, the number of first target associated entities can be determined according to the computing power of the electronic device, so that the expansion of the large model question and answer data can be flexibly realized, which not only ensures the expansion quantity, but also ensures the computing power matching, that is, ensures the efficiency of data generation.
[0097] In summary, the number of first target associated entities in the embodiments of the present disclosure can be one or more, and the number of first target associated entities is not limited here and can be selected according to actual conditions. The subsequent embodiments will take one first target associated entity as an example to illustrate the construction of the associated entity set, which is only an example and not a practical limitation.
[0098] As another optional way, after obtaining the first target associated entity, the embodiments of the present disclosure can repeatedly perform the above steps, that is, repeatedly perform the entity search operation to the associated entity set determination operation, but this time it is based on the first target associated entity, and the above is based on the seed entity.
[0099] For example, the search content is obtained by searching entity A (seed entity), and entity B (first target associated entity) associated with the seed entity is determined based on the search content. At this time, the embodiments of the present disclosure can obtain the second target associated entity (entity C) associated with entity B by using the same way.
[0100] That is, the embodiments of the present disclosure can search entity B through a search engine, and extract and splice the title and abstract related to the entity B to obtain the pure text abstract page of the entity B. Similar to entity A, by extracting the abstract content, the embodiments of the present disclosure can obtain the expansion triplets of entity B, that is, obtain the relationship information such as the entity list, the triplet relationship list and the attribute list corresponding to entity B. Through these relationship information, the embodiments of the present disclosure can select entity C as the second target associated entity from the multiple expansion entities of entity B.
[0101] The embodiments of the present disclosure can repeat the above steps to realize the extension of the relationship chain (retrieval generation chain) and obtain the basic data for constructing the large model question and answer data. Here, the relationship chain extension process can refer to Figure 8 , based on Figure 8It can be known that after the seed entity A (seed entity) is acquired, the embodiment of the disclosure can search the entity A through the search engine, and then extract and splice the title and the abstract A, so as to acquire the extended triples of the entity A through the abstract content, so as to obtain all related entities of A. Screening these entities can obtain entity B. Repeating the above steps can select entities C, entity D, entity E and entity F in the same way. Meanwhile, the embodiment of the disclosure can acquire the entity list, the triple relationship list and the attribute list of these entities.
[0102] In summary, the embodiment of the disclosure can repeatedly perform the following operations: screening a new entity that can be expanded most in the related object as a search word, searching public domain related information; acquiring a related entity list, an SPO triple and an entity attribute triple of the related entity according to the related information. For example, the embodiment of the disclosure can perform the above steps six times to obtain an associated entity set (entity A, entity B, entity C, entity D, entity E and entity F). On this basis, the embodiment of the disclosure can construct a knowledge network related to the seed entity according to the entity list, the SPO triple (triple relationship list) and the entity attribute triple (attribute list).
[0103] The embodiment of the disclosure can construct an SPO structure by searching the abstract, and then form an entity-relation multi-hop chain (search generation chain), so as to realize the construction of a dynamic knowledge graph. In addition, when screening the candidate entity, the embodiment of the disclosure can use an edge most priority strategy to dynamically extend the entity chain.
[0104] In step S130, a target question is generated according to the associated entity set, and large model question and answer data is generated from the target question and the seed entity.
[0105] As an optional way, after the associated entity set of the seed entity is acquired, the embodiment of the disclosure can generate a target question based on the associated entity set, and generate large model question and answer data (FAQ) from the target question and the seed entity. Here, the target question can be output by a large language model.
[0106] For example, the embodiment of the disclosure can acquire user demand information and analyze the demand information before generating a target question according to the associated entity set. The user demand is different, and the generation strategy adopted to generate the target question is also different. For example, when the user demand is to generate large model question and answer data for a specified evaluation set, the embodiment of the disclosure can generate a target question by using a first generation strategy. For another example, when the user demand is to generate a question for retrieval, the embodiment of the disclosure can generate a target question by using a second generation strategy.
[0107] In this embodiment of the disclosure, the designated evaluation set can also be a benchmark testing system, which can be a high-difficulty AI browsing ability evaluation benchmark. That is, the designated evaluation set can be used as a tool to automatically perform comparative tests and generate evaluation reports based on preset standards.
[0108] In some implementations, a first hint can be obtained when generating the target question. This first hint can be generated based on the test results of a large model trained on a specified evaluation set, where the large model is a model trained using large model question-answering data. Based on this, the target question is generated using the first hint, a set of associated entities, and the relationship information of each associated entity. Here, the relationship information of the associated entities can be a list of triplet relationships for each associated entity.
[0109] As described above, the target question can be generated using a large language model. The input data of this model can be a set of related entities, and the output data can be the target question. Furthermore, the input data varies depending on the generation strategy.
[0110] For example, when constructing a FAQ based on the features of a specified evaluation set, this embodiment of the disclosure can construct a multi-step retrieval system by combining the analysis of the features of the questions in the specified evaluation set with the retrieved entity relationships. This allows for obtaining answers through multi-step retrieval. In this process, the input data of the large language model can include not only the set of associated entities but also the relationship information of each associated entity within that set. Additionally, the input data can also include a list of seed entities, a list of triplet relationships, a list of attributes, and a list of relationships of the last associated entity in the set. This last associated entity can be the last entity (the final entity) in the retrieval generation chain. For example, Figure 8 The last associated entity is entity F.
[0111] To better illustrate the input and output data of the large language model at this time, the embodiments of this disclosure provide the following... Figure 9 The example diagram shown is obtained by... Figure 9 It can be seen that when constructing FAQ based on the features of the specified evaluation set, the input data of the large language model includes the associated entity set 901, the relation list information of each associated entity 902, the relation information of the candidate entities related to the seed entity 903, and the extended entity and SPO information of the end entity 904.
[0112] In addition, in this embodiment of the disclosure, the first prompt information obtained can be input into the large language model together with the above data. The first prompt information may include at least one of the following: question design requirements, question output format, generation example, difficulty strategy, question construction strategy prompt, and specific question generation strategy.
[0113] Among them, the topic design requirements include: conclusions are drawn based on the reverse path structure of the search generation chain, such as reasoning based on the path structure of F→E→D→C→B→A; the full name or abbreviation of each associated entity in the associated entity set cannot be directly used in the question; the answer needs to be one of the seed entities or candidate entities; it is prohibited to locate the answer from any one sentence with one hop, and the path needs to be reconstructed through multiple levels of hops; at least one hop needs to be based on time, text partial matching, attribute combination or structural logic reasoning; the expansion entity of the last entity (F) or the SPO information of F can be reasonably embedded as a confounding factor; the problem needs to have concealment, information fusion and uniqueness, and cannot be too redundant or directly indicate the answer direction; the uniqueness of the answer is strictly guaranteed to ensure that only one entity can be located after searching. These topic design requirements can be fixed or can be adaptively adjusted according to the performance of the large model.
[0114] The question output format can include a question, reasoning explanation and an answer, the question is a question based on a multi-hop relationship chain reasoning; the reasoning explanation is used to gradually explain how to backtrack to the answer from the back to the front with multiple hops; the answer is a seed entity or an expansion entity of the seed entity (one of the candidate entities).
[0115] The generation example (Few-shot) can be an example question input by the user according to the demand, which is mainly used to guide the generation style, that is, to guide the style of the generated target question.
[0116] The difficulty strategy includes partial text matching, structural jump, time reverse, concept alias, interference injection and combined limitation condition, among which, the target text matching can be used to indicate only the name feature or the split suffix; the structural jump is used to realize the jump through the intermediate organization / location / sub-organization; the time reverse is used to replace the specific year based on the year / relative time; the concept alias is used for familiar but not obvious answer alias; the interference injection is used to confuse the expansion entity; the combined limitation condition is used to cross-lock multiple attributes. The difficulty strategy is mainly used to specify which dimensions to generate high-quality target questions from.
[0117] As an example, in the process of generating a target question based on the associated entity set, if the design time information of the associated entity is detected, the specific year can be replaced by the year or the relative time. For example, a suspense novel was published by person A ten years ago.
[0118] Further, in the process of generating the target question, the embodiments of the present disclosure can detect whether the first prompt information involves an entity type construction question strategy recommendation, and if so, determine a target entity matching the plurality of associated entities. For example, the plurality of associated entities include a person entity, and the embodiments of the present disclosure can generate a target question using a person class entity strategy, wherein the person class entity strategy includes: person construction = ambiguous identity × (time point event chain + cross-domain background chain + interpersonal relationship chain) + uniqueness design (unique features). Wherein, the ambiguous identity is mainly used to generate a question without directly giving the name of the person, but only through the description to give the identity clues.
[0119] Optionally, when the plurality of entities include technical / system / concept class entities, the embodiments of the present disclosure can generate a question based on a technical construction strategy, wherein the technical construction strategy includes: function / definition × (structure evolution chain + application coverage chain + comparison feature chain) + uniqueness design.
[0120] As another optional way, in the process of generating a target question according to the associated entity set, the embodiments of the present disclosure can generate a target question based on the specific construction strategy described below.
[0121] Specifically, the embodiments of the present disclosure can obtain the influence of the seed entity, and determine the difficulty of generating the target question based on the influence. The greater the influence, the greater the difficulty of generating the target question. Conversely, the smaller the influence of the seed entity, the smaller the difficulty of generating the target question.
[0122] For example, after obtaining the influence of the seed entity, the embodiments of the present disclosure can generate a question based on the influence, starting from the ambiguity, where the ambiguity and the influence can be in a proportional relationship, that is, the greater the influence, the greater the strength of the ambiguity processing. Wherein, the ambiguity processing can be realized by at least one of the following ways: category / range substitution, attribute name substitution, time / digital ambiguity, calculation type features, summary statements, and attribute synthesis.
[0123] Optionally, the embodiments of the present disclosure can also generate a question through conditional combination (cut-in point), such as using direct conditional combination, conditional transformation and recombination, or conditional merging to generate a target question.
[0124] Optionally, the first prompt information can also include prohibited items, which can be updated in real time according to the training of the large language model. For example, the prohibited items include that no answer can be displayed in any SPO, and no SPO or entity is required.
[0125] It should be noted that the first prompt information can be updated continuously along with the training result of the large model on the specified evaluation set. For example, if the reply accuracy of the large model on the specified evaluation set is low at the first time, the next time the large model question and answer data is generated, the generation of the large model question and answer data related to the character type is emphasized.
[0126] As another optional way, in the process of generating the target question according to the set of associated entities, the second prompt information can also be obtained. The second prompt information can be used to instruct the large model to generate questions according to a specified strategy, which can be obtained based on the retrieval of the large model. On this basis, the target question is generated based on the second prompt information, the set of associated entities, and the summary content information of each associated entity.
[0127] Optionally, in the process of generating the target question based on the second prompt information, the set of associated entities, and the summary content information of each associated entity, the disclosure embodiment can obtain the complete content corresponding to the summary content; and the target question is generated based on the second prompt information, the set of associated entities, and the complete content of each associated entity.
[0128] In this process, the disclosure embodiment can determine whether the number of entities involved in the summary content exceeds the preset number. If the number of entities does not exceed the preset number, the disclosure embodiment can extract the complete content corresponding to the summary content, and on this basis, the target question is generated based on the complete content corresponding to the summary content.
[0129] In addition, in the process of determining whether the number of entities meets the standard, the disclosure embodiment can count the number of entities involved in the summary content of each search result. When the number of summaries whose entity number is less than the preset number is multiple, and the number of these summaries is greater than the preset threshold, the disclosure embodiment can obtain the complete content corresponding to all search results. Conversely, when the number of these summaries is less than the preset threshold, the disclosure embodiment can specifically obtain the complete content corresponding to the corresponding summary (the summary whose entity number does not meet the standard). In this way, the accuracy of the target question can be ensured while reducing the energy consumption of the electronic device.
[0130] In addition, when the large language model training is completed, and it is determined that the search question cannot obtain an accurate answer through the summary content during testing, the disclosure embodiment can prompt the model to obtain the complete content corresponding to the summary content, and generate the target question based on the second prompt information, the set of associated entities, and the complete content of each associated entity.
[0131] In the above manner, the meaning of multi-hop can be abstracted into a question construction strategy, that is, using the characteristics of the attention mechanism of the large language model, extracting information suitable for question setting from a large amount of retrieved information, and using the retrieved information to blur the given clues to construct questions that require analysis of entry points, multi-step jumping of clues, and multi-step retrieval to obtain answers.
[0132] When constructing questions (FAQ) based on the retrieved summary page, the embodiment of the disclosure can input the summary content information of each associated entity while inputting the associated entity set into the large language model. In order to better illustrate the input data and output data of the large language model at this time, the embodiment of the disclosure gives an example as shown in Figure 10 As shown in Figure 10 It can be seen that when constructing a FAQ based on the retrieved summary page, the input data of the large language model includes the associated entity set 901 and the summary page content 1002 returned by each associated entity.
[0133] In addition, the embodiment of the disclosure can input the obtained second prompt information into the large language model together with the above data, wherein the second prompt information can include at least one of a question setting target (question design requirement), a question output format, a recommended question setting strategy, a specific question setting strategy, and a forced constraint strategy. Among them, the recommended question setting strategy is similar to the difficulty strategy, which can include alias shielding strategy, structure jumping strategy, time interleaving strategy, implicit attribute combination strategy, intermediate information ambiguity guiding strategy, and alias and concept generalization. Other information is similar to the first prompt information described above and will not be described here.
[0134] After obtaining the large model question and answer data, the embodiment of the disclosure can perform reinforcement learning (RL) and SFT (Supervised Fine-Tuning) training on the large language model, so as to ensure that the answer can be obtained without model memory. Compared with constructing QA based on encyclopedic information, the embodiment of the disclosure can be searched in the public domain and does not rely on professional databases, so that the model can learn more in the RL training. It is search, rather than relying on solid knowledge reasoning. That is, by introducing the reinforcement learning mechanism, the embodiment of the disclosure can optimize the retrieval and reasoning strategy, effectively reduce the model hallucination rate and improve the factual correctness.
[0135] In one specific embodiment, as shown in Figure 11As shown, after obtaining the keyword entity A (seed entity), the embodiment of the disclosure can search and splice the summary of the entity A, and the extended triple of the entity A can be obtained by extracting the summary content, which can include an entity list, a triple list, an attribute list, and a summary, etc. On this basis, the embodiment of the disclosure can take the object with the most expandable edges as entity B, and obtain the extended triple of entity B in the same way. By repeatedly performing the above operations, the extended triple of entity C, the extended triple of entity D, the extended triple of entity E, and the extended triple of entity F can be obtained. On this basis, the embodiment of the disclosure can be based on the summary question or based on the specified test set feature question to realize the construction of the large model question and answer data set.
[0136] The embodiment of the disclosure can systematically solve the problems of relying on model solid-state knowledge, high illusion rate, and lack of reasoning depth through multi-hop public information chain construction, multi-summary fusion, and RL training adaptation mechanism. Specifically, through the controllable relationship chain generation mechanism, the retrievable facts can be mapped to high complexity problem structures, and through embedding unique constraints and multi-layer verification paths, the strategy evolution of the Agent (intelligent agent) in the real retrieval environment can be realized in cooperation with reinforcement learning. In addition, the scheme can be applied to education, knowledge base construction, intelligent customer service, and large model training, etc. multiple scenes, that is, the application scenarios are widely used.
[0137] Based on the same inventive concept, the disclosure also provides a large model question and answer data generation device, Figure 12 is a block diagram of a large model question and answer data generation device 1200 according to an example embodiment, as shown in Figure 12 As shown, the large model question and answer data generation device 1200 can include a search module 1210, a determination module 1220, and a generation module 1230.
[0138] The search module 1210 is configured to obtain a seed entity and search the seed entity to obtain summary content, the summary content being summary content of a plurality of search results related to the seed entity;
[0139] The determination module 1220 is configured to determine a set of associated entities of the seed entity based on the summary content, the set of associated entities including a plurality of associated entities;
[0140] The generation module 1230 is configured to generate a target question according to the set of associated entities, and generate the large model question and answer data from the target question and the seed entity.
[0141] Optionally, the plurality of associated entities includes a first target associated entity and other associated entities, and the determining module 1220 is further configured to extract a plurality of candidate entities related to the seed entity in the summary content, and determine relationship information between each of the candidate entities and the seed entity; determine the first target associated entity from the plurality of candidate entities according to the relationship information, and determine the other associated entities based on the first target associated entity, to obtain the set of associated entities.
[0142] Optionally, the relationship information includes an entity list, a triple relationship list, and an attribute list, and the determining module 1220 is further configured to construct the entity list based on the plurality of candidate entities; extract a relationship between each of the candidate entities and the seed entity, and construct the triple relationship list based on the relationship; obtain attribute information of each of the candidate entities, and construct the attribute list based on the attribute information.
[0143] Optionally, the determining module 1220 is further configured to construct a knowledge graph based on the triple relationship list, and determine the first target associated entity according to a number of edges of each of the candidate entities in the knowledge graph.
[0144] Optionally, the generating module 1230 is further configured to obtain first prompt information, the first prompt information being prompt information generated by testing a large model using a specified evaluation set and based on a test result; and generate the target question based on the first prompt information, the set of associated entities, and the relationship information of each of the associated entities.
[0145] Optionally, the generating module 1230 is further configured to obtain second prompt information, the second prompt information being used to instruct the large model to generate questions according to a specified strategy, the specified strategy being obtained based on a retrieval situation of the large model; and generate the target question based on the second prompt information, the set of associated entities, and the summary content information of each of the associated entities.
[0146] Optionally, the generating module 1230 is further configured to obtain complete content corresponding to the summary content; and generate the target question based on the second prompt information, the set of associated entities, and the complete content of each of the associated entities.
[0147] Optionally, the large model question and answer data generation apparatus 1200 can further include:
[0148] An updating module configured to update an entity library based on the seed entity and the set of associated entities, the entity library including an entity set, a relationship set, and an attribute set, the entity set including a plurality of benchmark entities, each of the benchmark entities being configured with an entity label.
[0149] Optionally, the entity label includes at least one of the following information:
[0150] entity's popularity information;
[0151] domain label;
[0152] task type label;
[0153] elimination label.
[0154] Optionally, the large model question and answer data generation apparatus 1200 can further include:
[0155] The setting module is configured to set the elimination label to a True value when the reference entity meets the elimination condition.
[0156] Optionally, the elimination condition includes at least one of:
[0157] The reference entity has ambiguity;
[0158] The reference entity is a fuzzy category entity;
[0159] The reference entity is a broad category generalization.
[0160] Optionally, the search module 1210 is further configured to select the seed entity from the plurality of reference entities based on the entity label, wherein the seed entity corresponds to the elimination label with a first specified numerical value.
[0161] Optionally, the large model question and answer data generation apparatus 1200 can further include:
[0162] The clarity processing module is configured to perform clarity processing on the reference entity with the elimination label having a second specified numerical value.
[0163] The embodiments of the present disclosure can realize logical reasoning question construction, that is, according to data requirements, not only the length of the entity chain can be adjusted, but also the relationship chain can be checked for duplication to avoid truncation caused by relationship chain loop. Through embedding attribute fusion, interference terms, concept aliases and other rules, the answer can be ensured to be unique. Moreover, the entities in the chain can be continuously expanded in the retrievable range to form an entity library according to the retrieved information, that is, without the need to continuously collect seed entities, the update of the entity library can be realized. At the same time, the large model question and answer data generated based on the embodiments of the present disclosure can improve the ability of the Agent in multi-hop reasoning and information aggregation in RL training, that is, it has the ability of intervention-feedback-optimization.
[0164] The embodiments of the modules in the large model question and answer data generation apparatus 1200 described above can refer to the related embodiments of the above method, and the present embodiment will not be repeated here.
[0165] The embodiment of the present disclosure further provides a computer readable medium, which has stored thereon a computer program, and the computer program is executed by a processing device to implement the steps of the method for generating large model question and answer data.
[0166] The embodiment of the present disclosure further provides a computer program product, which comprises a computer program, and the computer program is executed by a processor to implement the steps of the method for generating large model question and answer data.
[0167] The embodiment of the present disclosure further provides an electronic device, which comprises:
[0168] a storage device having stored thereon a computer program;
[0169] a processing device configured to execute the computer program in the storage device to implement the steps of the method for generating large model question and answer data.
[0170] Reference will be made to the following description of the embodiments of the present disclosure. Figure 13 which shows a structural schematic diagram of an electronic device 1300 suitable for implementing the embodiments of the present disclosure. The terminal device in the embodiments of the present disclosure can include, but is not limited to, a mobile terminal such as a mobile phone, a notebook computer, a digital broadcast receiver, a PDA (Personal Digital Assistant), a PAD (Tablet Personal Computer), a PMP (Portable Multimedia Player), a vehicle terminal (for example, a car navigation terminal), and the like, and a fixed terminal such as a digital TV, a desktop computer, and the like. Figure 13 The electronic device shown is only an example, and should not bring any limitation to the functions and use range of the embodiments of the present disclosure.
[0171] As shown in Figure 13 , the electronic device 1300 can include a processing device (for example, a central processing unit, a graphics processing unit, and the like) 1301, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1302 or a program loaded from a storage device 1308 into a random access memory (RAM) 1303. In the RAM 1303, various programs and data required for the operation of the electronic device 1300 are also stored. The processing device 1301, the ROM 1302, and the RAM 1303 are connected to each other through a bus 1304. An input / output (I / O) interface 1305 is also connected to the bus 1304.
[0172] In general, the following devices can be connected to the I / O interface 1305: input devices 1306, including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, and the like; output devices 1307, including, for example, a liquid crystal display (LCD), a speaker, a vibrator, and the like; storage devices 1308, including, for example, a magnetic tape, a hard disk, and the like; and communication devices 1309. The communication devices 1309 can allow the electronic device 1300 to communicate wirelessly or through a wire with other devices to exchange data. Although Figure 13 The electronic device 1300 is shown with various devices, but it is understood that not all of the shown devices are required to be implemented or present. More or fewer devices can alternatively be implemented or present.
[0173] In particular, the processes described above with reference to the flowcharts can be implemented as a computer software program according to embodiments of the present disclosure. For example, embodiments of the present disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for executing the methods illustrated by the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network through the communication devices 1309, or installed from the storage devices 1308, or installed from the ROM 1302. When the computer program is executed by the processing devices 1301, the above-described functions defined in the methods of embodiments of the present disclosure are performed.
[0174] It is noted that the aforementioned computer-readable medium of the present disclosure can be a computer-readable signal medium or a computer-readable storage medium or any combination thereof. The computer-readable storage medium can be, for example and without limitation, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the computer-readable storage medium can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In the present disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program used by or in connection with an instruction execution system, apparatus, or device. In the present disclosure, the computer-readable signal medium can include a computer-readable program code transmitted by a computer-readable signal medium, in a baseband or as a part of a carrier wave. Such a computer-readable signal medium can take a variety of forms, including but not limited to, electro-magnetic, optical, or any suitable combination of the foregoing. The computer-readable signal medium can also be any computer-readable medium that is not a computer-readable storage medium and that can be used to carry or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained in the computer-readable medium can be transmitted by any suitable medium, including but not limited to, wire, cable, RF, etc., or any suitable combination of the foregoing.
[0175] In some embodiments, the terminal device, the server can communicate using any currently known or future developed network protocol, such as HTTP (HyperText Transfer Protocol), and can be interconnected with digital data communication of any form or medium (e.g., a communication network). Examples of the communication network include a local area network ("LAN"), a wide area network ("WAN"), an internetwork (e.g., the Internet), and an end-to-end network (e.g., an ad hoc end-to-end network), as well as any currently known or future developed network.
[0176] The aforementioned computer-readable medium can be included in the aforementioned electronic device; or can exist separately from the electronic device and not be assembled into the electronic device.
[0177] The computer readable medium described above carries one or more programs, when the one or more programs are executed by the electronic device, cause the electronic device to: obtain a seed entity, and search the seed entity to obtain summary content, the summary content being summary content of a plurality of search results related to the seed entity; determine a set of associated entities of the seed entity based on the summary content, the set of associated entities including a plurality of associated entities; generate a target question according to the set of associated entities, and generate the large model question and answer data from the target question and the seed entity.
[0178] Computer program code for carrying out operations of the present disclosure can be written in one or more programming languages or combinations of languages including object oriented programming languages such as Java, Smalltalk, C++ as well as conventional procedural programming languages such as "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).
[0179] The flow diagrams and the block diagrams in the drawings are illustrations of possible architectures, functions, and operations for systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flow diagrams or block diagrams can represent a module, a segment, or a portion of code, which comprises one or more executable instructions for implementing the specified logical function(s). It should also be noted that in some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently or the blocks may
[0180] The modules involved in the embodiments of the present disclosure can be implemented in the form of software or in the form of hardware. Among them, the name of the module does not constitute a limitation to the module itself in some cases.
[0181] The functionality described above can be performed, at least in part, by one or more hardware logic components. For example, and without limitation, an exemplary type of hardware logic components that can be used include Field-programmable Gate Arrays (FPGAs), Application-specific Integrated Circuits (ASICs), Application-specific Standard Products (ASSPs), System-on-a-chip systems (SOCs), Complex Programmable Logic Devices (CPLDs), etc.
[0182] In the context of the present disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium will include one or more lines of electrical connections, portable computer disks, hard disk drives, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or Flash memory), optical fibers, portable compact disc read-only memories (CD-ROMs), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0183] The above description is only preferred embodiments of the present disclosure and the explanation of the applied technical principles. It should be understood by those skilled in the art that the disclosure range involved in the present disclosure is not limited to the technical solutions formed by the specific combinations of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosed concept. For example, the technical solutions formed by replacing the above features with the technical features disclosed in the present disclosure (but not limited to) having similar functions.
[0184] Further, although operations are depicted in a particular, chronological sequence, this should not be understood as requiring this particular order or sequence. On the contrary, some of the operations can be performed in a different order or concurrently with one another. Additionally, certain details of the implementation can have been left out in order to avoid obscuring the point of the present disclosure. Also, some of the features of the present disclosure can be utilized to their fullest effect in only some instances of the technology. Furthermore, it should be appreciated that the technology might be practiced in the absence of any element that is not specifically disclosed herein.
Claims
1. A method for generating large-scale question-and-answer data, characterized in that, The method includes: Obtain a seed entity and search the seed entity to obtain summary content, which is a summary of multiple search results related to the seed entity; Based on the summary content, the associated entity set of the seed entity is determined. The associated entity set includes multiple associated entities, including a first target associated entity and other associated entities. The step of determining the associated entity set of the seed entity based on the summary content includes: Extract multiple candidate entities related to the seed entity from the summary content, and determine the relationship information between each candidate entity and the seed entity. The relationship information includes an entity list, a triplet relationship list, and an attribute list. The first target associated entity is determined from the plurality of candidate entities based on the relationship information, and the other associated entities are determined based on the first target associated entity to obtain the associated entity set; A target question is generated based on the set of associated entities, and the large model question-and-answer data is generated from the target question and the seed entities. The answer to the target question in the large model question-and-answer data includes the seed entities. The step of generating the target problem based on the associated entity set includes: Obtain the first prompt information, which is a prompt information generated based on the test results of testing a large model using a specified evaluation set; The target question is generated based on the first prompt information, the set of associated entities, and the relationship information of each associated entity.
2. The method for generating large-scale model question-answering data according to claim 1, characterized in that, The determination of the relationship information between each candidate entity and the seed entity includes: The entity list is constructed based on multiple candidate entities; Extract the relationship between each candidate entity and the seed entity, and construct the triplet relationship list based on the relationship; Obtain the attribute information of each candidate entity, and construct the attribute list based on the attribute information.
3. The method for generating large-scale model question-answering data according to claim 2, characterized in that, The step of determining the first target associated entity from the plurality of candidate entities based on the relationship information includes: A knowledge graph is constructed based on the triplet relation list, and the first target associated entity is determined according to the number of edges of each candidate entity in the knowledge graph.
4. The method for generating large-scale question-and-answer data according to claim 1, characterized in that, The step of generating the target problem based on the associated entity set includes: Obtain a second prompt message, which is used to instruct the large model to generate questions according to a specified strategy, the specified strategy being obtained based on the retrieval results of the large model; The target question is generated based on the second prompt information, the set of associated entities, and the summary content information of each associated entity.
5. The method for generating large-scale model question-answering data according to claim 4, characterized in that, The step of generating the target question based on the second prompt information, the set of associated entities, and the summary content information of each associated entity includes: Obtain the complete content corresponding to the summary content; The target question is generated based on the second prompt information, the set of associated entities, and the complete content of each associated entity.
6. The method for generating large-scale question-and-answer data according to claim 1, characterized in that, The method further includes: The entity library is updated based on the seed entity and the associated entity set. The entity library includes an entity set, a relationship set, and an attribute set. The entity set includes multiple base entities, and each base entity is configured with an entity tag.
7. The method for generating large-scale model question-answering data according to claim 6, characterized in that, The entity tag includes at least one of the following information: Information about the entity's brand recognition; Domain tags; Task type tags; Remove the tags.
8. The method for generating large-scale model question-answering data according to claim 7, characterized in that, The method further includes: When the baseline entity is determined to meet the exclusion criteria, the exclusion label is set to the True value.
9. The method for generating large-scale question-and-answer data according to claim 8, characterized in that, The rejection criteria include at least one of the following: The reference entity is ambiguous; The benchmark entity is a fuzzy entity; The reference entity is a broad category.
10. The method for generating large-scale model question-answering data according to claim 7, characterized in that, The acquisition of the seed entity includes: The seed entity is selected from the plurality of benchmark entities based on the entity label, and the value of the rejection label corresponding to the seed entity is a first specified value.
11. The method for generating large-scale model question-answering data according to claim 7, characterized in that, The method further includes: The benchmark entity whose removal label value is a second specified value is subjected to sharpness processing.
12. A device for generating large-scale model question-and-answer data, characterized in that, The device includes: The search module is used to obtain seed entities and search the seed entities to obtain summary content, which is a summary of multiple search results related to the seed entities. A determining module is configured to: determine a set of associated entities of the seed entity based on the summary content, wherein the set of associated entities includes multiple associated entities, including a first target associated entity and other associated entities; extract multiple candidate entities related to the seed entity from the summary content, and determine the relationship information between each candidate entity and the seed entity, wherein the relationship information includes an entity list, a triplet relationship list, and an attribute list; determine the first target associated entity from the multiple candidate entities according to the relationship information, and determine the other associated entities based on the first target associated entity, thereby obtaining the set of associated entities; A generation module is configured to generate a target question based on the associated entity set, and generate the large model question-and-answer data from the target question and the seed entities, wherein the answer to the target question in the large model question-and-answer data includes the seed entities; the step of generating the target question based on the associated entity set includes: obtaining first prompt information, wherein the first prompt information is generated based on the test results of testing the large model using a specified evaluation set; and generating the target question based on the first prompt information, the associated entity set, and the relationship information of each associated entity.
13. A computer-readable medium having a computer program stored thereon, characterized in that, When executed by a processing device, the computer program performs the steps of the method described in any one of claims 1-11.
14. An electronic device, characterized in that, include: A storage device on which computer programs are stored; A processing device for executing the computer program in the storage device to implement the steps of the method according to any one of claims 1-11.
15. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 11.
Citation Information
Patent Citations
Data processing method and device based on knowledge graph, medium and electronic equipment
CN111552880A
Question and answer method and device, storage medium and computing equipment
CN118035409A