Information synthesis method and apparatus
By decomposing seed problems and combining meta-information, diversified synthetic data is generated, and the problems of poor data scalability, high cost and insufficient diversity in the prior art are solved, and larger-scale and more diverse data synthesis is achieved.
Patent Information
- Application Number
- CN202411580716.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-07
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2044-11-07
AI Technical Summary
The existing data information synthesis methods have poor data scalability, high data synthesis cost, and insufficient data diversity.
The seed problem is decomposed through a preset decomposition model to generate meta information related to the seed problem; a meta information relationship diagram is constructed based on the relationship between the meta information; a meta information node is combined to generate multiple types of meta information combinations; a question-and-answer generation model is used to generate synthetic questions and answers based on the meta information combination.
The scale of seed data is achieved, reducing dependence on the original seed data, increasing the diversity of synthetic data, and reducing the cost of synthesis.
Smart Images

Figure CN119514690B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and particularly to an information synthesis method and device. Background Art
[0002] Although large language models have demonstrated remarkable capabilities in various language tasks, they still have problems in understanding and solving complex problems (e.g., mathematics and programming). An effective way to solve this problem is for large language models to use large-scale high-quality synthetic data when understanding and solving complex problems. However, developing a low-cost and effective data synthesis method remains a challenge.
[0003] Existing data synthesis methods have three main drawbacks:
[0004] (1) Limited scalability: Existing methods are difficult to synthesize larger-scale data based on less seed data.
[0005] (2) High cost: Current data synthesis methods rely on the assistance of commercial models, resulting in a significant increase in synthesis cost.
[0006] (3) Similar to seed data: Due to the excessive dependence on seed data during the synthesis process, the newly generated data is very similar to the seed data, resulting in insufficient diversity of the generated data.
[0007] Regarding the technical problems of poor data scalability, high data synthesis cost, and insufficient diversity of the generated data in the existing data information synthesis methods described above, no effective solution has been proposed yet. Summary of the Invention
[0008] Embodiments of the present application provide an information synthesis method and device to at least solve the technical problems of poor data scalability, high data synthesis cost, and insufficient diversity of the generated data existing in the prior art.
[0009] According to one aspect of the embodiments of the present application, an information synthesis method is provided, including: decomposing a seed problem through a preset decomposition model to generate meta-information related to the seed problem; constructing a first meta-information relationship graph according to the meta-information relationships between the meta-information, where the meta-information relationship is used to indicate the association relationships between the meta-information in the same seed problem and the association relationships between the meta-information in different seed problems; combining the meta-information nodes according to the meta-information relationships between the meta-information nodes in the first meta-information relationship graph to generate meta-information combinations, where different meta-information combinations have different meta-information relationships; and generating synthetic problems and corresponding synthetic answers through a question-and-answer generation model according to the meta-information combinations and corresponding prompt words.
[0010] According to another aspect of the embodiments of the present application, an information synthesis device is further provided, including: an information generation module, configured to decompose a seed problem through a preset decomposition model to generate meta-information related to the seed problem; a relationship graph generation module, configured to construct a first meta-information relationship graph according to the meta-information relationships between the meta-information, where the meta-information relationship is used to indicate the association relationships between the meta-information in the same seed problem and the association relationships between the meta-information in different seed problems; a combination generation module, configured to combine the meta-information nodes according to the meta-information relationships between the meta-information nodes in the first meta-information relationship graph to generate meta-information combinations, where different meta-information combinations have different meta-information relationships; and an information synthesis module, configured to generate a synthesized problem and a corresponding synthesized answer through a question-and-answer generation model according to the meta-information combination and the corresponding prompt words.
[0011] According to another aspect of the embodiments of the present application, an information synthesis device is further provided, including: a processor; and a memory, connected to the processor, configured to provide instructions for the processor to perform the following processing steps: decompose a seed problem through a preset decomposition model to generate meta-information related to the seed problem; construct a first meta-information relationship graph according to the meta-information relationships between the meta-information, where the meta-information relationship is used to indicate the association relationships between the meta-information in the same seed problem and the association relationships between the meta-information in different seed problems; combine the meta-information nodes according to the meta-information relationships between the meta-information nodes in the first meta-information relationship graph to generate meta-information combinations, where different meta-information combinations have different meta-information relationships; and generate a synthesized problem and a corresponding synthesized answer through a question-and-answer generation model according to the meta-information combination and the corresponding prompt words.
[0012] In the embodiments of the present application, the computing device decomposes the seed problem, so that meta-information corresponding to the seed problem can be obtained in multiple aspects and dimensions as extended seed data, thereby realizing the expansion of the scale of the seed data in the technical solution, and obtaining a larger scale of seed data. And in the technical solution, by constructing the relationships between the meta-information, a meta-information relationship graph is generated, and then the meta-information nodes in the meta-information relationship graph are flexibly combined to generate various types of meta-information combinations corresponding to the meta-information relationships, and the corresponding questions and answers are generated by using the meta-information combinations, so as to generate more types of seed data, expand the scale of the seed data, reduce the dependence on the original seed data, and increase the diversity of the synthesized data. And most of the data models used in the data synthesis in the technical solution are open-source models, avoiding relying on the help of commercial models and reducing the synthesis cost. Thereby solving the technical problems of poor data scalability, high data synthesis cost, and insufficient diversity of the generated data in the data information synthesis method. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] The accompanying drawings described herein are used to provide a further understanding of the present application and form a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation of the present application. In the drawings:
[0014] Figure 1 is a hardware structure block diagram of a computing device for implementing the method according to Embodiment 1 of the present application;
[0015] Figure 2 is a schematic flowchart of the information synthesis method according to the first aspect of Embodiment 1 of the present application;
[0016] Figure 3 is a schematic diagram of the first meta-information relationship diagram according to Embodiment 1 of the present application;
[0017] Figure 4 is a schematic diagram of the second meta-information relationship diagram according to Embodiment 1 of the present application;
[0018] Figure 5 is a schematic diagram of the third meta-information relationship diagram according to Embodiment 1 of the present application;
[0019] Figure 6 is a schematic diagram of the fourth meta-information relationship diagram according to Embodiment 1 of the present application;
[0020] Figure 7 is a sequential flowchart of the information synthesis method according to the first aspect of Embodiment 1 of the present application;
[0021] Figure 8 is a schematic diagram of the information synthesis device according to Embodiment 2 of the present application; and
[0022] Figure 9 is a schematic diagram of the information synthesis device according to Embodiment 3 of the present application. Detailed Description of the Embodiments
[0023] In order to enable those skilled in the art to better understand the technical solutions of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0024] It should be noted that the terms "first", "second", etc. in the description, claims and the above-mentioned drawings of this application are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of this application described here can be implemented in an order different from those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0025] Embodiment 1
[0026] According to this embodiment, a method embodiment of an information synthesis method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that here.
[0027] The method embodiment provided by this embodiment can be executed in a mobile terminal, a computer terminal, a server or a similar computing device. Figure 1 A hardware structure block diagram of a computing device for implementing the information synthesis method is shown. As Figure 1 shown, the computing device may include one or more processors (the processor may include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory for storing data, and a transmission device for communication functions. In addition, it may further include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the I / O interface), a network interface, a power supply and / or a camera. Those of ordinary skill in the art can understand that Figure 1 the structure shown is only schematic and does not limit the structure of the above-mentioned electronic device. For example, the computing device may further include more or fewer components than Figure 1 shown, or have a different configuration from Figure 1 shown.
[0028] It should be noted that the one or more processors and / or other data processing circuits described above can generally be referred to as "data processing circuits" herein. The data processing circuit can be embodied in software, hardware, firmware, or any combination thereof, in whole or in part. In addition, the data processing circuit can be a single independent processing module, or be incorporated in whole or in part into any one of other elements in the computing device. As involved in the embodiments of the present application, the data processing circuit is a kind of processor control (for example, the selection of a variable resistor terminal path connected to an interface).
[0029] The memory can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the information synthesis method in the embodiments of the present application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, realizes the information synthesis method of the above application program. The memory can include high-speed random access memory, and can also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memories. In some instances, the memory can further include a memory remotely disposed relative to the processor, and these remote memories can be connected to the computing device through a network. Examples of the above network include but are not limited to the Internet, intranet, local area network, mobile communication network, and combinations thereof.
[0030] The transmission device is used to receive or send data via a network. Specific examples of the above network can include a wireless network provided by a communication provider of the computing device. In one instance, the transmission device includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices through a base station and thus can communicate with the Internet. In one instance, the transmission device can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0031] The display can be, for example, a touch-screen liquid crystal display (LCD), which enables a user to interact with the user interface of the computing device.
[0032] It should be noted here that in some alternative embodiments, the above Figure 1 illustrated computing device can include hardware elements (including circuits), software elements (including computer code stored on a computer-readable medium), or a combination of both hardware elements and software elements. It should be pointed out that Figure 1 is only an example of a specific specific instance, and is intended to illustrate the types of components that can exist in the above computing device.
[0033] Under the above operating environment, according to the first aspect of this embodiment, an information synthesis method is provided, which is implemented by the Figure 1 computing device shown in Figure 2 FIG. shows a schematic flow diagram of the method. Referring to Figure 3 shown, the method includes:
[0034] S202: Decompose the seed problem through a preset decomposition model to generate meta-information related to the seed problem;
[0035] S204: Construct a first meta-information relationship graph according to the meta-information relationships between meta-information, where the meta-information relationship is used to indicate the association relationships between various meta-information in the same seed problem and the association relationships between various meta-information in different seed problems;
[0036] S206: Combine the meta-information nodes according to the meta-information relationships between the meta-information nodes in the first meta-information relationship graph to generate meta-information combinations, where different meta-information combinations have different meta-information relationships; and
[0037] S208: Generate a synthesized question and a corresponding synthesized answer through a question-and-answer generation model according to the meta-information combination and the corresponding prompt words.
[0038] Specifically, the computing device collects seed data for data synthesis. Taking the question-and-answer system as an example, the computing device thus collects seed questions for data synthesis. Then the computing device inputs the seed questions and the corresponding prompt words into a preset decomposition model, and decomposes the seed questions through the decomposition model to output the meta-information corresponding to the seed questions. The prompt words are pre-set prompts for prompting the decomposition model to extract meta-information from the seed questions. And the decomposition model can be a natural language processing model.
[0039] The prompt words corresponding to the decomposition model are: "As a mathematics education expert, please analyze the given mathematics problem and its solution to extract specific mathematics knowledge points. These knowledge points will help teachers create similar exercises to help students understand and master key learning objectives.
[0040] Please follow the following requirements:
[0041] 1. Extract knowledge points: Identify relevant mathematics knowledge points from the given problem and its solution.
[0042] 2. Ensure relevance: Ensure that the extracted knowledge points are directly related, accurate, and concise to the question, and avoid vague concepts.
[0043] 3. Focus on key concepts: Focus on the specific concepts necessary to solve the problem and explain the solution.
[0044] 4. Provide a clear and concise list: Provide a clear and concise list of key knowledge points so that educators can design relevant exercises to help students focus on the key learning outcomes they need to master.
[0045] Mathematical problem: {Seed problem}
[0046] Solution: {Corresponding answer}
[0047] Limit the list to no more than ten key knowledge points, ensuring that the listed knowledge points are strictly relevant to solving the given mathematical problem and contribute to its conceptual understanding. Please output in the following format:
[0048] Related mathematical knowledge points: 1. 2. 3.
[0052] 4. (Continue listing as needed)
[0053] ”
[0054] For example, the seed problem is a mathematical problem, and the meta-information is the knowledge points involved in this mathematical problem, such as "calculus", "limit", or "derivative". Thus, the computing device inputs the mathematical problem (i.e., the seed problem) and the prompt into a preset decomposition model, and the decomposition model decomposes the mathematical problem (i.e., the seed problem) to output the knowledge points (i.e., the meta-information) corresponding to this mathematical problem.
[0055] It should be noted that in this embodiment, only taking knowledge points as meta-information is used to illustrate the process of data synthesis. In addition, the meta-information can also be a subject (such as "mathematics") or a topic (such as "algebra"), which is not specifically limited here.
[0056] Further, the computing device inputs the decomposed meta-information and the corresponding prompt into a preset large language model. The large language model filters the meta-information, for example, filters out knowledge points containing mathematical errors, incompleteness, or being too general, to obtain the filtered knowledge points (i.e., the meta-information). Then the computing device inputs the filtered meta-information into an embedding model. The embedding model clusters the filtered meta-information and stores the clustered meta-information in a meta-information library for storing various types of meta-information. For each cluster center, a large language model is used to select the most appropriate meta-information from it to represent this cluster center. The embedding model can be the bge-en-large model.
[0057] The prompt corresponding to the large language model is: "Given a list of similar mathematical knowledge points, select a knowledge point that best represents all these knowledge points. This knowledge point should be the most commonly known and used in mathematical discussions and ensure that it includes all the listed knowledge points.
[0058] Knowledge point list: {Knowledge point list}
[0059] How to select:
[0060] 1. Common usage: Knowledge points should be widely used in school and academic scenarios.
[0061] 2. Wide coverage: It should cover all key aspects and details of the listed similar knowledge points.
[0062] 3. Standard terms: The terms used should be standard and widely accepted in the mathematical community.
[0063] Please review each knowledge point and select a knowledge point that best meets these criteria. Provide a short explanation to justify your choice.
[0064] The format is as follows:
[0065] Best knowledge point: {The knowledge point you selected}
[0066] Reason: {Explanation}”
[0067] Furthermore, the computing device filters the meta-information that conforms to the meta-information relationship from the meta-information library according to the preset meta-information relationship between the meta-information. The meta-information relationship includes the association relationship between the meta-information in the same sub-problem and the association relationship between the meta-information in different sub-problems.
[0068] For example, the knowledge points (i.e., meta-information) in sub-problem 1 include: knowledge point A, knowledge point B, and knowledge point C; the knowledge points (i.e., meta-information) in sub-problem 2 include: knowledge point A and knowledge point D; the knowledge points (i.e., meta-information) in sub-problem 3 include: knowledge point D and knowledge point E; the knowledge points (i.e., meta-information) in sub-problem 4 include: knowledge point E and knowledge point F.
[0069] Thus, the computing device associates the knowledge points in the same sub-problem. For example, it associates knowledge point A, knowledge point B, and knowledge point C in sub-problem 1, associates knowledge point A and knowledge point D in sub-problem 2, associates knowledge point D and knowledge point E in sub-problem 3, and associates knowledge point E and knowledge point F in sub-problem 4. Then the computing device associates the knowledge points in different sub-problems. For example, it associates knowledge point A in sub-problem 1 with knowledge point E in sub-problem 3, associates knowledge point A in sub-problem 1 with knowledge point F in sub-problem 4, and so on. After that, the computing device uses the meta-information as the meta-information nodes for generating the meta-information relationship graph, and then connects the meta-information nodes corresponding to the knowledge points (i.e., meta-information) with an association relationship through connection edges to construct as Figure 3The shown meta-information relationship diagram (i.e., the first meta-information relationship diagram).
[0070] Furthermore, the computing device determines meta-information nodes with meta-information relationships according to the connection edges in the first meta-information relationship diagram. Then the computing device generates meta-information combinations from the meta-information nodes with meta-information relationships. For example, if knowledge point A, knowledge point B, and knowledge point C have meta-information relationships, the computing device combines knowledge point A, knowledge point B, and knowledge point C to generate a meta-information combination. Another example is that if knowledge point A and knowledge point E have meta-information relationships, the computing device combines knowledge point A and knowledge point E to generate a meta-information combination.
[0071] Furthermore, after the computing device generates multiple meta-information combinations from the first meta-information relationship diagram, it inputs each meta-information combination and the corresponding prompt words for prompting the generation of questions and answers into a preset question-answer generation model. The question-answer generation model synthesizes and processes the meta-information combinations to generate synthetic questions and corresponding synthetic answers. Then the computing device stores the synthetic questions and synthetic answers in a database, which are used as sample data when training the question-answer model.
[0072] As described in the background art, although large language models have demonstrated remarkable capabilities in various language tasks, they still have problems in understanding and solving complex problems (e.g., mathematics and programming). An effective way to solve this problem is for large language models to use large-scale high-quality synthetic data when understanding and solving complex problems. However, developing a low-cost and effective data synthesis method remains a challenge. Existing data synthesis methods have three main drawbacks: (1) Limited scalability: Existing methods are difficult to synthesize larger-scale data based on less seed data. (2) High cost: Current data synthesis methods rely on the help of commercial models, resulting in a significant increase in synthesis costs. (3) Similar to seed data: Due to the excessive dependence on seed data during the synthesis process, the newly generated data is very similar to the seed data, resulting in insufficient diversity of the generated data.
[0073] In view of the above technical problems, through the technical solution of the embodiments of the present application, the computing device decomposes the seed problem, so that meta-information corresponding to the seed problem can be obtained in multiple aspects and dimensions as extended seed data, thereby realizing the expansion of the scale of the seed data by this technical solution and obtaining a larger-scale seed data. And this technical solution generates a meta-information relationship graph by constructing the relationship between meta-information, and then flexibly combines each meta-information node in the meta-information relationship graph to generate various types of meta-information combinations corresponding to the meta-information relationship, and uses the meta-information combinations to generate corresponding questions and answers, and then generates more types of seed data, expands the scale of the seed data, reduces the dependence on the original seed data, and increases the diversity of the synthetic data. And most of the data models used in this technical solution for data synthesis are open-source models, avoiding relying on the help of commercial models and reducing the synthesis cost. Thus, the technical problems of poor data scalability, high data synthesis cost, and insufficient diversity of the generated data in the data information synthesis method are solved.
[0074] Optionally, the operation of constructing the first meta-information relationship graph according to the meta-information relationship between meta-information includes: selecting meta-information nodes in the same seed problem to form a single-hop relationship; connecting the meta-information nodes with a single-hop relationship through an explicit connection edge to construct a second meta-information relationship graph; and constructing the first meta-information relationship graph according to the second meta-information relationship graph and the meta-information relationship.
[0075] Specifically, the computing device screens the meta-information in the same seed problem from the meta-information library, so that the meta-information in the same seed problem forms a single-hop relationship between each other. For example, the computing device associates knowledge point A and knowledge point B in seed problem 1 to form a single-hop relationship; the computing device associates knowledge point A and knowledge point C in seed problem 1 to form a single-hop relationship; the computing device associates knowledge point B and knowledge point C in seed problem 1 to form a single-hop relationship. Another example is that the computing device associates knowledge point A and knowledge point D in seed problem 2 to form a single-hop relationship. And so on, so that the computing device forms a single-hop relationship for the knowledge points (i.e., meta-information) in all problem seeds.
[0076] Further, as shown in Figure 4 The computing device uses the meta-information with a single-hop relationship as the meta-information nodes for generating the second meta-information relationship graph, and connects them through an explicit connection edge (i.e., a solid line edge), thereby generating a meta-information relationship graph as shown in Figure 4 Shown (i.e., the second meta-information relationship graph). The single-hop relationship is an association relationship formed by connecting two meta-information nodes through an explicit connection edge.
[0077] Further, the computing device constructs a first meta-information relationship graph according to the second meta-information relationship graph and other meta-information relationships.
[0078] Thus, in this technical solution, by clearly defining the single-hop relationships between meta-information nodes within the same sub-problem, the direct associations between these meta-information can be accurately captured, ensuring the clarity and accuracy of the meta-information relationship graph. And in this technical solution, the meta-information with single-hop relationships is the information in the original sub-problem. Therefore, the correlation between this meta-information is relatively strong. Thus, when synthesizing data based on the meta-information with single-hop relationships, relatively core synthesized data can be obtained, making the synthesized data more reasonable.
[0079] Optionally, the operation of constructing a first meta-information relationship graph according to the second meta-information relationship graph and the meta-information relationships includes: screening two meta-information nodes connected by a common meta-information node in the second meta-information relationship graph to form a double-hop relationship; in the second meta-information relationship graph, connecting the meta-information nodes with double-hop relationships through a first implicit connection edge to construct a third meta-information relationship graph; and constructing a first meta-information relationship graph according to the third meta-information relationship graph and the meta-information relationships.
[0080] Specifically, after the computing device constructs the second meta-information relationship graph, it screens two meta-information nodes connected by a common meta-information node from the second meta-information relationship graph and associates the two meta-information nodes to form a double-hop relationship. Referring to the second meta-information relationship graph as shown in Figure 4 After that, the computing device screens from the second meta-information relationship graph that knowledge point A and knowledge point E are two meta-information nodes connected by a common meta-information node (i.e., knowledge point D). Therefore, knowledge point A and knowledge point E are associated to form a double-hop relationship. The meta-information forming the double-hop relationship is the meta-information in different sub-problems. Thus, the computing device constructs the double-hop relationships between all the meta-information nodes in the second meta-information relationship according to the method for generating the double-hop relationships between meta-information described above. Thus, the meta-information forming the double-hop relationship also includes: (1) knowledge point B and knowledge point D; (2) knowledge point C and knowledge point D; (3) knowledge point D and knowledge point E.
[0081] Further, the computing device connects the meta-information nodes with double-hop relationships in the second meta-information relationship graph through a first implicit connection edge (i.e., a long dashed line edge), thereby generating a meta-information relationship graph as shown in Figure 5 After that (i.e., the third meta-information relationship graph).
[0082] Further, the computing device constructs a first meta-information relationship graph according to the third meta-information relationship graph and other meta-information relationships.
[0083] Therefore, in this technical solution, by screening and associating two meta-information nodes connected by a common meta-information node in the second meta-information relationship graph, a double-hop relationship is formed, effectively connecting those indirectly but equally important meta-information. And in this technical solution, the double-hop relationship is the closest implicit relationship, indirectly connected through a node, so that a certain degree of relevance can be retained while expanding the scale of seed data.
[0084] Optionally, the operation of constructing the first meta-information relationship graph according to the third meta-information relationship graph and the meta-information relationship includes: determining, from the third meta-information relationship graph, other meta-information nodes connected to the core meta-information node through two common meta-information nodes; constituting a three-hop relationship according to the core meta-information node and the corresponding other meta-information nodes, where the core meta-information node is used to indicate a meta-information node whose explicit connection edge is greater than a preset threshold; in the third meta-information relationship graph, connecting the meta-information nodes with a three-hop relationship through a second implicit connection edge to construct a fourth meta-information relationship graph; and constructing the first meta-information relationship graph according to the fourth meta-information relationship graph and the meta-information relationship.
[0085] Specifically, after the computing device constructs the third meta-information relationship graph, it screens out the meta-information nodes in the third meta-information relationship graph whose connected explicit connection edges are greater than the preset threshold as the core meta-information nodes. The core meta-information nodes are usually meta-information nodes with more connected explicit connection edges in the meta-information relationship graph. For example, the preset threshold is 2, so the computing device screens out the meta-information nodes in the third meta-information relationship graph whose connected explicit connection edges are greater than 2 as the core meta-information nodes. Thus, the computing device screens out the core meta-information node with an explicit connection edge greater than 2 in the third meta-information graph as "Knowledge Point A".
[0086] After that, the computing device determines, from the third meta-information relationship graph, other meta-information nodes connected to the core meta-information node through two common meta-information nodes, that is, there is a path with an explicit connection edge of 3 between the core knowledge point and other knowledge points. For example, the core meta-information node "Knowledge Point A" is connected to another meta-information node "Knowledge Point F" through two common meta-information nodes (i.e., Knowledge Point D and Knowledge Point F), thereby associating the core meta-information node "Knowledge Point A" with the other meta-information node "Knowledge Point F" to form a three-hop relationship. The meta-information constituting the three-hop relationship is the meta-information in different seed questions.
[0087] Further, the computing device connects the meta-information nodes with a three-hop relationship in the third meta-information relationship graph through a second implicit connection edge (i.e., a short dashed line edge), thereby generating as Figure 6 shown in the meta-information relationship graph (i.e., the fourth meta-information relationship graph).
[0088] Further, the computing device constructs a first meta-information relationship graph according to the fourth meta-information relationship graph and the meta-information relationships.
[0089] Thus, in this technical solution, by identifying other meta-information nodes that are connected to the core meta-information node through two common meta-information nodes, meta-information with indirect associations across multiple levels can be recognized, thereby expanding the coverage of the meta-information network. And as the distance between knowledge points increases, the correlation gradually weakens. Therefore, this technical solution will search for core meta-information nodes in the meta-information relationship graph whose explicit connection edges exceed a preset threshold, thereby highlighting the key nodes in the meta-information relationship graph and ensuring that there is core data for meta-information with a three-hop relationship.
[0090] Optionally, the operation of constructing the first meta-information relationship graph according to the fourth meta-information relationship graph and the meta-information relationships includes: screening the pairwise-connected meta-information nodes in the fourth meta-information relationship graph where the number of meta-information nodes is equal to three to form a group relationship; and constructing the first meta-information relationship graph according to the fourth meta-information relationship graph and the group relationship.
[0091] Specifically, after the computing device constructs the third meta-information relationship graph, it screens from the fourth meta-information relationship graph the pairwise-connected meta-information nodes where the number of meta-information nodes is equal to three, and associates these meta-information nodes to form a group relationship. Referring to the fourth meta-information relationship graph as shown in Figure 6 The computing device screens from the fourth meta-information relationship graph that the number of meta-information nodes of knowledge point A, knowledge point B, and knowledge point C is equal to three and they are pairwise connected. Therefore, knowledge point A, knowledge point B, and knowledge point C are associated to form a group relationship. The meta-information nodes forming the group relationship can be meta-information in different seed problems or meta-information in the same seed problem.
[0092] Further, the computing device constructs the meta-information relationship graph (i.e., the first meta-information relationship graph) as shown in Figure 1 according to the fourth meta-information relationship graph and the meta-information combination having the group relationship.
[0093] Thus, the meta-information nodes with a group relationship in this technical solution are all strongly correlated with each other. Therefore, when synthesizing data based on the meta-information with a group relationship, relatively core synthesized data can be obtained, making the synthesized data more reasonable.
[0094] Optionally, according to the meta-information relationships between the meta-information nodes in the first meta-information relationship graph, the operation of combining the meta-information nodes to generate meta-information combinations includes: determining the meta-information nodes with single-hop relationships in the first meta-information relationship graph for combination to generate single-hop meta-information combinations; determining the meta-information nodes with double-hop relationships in the first meta-information relationship graph for combination to generate double-hop meta-information combinations; determining the meta-information nodes with triple-hop relationships in the first meta-information relationship graph for combination to generate triple-hop meta-information combinations; and determining the meta-information nodes with group relationships in the first meta-information relationship graph for combination to generate group meta-information combinations.
[0095] Specifically, the computing device determines the meta-information nodes with single-hop relationships according to the two meta-information nodes connected by the explicit connection edges in the first meta-information graph, such as knowledge point A and knowledge point B. Thus, the computing device combines the meta-information nodes with single-hop relationships to generate single-hop meta-information combinations. For example, combining knowledge point A and knowledge point B to generate a single-hop meta-information combination. Thus, the computing device combines all the meta-information nodes in the first meta-information relationship graph in the above manner to generate corresponding multiple single-hop meta-information combinations.
[0096] Further, the computing device determines the meta-information nodes with double-hop relationships according to the meta-information nodes connected by the first implicit connection edges in the first meta-information relationship graph, such as knowledge point A and knowledge point E. Thus, the computing device combines the meta-information nodes with double-hop relationships to generate double-hop meta-information combinations. For example, combining knowledge point A and knowledge point E to generate a double-hop meta-information combination. Thus, the computing device combines all the meta-information nodes in the first meta-information relationship graph in the above manner to generate corresponding multiple double-hop meta-information combinations.
[0097] Further, the computing device determines the meta-information nodes with triple-hop relationships according to the meta-information nodes connected by the second implicit connection edges in the first meta-information relationship graph, such as knowledge point A and knowledge point F. Thus, the computing device combines the meta-information nodes with triple-hop relationships to generate triple-hop meta-information combinations. For example, combining knowledge point A and knowledge point F to generate a triple-hop meta-information combination. Thus, the computing device combines all the meta-information nodes in the first meta-information relationship graph in the above manner to generate corresponding multiple triple-hop meta-information combinations.
[0098] Further, the computing device determines meta-information nodes with a group relationship based on three meta-information nodes that are pairwise connected using the explicit connection edges in the first meta-information relationship graph, such as knowledge point A, knowledge point B, and knowledge point C. Thus, the computing device combines the meta-information nodes with a group relationship to generate a group meta-information combination. For example, knowledge point A, knowledge point B, and knowledge point C are combined to generate a group meta-information combination. Thus, the computing device combines all the meta-information nodes in the first meta-information relationship graph in the above manner to generate corresponding multiple group meta-information combinations.
[0099] Thus, in this technical solution, the single-hop meta-information combination is used to synthesize variant questions of seed data. Therefore, using the single-hop meta-information combination for question synthesis can explore more question types. The double-hop meta-information combination and the triple-hop meta-information combination are used to synthesize distributed data, thereby increasing the diversity of the synthesized data. The group meta-information combination is used to generate comprehensive questions, thereby enriching the difficulty levels of the synthesized data. In summary, this technical solution greatly expands the scale of the seed data and improves the diversity of the seed data by generating various types of meta-information combinations.
[0100] Optionally, the method further includes: determining the number of times the meta-information nodes connected by each connection edge for connecting meta-information nodes in the first meta-information relationship graph appear in a seed question; and using the number of times as the weight of the corresponding connection edge.
[0101] Specifically, after generating the first meta-information relationship graph, the computing device determines the number of times the meta-information nodes connected by each connection edge for connecting meta-information nodes (i.e., including explicit connection edges and implicit connection edges) appear in a seed question.
[0102] For an explicit connection edge, for example, knowledge point A and knowledge point B are connected by an explicit connection edge. Thus, the computing device counts the number of times knowledge point A and knowledge point B appear in a seed question from all the collected seed questions. For example, if knowledge point A and knowledge point B appear in 3 seed questions at the same time, then the number of times knowledge point A and knowledge point B appear in a seed question is 3. After that, the computing device uses the number of times 3 as the weight of the explicit connection edge for connecting knowledge point A and knowledge point B.
[0103] For an implicit connection edge, the number of times the meta-information nodes connected by the implicit connection edge appear in a seed question is necessarily 0. Therefore, the weights of all implicit connection edges are 0.
[0104] Thus, this technical solution intuitively reflects the correlation degree and co-occurrence frequency between different meta-information nodes by counting the number of times the meta-information nodes connected by each connection edge appear in the same seed question and using this number of times as the weight of the connection edge.
[0105] Optionally, the operation of generating synthetic questions and corresponding synthetic answers by the question-and-answer generation model according to the meta-information combination and the corresponding prompting words includes: screening the meta-information combinations that meet the weight requirements as candidate meta-information combinations according to the weights of the connecting edges; generating prompting words for indicating the generation of synthetic questions by the prompting word generator according to the candidate meta-information combinations; generating synthetic questions by the question synthesis model according to the candidate meta-information combinations and the corresponding prompting words; assigning a difficulty level to the synthetic questions by the scoring model; and generating synthetic answers by the large language model corresponding to the difficulty level according to the corresponding synthetic questions.
[0106] Specifically, before data synthesis, the computing device determines the weights of each connecting edge (i.e., explicit connecting edges and implicit connecting edges) in the first meta-information relationship graph. Then the computing device screens the meta-information combinations in the first meta-information relationship graph that meet the weight requirements as candidate meta-information combinations. For example, if the computing device requires screening the connecting edges with a weight of 0, i.e., implicit connecting edges, then the computing device screens the meta-information combinations corresponding to the connecting edges with a weight of 0 in the first meta-information relationship graph as candidate meta-information combinations.
[0107] For another example, the computing device can also require the weight to be any natural number, that is, the computing device determines all meta-information combinations in the first meta-information relationship graph, including single-hop meta-information combinations with single-hop relationships, double-hop meta-information combinations with double-hop relationships, triple-hop meta-information combinations with triple-hop relationships, and group meta-information combinations with group relationships, so as to use the determined above-mentioned meta-information combinations as candidate meta-information combinations.
[0108] Furthermore, the computing device inputs the candidate meta-information combinations into a preset prompting word generator, and the prompting word generator generates prompting words for prompting the generation of synthetic questions. Then the computing device inputs the candidate meta-information combinations and the corresponding prompting words into the question-and-answer generation model. The question-and-answer generation model includes a question synthesis model, a scoring model, and a large language model.
[0109] Thus, the question-and-answer generation model performs a synthesis operation on the meta-information in each candidate meta-information combination according to the prompting words through the question synthesis model to generate corresponding synthetic questions. One candidate meta-information combination can generate one synthetic question. The question synthesis model can be the deepseek-math-rl model.
[0110] The prompting words corresponding to the question synthesis model are: "You are a math teacher. Now, you need to help your students learn the following math knowledge points. Using these knowledge points as a guide, please construct a new, original math question that requires understanding and application of all these knowledge points.
[0111] Ensure the following:
[0112] 1. The constructed problem must have no mathematical logic errors.
[0113] 2. The problem must combine all knowledge points.
[0114] 3. The problem should have sufficient difficulty and be logically coherent.
[0115] Knowledge point 1: {Knowledge point 1}
[0116] Knowledge point 2: {Knowledge point 2}
[0117] Knowledge point 3: {Knowledge point 3}
[0118] Please provide your answer in the following format:
[0119] New problem: {The new problem you created}”.
[0120] Furthermore, the computing device inputs each synthesized problem into a preset scoring model respectively, and the scoring model scores the difficulty of each synthesized problem. According to this score, a difficulty level is assigned to the synthesized problem. For example, a score of 0 - 60 represents a low difficulty level, a score of 60 - 80 represents a medium difficulty level, and a score of 81 - 100 represents a high difficulty level. The scoring model can be a bert model. The training data used to train this scoring model is mathematical Q&A data with difficulty ratings. Using the training data as supervision, let bert learn how to rate the difficulty of a mathematical problem as low, medium, or high.
[0121] Each of the low, medium, and high difficulty levels corresponds to a large - language model with a different parameter scale. For example, the large - language model corresponding to the low difficulty level has a smaller parameter scale, the large - language model corresponding to the high difficulty level has a larger parameter scale, and the large - language model corresponding to the medium difficulty level has a parameter scale between the two.
[0122] Furthermore, the computing device selects the corresponding large - language model according to the difficulty level of each synthesized problem, and then inputs the synthesized problem into the corresponding large - language model. The corresponding large - language model analyzes the synthesized problem and generates a corresponding answer (i.e., a synthesized answer).
[0123] The prompt word corresponding to this large - language model is: “
[0124] {Question}
[0125] Please think step by step to solve this question and put the final answer in \box{}.
[0126] Therefore, this technical solution screens out combinations of meta-information that meet specific conditions based on the weights of the connection edges as candidate objects, so that the required combinations of meta-information can be quickly located among a large amount of information, effectively reducing unnecessary consumption of computing resources and improving the operating efficiency of the entire system. And this technical solution matches the difficulty level of the synthesis problem with the parameter scale of the large language model, so that synthesis problems can be generated in multiple dimensions and at multiple levels.
[0127] In addition, after generating the synthesis problem and the synthesis answer, the computing device will also evaluate the synthesis problem and the synthesis answer. Specifically, to save costs, this technical solution only uses open-source evaluation models. To achieve the effect of using a closed-source model, this technical solution uses multiple evaluation models to jointly score and filter data. Experiments show that this strategy can achieve 94% of the effect of a closed-source model.
[0128] For the evaluation of the problem, this technical solution adopts a weighted scoring and filtering strategy. The synthesis problem is evaluated according to two criteria: (1) logical integrity (no mathematical errors and accurately related to the provided knowledge points); (2) expression integrity (clarity, integrity, and no hints or answers in the question). Each synthesis problem is scored by multiple evaluation models, and the score ranges from 0 to 1. Then, the computing device calculates the weighted score and filters out the questions with a weighted score lower than 0.85. For the evaluation of the synthesis answer, this technical solution adopts a single-vote veto strategy. The evaluation model requires that the synthesis answer has no mathematical errors and fully meets the requirements of the question. Each synthesis answer is scored 0 or 1 according to the consensus of the models. As long as one evaluation model deems the synthesis answer inappropriate, the synthesis answer will be eliminated.
[0129] Among them, the computing device inputs the synthesis problem and the corresponding prompt words into the evaluation model, and the evaluation model evaluates the synthesis problem.
[0130] Among them, the prompt words corresponding to the evaluation of the question are: "
[0131] Given these mathematical knowledge points:
[0132] Knowledge point 1: {Knowledge point 1}
[0133] Knowledge point 2: {Knowledge point 1}
[0134] I have constructed a new math problem as follows: {Newly constructed question}
[0135] Please evaluate whether this new math problem effectively covers all the knowledge points provided and identify any factual or logical errors in the problem. Please provide a floating-point number score between 0 and 1, where 1 indicates that the problem effectively covers both knowledge points and has no factual or logical errors, and 0 indicates that the problem fails to effectively cover any of the knowledge points or has factual or logical errors.
[0136] Please provide a reasonable score strictly in accordance with the above requirements. The format is as follows:
[0137] Score: {score}
[0138] Explanation: {explanation}”.
[0139] In addition, the computing device inputs the synthesized answer and the corresponding prompt words into the evaluation model, and the evaluation model evaluates the synthesized answer. The prompt words corresponding to the evaluation of the answer are:
[0140] You are given a math problem and its solution. Your task is to determine whether the provided solution is correct. Please read the problem and the solution carefully and consider the following criteria: “
[0141] 1. Computational accuracy: Check whether all numerical calculations are performed correctly.
[0142] 2. Logical consistency: Confirm whether the logical steps are carried out coherently and correctly in sequence.
[0143] 3. Solution completeness: Ensure that all parts of the problem are solved and whether the solution is comprehensive.
[0144] If any of these criteria is not met, you should answer "wrong".
[0145] Question: {question}
[0146] Solution: {solution}
[0147] Is the provided solution correct? Please answer in the following format:
[0148] Answer: {give your judgment, correct or wrong}
[0149] Explanation: {explanation}”.
[0150] Method effect: This technical solution uses the above method to synthesize 1.91 million math reasoning datasets GSDP-MATH with detailed solutions at a very low cost. The dataset covers various data distributions, multiple math fields, and multiple difficulty levels, providing high-quality math training data for large language models.
[0151] This technical solution uses GSDP-MATH to fine-tune on multiple baseline models (Qwen1.5-7B, LLaMA3-8B, and Mistral-7B), and verifies the effectiveness of this method on multiple mathematical reasoning metrics (GSM8K, MATH, GAOKAO-Math, and SVAMP). Using only GSDP-MATH, our model GSDP-7B (based on Mistral-7B) achieved an accuracy of 78.4% on GSM8K and 37.7% on MATH. The overall performance exceeded the baseline model by 26 points, defeating all competitors under the same conditions.
[0152] To further demonstrate the advantages of this method, we compared various data synthesis methods in terms of the expansion ratio (the ratio of the final synthesized data volume to the seed data volume) and the cost of synthesizing a single piece of data. As shown in Table 1, our method achieved the highest expansion ratio of 255, and at the same time, since our method does not use closed-source models, the cost of synthesizing a single piece of data (in units of 0.01 cents) is 5% or even 1% of other methods.
[0153] Table 1 Comparison of various methods in terms of expansion ratio and synthesis cost.
[0154]
[0155] In summary, as shown in Figure 7 the sequential method for a computing device to synthesize data based on seed problems is as follows:
[0156] S701: The computing device decomposes the seed problem through a preset decomposition model to generate meta-information related to the seed problem;
[0157] S702: The computing device selects meta-information nodes in the same seed problem to form single-hop relationships;
[0158] S703: The computing device connects the meta-information nodes with single-hop relationships through explicit connection edges to construct a second meta-information relationship graph;
[0159] S704: The computing device filters out two meta-information nodes in the second meta-information relationship graph that are connected by a common meta-information node to form double-hop relationships;
[0160] S705: The computing device connects the meta-information nodes with double-hop relationships through the first implicit connection edge in the second meta-information relationship graph to construct a third meta-information relationship graph;
[0161] S706: The computing device determines other meta-information nodes in the third meta-information relationship graph that are connected to the core meta-information node through two common meta-information nodes;
[0162] S707: The computing device forms a three-hop relationship based on the core meta-information node and the corresponding other meta-information nodes;
[0163] S708: In the third meta-information relationship graph, the computing device connects the meta-information nodes with a three-hop relationship through the second implicit connection edge to construct a fourth meta-information relationship graph;
[0164] S709: The computing device filters the meta-information nodes in the fourth meta-information relationship graph where the number of meta-information nodes is equal to three and are pairwise connected to form a group relationship;
[0165] S710: The computing device constructs a first meta-information relationship graph based on the fourth meta-information relationship graph and the group relationship;
[0166] S711: The computing device determines the number of times the meta-information nodes connected by each connection edge in the first meta-information relationship graph appear in a seed question, and uses the number of times as the weight of the corresponding connection edge;
[0167] S712: The computing device determines the meta-information nodes with a single-hop relationship in the first meta-information relationship graph for combination to generate a single-hop meta-information combination;
[0168] S713: The computing device determines the meta-information nodes with a double-hop relationship in the first meta-information relationship graph for combination to generate a double-hop meta-information combination;
[0169] S714: The computing device determines the meta-information nodes with a three-hop relationship in the first meta-information relationship graph for combination to generate a three-hop meta-information combination;
[0170] S715: The computing device determines the meta-information nodes with a group relationship in the first meta-information relationship graph for combination to generate a group meta-information combination;
[0171] S716: The computing device filters the meta-information combinations that meet the weight requirements as candidate meta-information combinations according to the weight of the connection edge;
[0172] S717: The computing device generates a prompt for indicating the generation of a synthetic question based on the candidate meta-information combination through a prompt generator;
[0173] S718: The computing device generates a synthetic question through a question synthesis model based on the candidate meta-information combination and the corresponding prompt;
[0174] S719: The computing device assigns a difficulty level to the synthetic question through a scoring model;
[0175] S720: The computing device generates a synthetic answer through a large language model corresponding to the difficulty level based on the corresponding synthetic question.
[0176] Thus, according to this embodiment, the computing device decomposes the seed problem, so that meta-information corresponding to the seed problem can be obtained in multiple aspects and dimensions as extended seed data. Therefore, this technical solution realizes the expansion of the scale of the seed data and obtains a larger scale of seed data. And this technical solution generates a meta-information relationship graph by constructing the relationships between meta-information. Then, each meta-information node in the meta-information relationship graph is flexibly combined to generate various types of meta-information combinations corresponding to the meta-information relationships, and corresponding questions and answers are generated using the meta-information combinations, thereby generating more types of seed data, expanding the scale of the seed data, reducing the dependence on the original seed data, and increasing the diversity of the synthetic data. And most of the data models used in this technical solution for data synthesis are open-source models, avoiding the dependence on commercial models and reducing the synthesis cost. Thus, the technical problems of poor data scalability, high data synthesis cost, and insufficient diversity of the generated data in the data information synthesis method are solved.
[0177] It should be noted that, for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present invention is not limited by the described action sequence, because according to the present invention, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present invention.
[0178] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal device (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods described in various embodiments of the present invention.
[0179] Embodiment 2
[0180] Figure 8 Shows an information synthesis device 800 according to this embodiment. The device 800 corresponds to the method described in the first aspect of Embodiment 1. Refer to Figure 8As shown, the device 800 includes: an information generation module 810, configured to decompose a seed problem through a preset decomposition model to generate meta-information related to the seed problem; a relationship graph generation module 820, configured to construct a first meta-information relationship graph according to the meta-information relationships between meta-information, where the meta-information relationships are used to indicate the association relationships between each meta-information in the same seed problem and the association relationships between each meta-information in different seed problems; a combination generation module 830, configured to combine meta-information nodes according to the meta-information relationships between the meta-information nodes in the first meta-information relationship graph to generate meta-information combinations, where different meta-information combinations have different meta-information relationships; and an information synthesis module 840, configured to generate a synthesized problem and a corresponding synthesized answer through a question-and-answer generation model according to the meta-information combination and the corresponding prompt words.
[0181] Optionally, the relationship graph generation module 820 includes: a first construction sub-module, configured to select meta-information nodes in the same seed problem to form a single-hop relationship; a second construction sub-module, configured to connect the meta-information nodes with a single-hop relationship through an explicit connection edge to construct a second meta-information relationship graph; and a third construction sub-module, configured to construct a first meta-information relationship graph according to the second meta-information relationship graph and the meta-information relationships.
[0182] Optionally, the third construction sub-module includes: a first construction unit, configured to screen two meta-information nodes connected by a common meta-information node in the second meta-information relationship graph to form a double-hop relationship; a second construction unit, configured to connect the meta-information nodes with a double-hop relationship through a first implicit connection edge in the second meta-information relationship graph to construct a third meta-information relationship graph; and a third construction unit, configured to construct a first meta-information relationship graph according to the third meta-information relationship graph and the meta-information relationships.
[0183] Optionally, the operation of constructing a first meta-information relationship graph according to the third meta-information relationship graph and the meta-information relationships includes: determining, from the third meta-information relationship graph, other meta-information nodes connected to a core meta-information node through two common meta-information nodes; forming a triple-hop relationship according to the core meta-information node and the corresponding other meta-information nodes, where the core meta-information node is used to indicate a meta-information node with an explicit connection edge greater than a preset threshold; connecting the meta-information nodes with a triple-hop relationship through a second implicit connection edge in the third meta-information relationship graph to construct a fourth meta-information relationship graph; and constructing a first meta-information relationship graph according to the fourth meta-information relationship graph and the meta-information relationships.
[0184] Optionally, the operation of constructing the first meta-information relationship graph according to the fourth meta-information relationship graph and the meta-information relationship includes: screening the pairwise-connected meta-information nodes in the fourth meta-information relationship graph where the number of meta-information nodes is equal to three to form a group relationship; and constructing the first meta-information relationship graph according to the fourth meta-information relationship graph and the group relationship.
[0185] Optionally, the operation of combining meta-information nodes according to the meta-information relationship between the meta-information nodes in the first meta-information relationship graph to generate meta-information combinations includes: determining the meta-information nodes with a single-hop relationship in the first meta-information relationship graph for combination to generate single-hop meta-information combinations; determining the meta-information nodes with a double-hop relationship in the first meta-information relationship graph for combination to generate double-hop meta-information combinations; determining the meta-information nodes with a triple-hop relationship in the first meta-information relationship graph for combination to generate triple-hop meta-information combinations; and determining the meta-information nodes with a group relationship in the first meta-information relationship graph for combination to generate group meta-information combinations.
[0186] Optionally, the apparatus 800 further includes: a first determination module, configured to determine the number of times the meta-information nodes connected by each connection edge for connecting meta-information nodes in the first meta-information relationship graph appear in a seed problem; and a second determination module, configured to use the number of times as the weight of the corresponding connection edge.
[0187] Optionally, the information synthesis module 840 includes: a first generation sub-module, configured to screen the meta-information combinations that meet the weight requirements as candidate meta-information combinations according to the weight of the connection edge; a second generation sub-module, configured to generate a prompt word for indicating the generation of a synthesis problem by a prompt word generator according to the candidate meta-information combination; a third generation sub-module, configured to generate a synthesis problem by a problem synthesis model according to the candidate meta-information combination and the corresponding prompt word; a fourth generation sub-module, configured to assign a difficulty level to the synthesis problem by a scoring model; and a fifth generation sub-module, configured to generate a synthesis answer by a large language model corresponding to the difficulty level according to the corresponding synthesis problem.
[0188] Thus, according to this embodiment, the computing device decomposes the seed problem, so that meta-information corresponding to the seed problem can be obtained in multiple aspects and dimensions as the extended seed data. Therefore, this technical solution realizes the expansion of the scale of the seed data and obtains a larger-scale seed data. And this technical solution generates a meta-information relationship graph by constructing the relationships between meta-information, and then flexibly combines each meta-information node in the meta-information relationship graph to generate various types of meta-information combinations corresponding to the meta-information relationship, and uses the meta-information combinations to generate corresponding questions and answers, thereby generating more types of seed data, expanding the scale of the seed data, reducing the dependence on the original seed data, and increasing the diversity of the synthetic data. And most of the data models used in this technical solution for data synthesis are open-source models, avoiding relying on the help of commercial models and reducing the synthesis cost. Thus, the technical problems of poor data scalability, high data synthesis cost, and insufficient diversity of the generated data in the data information synthesis method are solved.
[0189] Embodiment 3
[0190] Figure 9 Fig. shows an information synthesis device 900 according to this embodiment, and this device 900 corresponds to the method according to the first aspect of Embodiment 1. Refer to Figure 9 As shown, this device 900 includes: a processor 910; and a memory 920, connected to the processor 910, for providing instructions for the processor 910 to perform the following processing steps: decomposing the seed problem through a preset decomposition model to generate meta-information related to the seed problem; constructing a first meta-information relationship graph according to the meta-information relationship between meta-information, where the meta-information relationship is used to indicate the association relationship between each meta-information in the same seed problem and the association relationship between each meta-information in different seed problems; combining the meta-information nodes according to the meta-information relationship between the meta-information nodes in the first meta-information relationship graph to generate meta-information combinations, where different meta-information combinations have different meta-information relationships; and generating synthetic questions and corresponding synthetic answers through a question-and-answer generation model according to the meta-information combinations and corresponding prompt words.
[0191] Optionally, the operation of constructing a first meta-information relationship graph according to the meta-information relationship between meta-information includes: selecting meta-information nodes in the same seed problem to form a single-hop relationship; connecting the meta-information nodes with a single-hop relationship through an explicit connection edge to construct a second meta-information relationship graph; and constructing a first meta-information relationship graph according to the second meta-information relationship graph and the meta-information relationship.
[0192] Optionally, the operation of constructing the first meta-information relationship graph according to the second meta-information relationship graph and the meta-information relationship includes: screening two meta-information nodes connected by a common meta-information node in the second meta-information relationship graph to form a double-hop relationship; in the second meta-information relationship graph, connecting the meta-information nodes with a double-hop relationship through a first implicit connection edge to construct a third meta-information relationship graph; and constructing the first meta-information relationship graph according to the third meta-information relationship graph and the meta-information relationship.
[0193] Optionally, the operation of constructing the first meta-information relationship graph according to the third meta-information relationship graph and the meta-information relationship includes: determining, from the third meta-information relationship graph, other meta-information nodes connected to the core meta-information node through two common meta-information nodes; forming a triple-hop relationship according to the core meta-information node and the corresponding other meta-information nodes, where the core meta-information node is used to indicate a meta-information node with an explicit connection edge greater than a preset threshold; in the third meta-information relationship graph, connecting the meta-information nodes with a triple-hop relationship through a second implicit connection edge to construct a fourth meta-information relationship graph; and constructing the first meta-information relationship graph according to the fourth meta-information relationship graph and the meta-information relationship.
[0194] Optionally, the operation of constructing the first meta-information relationship graph according to the fourth meta-information relationship graph and the meta-information relationship includes: screening, in the fourth meta-information relationship graph, pairwise-connected meta-information nodes with the number of meta-information nodes equal to three to form a group relationship; and constructing the first meta-information relationship graph according to the fourth meta-information relationship graph and the group relationship.
[0195] Optionally, the operation of combining meta-information nodes according to the meta-information relationship between meta-information nodes in the first meta-information relationship graph to generate a meta-information combination includes: determining and combining meta-information nodes with a single-hop relationship in the first meta-information relationship graph to generate a single-hop meta-information combination; determining and combining meta-information nodes with a double-hop relationship in the first meta-information relationship graph to generate a double-hop meta-information combination; determining and combining meta-information nodes with a triple-hop relationship in the first meta-information relationship graph to generate a triple-hop meta-information combination; and determining and combining meta-information nodes with a group relationship in the first meta-information relationship graph to generate a group meta-information combination.
[0196] Optionally, the memory 920 is further configured to provide instructions for the processor 910 to process the following processing steps: determining the number of times the meta-information nodes connected by each connection edge for connecting meta-information nodes in the first meta-information relationship graph appear in a seed problem; and using the number of times as the weight of the corresponding connection edge.
[0197] Optionally, the operation of generating synthetic questions and corresponding synthetic answers by the question-and-answer generation model based on the meta-information combination and the corresponding prompt words includes: screening the meta-information combinations that meet the weight requirements as candidate meta-information combinations according to the weights of the connection edges; generating prompt words for indicating the generation of synthetic questions by the prompt word generator according to the candidate meta-information combinations; generating synthetic questions by the question synthesis model according to the candidate meta-information combinations and the corresponding prompt words; assigning a difficulty level to the synthetic questions by the scoring model; and generating synthetic answers by the large language model corresponding to the difficulty level according to the corresponding synthetic questions.
[0198] Thus, according to this embodiment, the computing device decomposes the seed question, so that meta-information corresponding to the seed question can be obtained in multiple aspects and dimensions as extended seed data. Thus, the technical solution realizes the expansion of the scale of the seed data and obtains a larger scale of seed data. And this technical solution generates a meta-information relationship graph by constructing the relationship between meta-information, and then flexibly combines each meta-information node in the meta-information relationship graph to generate various types of meta-information combinations corresponding to the meta-information relationship, and uses the meta-information combinations to generate corresponding questions and answers, thereby generating more types of seed data, expanding the scale of the seed data, reducing the dependence on the original seed data, and increasing the diversity of the synthetic data. And most of the data models used in this technical solution for data synthesis are open-source models, avoiding relying on the help of commercial models and reducing the synthesis cost. Furthermore, the technical problems of poor data scalability, high data synthesis cost, and insufficient diversity of the generated data in the data information synthesis method are solved.
[0199] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.
[0200] In the above embodiments of the present invention, the descriptions of the respective embodiments have their own emphases. For parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.
[0201] In several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection may be through some interfaces, and the indirect coupling or communication connection of the units or modules may be in an electrical or other form.
[0202] The unit described as a separation component may or may not be physically separated. The component displayed as a unit may or may not be a physical unit, that is, it may be located in one place or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0203] In addition, each functional unit in various embodiments of the present invention may be integrated in a processing unit, may exist separately as individual physical units, or two or more units may be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0204] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: USB flash drives, read-only memory (ROM), random access memory (RAM), mobile hard disks, magnetic disks, or optical discs and other various media that can store program codes.
[0205] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. An information synthesis method, characterized in that: include: Decomposing the seed problem by using a preset decomposition model to generate meta-information related to the seed problem; Constructing a first meta information relationship graph according to the meta information relationship between the meta information, wherein the meta information relationship is used to indicate the association relationship between each meta information in the same seed question and the association relationship between each meta information in different seed questions; According to the meta information relationship between the meta information nodes in the first meta information relationship graph, the meta information nodes are combined to generate meta information combinations, wherein different meta information combinations have different meta information relationships; as well as A synthetic question and a corresponding synthetic answer are generated by a question-answering generation model according to the meta-information combination and the corresponding prompt words.
2. The method according to claim 1, characterized in that The operation of constructing a first meta information relationship graph according to the meta information relationship between the meta information includes: Select the meta-information nodes in the same seed problem to form a single-hop relationship; Connecting the meta information nodes having a single-hop relationship through explicit connection edges to construct a second meta information relationship graph; and The first meta information relationship graph is constructed according to the second meta information relationship graph and the meta information relationship.
3. The method according to claim 2, characterized in that The operation of constructing the first meta information relationship graph according to the second meta information relationship graph and the meta information relationship includes: In the second meta-information relationship graph, two meta-information nodes connected by a common meta-information node form a double-hop relationship; In the second meta-information relationship graph, the meta-information nodes having a double-hop relationship are connected through a first implicit connection edge to construct a third meta-information relationship graph; and The first meta-information relationship graph is constructed according to the third meta-information relationship graph and the meta-information relationship.
4. The method according to claim 3, characterized in that The operation of constructing the first meta information relationship graph according to the third meta information relationship graph and the meta information relationship includes: Determine, from the third meta information relationship graph, other meta information nodes connected to the core meta information node via two common meta information nodes; A three-hop relationship is formed according to the core meta-information node and corresponding other meta-information nodes, wherein the core meta-information node is used to indicate the meta-information node whose explicit connection edge is greater than a preset threshold; In the third metadata relationship graph, metadata nodes having a three-hop relationship are connected through second implicit connection edges to construct a fourth metadata relationship graph; and The first meta-information relationship graph is constructed according to the fourth meta-information relationship graph and the meta-information relationship.
5. The method according to claim 4, characterized in that The operation of constructing the first meta information relationship graph according to the fourth meta information relationship graph and the meta information relationship includes: Filtering the meta-information nodes connected in pairs, in which the number of the meta-information nodes is equal to three, in the fourth meta-information relationship graph to form a group relationship; and The first meta information relationship graph is constructed according to the fourth meta information relationship graph and the group relationship.
6. The method according to claim 5, characterized in that The operation of combining the meta information nodes according to the meta information relationship between the meta information nodes in the first meta information relationship graph to generate a meta information combination includes: Determine the meta information nodes having the single-hop relationship from the first meta information relationship graph and combine them to generate a single-hop meta information combination; Determine the meta information nodes having the double-hop relationship from the first meta information relationship graph and combine them to generate a double-hop meta information combination; Determine the meta information nodes having the three-hop relationship from the first meta information relationship graph and combine them to generate a three-hop meta information combination; and Meta information nodes having the group relationship are determined from the first meta information relationship graph and combined to generate a group meta information combination.
7. The method according to claim 1, characterized in that Also includes: Determine the number of times the meta-information nodes connected by the connection edges for connecting the meta-information nodes in the first meta-information relationship graph appear in a seed question; as well as The number is used as the weight of the corresponding connecting edge.
8. The method according to claim 7, characterized in that The operation of generating a synthetic question and a corresponding synthetic answer according to the meta-information combination and the corresponding prompt words by the question-answer generation model includes: According to the weight of the connection edge, selecting a meta information combination that meets the weight requirement as a candidate meta information combination; Generate a prompt word for indicating generation of the synthetic question according to the candidate meta-information combination by a prompt word generator; Generate the synthetic question according to the candidate meta-information combination and the corresponding prompt words through the question synthesis model; assigning a difficulty level to the synthetic question by a scoring model; and The synthetic answer is generated according to the corresponding synthetic question through a large language model corresponding to the difficulty level.
9. An information synthesis device, characterized in that: include: An information generation module, used to decompose the seed problem through a preset decomposition model to generate meta-information related to the seed problem; A relationship graph generating module, used for constructing a first meta information relationship graph according to the meta information relationship between the meta information, wherein the meta information relationship is used for indicating the association relationship between each meta information in the same seed question and the association relationship between each meta information in different seed questions; a combination generating module, used for combining the meta information nodes according to the meta information relationship between the meta information nodes in the first meta information relationship graph to generate a meta information combination, wherein different meta information combinations have different meta information relationships; as well as The information synthesis module is used to generate a synthetic question and a corresponding synthetic answer according to the meta-information combination and the corresponding prompt words through a question-answer generation model.
10. An information synthesis device, characterized in that: include: processor; as well as A memory, connected to the processor, configured to provide the processor with instructions for processing the following processing steps: Decomposing the seed problem by using a preset decomposition model to generate meta-information related to the seed problem; Constructing a first meta information relationship graph according to the meta information relationship between the meta information, wherein the meta information relationship is used to indicate the association relationship between each meta information in the same seed question and the association relationship between each meta information in different seed questions; According to the meta information relationship between the meta information nodes in the first meta information relationship graph, the meta information nodes are combined to generate meta information combinations, wherein different meta information combinations have different meta information relationships; as well as A synthetic question and a corresponding synthetic answer are generated by a question-answering generation model according to the meta-information combination and the corresponding prompt words.
Citation Information
Patent Citations
Question and answer system providing indications of information gaps
CN103778471A
Analytical processing system supporting natural language analytic questions
CN111382171A