A method, apparatus, and device for generating large-scale data

Through LLM expansion and Monte Carlo search tree optimization, combined with asynchronous parallelization architecture and dynamic resource management, the problems of inefficiency and insufficient diversity in the data generation method are solved, and the balance of efficiency, diversity and goal-orientedness is achieved, and the quality and efficiency of data generation are improved.

CN119883657BActive Publication Date: 2025-07-04XIAMEN YUANTING INFORMATION TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510370436.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-07-04
Estimated Expiration
2045-03-27

AI Technical Summary

Technical Problem

Existing data generation methods are difficult to meet the needs of high efficiency, strong goal orientation and result diversity at the same time, and the computing resource utilization rate is insufficient, resulting in inefficient generation efficiency and increased costs.

Method used

LLM extension is used to generate multiple candidate content, through multi-dimensional evaluation and Monte Carlo search tree optimization, combining asynchronous parallelization architecture and dynamic resource management, dynamic parameters and exploration rewards are adjusted to form a self-iteration optimization mechanism.

Benefits of technology

It improves the efficiency and quality of data generation, ensures the rational use of resources, and achieves a balance of efficiency, diversity and goal-orientedness, and is suitable for large-scale data generation tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119883657B_ABST
    Figure CN119883657B_ABST
Patent Text Reader

Abstract

The present invention discloses a method, apparatus, and device for generating large-scale data, including: using an LLM to expand the text to be processed input by a user to generate multiple candidate contents, where each candidate content serves as an expansion node; performing multi-dimensional evaluation and recommendation evaluation on the candidate content corresponding to each expansion node through the LLM to obtain evaluation scores and optimization suggestions; creating a Monte Carlo search tree for each expansion node to obtain multiple expansion subtrees, where each expansion node serves as the root node of the expansion subtree; performing strategy optimization including selection, expansion, simulation, and backtracking on each expansion subtree based on the evaluation scores and the optimization suggestions to obtain large-scale data. It can meet the requirements of data generation efficiency, goal orientation, and diversity, and improve the quality and efficiency of data generation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to a method, apparatus, and device for generating large-scale data. Background Art

[0002] In the fields of artificial intelligence and data generation technology, with the continuous growth of data requirements, the limitations of traditional data generation methods have gradually emerged. On the one hand, although large models such as GPT and Stable Diffusion can generate rich and diverse data, in practical applications, due to the lack of an effective target-oriented mechanism, it is often necessary to manually adjust the prompt repeatedly to obtain data that meets the requirements. This process is not only inefficient but also difficult to ensure the stable quality of the generated data. On the other hand, the Monte Carlo Tree Search (MCTS), as an algorithm for optimizing decision-making paths, can provide certain optimization effects in data generation tasks, but its computational complexity is high, and the calculation time for a single node is too long to meet the speed requirements of large-scale data generation tasks.

[0003] In actual data generation tasks, the efficiency of data generation, the diversity of results, and the precise target orientation are three key requirements. However, the existing data generation methods have obvious deficiencies in these three aspects. Some methods often ignore the diversity and target orientation of data when pursuing high generation speed, resulting in a large amount of generated data, but the quality and practicality are poor. While other methods emphasize the diversity of results, but it is difficult to balance the generation efficiency and target orientation, making the data generation process time-consuming and unable to meet the timeliness requirements in practical applications. In addition, some methods lack effective integration and optimization of resources during the process of data generation, resulting in insufficient utilization of computing resources such as GPUs and increasing the cost of data generation. Summary of the Invention

[0004] In view of this, the purpose of the present invention is to propose a method, apparatus, and device for generating large-scale data, aiming to solve the problems that traditional data generation methods cannot simultaneously meet the requirements of high efficiency, strong target orientation, and result diversity.

[0005] To achieve the above object, the present invention provides a method for generating large-scale data, the method comprising:

[0006] Using an LLM to expand the text to be processed input by the user to generate a plurality of candidate contents, wherein each candidate content serves as an expansion node;

[0007] Performing multi-dimensional evaluation and recommendation evaluation on the candidate content corresponding to each expansion node through an LLM to obtain an evaluation score and an optimization recommendation;

[0008] Create a Monte Carlo search tree for each of the extended nodes, obtaining multiple extended subtrees, where each of the extended nodes serves as the root node of the extended subtree;

[0009] Perform policy optimization including selection, expansion, simulation, and backtracking on each of the extended subtrees based on the evaluation scores and the optimization suggestions to obtain large-scale data.

[0010] Preferably, the use of the LLM to expand the text to be processed input by the user to generate multiple candidate contents includes:

[0011] According to N threads = N cpu * (1 + W / E) for calculation to obtain the initial number of threads, where N cpu represents the number of CPU logical cores, and W / E represents the ratio of task waiting time to calculation time;

[0012] Call the corresponding threads to request the LLM according to the initial number of threads, and dynamically adjust the temperature parameter and the seed parameter to expand the text to be processed to obtain the candidate contents corresponding to the multiple extended nodes.

[0013] Preferably, the multi-dimensional evaluation of the candidate content corresponding to each of the extended nodes by the LLM to obtain evaluation scores includes:

[0014] Perform multi-dimensional evaluation of the candidate content corresponding to each of the extended nodes by the LLM according to the preset grammar correctness, context rationality, and domain relevance to obtain the grammar evaluation result, the context evaluation result, and the professionalism evaluation result;

[0015] Perform weighted summation on the grammar evaluation result, the context evaluation result, and the professionalism evaluation result to obtain the evaluation score.

[0016] Preferably, the policy optimization including selection, expansion, simulation, and backtracking on each of the extended subtrees based on the evaluation scores and the optimization suggestions includes:

[0017] Calculate the scores of the child nodes of each of the extended subtrees, and select the child node with the highest score as the target extended node;

[0018] Use the optimization suggestion as a prompt and input it into the LLM to expand the target extended node to generate the content to be evaluated corresponding to multiple branch nodes;

[0019] Perform multi-dimensional evaluation and recommended evaluation on the content to be evaluated corresponding to each of the branch nodes through the LLM to obtain the current evaluation score and the current optimization recommendation, where the current evaluation score is used to retrospectively update the score of the child node, and the current optimization recommendation is used as a prompt for the next expansion.

[0020] Preferably, the calculating the scores of the child nodes of each of the expanded subtrees includes:

[0021] Calculate according to UCB(S,a)=Q(S,a) / N(S,a)+C * sqrt(ln(N(S)) / N(S,a)) to obtain the scores of each child node; in the formula, Q(S,a) / N(S,a) represents the average reward obtained by taking action a from state S, C * sqrt(ln(N(S)) / N(S,a)) represents the exploration reward, C represents the exploration coefficient, Q represents the cumulative reward, and N represents the number of visits.

[0022] Preferably, the method further includes:

[0023] Monitor the child nodes expanded by the expanded subtree within a preset period. When the average reward of the child node is lower than the preset value and the number of visits exceeds the preset number, the child node will be deleted and the resources corresponding to the child node will be released.

[0024] To achieve the above object, the present invention also provides a large-scale data generation device, the device includes:

[0025] An expansion unit, configured to use the LLM to expand the text to be processed input by the user to generate a plurality of candidate contents, where each candidate content serves as an expansion node;

[0026] An evaluation unit, configured to perform multi-dimensional evaluation and recommended evaluation on the candidate content corresponding to each of the expansion nodes through the LLM to obtain an evaluation score and an optimization recommendation;

[0027] A creation unit, configured to create a Monte Carlo search tree for each of the expansion nodes to obtain a plurality of expanded subtrees, where each of the expansion nodes serves as the root node of the expanded subtree;

[0028] An optimization unit, configured to perform policy optimization including selection, expansion, simulation, and backtracking on each of the expanded subtrees based on the evaluation score and the optimization recommendation to obtain large-scale data.

[0029] Preferably, the device further includes:

[0030] The monitoring unit is used to monitor the child nodes expanded by the extended subtree within a preset period. When the average reward of a child node is lower than a preset value and the access times exceed a preset number, the child node is deleted and the resources corresponding to the child node are released.

[0031] To achieve the above object, the present invention also provides a large-scale data generation device, including a processor, a memory, and a computer program stored in the memory. The computer program is executed by the processor to implement the steps of a large-scale data generation method as described in the above embodiments.

[0032] To achieve the above object, the present invention also provides a computer-readable storage medium, on which a computer program is stored. The computer program is executed by a processor to implement the steps of a large-scale data generation method as described in the above embodiments.

[0033] To achieve the above object, the present invention also provides a computer program product, including a computer program / instructions. When the computer program / instructions are executed by a processor, the steps of a large-scale data generation method as described in the above embodiments are implemented.

[0034] Beneficial effects:

[0035] In the above solution, by constructing an asynchronous parallel architecture, decoupling expansion and evaluation, using multi-threading and queue mechanisms to improve the search efficiency, using the evaluation suggestions generated by the LLM as the optimized Prompt for the next round of expansion to form a dynamic feedback loop and self-iteration, continuously optimizing the generation result, and regarding each expanded node as an independent MCTS subtree for parallel search, and finally merging all subtree paths to achieve the global optimum, which not only ensures the diversity of the results but also can find the optimal generation path, improving the large-scale data generation efficiency and quality. In this case, by integrating the LLM and Monte Carlo tree search, it simultaneously meets the requirements of data generation efficiency, goal orientation, and diversity, improves the data generation quality and efficiency, and is applicable to large-scale data generation tasks.

[0036] Based on the ratio of the number of CPU cores to the task waiting time and the calculation time (W / E), the number of threads is dynamically calculated, realizing the reasonable allocation of thread resources, ensuring that computing resources can be efficiently utilized in different hardware environments, and improving the efficiency of expanded nodes; during the expansion process, the temperature parameter and the seed parameter are dynamically adjusted, which can further improve the generation speed and reduce resource consumption on the premise of ensuring the generation quality.

[0037] Adopt a multi-dimensional scoring mechanism, which integrates grammatical correctness, context rationality, and domain relevance for weighted evaluation, comprehensively measures the quality of candidate content, ensures that the generated data meets the requirements in terms of grammar, logic, and professionalism, provides a reliable basis for subsequent strategy optimization, and improves the overall quality of the generated data; The evaluation score is obtained by weighted summation, and the weights of each dimension can be flexibly adjusted according to actual needs, making the evaluation more targeted and scientific, and further improving the accuracy and practicality of the evaluation results.

[0038] During the strategy optimization process, the target expansion node is selected by calculating the scores of child nodes, and the optimization suggestions are used as prompts for expansion, realizing the self-optimization and improvement of the generation process, continuously adjusting the generation direction, and improving the accuracy and compliance of the generation results; The multi-dimensional evaluation and suggestion evaluation are carried out on the content of the newly generated branch nodes, further optimizing the data generation process. At the same time, the current evaluation score is used to backtrack and update the scores of child nodes, and the current optimization suggestions are used as prompts for the next expansion, forming a closed-loop optimization mechanism to continuously improve the quality and diversity of the generated data.

[0039] By using the UCB formula to calculate the scores of child nodes, considering both the average reward and the exploration reward, balancing exploration and exploitation, taking into account both historical performance and potential exploration value when selecting expansion nodes, which helps to discover better generation paths, avoid falling into local optima, improve the search efficiency and the quality of the generated data, and enhance the global search ability.

[0040] Through real-time monitoring and pruning, the inefficient branches are removed, optimizing the resource allocation, avoiding wasting computing resources on invalid or inefficient paths, further improving the efficiency and quality of large-scale data generation, and ensuring that resources can be concentrated on exploring effective paths. The thread resources released by pruning are automatically allocated to high-potential subtrees, realizing the dynamic optimization allocation of resources, improving the overall search efficiency, and accelerating the process of large-scale data generation. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0042] Figure 1 It is a schematic flowchart of a method for generating large-scale data provided by an embodiment of the present invention.

[0043] Figure 2Schematic diagram of the overall process for generating large-scale data provided by an embodiment of the present invention.

[0044] Figure 3 Schematic diagram of the structure of a device for generating large-scale data provided by an embodiment of the present invention.

[0045] The realization of the invention objective, functional features and advantages will be further described with reference to the embodiments and the accompanying drawings. Detailed implementation manners

[0046] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed present invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

[0047] The content of the present invention will be elaborated in detail below with reference to the embodiments.

[0048] Refer to Figure 1 The flowchart of a method for generating large-scale data provided by an embodiment of the present invention is shown as follows.

[0049] In this embodiment, the method includes:

[0050] S11. Use the LLM to expand the text to be processed input by the user to generate multiple candidate contents, where each candidate content serves as an expansion node.

[0051] Further, in step S11, the use of the LLM to expand the text to be processed input by the user to generate multiple candidate contents includes:

[0052] S11-1. Calculate according to N threads = N cpu * (1 + W / E) to obtain the initial number of threads, where N cpu represents the number of CPU logical cores, and W / E represents the ratio of the task waiting time to the calculation time;

[0053] S11-2. According to the initial number of threads, call the corresponding threads to request the LLM respectively, and dynamically adjust the temperature parameter and the seed parameter to expand the text to be processed, so as to obtain candidate contents corresponding to multiple said expanded nodes.

[0054] In this embodiment, this method can be applied to the form of question-answer pairs with an inference process to generate high-quality data for model training and fine-tuning. Refer to Figure 2 As shown, create an expanded node task queue according to the input text to be processed, and generate tasks for child nodes through the LLM for asynchronous node expansion. Specifically, according to the formula N threads =N cpu * (1 + W / E) to determine the initial number of threads. In the formula, N cpu represents the number of CPU logical cores, and W / E represents the ratio of task waiting time to computing time. Assume that it takes 50 ms of computing time and 150 ms of IO waiting time to generate each node, and the CPU has 4 cores, then the initial number of threads = 4 * (1 + 150 / 50) = 16 threads.

[0055] Use the LLM to batch generate K1 expanded nodes through the prompt, and its inference of the LLM is realized through asynchronous calls. Specifically, declare a queue for expanded node tasks, store the tasks of expanded nodes in this queue, and call threads respectively to request the LLM to generate corresponding expanded nodes; and initialize the generated expanded nodes, including: recording the parent node, the number of accesses (N, initially 0, incremented by 1 each time it is accessed), the cumulative reward (Q), the evaluation score (W), optimization suggestions, etc. When calling the LLM for inference, by dynamically adjusting the parameters, make the generated results (K1 expanded nodes) have greater differences. The dynamically adjusted parameters include: Temperature (temperature parameter), Temperature changes periodically between 0.3 and 0.8 to achieve a balanced exploration effect, and adjust the seed parameter, which is a time-based random seed to avoid pattern fixation and avoid too high a repetition rate of K1 expanded nodes. When expanding child nodes for the first time, the prompt is the user's input, and in subsequent expansions, the previous optimization suggestion is used as the prompt for feedback to achieve optimized expansion.

[0056] S12. Use the LLM to perform multi-dimensional evaluation and recommendation evaluation on the candidate content corresponding to each said expanded node to obtain an evaluation score and an optimization suggestion.

[0057] Further, in step S12, the use of the LLM to perform multi-dimensional evaluation on the candidate content corresponding to each said expanded node to obtain an evaluation score includes:

[0058] S12-1. Use the LLM to perform multi-dimensional evaluations on the candidate content corresponding to each of the extended nodes based on preset grammar correctness, context rationality, and domain relevance to obtain a grammar evaluation result, a context evaluation result, and a professionalism evaluation result.

[0059] S12-2. Perform a weighted sum of the grammar evaluation result, the context evaluation result, and the professionalism evaluation result to obtain the evaluation score.

[0060] In this embodiment, the extended nodes are given to the LLM through prompts for evaluation based on grammar correctness, context rationality, and domain relevance. The evaluation obtains the scores and optimization suggestions for the corresponding nodes. When expanding the next node, use this optimization suggestion as a prompt to achieve the optimization effect. The evaluation tasks are also added to the evaluation queue, and the producer-consumer pattern is used (the producer generates tasks into the queue, and the consumer fetches tasks from the queue to execute. That is, the generated evaluation tasks are placed in the queue, and when the thread is idle, the evaluation tasks are fetched from the queue for evaluation.) to achieve asynchronous evaluation. That is, the expanded nodes (without waiting for all nodes to be expanded, directly perform the next operation for the expanded nodes) are asynchronously evaluated by the LLM through prompts, and the tasks to be evaluated are put into the evaluation queue to achieve asynchronous evaluation. The LLM evaluation includes scoring evaluation and suggestion evaluation. The scoring evaluation backtracks and updates the corresponding Q value and the access count N, and the suggestion evaluation is used to inform the large model of the points to be optimized during the next expansion. For example:

[0061] # Sentence Quality Evaluation Prompt

[0062] [Input Requirements]

[0063] Please perform multi-dimensional scoring on the following text and give optimization opinions:

[0064] {{Sentence to be evaluated}}

[0065] * The requirement proposed by the user is: {The text initially input by the user, for example: Generate a medical conversation for me}

[0066] [Scoring Dimensions]

[0067] Weight Allocation:

[0068] - Grammar Correctness (30%): Check the subject-predicate-object structure, punctuation usage, and sentence pattern normality

[0069] - Context Rationality (35%): Evaluate logical coherence, scene suitability, and information completeness

[0070] - Domain Relevance (35%): Determine the domain based on the requirement initially proposed by the user

[0071] Output Format

[0072] {

[0073] "Syntax Evaluation": {

[0074] "score": (0 - 30),

[0075] "issues": ["List of specific syntax problems"],

[0076] "suggestions": ["Passive voice conversion", "Professional sentence restructuring"]

[0077] },

[0078] "Context Evaluation": {

[0079] "score": (0 - 35),

[0080] "logic_flow": "Excellent / Qualified / Broken",

[0081] },

[0082] "Professionalism Evaluation": {

[0083] "score": (0 - 35),

[0084] "term_accuracy": ["List of terms to be corrected"],

[0085] "compliance_check": "Pass / Partially Pass / Fail"

[0086] },

[0087] "Total Score": (Weighted calculated value), "Optimization Suggestions": "Places that need to be optimized and improved",

[0088] }

[0089] In the above, Suggestions are the suggestions to be given in the syntax evaluation. For example, passive voice conversion is to change an active sentence to a passive sentence to weaken the subjective text; professional sentence restructuring is to replace colloquial expressions with domain terms. Logic_flow: is the rating standard. Compliance_check: Compliance check, whether it complies with industry regulations, and if it passes, it means compliance. With these guidelines, the output of optimization suggestions will be more specific. That is, by guiding the LLM to score from angles such as syntax, context, and professionalism, and then corresponding optimization suggestions are put forward based on the scores of these angles.

[0090] S13. Create a Monte Carlo search tree for each of the extended nodes to obtain multiple extended subtrees, where each of the extended nodes serves as the root node of the extended subtree.

[0091] S14. Based on the evaluation scores and the optimization suggestions, perform policy optimization including selection, expansion, simulation, and backtracking on each of the extended subtrees to obtain large-scale data.

[0092] Further, in step S14, the performing policy optimization including selection, expansion, simulation, and backtracking on each of the extended subtrees based on the evaluation scores and the optimization suggestions includes:

[0093] S14-1. Calculate the scores of the child nodes of each of the extended subtrees, and select the child node with the highest score as the target extended node.

[0094] S14-2. Use the optimization suggestion as a prompt and input it into the LLM to expand the target extended node, generating the content to be evaluated corresponding to multiple branch nodes.

[0095] S14-3. Through the LLM, perform multi-dimensional evaluation and recommended evaluation on the content to be evaluated corresponding to each of the branch nodes to obtain the current evaluation score and the current optimization suggestion, where the current evaluation score is used to update the scores of the child nodes by backtracking, and the current optimization suggestion is used as a prompt for the next expansion.

[0096] Further, in step S14-1, the calculating the scores of the child nodes of each of the extended subtrees includes:

[0097] Calculate according to UCB(S,a)=Q(S,a) / N(S,a)+C * sqrt(ln(N(S)) / N(S,a)) to obtain the scores of each child node; in the formula, Q(S,a) / N(S,a) represents the average reward obtained by taking action a from state S, C * sqrt(ln(N(S)) / N(S,a)) represents the exploration reward, C represents the exploration coefficient, Q represents the cumulative reward, and N represents the number of visits; N(S) represents the total number of times state S has been visited. For example, in Monte Carlo tree search (MCTS), each time an action is selected starting from state S, N(S) accumulates all actions; N(S,a) represents the number of times action a has been selected in state S, and Q(S,a) represents the sum of all rewards obtained historically after selecting action a in state S.

[0098] In this embodiment, the extended nodes obtained by the above extension and initialization are used as the root nodes of the subtrees to be divided. Each subtree is used as an independent Monte Carlo search tree, and each subtree is optimized by including selection, expansion, simulation, and backtracking, that is, each subtree executes the complete selection → expansion → simulation → backtracking process to explore diverse paths in parallel. Specifically:

[0099] Phase 1: Selection (Path Selection):

[0100] Based on the UCB formula (balancing exploration and exploitation), calculate the scores of each child node (the higher the score, the greater the potential of this child node). That is,

[0101] Use the formula UCB(S,a)=Q(S,a) / N(S,a)+C * sqrt(ln(N(S)) / N(S,a)) to calculate the score, and select the child node corresponding to the most promising target path from all current child nodes according to the score. Among them,

[0102] Q(s,a) / N(s,a): Cumulative reward / number of visits, that is, the historical average reward of this path (biased towards high-reward paths, from the Simulation phase);

[0103] C * sqrt(ln(N(S)) / N(s,a)): Exploration reward (exploration term, encouraging paths with fewer visits, that is, when the number of visits is small, the larger the value, the more likely to be visited first);

[0104] Cumulative reward (Q): The total reward value obtained by all paths starting from this child node;

[0105] Number of visits (N): The number of times this child node is selected;

[0106] C represents the exploration coefficient (hyperparameter, increasing C at the beginning to promote exploration and decreasing it later).

[0107] By dynamically adjusting the search direction, select the appropriate target path, including: when the cumulative reward Q of the current path is high but the number of visits is small, this current path will be preferentially selected (when the cumulative reward Q is high and the number of visits N is small, the UCB value calculated according to the formula will be large, and when the UCB value is large, it will be preferentially selected) as the target path; when the number of visits of the current path is large and the cumulative reward tends to be stable, this current path will be more inclined to be selected as the target path.

[0108] Phase 2: Expansion (LLM Core Role in Expanding Nodes):

[0109] By invoking the optimized prompt (the optimized suggestions obtained previously are the optimized prompt), the LLM generates K2 new nodes (sources of diversity). That is, the optimized suggestions obtained from the previous evaluation are used as the new prompt to expand new branch nodes, achieving the effect of feedback optimization and self-correction. The tasks to be expanded are placed in the expansion node queue for asynchronous expansion. And by controlling the number of branches, the exploration efficiency is balanced. When it is shallow, for example, within 5 layers, the branch exploration is increased. When it is deep, the focus is on optimization and the number of branches is reduced. The termination conditions include: a relatively high feedback score value from the large model (the score evaluated by the large model) and no optimization content; reaching the preset exploration path depth; the number of nodes exceeding the preset maximum number.

[0110] Phase 3: Simulation (simulation-quality assessment):

[0111] The above-mentioned evaluation prompt is given to the large model to asynchronously evaluate each expanded branch node, obtaining 2 results, including the evaluation score and optimization suggestions, so as to achieve the quality assessment of the content to be evaluated corresponding to the evaluated branch node (instead of manual annotation). The evaluation score W is used to backtrack and calculate the UCB value (the cumulative reward Q in UCB = Q + W. When backtracking to the parent node, Q + W and the access count N + 1. At this time, by recalculating the UCB value, the UCB value is updated backtrackingly). The optimization suggestions are used as the prompt for the next expansion for optimization, realizing the closed loop of the LLM.

[0112] Phase 4: Backpropagation (backtracking-strategy update):

[0113] Operation: Backpropagate and update the corresponding path weights with the reward value of this round (the score evaluated by the large model);

[0114] Effect: The Q value of the high-frequency and high-reward path is increased, and it is more likely to be selected in subsequent iterations.

[0115] Furthermore, the method further includes:

[0116] S15, monitor the child nodes expanded by the expanded subtree within a preset period. When the average reward of the child nodes is lower than the preset value and the access count exceeds the preset number of times, the child nodes are deleted and the resources corresponding to the child nodes are released.

[0117] By setting up a monitoring service to monitor the change of Q value corresponding to each subtree node, deleting low-value branches according to the change of real-time Q value, and optimizing resource allocation. That is, it is monitored once every preset period of s seconds. If the average reward of a certain child node (the historical average reward of the corresponding path) ranks continuously below the threshold (such as the nodes with Q value ranking in the last 20% and the number of visits ≥ 5), and the number of visits N exceeds the preset number (such as N is greater than or equal to 5), then pruning is performed, and the thread resources released by pruning are automatically allocated to high-potential subtrees to improve the overall search efficiency. For example, if a parent node has 10 child nodes, and a certain child node ranks 9th or 10th in the average reward after multiple simulations, then it is pruned.

[0118] Refer to Figure 3 The structure diagram of a large-scale data generation device provided by an embodiment of the present invention is shown as follows.

[0119] In this embodiment, the device 20 includes:

[0120] An expansion unit 21, configured to expand the to-be-processed text input by the user by using an LLM to generate multiple candidate contents, where each candidate content serves as an expansion node;

[0121] An evaluation unit 22, configured to perform multi-dimensional evaluation and recommended evaluation on the candidate content corresponding to each expansion node through an LLM to obtain an evaluation score and an optimization recommendation;

[0122] A creation unit 23, configured to create a Monte Carlo search tree for each expansion node to obtain multiple expanded subtrees, where each expansion node serves as the root node of the expanded subtree;

[0123] An optimization unit 24, configured to perform strategy optimization including selection, expansion, simulation, and backtracking on each expanded subtree based on the evaluation score and the optimization recommendation to obtain large-scale data.

[0124] In another embodiment, the device 20 further includes:

[0125] A monitoring unit, configured to monitor the child nodes expanded by the expanded subtree within a preset period. When the average reward of the child node is lower than the preset value and the number of visits exceeds the preset number, the child node is deleted and the resources corresponding to the child node are released.

[0126] Each unit module of the device 20 can respectively execute the corresponding steps in the above method embodiment, so the unit modules will not be described in detail here. For details, please refer to the description of the above corresponding steps.

[0127] An embodiment of the present invention further provides a large-scale data generation device, which includes the large-scale data generation device as described above. Among them, the large-scale data generation device can adoptFigure 3 The structure of the embodiment, correspondingly, can execute Figure 1 The technical solution of the method embodiment shown, whose implementation principle and technical effects are similar. For details, reference can be made to the relevant records in the above embodiments, and will not be elaborated here.

[0128] The device includes: devices with a photographing function such as mobile phones, digital cameras or tablet computers, or devices with an image processing function, or devices with an image display function. The device may include components such as a memory, a processor, an input unit, a display unit, and a power supply.

[0129] Among them, the memory can be used to store software programs and modules. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory. The memory mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as an image playback function, etc.); the data storage area can store data created according to the use of the device, etc. In addition, the memory can include a high-speed random access memory, and can also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices. Correspondingly, the memory can also include a memory controller to provide access to the memory by the processor and the input unit.

[0130] The input unit can be used to receive input digital or character or image information, and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function controls. Specifically, the input unit of this embodiment, in addition to including a camera, can also include a touch-sensitive surface (such as a touch display screen) and other input devices.

[0131] The display unit can be used to display information input by the user or information provided to the user and various graphical user interfaces of the device. These graphical user interfaces can be composed of graphics, text, icons, videos and any combination thereof. The display unit can include a display panel. Optionally, the display panel can be configured in the form of an LCD (Liquid Crystal Display), an OLED (Organic Light-Emitting Diode), etc. Further, the touch-sensitive surface can cover the display panel. When the touch-sensitive surface detects a touch operation on or near it, it is transmitted to the processor to determine the type of touch event, and then the processor provides a corresponding visual output on the display panel according to the type of touch event.

[0132] An embodiment of the present invention further provides a computer-readable storage medium, which may be the computer-readable storage medium included in the memory in the above embodiment; or it may exist alone and be a computer-readable storage medium not assembled into the device. At least one instruction is stored in the computer-readable storage medium, and the instruction is loaded and executed by a processor to implement Figure 1 the method for generating large-scale data shown. The computer-readable storage medium may be a read-only memory, a magnetic disk, an optical disc, etc.

[0133] An embodiment of the present invention further provides a computer program product, including a computer program / instructions, and the computer program / instructions are loaded and executed by a processor to implement Figure 1 the method for generating a large-scale data shown.

[0134] It should be noted that the embodiments in this specification are all described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other. For the device embodiments, equipment embodiments, and storage medium embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiments.

[0135] Moreover, in this article, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without more limitations, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, article or device including the element.

[0136] The above description shows and describes the preferred embodiments of the present invention. It should be understood that the present invention is not limited to the form disclosed herein, and should not be regarded as excluding other embodiments, but can be used in various other combinations, modifications and environments, and can be changed within the scope of the inventive concept herein through the above teachings or the technology or knowledge in the relevant field. And the changes and modifications made by those skilled in the art without departing from the spirit and scope of the present invention should all be within the protection scope of the appended claims of the present invention.

Claims

1. A method for generating large-scale data, characterized in that, The method includes: Using an LLM to expand the text to be processed input by the user to generate multiple candidate contents, where each candidate content serves as an expansion node; Performing multi-dimensional evaluation and recommendation evaluation on the candidate content corresponding to each expansion node through the LLM to obtain evaluation scores and optimization suggestions; Creating a Monte Carlo search tree for each expansion node to obtain multiple expansion subtrees, where each expansion node serves as the root node of the expansion subtree; Performing strategy optimization including selection, expansion, simulation, and backtracking on each expansion subtree based on the evaluation scores and the optimization suggestions to obtain large-scale data; Further, the performing strategy optimization including selection, expansion, simulation, and backtracking on each expansion subtree based on the evaluation scores and the optimization suggestions includes: Calculating the scores of the child nodes of each expansion subtree and selecting the child node with the highest score as the target expansion node; Using the optimization suggestion as a prompt and inputting it into the LLM to expand the target expansion node to generate the content to be evaluated corresponding to multiple branch nodes; Performing multi-dimensional evaluation and recommendation evaluation on the content to be evaluated corresponding to each branch node through the LLM to obtain the current evaluation score and the current optimization suggestion, where the current evaluation score is used to update the score of the child node by backtracking, and the current optimization suggestion is used as a prompt for the next expansion; Further, the calculating the scores of the child nodes of each expansion subtree includes: Calculating according to UCB(S,a)=Q(S,a) / N(S,a)+C * sqrt(ln(N(S)) / N(S,a)) to obtain the scores of each child node; in the formula, Q(S,a) / N(S,a) represents the average reward obtained by taking action a from state S, C * sqrt(ln(N(S)) / N(S,a)) represents the exploration reward, C represents the exploration coefficient, Q represents the cumulative reward, N represents the number of visits, N(S) represents the total number of times state S is visited, N(S,a) represents the number of times action a is selected in state S, and Q(S,a) represents the sum of all rewards obtained historically after selecting action a in state S.

2. The method for generating a large-scale data according to claim 1, wherein The using an LLM to expand the text to be processed input by the user to generate multiple candidate contents includes: Calculate according to N threads = N cpu * (1 + W / E) to obtain the initial number of threads, where N cpu represents the number of CPU logical cores, and W / E represents the ratio of task waiting time to computing time; Respectively calling corresponding threads to request the LLM according to the initial number of threads, and dynamically adjusting the temperature parameter and the seed parameter to expand the text to be processed to obtain the candidate contents corresponding to multiple expansion nodes.

3. A method for generating large-scale data according to claim 1, characterized in that, The performing multi-dimensional evaluation on the candidate content corresponding to each expansion node through the LLM to obtain an evaluation score includes: Performing multi-dimensional evaluation on the candidate content corresponding to each expansion node through the LLM according to the preset grammar correctness, context rationality, and domain relevance to obtain the grammar evaluation result, the context evaluation result, and the professionalism evaluation result; Performing weighted summation on the grammar evaluation result, the context evaluation result, and the professionalism evaluation result to obtain the evaluation score.

4. A method for generating large-scale data according to claim 1, characterized in that, The method further includes: Monitor the child nodes expanded by the extended subtree within a preset period. When the average reward of a child node is lower than a preset value and the number of visits exceeds a preset number, delete the child node and release the resources corresponding to the child node.

5. A generating device for large-scale data, characterized in that, The device includes: An expansion unit, configured to use an LLM to expand the text to be processed input by a user, generating multiple candidate contents, where each candidate content serves as an expansion node; An evaluation unit, configured to perform multi-dimensional evaluation and recommendation evaluation on the candidate content corresponding to each expansion node through an LLM, obtaining an evaluation score and an optimization recommendation; A creation unit, configured to create a Monte Carlo search tree for each expansion node, obtaining multiple extended subtrees, where each expansion node serves as the root node of the extended subtree; An optimization unit, configured to perform strategy optimization including selection, expansion, simulation, and backtracking on each extended subtree based on the evaluation score and the optimization recommendation, obtaining large-scale data; Further, the optimization unit is further configured to: Calculate the scores of the child nodes of each extended subtree and select the child node with the highest score as the target expansion node; Use the optimization recommendation as a prompt and input it into the LLM to expand the target expansion node, generating the content to be evaluated corresponding to multiple branch nodes; Perform multi-dimensional evaluation and recommendation evaluation on the content to be evaluated corresponding to each branch node through the LLM, obtaining the current evaluation score and the current optimization recommendation, where the current evaluation score is used to update the score of the child node by backtracking, and the current optimization recommendation is used as a prompt for the next expansion; Further, the calculating the scores of the child nodes of each extended subtree includes: Calculating according to UCB(S,a)=Q(S,a) / N(S,a)+C * sqrt(ln(N(S)) / N(S,a)) to obtain the scores of each child node; in the formula, Q(S,a) / N(S,a) represents the average reward obtained by taking action a from state S, C * sqrt(ln(N(S)) / N(S,a)) represents the exploration reward, C represents the exploration coefficient, Q represents the cumulative reward, N represents the number of visits, N(S) represents the total number of times state S is visited, N(S,a) represents the number of times action a is selected in state S, and Q(S,a) represents the sum of all rewards obtained historically after selecting action a in state S.

6. The generating device for large-scale data according to claim 5, characterized in that, The device further includes: A monitoring unit, configured to monitor the child nodes expanded by the extended subtree within a preset period. When the average reward of a child node is lower than a preset value and the number of visits exceeds a preset number, delete the child node and release the resources corresponding to the child node.

7. A generating device for large-scale data, characterized in that, It includes a processor, a memory, and a computer program stored in the memory. The computer program is executed by the processor to implement the steps of a method for generating large-scale data according to any one of claims 1 to 4.

8. A computer program product comprising computer programs / instructions which, when executed by a processor, implement the steps of a method for generating large-scale data according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Knowledge graph question and answer retrieval method based on large language model and MCTS algorithm

    CN118296114A