Case generation method and related apparatus

By automatically generating and evaluating cases through a large language model, the problem of low efficiency in manually constructing cases is solved, and efficient and accurate case generation is achieved.

WO2025190144A1PCT designated stage Publication Date: 2025-09-18HUAWEI TECH CO LTD

Patent Information

Application Number
PCT/CN2025/080973
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-13
Filing Date
2025-03-06
Publication Date
2025-09-18

AI Technical Summary

Technical Problem

In the existing technology, case construction in professional fields mainly relies on manual construction, which is inefficient and consumes a lot of human resources.

Method used

Use a large language model to generate cases, evaluate case accuracy from different dimensions through the large language model, set modification goals and obtain modification opinions, and cyclically optimize the generation process to ensure case accuracy.

Benefits of technology

It improves the efficiency of case construction, ensures the accuracy of generated cases in all dimensions, reduces manual intervention, and improves case quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025080973_18092025_PF_FP_ABST
    Figure CN2025080973_18092025_PF_FP_ABST
Patent Text Reader

Abstract

A case generation method, used for generating cases in professional fields such as the field of law, the field of medicine or the field of economy. In the method, on the basis of using existing cases as reference cases, similar new cases can be generated by means of using a large language model, and after the new cases are generated, the accuracy of the new cases can be evaluated from different dimensions by means of a large language model; and when the accuracy of the new cases satisfy requirements, the generated new cases are output. The present application uses knowledge repositories and semantic comprehension capabilities of large language models to instruct the large language models to generate new cases on the basis of existing cases, and further evaluate the generated new cases from different dimensions, thus solving the hallucination problem of large language models, and ensuring the accuracy of generated new cases. Moreover, the present solution uses large language models to automatically generate cases without requiring manual construction of cases by people, thereby effectively improving the case construction efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

A case generation method and related device

[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office on March 13, 2024, with application number 202410292519.5 and invention name “A case generation method and related device”, the entire contents of which are incorporated by reference into this application. Technical Field

[0002] The present application relates to the field of artificial intelligence (AI) technology, and in particular to a case generation method and related devices. Background Art

[0003] Case studies typically refer to typical instances within specific fields, particularly law, economics, and medicine, that possess universal, representative, and practical significance. Case studies reflect the development and progression of problems and issues. By analyzing these cases, solutions and ideas can be proposed. Based on existing cases, researchers can conduct in-depth research and analysis of related issues, uncovering patterns and universal elements within them. This is one of the fastest and most accurate research methods in applied disciplines.

[0004] Currently, case studies in most fields are manually constructed by professionals in the field. For example, in the legal field, case studies often include past legal cases and their interpretations. Therefore, it is necessary for practitioners proficient in law to carefully annotate past legal cases from the perspective of laws and regulations to construct high-quality case studies.

[0005] However, this method of manually building cases by professionals often consumes a lot of human resources and the efficiency of building cases is too low. Summary of the Invention

[0006] This application provides a case generation method that does not require manual case construction and can effectively improve the efficiency of case construction.

[0007] The first aspect of the present application provides a case generation method for generating cases in professional fields such as law, medicine, or economics. In this method, it includes: first obtaining a first prompt word, the first prompt word includes a first case, and the first prompt word is used to indicate the generation of other cases with reference to the first case. The first case includes a historical event and a reasoning process for the historical event, and the first case is, for example, a case manually constructed by a professional. The first prompt word can be, for example: Please generate a similar case based on the following case; the case content includes "the case process is XXX, and the case reasoning process is XXX".

[0008] Then, the first prompt word is input into the first model, resulting in the output of the second example. The second example is generated by the first model with reference to the first example, as the first model is a large language model. By inputting the first prompt word into the first model, the first model can recognize the semantics expressed by the first prompt word, thereby understanding the task the user expects the first model to complete and outputting the second example.

[0009] Next, the second model is used to evaluate the accuracy of the second case from different dimensions, obtaining scores for the second case in these dimensions. That is, after obtaining the second case generated by the large language model, the large language model is further used as a judge to evaluate the accuracy of the generated second case from different dimensions, thereby ensuring that the evaluation of the second case in each dimension is accurate and ultimately guaranteeing the quality of the resulting new case. The second model can be a large language model.

[0010] When the scores of the second case in different dimensions meet the target conditions, it means that the accuracy of the second case has reached the requirement, so the second case is output.

[0011] In this solution, based on existing cases as reference cases, a large language model is used to generate similar new cases, and after the new cases are generated, the accuracy of the new cases is evaluated from different dimensions by the large language model. When the accuracy of the new cases meets the requirements, the generated new cases are output. This solution uses the powerful knowledge storage and semantic understanding capabilities of the existing large language model to instruct the large language model to generate new cases with reference to existing cases, and further evaluate the generated new cases from different dimensions to ensure the accuracy of the generated new cases, overcome the illusion problem of the large language model itself (that is, when generating cases, too much content is associated with it, resulting in inaccurate generated cases), and ultimately ensure that the content and format of the generated new cases can be aligned with the existing cases. In addition, this solution uses a large language model to automatically generate cases, without the need for manual case construction, which effectively improves the efficiency of case construction.

[0012] In one possible implementation, when the scores of the second case in different dimensions do not meet the target conditions, a modification target is set for the first model, and the second case is modified according to the modification target through the first model, so that the first model can output a modified case whose scores in different dimensions meet the target conditions.

[0013] In other words, if the second case does not meet the target conditions, it means that the second case generated by the first model is incorrect and needs to be further modified. Therefore, by setting a modification target for the first model, you can instruct the first model to continuously modify the second case until the scores of the modified case in different dimensions meet the target conditions.

[0014] In this solution, when the cases output by the large language model do not meet the requirements, a modification target is set for the large language model, prompting the large language model to continuously modify the cases previously output until the accuracy of the cases finally output can meet the requirements. This ensures the accuracy of the cases output by the large language model and effectively utilizes the cases initially output by the large language model, thereby improving the efficiency of the large language model in generating valid cases.

[0015] In one possible implementation, when the second model evaluates the accuracy of the second case from different dimensions, modification suggestions for the second case can also be obtained. The modification suggestions are obtained when the second model evaluates the accuracy of the second case from different dimensions. That is, when the second model evaluates the accuracy of the second case from different dimensions, the second model is also required to provide corresponding modification suggestions for the second case.

[0016] In this solution, when evaluating the accuracy of a case through a large language model, the large language model is also required to output modification suggestions for the case at the same time, so that the case can be modified in a targeted manner based on the modification suggestions given by the large language model, thereby improving the modification efficiency of the case and ensuring that the modified case can effectively improve the accuracy and meet the needs of users.

[0017] In a possible implementation, setting a modification target for the first model and modifying the second case according to the modification target through the first model may be implemented by executing a plurality of steps in a loop.

[0018] First, based on the modification suggestions corresponding to the case output by the first model, the first model modifies the case output by the first model to obtain a new case. The second case is the case output by the first model for the first time, so the second case is also the first version of the case that needs to be modified by the first model.

[0019] Then, the second model is used to evaluate the accuracy of the new case from different dimensions to obtain the scores of the new case in different dimensions.

[0020] In this way, the above steps are executed cyclically until the scores of the cases output by the first model in different dimensions meet the target conditions, so as to obtain the cases output by the first model for the last time.

[0021] In this solution, by cyclically executing the process in which the second model, acting as a judge, outputs modification opinions, and the first model, acting as a case generator, continuously modifies the generated cases based on the modification opinions, the first model can continuously and specifically improve the initially generated cases, so that the cases can be automatically optimized in the cycle, which is conducive to the batch generation of high-quality cases.

[0022] In one possible implementation, the second model includes multiple different large language models. To evaluate the accuracy of the second case from different dimensions, the second case can be input into each of the multiple different large language models, and scores output by the multiple different large language models can be obtained. The multiple different large language models may refer to large language models with different structures or training processes, such as those built by different manufacturers or companies.

[0023] In this solution, the accuracy of cases is comprehensively evaluated through multiple different large language models, which can ensure that the output cases are recognized by most large language models, avoid the overly one-sided situation when using a single large language model for evaluation, and effectively guarantee the accuracy of the final output cases.

[0024] In one possible implementation, after obtaining the second case output by the first model, multiple prompt words can be constructed based on the second case. Each of the multiple prompt words includes the second case, and each of the multiple prompt words indicates how the accuracy of the second case should be evaluated from different perspectives, each of which is related to the field to which the second case belongs. The multiple prompt words are then input into the second model to obtain scores for the second case from different evaluation perspectives.

[0025] In this solution, by setting different field-related evaluation angles to evaluate the same generated case according to the field to which the case belongs, it is possible to evaluate the accuracy of the case separately from different focus points, improve the accuracy of the large language model in evaluating the case, and thus ensure that the final output case has high accuracy from various angles.

[0026] In one possible implementation, the target condition includes that a weighted average of scores of the case in different dimensions is greater than or equal to a target threshold.

[0027] In one possible implementation, after outputting the second case, a second prompt word can also be obtained. The second prompt word is obtained by modifying the first prompt word based on the second case. The second prompt word is then input into the first model to obtain a third case output by the first model. The second model then evaluates the accuracy of the third case from different dimensions, obtaining scores for the third case in each dimension. If the scores for the third case in each dimension meet the target conditions, the third case is output.

[0028] In this solution, after the large language model outputs a case with guaranteed accuracy, it is further verified by humans to ensure the accuracy of the final case, and the prompt words adjusted by humans based on the problems found when verifying the case are obtained. The adjusted prompt words are then used to generate higher-quality cases, further improving the quality of subsequently generated cases.

[0029] In one possible implementation, the second prompt word is generated by adding guidance information to the first prompt word based on the second case. The guidance information is used to indicate how to generate a new case with reference to the first case. That is, after the professional verifies the second case and finds some inaccuracies in the second case, they can add some guidance information to the original prompt word to help the large language model generate a better case. This allows the large language model to generate a more accurate case based on the guidance of the guidance information.

[0030] In this solution, by having professionals in the field to which the case belongs add guidance information to the prompt words, more professional background knowledge can be integrated into the prompt words, which is conducive to guiding the large language model to generate higher quality cases.

[0031] In one possible implementation, the second prompt is generated by modifying the first case in the first prompt based on the second case. For example, unclear explanations or reasoning in the first case can be further explained to help the large language model better understand the reference case in the first prompt. For another example, if the second case is of higher quality, the first case in the first prompt can be replaced with the second case, allowing the newly generated second case to serve as a reference for generating further cases.

[0032] The second aspect of the present application provides a case generation device, including: an acquisition module, used to obtain a first prompt word, the first prompt word includes a first case, and the first prompt word is used to indicate the generation of other cases with reference to the first case; a processing module, used to input the first prompt word into a first model to obtain a second case output by the first model, and the second case is generated by the first model with reference to the first case; the processing module is also used to evaluate the accuracy of the second case from different dimensions through the second model, and obtain the scores of the second case in different dimensions; the processing module is also used to output the second case when the scores of the second case in different dimensions meet the target conditions.

[0033] In a possible implementation, both the first model and the second model are large language models.

[0034] In one possible implementation, the processing module is also used to: when the score of the second case in different dimensions does not meet the target conditions, set a modification target for the first model, and modify the second case according to the modification target through the first model, so that the first model can output a modified case whose scores in different dimensions meet the target conditions.

[0035] In one possible implementation, the acquisition module is also used to obtain modification opinions for the second case, which are obtained by evaluating the accuracy of the second case from different dimensions through the second model; the modification target includes modifying the second case with reference to the modification opinions so that the scores of the modified case in different dimensions meet the target conditions.

[0036] In one possible implementation, the processing module is further used to: based on the modification suggestions corresponding to the case last output by the first model, modify the case last output by the first model to obtain a new case, wherein the second case is the case output by the first model for the first time; evaluate the accuracy of the new case from different dimensions through the second model to obtain the scores of the new case in different dimensions; and repeatedly perform the above steps until the scores of the case output by the first model in different dimensions meet the target conditions to obtain the case output by the first model for the last time.

[0037] In one possible implementation, the second model includes multiple different large language models, and the processing module is further used to: input the second case into the multiple different large language models respectively, and obtain the scores output by the multiple different large language models respectively.

[0038] In one possible implementation, the processing module is further used to: construct multiple prompt words based on the second case, the multiple prompt words all include the second case, and the multiple prompt words are respectively used to indicate the evaluation of the accuracy of the second case from different evaluation perspectives, and the evaluation perspectives are all related to the field to which the second case belongs; input the multiple prompt words into the second model respectively to obtain the scores of the second case under different evaluation perspectives.

[0039] In one possible implementation, the target condition includes that a weighted average of scores of the case in different dimensions is greater than or equal to a target threshold.

[0040] In one possible implementation, after outputting the second case, the acquisition module is further used to obtain a second prompt word, where the second prompt word is obtained by the user modifying the first prompt word based on the second case; the processing module is further used to: input the second prompt word into the first model to obtain a third case output by the first model; evaluate the accuracy of the third case from different dimensions through the second model to obtain the scores of the third case in different dimensions; and output the third case when the scores of the third case in different dimensions meet the target conditions.

[0041] In a possible implementation, the second prompt word is obtained by adding guidance information to the first prompt word based on the second case, where the guidance information is used to instruct how to generate a new case with reference to the first case.

[0042] In a possible implementation, the second prompt word is obtained by modifying the first case in the first prompt word based on the second case.

[0043] A third aspect of the present application provides a computing device cluster, comprising at least one computing device, each computing device including a processor and memory. The processor of at least one computing device is configured to execute instructions stored in the memory of at least one computing device, causing the computing device cluster to perform the method described in the first aspect or any of the implementations of the first aspect. For details regarding the steps in each possible implementation of the first aspect performed by the computing device cluster, please refer to the first aspect and will not be repeated here.

[0044] In a fourth aspect, the present application provides a computer-readable storage medium having instructions stored therein. When the instructions are executed on a computer, the computer can execute any of the above methods.

[0045] A fifth aspect of the present application provides a computer program product comprising instructions, which, when executed on a computer, enables the computer to execute any of the methods described above.

[0046] In a sixth aspect, the present application provides a chip comprising a processor and a communication interface, wherein the communication interface is used to communicate with modules outside the chip, and the processor is used to run computer programs or instructions so that a device in which the chip is installed can execute any of the methods described above.

[0047] Among them, the technical effects brought about by any design method in the second to sixth aspects can refer to the technical effects brought about by different implementation methods in the above-mentioned first aspect, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] FIG1 is a schematic diagram of a system architecture 100 provided in an embodiment of the present application;

[0049] FIG2 is a flow chart of a case generation method provided in an embodiment of the present application;

[0050] FIG3 is a schematic diagram of a case generation based on a large language model provided in an embodiment of the present application;

[0051] FIG4 is a schematic diagram of an evaluation case based on multiple large language models provided in an embodiment of the present application;

[0052] FIG5 is a schematic diagram of implementing case evaluation based on multiple prompt words constructed from different evaluation perspectives, provided by an embodiment of the present application;

[0053] FIG6 is a schematic diagram of modifying a case by setting a modification target according to an embodiment of the present application;

[0054] FIG7 is a schematic diagram of a method for modifying a case by setting a modification target and combining modification opinions, provided by an embodiment of the present application;

[0055] FIG8 is a schematic diagram of a process for generating cases using a large language model according to an embodiment of the present application;

[0056] FIG9 is a schematic structural diagram of a case generation device provided in an embodiment of the present application;

[0057] FIG10 is a schematic structural diagram of a chip provided in an embodiment of the present application;

[0058] FIG11 is a schematic diagram of the structure of a computing device 1100 provided in an embodiment of the present application;

[0059] FIG12 is a schematic diagram of the structure of a computing device cluster provided in an embodiment of the present application;

[0060] FIG13 is a schematic diagram of the structure of another computing device cluster provided in an embodiment of the present application;

[0061] FIG14 is a schematic diagram of the structure of a computer-readable storage medium provided in an embodiment of the present application. DETAILED DESCRIPTION

[0062] In order to make the purpose, technical solutions and advantages of this application more clear, the embodiments of this application are described below in conjunction with the accompanying drawings. Obviously, the described embodiments are only embodiments of a part of this application, rather than all embodiments. It is known to those skilled in the art that with the emergence of new application scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.

[0063] The terms "first", "second", etc. in the specification and claims of this application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the descriptions used in this way can be interchangeable where appropriate so that the embodiments can be implemented in a sequence other than that illustrated or described in this application. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or modules is not necessarily limited to those steps or modules clearly listed, but may include other steps or modules that are not clearly listed or that are inherent to these processes, methods, products or devices. The naming or numbering of steps in this application does not mean that the steps in the method flow must be executed in the time / logical sequence indicated by the naming or numbering. The named or numbered process steps can change the execution order according to the technical purpose to be achieved, as long as the same or similar technical effects can be achieved. The division of units in this application is a logical division. In actual application, there may be other division methods. For example, multiple units can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between each other shown or discussed can be through some interfaces, and the indirect coupling or communication connection between units can be electrical or other similar forms, which are not limited in this application. Moreover, the units or sub-units described as separate components may or may not be physically separated, may or may not be physical units, or may be distributed into multiple circuit units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this application.

[0064] To facilitate understanding, some technical terms involved in the embodiments of this application are first introduced below.

[0065] (1) Case

[0066] Case studies typically refer to typical instances within a specific field, particularly law, economics, or medicine, that are universal, representative, and of practical significance. Case studies reflect the development and progression of issues and work, and through analysis of these cases, solutions and ideas can be proposed.

[0067] For example, in medicine, a case is a documented record of the diagnosis and treatment of a condition, ensuring a verifiable basis. In education, case studies utilize real, typical, and representative examples to help students master problem-solving methods and thinking, strengthening their understanding and ability to address real-world issues. In business, cases demonstrate the importance of management practices and the commercial value of successful cases, helping companies develop competitive strategies. Cases can also be literary accounts of real or fictional events, often involving individuals or specific groups, to convey situations and scenarios.

[0068] Using case studies, people can conduct in-depth research and analysis on relevant issues, uncovering patterns and universal elements. This is one of the fastest and most accurate research methods in applied disciplines. Case analysis has long been a highly effective research tool in disciplines such as psychology, management, education, medicine, and law.

[0069] (2) Large language model (LLM)

[0070] Large language models are deep learning models trained using large amounts of text data. They can generate natural language text or understand the meaning of text. Large language models can handle a variety of natural language tasks, such as text classification, question-answering, and conversation, and are an important path to artificial intelligence.

[0071] Specifically, large language models are a technology that has emerged in recent years. Because they undergo sophisticated data engineering and training processes, their parameters already incorporate a wealth of existing natural language processing knowledge. This knowledge can already replace humans in many language-related tasks, such as having large language models write code or perform text summarization.

[0072] (3) Prompt

[0073] Prompt originated from a form of input designed by researchers for downstream tasks. Its function is to help the pre-trained model "recall" what it has "learned" during pre-training, so it can also be called a prompt word. For large language models, Prompt is the user's input, which is used to instruct the large language model on the task to be completed. Prompt can be a simple question, a longer text, or a set of instructions, depending on the specific needs of the user. Generally speaking, Prompt is usually a short text string that can provide context and task-related information to help the model better understand the requirements and generate correct output. For example, in a question-answering task, Prompt may contain a description of the question or topic to help the large language model generate the correct answer. In addition, Prompt is usually designed by humans to help large language models better understand specific tasks or fields.

[0074] When the large language model generates content, it first processes the prompt and then generates output based on its understanding of it. The large language model works by predicting the probability of the next word appearing based on the previous context of the user's input, thereby generating the following text word by word. Therefore, differences in user-entered prompts can directly affect the quality of the large language model's output. In some cases, even if the user's input prompt differs by just a few words, the large language model may still generate significantly different content.

[0075] Currently, most cases in specialized fields are manually constructed by professionals in the field. Although the cases constructed by professionals are of high quality, this case construction method often consumes a lot of human resources and is too inefficient to effectively generate large quantities of cases.

[0076] Based on this, this application provides a case generation method. Based on existing cases as reference cases, a large language model is used to generate similar new cases. After generating the new cases, the large language model is used to evaluate the accuracy of the new cases from different dimensions. If the accuracy of the new cases meets the requirements, the generated new cases are output.

[0077] This solution leverages the powerful knowledge and semantic understanding capabilities of existing large language models, instructing them to generate new cases based on existing ones. Furthermore, the generated cases are evaluated from various perspectives to ensure their accuracy. This overcomes the inherent illusion problem inherent in large language models (i.e., the problem of generating inaccurate cases due to excessive associations when generating cases). Ultimately, the content and format of the generated new cases are aligned with existing ones. Furthermore, this solution utilizes large language models to automatically generate cases, eliminating the need for manual case construction, effectively improving case construction efficiency.

[0078] Specifically, because large language models learn so much knowledge, when generating new cases based on reference cases, they often associate the reference cases with a large amount of knowledge they have already learned (this is the phantom problem of large language models). This can ultimately lead to errors in the cases generated by the large language models. Therefore, in addition to using large language models as content producers, this solution also uses them as judges, evaluating the accuracy of the generated new cases from different dimensions. This ensures that the evaluation of new cases in each dimension is accurate, ultimately guaranteeing the quality of the resulting new cases.

[0079] Please refer to Figure 1, which is a schematic diagram of a system architecture 100 provided in an embodiment of the present application. As shown in Figure 1, in this system architecture 100, the execution device 110 can be implemented by at least one computing instance of a physical host (computing device), a virtual machine, or a container. When the execution device 110 is implemented by a virtual machine or a container, the execution device 110 actually exists in the form of a cloud computing product and can provide cloud services.

[0080] Optionally, the execution device 110 cooperates with other computing devices, such as data storage devices, load balancers, and other devices; the execution device 110 can be deployed on one physical site or distributed on multiple physical sites.

[0081] Optionally, in order to perform persistent storage of data, a data storage system 120 is further provided in the system architecture 100. The data storage system 120 may be located outside the execution device 110 (as shown in FIG1 ), and exchange data with the execution device 110 through a network. Optionally, in the case where the execution device 110 is a physical host, the data storage system 120 may also be located inside the execution device 110, such as when the data storage system 120 exchanges data with the processor through a bus. In this case, the data storage system 120 appears as a hard disk. In the case of having a data storage system 120, the execution device 110 may use the data in the data storage system 120, or call the program code in the data storage system 120 to implement the case generation method provided in the embodiment of the present application.

[0082] Optionally, users can operate their respective user devices (such as local device 101 and local device 102) to interact with execution device 110. Each local device can represent any computing device, such as a personal computer, a computer workstation, a smart phone, a tablet computer, a laptop computer, and a smart car.

[0083] Each user's local device can interact with the execution device 110 through a communication network of any communication mechanism / communication standard. The communication network can be a wide area network, a local area network, a point-to-point connection, etc., or any combination thereof.

[0084] In one implementation, the execution device 110 is used to implement the case generation method provided by the embodiment of the present application, thereby constructing a new case. Optionally, in the process of the execution device 110 implementing the case generation method, the local device 101 and the local device 102 can provide the execution device 110 with corresponding cases or prompt words, so that the execution device 110 can realize the generation of the case. Moreover, in the process of the local device 101 and the local device 102 needing to generate a new case, the execution device 110 can process the task data input by the user (such as the case or prompt word provided by the user) based on the case generation method to obtain a new case, and then return the generated new case to the local device 101 and the local device 102. In another implementation, one aspect or multiple aspects of the execution device 110 can be implemented by each local device, for example, the local device 101 can provide local data or feedback calculation results for the execution device 110, or execute the case generation method provided by the embodiment of the present application.

[0085] In general, the case generation method provided in the embodiments of the present application can be applied to electronic devices, such as the above-mentioned execution device 110, local device 101 or local device 102.

[0086] Please refer to Figure 2, which is a flow chart of a case generation method provided in an embodiment of the present application. As shown in Figure 2, the case generation method provided in an embodiment of the present application includes the following steps 201-204.

[0087] Step 201: Obtain a first prompt word, where the first prompt word includes a first case, and the first prompt word is used to indicate generating other cases with reference to the first case.

[0088] In this embodiment, before generating a new case, an existing case can be obtained as a reference case. For example, a case manually constructed by a professional can be obtained as a reference case to ensure the accuracy of the obtained reference case. In this way, the prompt word can be constructed based on the obtained case.

[0089] Specifically, the first prompt word obtained in this embodiment is constructed based on the first case that has already been obtained. Furthermore, the first prompt word is used in the input model to instruct the model to refer to the first case in the first prompt word to generate other cases. The first case included in the first prompt word is a case that has already been obtained and has guaranteed accuracy, and is used as a reference use case.

[0090] The first case includes a historical event and a reasoning process based on the historical event, wherein both the historical event and the reasoning process based on the historical event are closely related to the field to which the first case belongs.

[0091] For example, when the first case is a case in the legal field, the historical event can be, for example, a legal case that occurred in the past (such as a criminal case, a civil case, or other legal case), and the reasoning process can be a legal reasoning process for the legal case, such as judging the applicable legal provisions of the legal case based on the circumstances of the legal case and determining the judgment result of the legal case based on the applicable legal provisions of the legal case.

[0092] For another example, when the first case is a case in the medical field, the historical event may be, for example, a patient's medical record under the historical event, and the reasoning process may be a medical reasoning process for the patient's symptoms, such as asking the patient whether he has other related symptoms based on some of the patient's symptoms, judging the root cause of the symptoms based on all the patient's symptoms, and determining the medication method for the patient based on the root cause of the patient's symptoms.

[0093] The first prompt word can be further generated based on the first case to instruct the model to generate a case similar to the first case. For example, please refer to Figure 3, which is a schematic diagram of a case generation based on a large language model provided in an embodiment of the present application. As shown in Figure 3, for the first prompt word that needs to be input into the first model, the first prompt word can specifically be: Please generate a similar case based on the following case; the case content includes "the case process is XXX, and the case reasoning process is XXX".

[0094] Step 202: Input the first prompt word into the first model to obtain a second case output by the first model. The second case is generated by the first model with reference to the first case.

[0095] In this embodiment, the first model can be a large language model. By inputting the first prompt word into the first model, the first model can recognize the semantics expressed by the first prompt word, thereby understanding the task the user expects the first model to complete and outputting a second case. In other words, when processing the first prompt word, the first model will refer to the first case to generate a corresponding second case.

[0096] Step 203: Using the second model, the accuracy of the second case is evaluated from different dimensions to obtain scores of the second case in different dimensions.

[0097] Among them, the second model is specifically a large language model. It is understandable that since the large language model has learned too much knowledge, when the large language model generates a new case based on the reference case, it often associates the reference case with a large amount of knowledge content that has been learned (that is, the illusion problem of the large language model), which may eventually cause the case generated by the large language model to be prone to errors. Therefore, in this step, after obtaining the second case generated by the large language model, the large language model is further used as a judge to evaluate the accuracy of the generated second case from different dimensions, thereby ensuring that the evaluation of the second case in each dimension is accurate, and ultimately ensuring the quality of the new case obtained.

[0098] As shown in Figure 3, based on the large language model, we can construct multiple evaluation modules for different dimensions, such as an evaluation module for dimension 1, an evaluation module for dimension 2, and so on. This allows us to evaluate the same second case using evaluation modules across different dimensions, thereby obtaining scores for the second case across different dimensions.

[0099] Step 204: When the scores of the second case in different dimensions meet the target conditions, the second case is output.

[0100] In this embodiment, the target condition can be determined based on the actual application scenario (e.g., the number of dimensions and the accuracy requirements for the case) to measure the accuracy of the case. When the scores of the second case in different dimensions meet the target condition, it means that the accuracy of the second case has met the requirements, and thus the second case can be output. When the scores of the second case in different dimensions do not meet the target condition, it means that the accuracy of the second case has not met the requirements, and thus the second case may not be output.

[0101] For example, the target condition can be that the weighted average of the case's scores across different dimensions is greater than or equal to a target threshold. For example, for the second case, the second case has a score across multiple dimensions. Therefore, the scores across these dimensions can be weighted and summed, and the average of the summed values ​​can be calculated to ultimately obtain a weighted average. Thus, by comparing the calculated weighted average with the target threshold, it can be determined whether the second case meets the target condition.

[0102] As shown in Figure 3, after the evaluation modules based on different dimensions score the new cases output by the first model, scores for different dimensions can be obtained, such as score 1 for dimension 1, score 2 for dimension 2, and score n for dimension n. Then, for the scores for different dimensions, a weighted average of the scores for multiple dimensions can be calculated based on the formula (score 1 + score 2 + ... + score n) / n. The calculated weighted average is then compared to see if it is greater than or equal to a target threshold to determine whether to output a new case.

[0103] In this embodiment, there are multiple ways to evaluate the accuracy of the second case from different dimensions.

[0104] In one possible implementation, the second model may specifically include multiple different large language models. After obtaining the second case output by the first model, the second case is input into multiple different large language models respectively to obtain scores output by multiple different large language models. Among them, multiple different large language models may refer to large language models with different structures or different training processes, such as large language models constructed by different manufacturers or companies. This embodiment does not limit the source of the large language model. Exemplarily, multiple different large language models may be, for example, the Pangu large language model, the Qwen large language model, and the GPT-4 large language model.

[0105] In addition, the process of inputting the second case into multiple different large language models can specifically be to first generate a prompt word in combination with the second case, the prompt word includes the second case, and the prompt word is used to instruct the large language model to evaluate the accuracy of the second case. For example, the prompt word can be specifically: "In the XX scenario, the second case is specifically XXX. Please analyze the second case to determine whether the second case is accurate, and score the second case based on the accuracy of the second case. The scoring rule is specifically XXX." For each of the multiple different large language models, the same prompt word can be input, so that each large language model scores the second case separately. Ultimately, the scores output by different large language models can be used as scores of the second case in different dimensions.

[0106] Due to differences in structure or training processes, different large language models often have different performance. Therefore, when using different large language models to evaluate the accuracy of the second case, if multiple large language models output relatively high scores, it means that each large language model recognizes the accuracy of the second case, and therefore the accuracy of the second case can be considered to meet the requirements. In general, using multiple different large language models to comprehensively evaluate the accuracy of a case ensures that the output case is recognized by the majority of large language models, avoiding the overly biased evaluation that can occur when using a single large language model, and effectively ensuring the accuracy of the final output case.

[0107] For example, please refer to Figure 4, which is a schematic diagram of a case evaluation based on multiple large language models provided in an embodiment of the present application. As shown in Figure 4, after the first model outputs a new case based on the input prompt word, a prompt word can be constructed based on the new case to indicate that the new case is to be evaluated. Then, the prompt words constructed based on the new case are respectively input into the large language model 1, the large language model 2... the large language model n, so as to obtain the score 1 output by the large language model 1 (i.e., the score under dimension 1), the score 2 output by the large language model 2 (i.e., the score under dimension 2)... the score n output by the large language model n (i.e., the score under dimension n). Finally, based on the scores under n dimensions output by the n large language models, it can be determined whether the score for the new case meets the requirements (i.e., whether it meets the above-mentioned target conditions), thereby determining whether to output a new case.

[0108] In another possible implementation, after obtaining the second case output by the first model, multiple prompt words are first constructed based on the second case. These multiple prompt words all include the second case, and each prompt word is used to indicate how the accuracy of the second case is evaluated from different evaluation perspectives, each of which is related to the field to which the second case belongs. In this way, for each of the multiple prompt words constructed based on the second case, the prompt word instructs the large language model to evaluate the accuracy of the second case from one of the evaluation perspectives, and different prompt words indicate different evaluation perspectives.

[0109] For example, if the second case is a case in the legal field, the multiple different evaluation angles used to evaluate the second case may include evaluation angles such as corporate law, tax law, labor law, contract law, tax treatment, legal applicability, and reasoning of the resulting judgment. For another example, if the second case is a case in the medical field, the multiple different evaluation angles used to evaluate the second case may include evaluation angles such as symptom speculation, medical knowledge matching, clinical medicine, and pharmacology.

[0110] Then, the multiple prompt words are input into the second model to obtain scores for the second case from different evaluation perspectives. The second model used in this step can be the same model as the first model, or a different large language model. That is, the multiple prompt words are input into the same large language model to obtain scores for the second case from different evaluation perspectives.

[0111] For example, please refer to Figure 5, which is a schematic diagram of a case evaluation based on multiple prompt words constructed from different evaluation perspectives provided by an embodiment of the present application. As shown in Figure 5, after the first model outputs a new case based on the input prompt words, multiple prompt words (i.e., prompt word 1, prompt word 2... prompt word n ​​in Figure 5) can be constructed for the new case from different evaluation perspectives. Among them, multiple prompt words are used to indicate that the new case is evaluated from different evaluation perspectives. Then, the prompt words 1, prompt word 2... prompt word n ​​constructed based on the new case are respectively input into the large language model, thereby obtaining the score 1 (i.e., the score under dimension 1), score 2 (i.e., the score under dimension 2)... score n (i.e., the score under dimension n) output by the large language model for different prompt words. Finally, based on the scores under the n dimensions output by the large language model, it can be determined whether the score for the new case meets the requirements (i.e., whether the above-mentioned target conditions are met), thereby determining whether to output the new case.

[0112] In this solution, by setting different field-related evaluation angles to evaluate the same generated case according to the field to which the case belongs, it is possible to evaluate the accuracy of the case separately from different focus points, improve the accuracy of the large language model in evaluating the case, and thus ensure that the final output case has high accuracy from various angles.

[0113] It should be noted that in some cases, the two implementation methods introduced above can be combined. That is, after constructing multiple prompt words from different evaluation perspectives, the multiple prompt words are input into each large language model in multiple different large language models respectively, so as to obtain the score of the second case output by each large language model under different evaluation perspectives. Finally, the scores output by multiple large language models are combined to determine whether the accuracy of the second case meets the target conditions.

[0114] The above describes the process of scoring the cases output by the first model, and outputting the cases as generated cases if the case scores meet the target conditions. However, in some cases, the cases output by the first model may not meet the target conditions. Therefore, this embodiment also provides a method for modifying cases, which instructs the first model to modify the output cases so that the modified cases meet the target conditions.

[0115] Exemplarily, when the scores of the second case output by the first model in different dimensions do not meet the target conditions, a modification target can be set for the first model, and the second case can be modified according to the modification target through the first model, so that the first model can output a modified case whose scores in different dimensions meet the target conditions.

[0116] In other words, if the second case does not meet the target conditions, it means that the second case generated by the first model is incorrect and needs to be further modified. Therefore, by setting a modification target for the first model, you can instruct the first model to continue modifying the second case until the scores of the modified case in different dimensions meet the target conditions. In other words, the modification target set for the first model is to have the first model continue modifying the second case until the scores of the modified case in different dimensions meet the target conditions.

[0117] For example, for the second case output by the first model, the following prompt word can be generated: "Please modify the following second case so that the modified case can obtain a higher score when evaluating accuracy", and the generated prompt word is input into the first model, thereby instructing the first model to modify the second case. Of course, if the score of the modified case in different dimensions after the first model executes it once still does not meet the target conditions, the modified case will continue to be re-input into the first model for modification, that is, the case output by the first model is cyclically modified through the first model until the score of the modified case in different dimensions can meet the target conditions.

[0118] For example, please refer to Figure 6, which is a schematic diagram of a method for modifying a case by setting a modification target provided by an embodiment of the present application. As shown in Figure 6, after the first model outputs a new case based on the input prompt word and scores the new case from different dimensions, if the score of the new case in different dimensions does not meet the requirements (i.e., does not meet the above-mentioned target conditions), a modification target is set for the first model, and the new case output by the first model is further input into the first model, and the first model modifies the new case previously output until the score of the case modified by the first model in different dimensions meets the requirements.

[0119] In this solution, when the cases output by the large language model do not meet the requirements, a modification target is set for the large language model, prompting the large language model to continuously modify the cases previously output until the accuracy of the cases finally output can meet the requirements. This ensures the accuracy of the cases output by the large language model and effectively utilizes the cases initially output by the large language model, thereby improving the efficiency of the large language model in generating valid cases.

[0120] Optionally, when the large language model evaluates the accuracy of the second case from different dimensions, modification suggestions for the second case can also be obtained. The modification suggestions are obtained when the second model evaluates the accuracy of the second case from different dimensions. That is, when the second model evaluates the accuracy of the second case from different dimensions, the second model is also required to provide corresponding modification suggestions for the second case.

[0121] For example, before evaluating the second case through the second model, the following prompt words can be constructed based on the second case: "In the XX scenario, the second case is specifically XXX. Please analyze the second case to determine whether the second case is accurate, and score the second case based on the accuracy of the second case. The scoring rule is specifically XXX; and, if the second case is inaccurate, give modification suggestions for the second case."; then, the constructed prompt words are input into the second model, thereby instructing the second model to evaluate the second case and output modification suggestions for the second case.

[0122] Thus, in the case of a modification suggestion for the second case, the modification goal set for the first model can include modifying the second case with reference to the modification suggestion, so that the scores of the modified case in different dimensions meet the target conditions. For example, for the second case output by the first model, the following prompt words can be generated: "Please modify case XXX according to the modification suggestion XXX, so that the modified case can receive a higher score when evaluating accuracy", and the generated prompt words can be input into the first model, thereby instructing the first model to modify the second case according to the modification suggestion.

[0123] Exemplarily, please refer to Figure 7, which is a schematic diagram of a method for modifying a case based on setting a modification target and combining modification opinions provided by an embodiment of the present application. As shown in Figure 7, after the first model outputs a new case based on the input prompt word and scores the new case from different dimensions, each large language model participating in the scoring (i.e., large language model 1-large language model n) can output corresponding modification opinions. In this way, if the score of the new case in different dimensions does not meet the requirements (i.e., does not meet the above-mentioned target conditions), a modification target is set for the first model, and the new case output by the first model and the modification opinions given by the large language model 1-large language model n are continued to be input into the first model, and the new case output by the first model and the modification opinions given by the large language model 1-large language model n are modified by the first model with reference to the modification opinions given, until the score of the case modified by the first model in different dimensions can meet the requirements.

[0124] In this solution, when evaluating the accuracy of a case through a large language model, the large language model is also required to output modification suggestions for the case at the same time, so that the case can be modified in a targeted manner based on the modification suggestions given by the large language model, thereby improving the modification efficiency of the case and ensuring that the modified case can effectively improve the accuracy and meet the needs of users.

[0125] Optionally, a modification target is set for the first model, and the second case is modified according to the modification target through the first model, which may specifically include the following cyclic process.

[0126] Step 1: Based on the modification suggestions corresponding to the case output by the first model, the first model modifies the case output by the first model to obtain a new case. The second case is the first case output by the first model, and therefore the second case is also the first version of the case that needs to be modified by the first model.

[0127] Step 2: Use the large language model to evaluate the accuracy of the new case from different dimensions and obtain the scores of the new case in different dimensions.

[0128] In this way, the above steps can be executed cyclically until the scores of the cases output by the first model in different dimensions meet the target conditions, so as to obtain the cases output by the first model for the last time.

[0129] That is to say, if the scores of the new case (i.e., the case modified by the first model) in different dimensions meet the target conditions, the above loop can be terminated to output the case obtained by the last modification of the first model; if the scores of the new case in different dimensions still do not meet the target conditions, the above loop is continued, so that the first model continues to modify the case based on the modification suggestions fed back by the large language model until the scores of the modified case in different dimensions meet the target conditions.

[0130] In this solution, by cyclically executing the process in which the large language model serving as the judge outputs modification opinions, and the first model serving as the case generator continuously modifies the generated cases based on the modification opinions, the first model can continuously and specifically improve the initially generated cases, so that the cases can be automatically optimized in the cycle, which is conducive to the batch generation of high-quality cases.

[0131] The above describes the process of automatically generating and optimizing examples based on a large language model, ultimately enabling the large language model to output examples with a certain degree of accuracy. In some cases, to ensure the accuracy of the examples, professionals can also be employed to verify the examples output by the large language model to ensure that they are accurate and error-free. Furthermore, after verification of the examples output by the large language model, the professionals can provide feedback on the verified examples to improve the quality of subsequently generated examples.

[0132] Optionally, after the second case's scores across different dimensions meet the target criteria and the second case is generated, a second prompt word can be obtained. This second prompt word is obtained by modifying the first prompt word based on the second case. In other words, after verifying the second case generated based on the first prompt word, the user may identify deficiencies in the second case and further modify the first prompt word to generate a more optimal case based on the modified second prompt word.

[0133] Then, the second prompt word is input into the first model, resulting in the third case output by the first model. Furthermore, the large language model evaluates the accuracy of the third case from different dimensions, obtaining scores for the third case across these dimensions. If the scores for the third case across these dimensions meet the target conditions, the third case is output.

[0134] The process of obtaining the third case and evaluating the accuracy of the third case is similar to the process of obtaining the third case and evaluating the accuracy of the third case described above. Please refer to the description of the above embodiment for details, and will not be repeated here.

[0135] In this solution, after the large language model outputs a case with guaranteed accuracy, it is further verified by humans to ensure the accuracy of the final case, and the prompt words adjusted by humans based on the problems found when verifying the case are obtained. The adjusted prompt words are then used to generate higher-quality cases, further improving the quality of subsequently generated cases.

[0136] In one possible implementation, the second prompt word is generated by adding guidance information to the first prompt word based on the second case, where the guidance information is used to indicate how to generate a new case with reference to the first case. That is, after the professional verifies the second case and finds some inaccuracies in the second case, they can add some guidance information to the original prompt word to help the large language model generate a better case. This allows the large language model to generate a more accurate case based on the guidance of the guidance information.

[0137] Guidance information can include explanations of some of the content involved in the case generation process. For example, in the legal field, further explanations of some of the legal provisions involved in the case, or further explanations of the first case in the first prompt. Another example is in the medical field, further explanations of some of the medical terms involved in the case.

[0138] Alternatively, the guidance information may include constraints imposed on the case generation process, i.e., the cases generated by the first model should satisfy certain specific constraints. For example, in the legal field, the first prompt word may specify the type of law from which the generated case should be selected, or the legal provisions on which the generated case should be based. For another example, in the medical field, the first prompt word may specify which medical texts should be referenced for the generated case, or which medications should not be used for certain symptoms, etc.

[0139] In this solution, by having professionals in the field to which the case belongs add guidance information to the prompt words, more professional background knowledge can be integrated into the prompt words, which is conducive to guiding the large language model to generate higher quality cases.

[0140] In another possible implementation, the second prompt is generated by modifying the first case in the first prompt based on the second case. For example, unclear explanations or reasoning in the first case can be further explained to help the large language model better understand the reference case in the first prompt. For another example, if the second case is of higher quality, the first case in the first prompt can be replaced with the second case, allowing the newly generated second case to serve as a reference for generating further cases.

[0141] For ease of understanding, the specific implementation process of the case generation method provided in this embodiment will be described in detail below with reference to specific examples. Please refer to Figure 8, which is a schematic diagram of the process of generating cases using a large language model provided in this embodiment. As shown in Figure 8, the process of generating cases using a large language model can specifically include the following steps 1-7.

[0142] Step 1: Obtain prompt words generated based on the initial case.

[0143] First, we can obtain an initial case that has been constructed by professionals from the case library. The initial case includes the case and reasoning process. Then, we generate corresponding prompt words based on the initial case.

[0144] For example, a tax law expert interprets a tax scenario, reasoning step-by-step based on relevant legal provisions to arrive at a final tax treatment conclusion, thus forming an initial case. Based on this initial case, a prompt like "Please generate a similar case based on case XXX, subject to the following constraints: XXX" can be constructed.

[0145] Step 2: Select a large language model as the case generator, input the prompt word into the large language model, and the large language model generates a new case (hereinafter referred to as generated case) with reference to the initial case in the prompt word.

[0146] Among them, the new cases generated by the large language model also include cases and reasoning processes.

[0147] Step 3: Select multiple large language models as judges, and input the generated cases into the multiple large language models as judges respectively, so that different judges can score the generated cases.

[0148] Furthermore, different prompt words can be generated for each generated case from different evaluation perspectives to indicate how the generated case should be scored from these different evaluation perspectives. This allows each large language model acting as a judge to receive multiple different prompt words, thereby scoring the same generated case from multiple different evaluation perspectives. This allows different judges to output scores for the generated case from different evaluation perspectives.

[0149] Step 4: Combine the scores output by different judges to determine whether the generated case meets the accuracy requirements.

[0150] Specifically, since each judge may output one or more scores (such as scores from multiple different evaluation perspectives), a weighted average of the multiple scores output by all judges can be calculated to obtain the average score corresponding to the generated case.

[0151] If the average score corresponding to the generated case is greater than or equal to the target threshold, it means that multiple judges agree with the generated case, so the generated case is output.

[0152] If the average corresponding to the generated case is less than the target threshold, it means that multiple judges have rejected the generated case and the generated case needs to be further optimized.

[0153] It should be noted that when the judges score the generated cases, they will also identify errors in the generated cases and provide corresponding modification suggestions.

[0154] In step 5, if the average score of the generated cases is less than the target threshold, the large language model serving as the case generator is instructed to modify the generated cases based on the modification suggestions given by the judges.

[0155] When instructing the large language model (the case generator) to modify a generated case based on modification suggestions, the modification target of the large language model (the case generator) can be specified. For example, after receiving modification suggestions, the following prompt words can be generated based on the generated case and the modification suggestions: "For case XXX, the following error XXX exists; please modify this case to gain more approval from the judges." The generated prompt words are then input into the large language model (the case generator), thereby modifying the generated case.

[0156] Step 6: Repeat steps 3 to 5 above until the generated case meets the accuracy requirement, thereby outputting the final generated case.

[0157] In other words, after the generated example is modified in step 5, the modified example is continued to be scored by multiple large language models acting as judges to determine whether the modified example meets the accuracy requirements. If the accuracy of the modified example still does not meet the accuracy requirements, further modifications are made to the modified example based on the modification suggestions until the final modified example meets the accuracy requirements.

[0158] Step 7: User verification.

[0159] For generated cases whose scores meet accuracy requirements, users can perform a final review and verification. If the user is still dissatisfied with the generated case after verification, they can manually modify the errors in the generated case and substitute them as a new reference case to generate a new case in step 1. Alternatively, users can provide guidance information to construct new prompt words, thereby guiding the large language model to generate higher-quality cases. Specific guidance information may include explanations of laws and regulations, or restrictions on the content generated by the large language model. This information depends on the actual scenario and is not specifically limited here.

[0160] The above describes in detail the method provided by the embodiment of the present application. Next, the device provided by the embodiment of the present application for executing the above method will be introduced.

[0161] Please refer to Figure 9, which is a schematic diagram of the structure of a case generation device provided in an embodiment of the present application. As shown in Figure 9, the case generation device provided in an embodiment of the present application includes: an acquisition module 901, which is used to obtain a first prompt word, the first prompt word includes a first case, and the first prompt word is used to indicate that other cases are generated with reference to the first case; a processing module 902, which is used to input the first prompt word into the first model to obtain a second case output by the first model, and the second case is generated by the first model with reference to the first case; the processing module 902 is also used to evaluate the accuracy of the second case from different dimensions through the second model, and obtain the scores of the second case in different dimensions; the processing module 902 is also used to output the second case when the scores of the second case in different dimensions meet the target conditions.

[0162] In a possible implementation, both the first model and the second model are large language models.

[0163] In one possible implementation, the processing module 902 is further used to: when the score of the second case in different dimensions does not meet the target conditions, set a modification target for the first model, and modify the second case according to the modification target through the first model, so that the first model can output a modified case whose scores in different dimensions meet the target conditions.

[0164] In one possible implementation, the acquisition module 901 is also used to obtain modification opinions for the second case, which are obtained by evaluating the accuracy of the second case from different dimensions through the second model; the modification target includes modifying the second case with reference to the modification opinions so that the scores of the modified case in different dimensions meet the target conditions.

[0165] In one possible implementation, the processing module 902 is further used to: based on the modification suggestions corresponding to the case last output by the first model, modify the case last output by the first model to obtain a new case, wherein the second case is the case output by the first model for the first time; evaluate the accuracy of the new case from different dimensions through the second model to obtain the scores of the new case in different dimensions; and repeatedly perform the above steps until the scores of the case output by the first model in different dimensions meet the target conditions to obtain the case output by the first model for the last time.

[0166] In one possible implementation, the second model includes multiple different large language models, and the processing module 902 is further configured to input the second case into the multiple different large language models respectively to obtain scores outputted by the multiple different large language models respectively.

[0167] In one possible implementation, the processing module 902 is further used to: construct multiple prompt words based on the second case, where the multiple prompt words all include the second case, and the multiple prompt words are respectively used to indicate how to evaluate the accuracy of the second case from different evaluation perspectives, and the evaluation perspectives are all related to the field to which the second case belongs; and input the multiple prompt words into the second model respectively to obtain scores for the second case from different evaluation perspectives.

[0168] In one possible implementation, the target condition includes that a weighted average of scores of the case in different dimensions is greater than or equal to a target threshold.

[0169] In one possible implementation, after outputting the second case, the acquisition module 901 is further used to obtain a second prompt word, where the second prompt word is obtained by the user modifying the first prompt word based on the second case; the processing module 902 is further used to: input the second prompt word into the first model to obtain a third case output by the first model; evaluate the accuracy of the third case from different dimensions through the second model to obtain the scores of the third case in different dimensions; and output the third case when the scores of the third case in different dimensions meet the target conditions.

[0170] In a possible implementation, the second prompt word is obtained by adding guidance information to the first prompt word based on the second case, where the guidance information is used to instruct how to generate a new case with reference to the first case.

[0171] In a possible implementation, the second prompt word is obtained by modifying the first case in the first prompt word based on the second case.

[0172] The acquisition module 901 and the processing module 902 can be implemented by software or hardware. For example, the implementation of the processing module 902 will be described below using the processing module 902 as an example. Similarly, the implementation of the acquisition module 901 can refer to the implementation of the processing module 902.

[0173] The processing module 902 is an example of a software functional unit. The processing module 902 may include code running on a computing instance. The computing instance may include at least one of a physical host (computing device), a virtual machine, and a container. Furthermore, the computing instance may be one or more. For example, the processing module 902 may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers used to run the code may be distributed in the same region or in different regions. Furthermore, the multiple hosts / virtual machines / containers used to run the code may be distributed in the same availability zone (AZ) or in different AZs, each AZ including one data center or multiple geographically close data centers. Typically, a region may include multiple AZs.

[0174] Similarly, multiple hosts / virtual machines / containers running the code can be distributed within the same virtual private cloud (VPC) or across multiple VPCs. Typically, a VPC is set up within a region. Cross-region communication between two VPCs within the same region, or between VPCs in different regions, requires a communication gateway within each VPC to interconnect the VPCs.

[0175] As an example of a hardware functional unit, processing module 902 may include at least one computing device, such as a server. Alternatively, processing module 902 may be implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD may be implemented using a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.

[0176] The multiple computing devices included in processing module 902 can be distributed in the same region or in different regions. The multiple computing devices included in processing module 902 can be distributed in the same AZ or in different AZs. Similarly, the multiple computing devices included in processing module 902 can be distributed in the same VPC or in multiple VPCs. The multiple computing devices can be any combination of servers, ASICs, PLDs, CPLDs, FPGAs, GALs, and other computing devices.

[0177] It should be noted that the information interaction, implementation process, etc. between the modules / units of the above-mentioned device are based on the same concept as the method embodiment of the present application, and the technical effects they bring are the same as those of the method embodiment of the present application. For specific contents, please refer to the description in the method embodiment shown above in the embodiment of the present application, and no further details will be given here.

[0178] The case generation device provided in the embodiment of the present application can specifically be a chip, and the chip includes: a processing unit and a communication unit. The processing unit can be, for example, a processor, and the communication unit can be, for example, an input / output interface, a pin, or a circuit. The processing unit can execute the computer-executable instructions stored in the storage unit to enable the chip in the electronic device to execute the method described in the above embodiment. Optionally, the storage unit is a storage unit in the chip, such as a register, a cache, etc. The storage unit can also be a storage unit located outside the chip in the wireless access device, such as a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM), etc.

[0179] Specifically, see Figure 10, which is a schematic diagram of the structure of a chip provided in an embodiment of the present application. The chip can be represented as a neural network processor NPU 1000. NPU 1000 is mounted on the host CPU (host CPU) as a coprocessor and is assigned tasks by the host CPU. The core of the NPU is arithmetic circuit 1003, which is controlled by controller 1004 to extract matrix data from memory and perform multiplication operations.

[0180] In some implementations, the arithmetic circuit 1003 includes multiple processing units (PEs). In some implementations, the arithmetic circuit 1003 is a two-dimensional systolic array. The arithmetic circuit 1003 can also be a one-dimensional systolic array or other electronic circuit capable of performing mathematical operations such as multiplication and addition. In some implementations, the arithmetic circuit 1003 is a general-purpose matrix processor.

[0181] For example, assume there are input matrix A, weight matrix B, and output matrix C. The arithmetic circuit retrieves the corresponding data of matrix B from weight memory 1002 and caches it on each PE in the arithmetic circuit. The arithmetic circuit retrieves the data of matrix A from input memory 1001 and performs a matrix operation on matrix B. The partial or final matrix result is stored in accumulator 1008.

[0182] Unified memory 1006 is used to store input and output data. Weight data is directly transferred to weight memory 1002 through the Direct Memory Access Controller (DMAC) 1005. Input data is also transferred to unified memory 1006 through the DMAC.

[0183] BIU stands for Bus Interface Unit, i.e., bus interface unit 1010 , which is used for interaction between the AXI bus, DMAC, and instruction fetch buffer (IFB) 1009 .

[0184] The bus interface unit 1010 (BIU) is used for the instruction fetch memory 1009 to obtain instructions from the external memory, and is also used for the storage unit access controller 1005 to obtain the original data of the input matrix A or the weight matrix B from the external memory.

[0185] DMAC is mainly used to transfer input data in the external memory DDR to the unified memory 1006 or transfer weight data to the weight memory 1002 or transfer input data to the input memory 1001.

[0186] The vector calculation unit 1007 includes multiple operation processing units. When necessary, it further processes the output of the operation circuit 1003, such as vector multiplication, vector addition, exponential operation, logarithmic operation, size comparison, etc. It is mainly used for non-convolutional / fully connected layer network calculations in neural networks, such as batch normalization, pixel-level summation, and upsampling of feature planes.

[0187] In some implementations, the vector calculation unit 1007 can store the processed output vector to the unified memory 1006. For example, the vector calculation unit 1007 can apply a linear function or a nonlinear function to the output of the operation circuit 1003, such as linear interpolation of the feature plane extracted by the convolution layer, or accumulate a vector of values ​​to generate an activation value. In some implementations, the vector calculation unit 1007 generates a normalized value, a pixel-level summed value, or both. In some implementations, the processed output vector can be used as an activation input to the operation circuit 1003, for example, for use in subsequent layers in a neural network.

[0188] An instruction fetch buffer 1009 connected to the controller 1004 is used to store instructions used by the controller 1004;

[0189] Unified memory 1006, input memory 1001, weight memory 1002, and instruction fetch memory 1009 are all on-chip memories. External memories are private to the NPU hardware architecture.

[0190] The processor mentioned in any of the above places can be a general-purpose central processing unit, a microprocessor, an ASIC, or one or more integrated circuits for controlling the execution of the above program.

[0191] An embodiment of the present application also provides a computing device 1100. Please refer to Figure 11, which is a schematic diagram of the structure of a computing device 1100 provided in an embodiment of the present application. As shown in Figure 11, computing device 1100 includes: a bus 1102, a processor 1104, a memory 1106, and a communication interface 1108. The processor 1104, the memory 1106, and the communication interface 1108 communicate with each other via bus 1102. Computing device 1100 can be a server or a terminal device. It should be understood that this application does not limit the number of processors and memories in computing device 1100.

[0192] Bus 1102 may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, among others. Buses may be classified as address buses, data buses, control buses, and the like. For ease of illustration, FIG11 illustrates a single bus line, but this does not imply a single bus or type of bus. Bus 1102 may include a path for transmitting information between various components of computing device 1100 (e.g., memory 1106, processor 1104, and communication interface 1108).

[0193] The processor 1104 may include any one or more processors such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).

[0194] The memory 1106 may include volatile memory, such as random access memory (RAM). The processor 1104 may also include non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid state drive (SSD).

[0195] The memory 1106 stores executable program code, and the processor 1104 executes the executable program code to implement the functions of the aforementioned receiving module and processing module, thereby implementing the above-mentioned case generation method. In other words, the memory 1106 stores instructions for executing the case generation method.

[0196] The communication interface 1108 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the computing device 1100 and other devices or a communication network.

[0197] Embodiments of the present application also provide a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smartphone.

[0198] Please refer to Figure 12, which is a schematic diagram of the structure of a computing device cluster provided in an embodiment of the present application. As shown in Figure 12, the computing device cluster includes at least one computing device 1100. The memory 1106 of one or more computing devices 1100 in the computing device cluster may store the same instructions for executing the case generation method.

[0199] In some possible implementations, the memory 1106 of one or more computing devices 1100 in the computing device cluster may also store partial instructions for executing the case generation method. In other words, the combination of one or more computing devices 1100 can jointly execute the instructions for executing the case generation method.

[0200] It should be noted that the memory 1106 in different computing devices 1100 in the computing device cluster can store different instructions, each for executing a portion of the functions of the data processing apparatus. In other words, the instructions stored in the memory 1106 in different computing devices 1100 can implement the functions of one or more of the aforementioned receiving module and processing module.

[0201] In some possible implementations, one or more computing devices in a computing device cluster may be connected via a network. The network may be a wide area network or a local area network, etc. FIG13 shows a possible implementation. FIG13 is a schematic structural diagram of another computing device cluster provided in an embodiment of the present application. As shown in FIG13 , in a computing device cluster 1300, two computing devices 1100A and 1100B are connected via a network. Specifically, the network is connected via a communication interface in each computing device. In this type of possible implementation, the memory 1106 in the computing device 1100A stores instructions for executing the functions of the receiving module. At the same time, the memory 1106 in the computing device 1100B stores instructions for executing the functions of the processing module.

[0202] It should be understood that the functionality of the computing device 1100A shown in FIG13 may also be implemented by multiple computing devices 1100. Similarly, the functionality of the computing device 1100B may also be implemented by multiple computing devices 1100.

[0203] Please refer to Figure 14, which is a schematic diagram of the structure of a computer-readable storage medium provided in an embodiment of the present application. The present application also provides a computer-readable storage medium. In some embodiments, the workflow executed by the above-mentioned database system can be implemented as computer program instructions encoded in a machine-readable format on a computer-readable storage medium or on other non-transitory media or products.

[0204] 14 schematically illustrates a conceptual partial view of an example computer-readable storage medium including a computer program for executing a computer process on a computing device, arranged in accordance with at least some embodiments presented herein.

[0205] In one embodiment, the computer-readable storage medium 1400 is provided using a signal-bearing medium 1401. The signal-bearing medium 1401 may include one or more program instructions 1402, which when executed by one or more processors may provide the functions or part of the functions described above for the database system.

[0206] In some examples, signal bearing medium 1401 may include computer readable medium 1403 such as, but not limited to, a hard drive, compact disk (CD), digital video disk (DVD), digital tape, memory, ROM or RAM, and the like.

[0207] In some embodiments, the signal-bearing medium 1401 may include a computer-recordable medium 1404, such as, but not limited to, a memory, a read / write (R / W) CD, a R / W DVD, or the like. In some embodiments, the signal-bearing medium 1401 may include a communication medium 1405, such as, but not limited to, a digital and / or analog communication medium (e.g., a fiber optic cable, a waveguide, a wired communication link, a wireless communication link, or the like). Thus, for example, the signal-bearing medium 1401 may be communicated via a wireless form of the communication medium 1405 (e.g., a wireless communication medium conforming to the IEEE 802.X standard or other transmission protocol).

[0208] The one or more program instructions 1402 may be, for example, computer-executable instructions or logic-implemented instructions. In some examples, the computing device may be configured to provide various operations, functions, or actions in response to the program instructions 1402 communicated to the computing device via one or more of computer-readable media 1403, computer-recordable media 1404, and / or communication media 1405.

[0209] The present application also provides a computer program product containing instructions. This computer program product can be software or a program product containing instructions that can be run on a computing device or stored on any available medium. When the computer program product is run on at least one computing device, it causes the at least one computing device to execute the case generation method described in the above embodiment.

[0210] Through the description of the above embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software plus necessary general hardware, and of course can also be implemented by special hardware including application-specific integrated circuits, special CPUs, special memories, special components, etc. In general, all functions performed by computer programs can be easily implemented with corresponding hardware, and the specific hardware structures used to implement the same function can also be various, such as analog circuits, digital circuits or special circuits, etc. However, for the present application, software program implementation is a better implementation method in most cases. Based on such an understanding, the technical solution of the present application is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a readable storage medium, such as a computer's floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk or optical disk, etc., and includes a number of instructions to enable a computer device (which can be a personal computer, training equipment, or network equipment, etc.) to execute the methods of each embodiment of the present application.

[0211] In the above embodiments, all or part of the embodiments may be implemented by software, hardware, firmware, or any combination thereof. When implemented by software, all or part of the embodiments may be implemented in the form of a computer program product.

Claims

1. A case generation method, characterized in that: include: Obtaining a first prompt word, where the first prompt word includes a first case, and the first prompt word is used to indicate generating other cases with reference to the first case; Inputting the first prompt word into a first model to obtain a second case output by the first model, where the second case is generated by the first model with reference to the first case; Using the second model, evaluating the accuracy of the second case from different dimensions to obtain scores of the second case in different dimensions; When the scores of the second case in different dimensions meet the target conditions, the second case is output.

2. The method according to claim 1, characterized in that Both the first model and the second model are large language models.

3. The method according to claim 1 or 2, characterized in that The method further comprises: When the scores of the second case in different dimensions do not meet the target conditions, a modification target is set for the first model, and the second case is modified according to the modification target through the first model, so that the first model can output a modified case whose scores in the different dimensions meet the target conditions.

4. The method according to claim 3, characterized in that The method further comprises: Obtaining modification opinions for the second case, where the modification opinions are obtained by evaluating the accuracy of the second case from different dimensions using the second model; The modification target includes modifying the second case with reference to the modification opinion so that the scores of the modified case in the different dimensions meet the target conditions.

5. The method according to claim 4, characterized in that The step of setting a modification target for the first model and modifying the second case according to the modification target using the first model includes: Based on the modification suggestions corresponding to the case last output by the first model, modify the case last output by the first model to obtain a new case, wherein the second case is the case first output by the first model; Using the second model, evaluating the accuracy of the new case from the different dimensions to obtain scores of the new case in the different dimensions; The above steps are executed cyclically until the scores of the cases output by the first model in the different dimensions meet the target conditions, so as to obtain the cases output by the first model for the last time.

6. The method according to any one of claims 1 to 5, characterized in that The second model includes a plurality of different large language models. The second model is used to evaluate the accuracy of the second case from different dimensions, including: The second case is input into the multiple different large language models respectively to obtain scores output by the multiple different large language models respectively.

7. The method according to any one of claims 1 to 6, characterized in that The second model is used to evaluate the accuracy of the second case from different dimensions, including: Based on the second case, construct a plurality of prompt words, each of which includes the second case, and each of which is used to indicate evaluating the accuracy of the second case from different evaluation perspectives, each of which is related to the field to which the second case belongs; The multiple prompt words are input into the second model respectively to obtain scores of the second case from different evaluation perspectives.

8. The method according to any one of claims 1 to 7, characterized in that The target condition includes that the weighted average of the scores of the cases in different dimensions is greater than or equal to the target threshold.

9. The method according to any one of claims 1 to 8, characterized in that After outputting the second case, the method further includes: Obtaining a second prompt word, where the second prompt word is obtained by the user modifying the first prompt word based on the second case; Inputting the second prompt word into the first model to obtain a third case output by the first model; Using the second model, evaluating the accuracy of the third case from different dimensions to obtain scores of the third case in different dimensions; When the scores of the third case in different dimensions meet the target condition, the third case is output.

10. The method according to claim 9, characterized in that The second prompt word is obtained by adding guidance information to the first prompt word based on the second case, and the guidance information is used to indicate how to generate a new case with reference to the first case.

11. The method according to claim 9, characterized in that The second prompt word is obtained by modifying the first case in the first prompt word based on the second case.

12. A case generation device, characterized in that: include: An acquisition module, configured to acquire a first prompt word, wherein the first prompt word includes a first case, and the first prompt word is used to indicate generating other cases with reference to the first case; a processing module, configured to input the first prompt word into a first model to obtain a second case output by the first model, where the second case is generated by the first model with reference to the first case; The processing module is further configured to evaluate the accuracy of the second case from different dimensions using a second model to obtain scores of the second case in different dimensions; The processing module is further configured to output the second case when the scores of the second case in different dimensions meet target conditions.

13. The device according to claim 12, characterized in that Both the first model and the second model are large language models.

14. The device according to claim 12 or 13, characterized in that The processing module is further configured to: When the scores of the second case in different dimensions do not meet the target conditions, a modification target is set for the first model, and the second case is modified according to the modification target through the first model, so that the first model can output a modified case whose scores in the different dimensions meet the target conditions.

15. The device according to claim 14, characterized in that The acquisition module is further configured to acquire modification opinions for the second case, where the modification opinions are obtained by evaluating the accuracy of the second case from different dimensions using the second model; The modification target includes modifying the second case with reference to the modification opinion so that the scores of the modified case in the different dimensions meet the target conditions.

16. The device according to claim 15, characterized in that The processing module is further configured to: Based on the modification suggestions corresponding to the case last output by the first model, modify the case last output by the first model to obtain a new case, wherein the second case is the case first output by the first model; Using the second model, evaluating the accuracy of the new case from the different dimensions to obtain scores of the new case in the different dimensions; The above steps are executed cyclically until the scores of the cases output by the first model in the different dimensions meet the target conditions, so as to obtain the cases output by the first model for the last time.

17. The device according to any one of claims 12 to 16, characterized in that: The second model includes a plurality of different large language models, and the processing module is further configured to: The second case is input into the multiple different large language models respectively to obtain scores output by the multiple different large language models respectively.

18. The device according to any one of claims 12 to 17, characterized in that: The processing module is further configured to: Based on the second case, construct a plurality of prompt words, each of which includes the second case, and each of which is used to indicate evaluating the accuracy of the second case from different evaluation perspectives, each of which is related to the field to which the second case belongs; The multiple prompt words are input into the second model respectively to obtain scores of the second case from different evaluation perspectives.

19. The device according to any one of claims 12 to 18, characterized in that The target condition includes that the weighted average of the scores of the cases in different dimensions is greater than or equal to the target threshold.

20. The device according to any one of claims 12 to 19, characterized in that After outputting the second case, the acquisition module is further configured to acquire a second prompt word, where the second prompt word is obtained by the user modifying the first prompt word based on the second case; The processing module is further configured to: Inputting the second prompt word into the first model to obtain a third case output by the first model; Using the second model, evaluating the accuracy of the third case from different dimensions to obtain scores of the third case in different dimensions; When the scores of the third case in different dimensions meet the target condition, the third case is output.

21. The device according to claim 20, characterized in that The second prompt word is obtained by adding guidance information to the first prompt word based on the second case, and the guidance information is used to indicate how to generate a new case with reference to the first case.

22. The device according to claim 20, characterized in that The second prompt word is obtained by modifying the first case in the first prompt word based on the second case.

23. A computing device cluster, characterized in that: comprising at least one computing device, each computing device including a processor and a memory; The processor of the at least one computing device is configured to execute instructions stored in a memory of the at least one computing device, so that the computing device cluster performs the operating steps of the method according to any one of claims 1 to 11.

24. A computer program product comprising instructions, characterized in that When the instructions are executed by a computing device cluster, the computing device cluster is caused to perform the operation steps of the method according to any one of claims 1 to 11.

25. A computer-readable storage medium, characterized in that The method comprises computer program instructions. When the computer program instructions are executed by a computing device cluster, the computing device cluster performs the operation steps of the method according to any one of claims 1 to 11.

Citation Information

Patent Citations

  • Case generation method and related device

    CN120654688A

  • Method and system for interactive automobile driving case teaching

    CN101561972A

  • Enterprise basic law intelligent consulting terminal, system and method

    CN111125319A

  • Operation planning scheme generation method, system and device and medium

    CN117076655A

  • Automatic optimization method and device for cue word, equipment and storage medium

    CN117520507A

Cited By

  • Automatic data model extraction method and system based on large model

    CN120892707A

  • Prompt word optimization method, electronic equipment, storage medium and program product

    CN121480506A

  • Prompt word optimization method, device and equipment based on AI, RPA, LLM and AI Agent

    CN121525881A

  • Content generation method and device based on man-machine mixed feedback, equipment and medium

    CN121638305A

  • Pumping well production condition analysis method based on cue word self-adaptive generation

    CN121882220A