Mathematical problem purpose generation method and device, equipment and storage medium
By using large language models to strategically modify the benchmark mathematical problems, the problem of lack of high-quality mathematical problems in the existing technology is solved, and an effective assessment of the mathematical reasoning ability and robustness of natural language systems is achieved.
Patent Information
- Application Number
- CN202311584948.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-24
- Publication Date
- 2025-05-27
AI Technical Summary
The lack of high-quality and large-scale confrontational mathematical problems in the prior art makes it difficult to effectively evaluate the mathematical reasoning ability of natural language systems.
By obtaining benchmark mathematical problems, determining attack strategies for natural language systems, and using large language models to modify benchmark mathematical problems based on these strategies, generating adversarial mathematical problems to evaluate the mathematical reasoning ability of natural language systems.
The generated adversarial mathematical problems can effectively evaluate the mathematical reasoning ability and robustness of natural language systems, and provide a method to evaluate the anti-interference ability of natural language systems.
Smart Images

Figure CN120046744A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence, and particularly to a method, apparatus, device, and storage medium for generating mathematical problems. Background Art
[0002] With the development of natural language systems, their performance in mathematical reasoning ability has also been significantly improved. Evaluating whether a natural language system truly masters mathematical reasoning ability is a hot research field currently.
[0003] In related technologies, the mathematical reasoning ability of natural language systems is evaluated by combining benchmark mathematical problems and adversarial mathematical problems. Adversarial mathematical problems are generated by modifying the input word level of benchmark mathematical problems, and the anti-interference ability in natural language systems is tested using adversarial mathematical problems, so as to evaluate whether the natural language system masters mathematical reasoning ability.
[0004] However, there is a lack of high-quality and large-scale adversarial mathematical problems in related technologies. Therefore, how to generate a large number of reliable adversarial mathematical problems to evaluate the mathematical reasoning ability of natural language systems is an urgent problem to be solved currently. Summary of the Invention
[0005] This application provides a method, apparatus, device, and storage medium for generating mathematical problems, and the technical solutions are as follows:
[0006] According to one aspect of this application, a method for generating mathematical problems is provided, and the method includes:
[0007] Obtain a benchmark mathematical problem, where the benchmark mathematical problem is used to evaluate the mathematical reasoning ability of a natural language system;
[0008] Determine an attack strategy for the natural language system;
[0009] Modify the benchmark mathematical problem based on the attack strategy through a large language model to obtain an adversarial mathematical problem, where the adversarial mathematical problem is used to evaluate the mathematical reasoning ability of the natural language system using the attack strategy.
[0010] According to another aspect of this application, a device for generating mathematical problems is provided, and the device includes:
[0011] An obtaining module, configured to obtain a benchmark mathematical problem, where the benchmark mathematical problem is used to evaluate the mathematical reasoning ability of a natural language system;
[0012] A determining module, configured to determine an attack strategy for the natural language system;
[0013] A processing module, configured to modify the benchmark math problem based on the attack strategy through a large language model to obtain an adversarial math problem, where the adversarial math problem is used to evaluate the math reasoning ability of the natural language system using the attack strategy.
[0014] According to another aspect of the present application, there is provided a computer device, including a processor and a memory, where a computer program is stored in the memory, and the computer program is loaded and executed by the processor to implement the method for generating math problems as described above.
[0015] According to another aspect of the present application, there is provided a computer-readable storage medium, in which executable instructions are stored, and the executable instructions are loaded and executed by a processor to implement the method for generating math problems as described above.
[0016] According to another aspect of the present application, there is provided a computer program product, including computer instructions, where the computer instructions are stored in a computer-readable storage medium, and a processor reads and executes the computer instructions from the computer-readable storage medium to implement the method for generating math problems as described above.
[0017] The beneficial effects brought by the technical solution provided by the present application at least include:
[0018] The embodiment of the present application provides a method for generating math problems. First, a benchmark math problem is obtained, which is used to evaluate the math reasoning ability of a natural language system; an attack strategy for the natural language system is determined; the benchmark math problem is modified based on the attack strategy through a large language model to obtain an adversarial math problem, and the adversarial math problem is used to evaluate the math reasoning ability of the natural language system using the attack strategy. In the present application, by selecting an attack strategy for the benchmark math problem, the benchmark math problem and the selected attack strategy are jointly input into the large language model, and the large language model modifies the benchmark math problem according to the selected attack strategy, and an adversarial math problem can be obtained. The generated adversarial math problem can be used to evaluate the math reasoning ability and robustness of the natural language system. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0020] Figure 1 Shows a schematic diagram of the architecture of a computer system provided by an exemplary embodiment of the present application;
[0021] Figure 2 Shows a schematic diagram of a method for generating math problems provided by an exemplary embodiment of the present application;
[0022] Figure 3 Shows a flowchart of a method for generating math problems provided by an exemplary embodiment of the present application;
[0023] Figure 4 Shows a flowchart of a method for generating math problems provided by an exemplary embodiment of the present application;
[0024] Figure 5 Shows a flowchart of a method for generating math problems provided by an exemplary embodiment of the present application;
[0025] Figure 6 Shows a schematic diagram of a method for generating math problems provided by an exemplary embodiment of the present application;
[0026] Figure 7 Shows a flowchart of a method for generating math problems provided by an exemplary embodiment of the present application;
[0027] Figure 8 Shows a flowchart of a method for generating math problems provided by an exemplary embodiment of the present application;
[0028] Figure 9 Shows a flowchart of a method for generating math problems provided by an exemplary embodiment of the present application;
[0029] Figure 10 Shows a flowchart of a method for generating math problems provided by an exemplary embodiment of the present application;
[0030] Figure 11 Shows a schematic diagram of a method for generating math problems provided by an exemplary embodiment of the present application;
[0031] Figure 12 Shows a block diagram of the structure of a device for generating math problems provided by an exemplary embodiment of the present application;
[0032] Figure 13 Shows a schematic diagram of the structure of a computer device provided by an exemplary embodiment of the present application. Detailed implementation
[0033] To make the objectives, technical solutions, and advantages of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the accompanying drawings.
[0034] Exemplary embodiments will be described in detail herein, and examples thereof are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.
[0035] The terms used in the present application are for the purpose of describing specific embodiments only and are not intended to limit the present application. The singular forms "a", "the", and "said" used in the present application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.
[0036] It should be noted that before and during the process of collecting relevant data of the user, the present application can display a prompt interface, a pop-up window or a text box prompt message, which is used to prompt the user that their relevant data is being collected at present, so that the present application only starts to execute the relevant steps of obtaining the user's relevant data after obtaining the confirmation operation issued by the user for the prompt interface or the pop-up window. Otherwise (that is, when the confirmation operation issued by the user for the prompt interface or the pop-up window is not obtained), the relevant steps of obtaining the user's relevant data are ended, that is, the relevant data of the user is not obtained. In other words, all user data collected by the present application is collected with the consent and authorization of the user, and the collection, use and processing of the relevant user data need to comply with the relevant laws, regulations and standards of the relevant countries and regions.
[0037] First, the relevant terms involved in the present application are introduced:
[0038] 1) Large Language Model (LLM): It is a natural language processing model based on deep learning. It can learn the grammar and semantics of natural language, so as to generate human-readable text, and has excellent dialogue understanding and logical reasoning capabilities. LLM can be used as a decision-making component in complex scenarios that require human common sense understanding, and can handle various natural language processing tasks, such as dialogue systems, intelligent customer service, natural language generation, text classification, text summarization, machine translation, speech recognition, etc.
[0039] 2) Natural Language Processing (NLP): It is an important direction in the fields of computer science and artificial intelligence. It studies various theories and methods that can achieve effective communication between humans and computers using natural language. Natural language processing involves natural language, that is, the language people use in daily life, and is closely related to linguistic research; moreover, natural language processing also involves fields such as computer science and mathematics.
[0040] 3) Natural language system: It refers to a computer system that can understand and generate human natural language. A natural language system can process and interpret the meaning, grammar, and structure of human language, and can answer questions, generate text, etc. A natural language system can be a rule-based system, a statistical model, or a deep learning model, aiming to simulate human language understanding and generation capabilities.
[0041] 4) Adversarial math problems: Adversarial math problems are generated by making small, targeted perturbations to benchmark math problems. Exemplarily, an adversarial math problem can be a modified form of a benchmark math problem generated by a large language model based on an attack strategy. For example, if the benchmark math problem is "There is a rectangle with a length of 5 and a width of 3, find its area", a numerical modification to the benchmark math problem gives the adversarial math problem "There is a rectangle with a length of 6 and a width of 4, find its area".
[0042] 5) Robustness: It refers to the stability and robustness of a natural language system when facing different types of interference, attacks, or changes. In the field of artificial intelligence, robustness refers to the resistance of a natural language system to changes, perturbations, or attacks in input data. Exemplarily, when evaluating the mathematical reasoning ability of a natural language system, robustness means that the natural language system can maintain good performance and output correct results when facing adversarial math problem inputs.
[0043] 6) Adversarial training: Adversarial training is a method for defending against adversarial math problem attacks. Training with both adversarial math problems and benchmark math problems is an effective regularization method, aiming to improve the robustness and adversarial attack ability of a natural language system. By introducing adversarial math problems to train a natural language system, the natural language system can maintain stable performance when facing interference, perturbations, or attacks.
[0044] 7) Artificial Intelligence (AI) is a theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable machines to have the functions of perception, reasoning, and decision-making.
[0045] Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, pre-trained model technology, operation / interaction systems, mechatronics, etc. Among them, the pre-trained model, also known as the large model or the foundation model, can be widely applied to downstream tasks in various directions of artificial intelligence after fine-tuning. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0046] Figure 1 The structural block diagram of a computer system 100 provided by an exemplary embodiment of the present application is shown. The computer system 100 includes: a terminal device 110 and a server 120, and a client of a human-computer interaction program 130 is installed in the terminal device 110.
[0047] The terminal device 110 can be an electronic device such as a mobile phone, a tablet computer, an in-vehicle terminal (carputer), a wearable device, a PC (Personal Computer), an unmanned reservation terminal, etc. A client of a target application program can be installed and run in the terminal device 110, and the target application program can be the human-computer interaction program 130. Optionally, the human-computer interaction program 130 can be an application program for training and / or using a large language model, or other application programs that provide functions for training and / or using a large language model. The present application does not make any limitations in this regard. In addition, the form of the human-computer interaction program 130 is not limited in the present application, including but not limited to an App (Application) installed in the terminal device 110, a small program, etc., and can also be in the form of a web page.
[0048] Exemplarily, the method for generating math problems provided by this application can be executed by the client of the human-computer interaction program 130 in the terminal device 110. The human-computer interaction program 130 is installed on the terminal device 110. When it is necessary to use the method for generating math problems provided by the embodiments of this application, the client of the human-computer interaction program 130 can run the human-computer interaction program to execute the method for generating math problems.
[0049] The server 120 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud computing services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms. The server 120 can be the background server of the above-mentioned human-computer interaction program 130, and is used to provide background services for the client of the human-computer interaction program 130.
[0050] Among them, cloud technology refers to a hosting technology that unifies a series of resources such as hardware, software, and networks within a wide area network or local area network to achieve data computing, storage, processing, and sharing. Cloud technology is the general term for network technology, information technology, integration technology, management platform technology, application technology, etc. based on the cloud computing business model, and can form a resource pool, which can be used on demand, flexibly and conveniently. Cloud computing technology will become an important support. The background services of the technical network system require a large amount of computing and storage resources, such as video websites, image-based websites, and more portal websites. With the highly developed application of the Internet industry, in the future, each item may have its own identification mark and needs to be transmitted to the background system for logical processing. Data at different levels will be processed separately, and various industry data requires a powerful system back-end support, which can only be achieved through cloud computing.
[0051] In some embodiments, the above-mentioned server 120 can also be implemented as a node in a blockchain system. Blockchain is a new application mode of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, and encryption algorithms. Blockchain, essentially a decentralized database, is a string of data blocks generated by using cryptographic methods. Each data block contains information about a batch of network transactions, which is used to verify the validity (anti-counterfeiting) of the information and generate the next block. Blockchain can include the blockchain underlying platform, the platform product service layer, and the application service layer.
[0052] Server 120 is used to provide background services for the client of the human-computer interaction program 130 of the terminal device 110. Optionally, Server 120 undertakes the main computing work, and the terminal device 110 undertakes the secondary computing work; or, Server 120 undertakes the secondary computing work, and the terminal device 110 undertakes the main computing work; or, a distributed computing architecture is adopted between Server 120 and the terminal device 110 for collaborative computing.
[0053] In an alternative embodiment, the terminal device 110 and the server 120 can be connected to each other through a wired or wireless network.
[0054] Exemplarily, Figure 2 The schematic diagram of the method for generating a math problem provided by an exemplary embodiment of the present application is shown. Taking the method being executed by the server 120 as an example for illustration.
[0055] To evaluate the mathematical reasoning ability in a natural language system, based on the benchmark math problem 10, one or more attack strategies 20 are selected for the benchmark math problem 10, the benchmark math problem 10 combined with the attack strategies 20 is input into the large language model 30, the large language model 30 generates candidate adversarial math problems 40 based on the attack strategies 20, the quality of the candidate adversarial math problems 40 is evaluated from multiple evaluation dimensions 50, through the large language model 30, the candidate adversarial math problems 40 that meet the evaluation conditions are output as adversarial math problems 60, and the mathematical logical reasoning ability in the natural language system is tested through the adversarial math problems 60.
[0056] In the embodiments of the present application, taking the benchmark problem 10 as a geometric problem and the attack strategy 20 selecting numerical modification as an example for illustration.
[0057] A brief description of the steps of the method for generating a math problem in the embodiments of the present application is as follows:
[0058] (1) Determine the first input prompt according to the benchmark math problem 10 and the attack strategy 20;
[0059] Among them, the benchmark math problem 10 is used to evaluate the mathematical reasoning ability of the natural language system. The attack strategy 20 is used to modify or vary the benchmark math problem 10.
[0060] Optionally, the benchmark math problem 10 is usually a math problem, has a known answer, and can be described in natural language. For example, the benchmark math problem 10 is a geometric problem, such as "There is a rectangle with a length of 5 and a width of 3, find its area".
[0061] The attack strategy 20 includes at least one of numerically modifying the benchmark math problem 10, expanding the problem of the benchmark math problem 10, performing an identical transformation of the formula for the benchmark math problem 10, semantically rewriting the benchmark math problem 10, and injecting interference into the benchmark math problem 10.
[0062] Strategy 1 (numerical modification): The benchmark math problem 10 can be changed to "There is a rectangle with a length of 6 and a width of 4. Find its area." Strategy 2 (problem expansion): The benchmark math problem 10 can be changed to "There is a rectangle with a length of 6 and a width of 4. What is the result of adding 2 to its area?" Strategy 3 (identical formula transformation): The benchmark math problem 10 can be changed to "There is a rectangle with an area of 15 and a width of 3. Find its length." Strategy 4 (semantic rewriting): The benchmark math problem 10 can be changed to "A rectangle has side lengths of 5 and 3. Find its area." Strategy 5 (interference injection): The benchmark math problem 10 can be changed to "In a rectangular garden, the length of the rectangle is 5 and the width is 3. Find the area of the garden."
[0063] The first input prompt refers to the text fragment provided as the starting input to the large language model when generating adversarial math problems. The first input prompt can be determined based on the benchmark math problem 10 and the attack strategy 20, with the aim of guiding the large language model 30 to generate candidate adversarial math problems 40.
[0064] (2) Input the first input prompt into the large language model 30, and the large language model 30 generates candidate adversarial math problems 40 according to the first input prompt;
[0065] In an optional example, select one or more attack strategies 20 for the benchmark math problem 10, and input the selected attack strategies 20 and the content of the benchmark math problem 10 as the first input prompt into the large language model 30. The large language model 30 can generate candidate adversarial math problems 40 according to the first input prompt. Optionally, the large language model 30 numerically modifies the benchmark math problem 10 based on the attack strategy 20, and the resulting candidate adversarial math problem 40 is "There is a rectangle with a length of 6 and a width of 4. Find its area."
[0066] (3) Evaluate the candidate adversarial math problems 40 based on the evaluation dimension 50;
[0067] After the large language model 30 generates the candidate adversarial math problems 40, it will evaluate the candidate adversarial math problems 40. The large language model 30 includes multiple evaluation dimensions 50.
[0068] The evaluation dimension 50 is a criterion or metric used to evaluate and measure the characteristics or performance of the candidate adversarial math problem 40 in different aspects. The evaluation dimension 50 includes at least one of: difficulty consistency, problem-solving effectiveness, expected attack effect, and attack strategy diversity. Among them, difficulty consistency is used to evaluate whether the mathematical ability required to solve the candidate adversarial math problem 40 exceeds the mathematical ability required to solve the benchmark math problem 10; problem-solving effectiveness is used to evaluate whether the candidate adversarial math problem 40 has a unique and effective solution; the expected attack effect is used to evaluate whether the candidate adversarial math problem 40 is more likely to cause the large language model 30 to make mistakes than the benchmark math problem 10; and attack strategy diversity is used to evaluate whether the candidate adversarial math problem 40 contains the attack strategy 20.
[0069] The purpose of evaluating the candidate adversarial math problem 40 based on the evaluation dimension 50 is to guide the large language model 30 to generate the score of the candidate adversarial math problem 40.
[0070] In a possible implementation, the large language model 30 scores the candidate adversarial math problem 40 based on the evaluation dimension 50, and filters out the adversarial math problems 60 that meet the preset requirements according to the scored values.
[0071] Exemplarily, a preset scoring mechanism is used to score the candidate adversarial math problem 40, and a scoring range is given, for example, from 1 to 5 points. When the candidate adversarial math problem 40 does not meet any of the dimensions in the evaluation dimension 50, the score of the candidate adversarial math problem 40 is 1 point; when the candidate adversarial math problem 40 meets one dimension in the candidate adversarial sample evaluation dimension 50, the score of the candidate adversarial math problem 40 is 2 points; when the candidate adversarial math problem 40 meets two dimensions in the evaluation dimension 50, the score of the candidate adversarial math problem 40 is 3 points; when the candidate adversarial math problem 40 meets three dimensions in the evaluation dimension 50, the score of the candidate adversarial math problem 40 is 4 points; when the candidate adversarial math problem 40 meets all dimensions in the evaluation dimension 50, the score of the candidate adversarial math problem 40 is 5 points. The large language model 30 scores the candidate adversarial math problem 40, and only the candidate adversarial math problems 40 with a score of 4 points or 5 points are considered qualified. This preset scoring mechanism can help filter out the candidate adversarial math problems 40 that meet the requirements of multiple evaluation dimensions.
[0072] (4) When the score of the candidate adversarial math problem 40 meets the score threshold, output it as an adversarial math problem;
[0073] Among them, the score threshold refers to a criterion or boundary for measuring the quality of the candidate adversarial math problem 40. The score threshold is a preset value or condition used to determine whether the candidate adversarial math problem 40 meets the required quality level. Exemplarily, in the above scoring mechanism, the score threshold refers to the case where the score of the candidate adversarial math problem 40 is 4 points.
[0074] Optionally, when the quality of the candidate adversarial math problem 40 reaches the score threshold, that is, when the candidate adversarial math problem 40 satisfies at least three evaluation dimensions, the candidate adversarial math problem 40 is output as the adversarial math problem 60.
[0075] (5) When the score of the candidate adversarial math problem 40 does not meet the score threshold, reselect the attack strategy 20 to generate a new candidate adversarial math problem 40;
[0076] Optionally, when the quality of the candidate adversarial math problem 40 does not reach the score threshold, that is, when the candidate adversarial math problem 40 satisfies at most two evaluation dimensions, reselect the attack strategy 20 and generate a new candidate adversarial math problem 40.
[0077] (6) Input the adversarial math problem 60 and the benchmark math problem 10 into the natural language system for adversarial training.
[0078] Optionally, in order to evaluate the mathematical reasoning ability of the natural language system, the adversarial math problem 60 and the benchmark math problem 10 can be used for adversarial training. Input the adversarial math problem 60 and the benchmark math problem 10 into the natural language system for adversarial training together. By introducing the adversarial math problem 60 to train the natural language system, the natural language system can maintain stable performance when facing interference, perturbation, or attack. The adversarial math problem 60 has minor and targeted changes compared to the benchmark math problem 10, aiming to make the natural language system produce incorrect outputs or misjudgments. By conducting adversarial training with the adversarial math problem 60 and the benchmark math problem 10 together, the robustness and attack resistance ability of the natural language system can be improved.
[0079] Figure 3 Shows a flowchart of a method for generating math problems provided by an exemplary embodiment of the present application. Taking the method applied to Figure 1 the shown terminal device 110 as an example, the method includes the following steps 210, step 220, and step 230:
[0080] Step 210: Obtain a benchmark math problem;
[0081] Among them, the benchmark math problem is used to evaluate the mathematical reasoning ability of the natural language system.
[0082] Optionally, the benchmark math question is a question used to evaluate the mathematical reasoning ability of the natural language system. The benchmark math question is usually a math question with a known answer and can be described in natural language. Exemplarily, the benchmark math question can be a math question in a math question bank or other resources. The benchmark math question can be a question type in at least one of the fields of arithmetic operations, algebraic equations, and geometry problems, and can cover different difficulty levels.
[0083] For example, the benchmark math question is a geometry problem, such as “a rectangle has a length of 5 and a width of 3, find its area”.
[0084] In some embodiments, the benchmark math problem is used to evaluate the mathematical reasoning ability of the natural language system when solving the benchmark math problem. Mathematical reasoning ability refers to the ability to perform logical analysis and reasoning on math problems. Optionally, mathematical reasoning ability includes understanding mathematical concepts and using mathematical principles for reasoning.
[0085] Step 220: Determine an attack strategy against the natural language system;
[0086] Attack strategies refer to methods used to generate adversarial math problems. Attack strategies are used to modify or change benchmark math problems. The purpose of attack strategies is to increase the difficulty or diversity of benchmark math problems in order to test the mathematical reasoning ability and robustness of natural language systems.
[0087] In an optional example, the attack strategy includes at least one of the following strategies:
[0088] · Make numerical modifications to benchmark math problems;
[0089] Expand the basic math questions;
[0090] ·Perform formula identity transformation on benchmark math problems;
[0091] ·Semantic rewriting of benchmark math problems;
[0092] ·Inject interference into benchmark math problems.
[0093] Exemplarily, the benchmark math question is a geometry problem, such as “a rectangle has a length of 5 and a width of 3, find its area”.
[0094] Strategy 1 (Numerical modification): Numerical modification of the basic math problem means changing the numerical value in the basic math problem while keeping the original content of the basic math problem as much as possible; for example, the basic math problem can be changed to "There is a rectangle with a length of 6 and a width of 4. Find its area."
[0095] Strategy 2 (Problem Expansion): Expanding the problem of the benchmark math problem is to modify the constraints of the benchmark math problem by incorporating additional arithmetic operations into the benchmark math problem. Optionally, the allowed arithmetic operations include addition, subtraction, multiplication, and division; in this attack strategy, the constraints of the benchmark math problem are modified. For example, the benchmark math problem can be changed to "There is a rectangle with a length of 6 and a width of 4. What is the result of adding 2 to its area?".
[0096] Strategy 3 (Formula Identity Transformation): Performing a formula identity transformation on the benchmark math problem, without adding additional constraints, modifies the description of the benchmark math problem. Optionally, convert known conditions to unknown conditions. For example, the benchmark math problem can be changed to "There is a rectangle with an area of 15 and a width of 3. What is its length?".
[0097] Strategy 4 (Semantic Rewriting): Semantically rewriting the benchmark math problem means changing the expression form of the benchmark math problem to generate a new math problem without changing the solution of the benchmark math problem. For example, the benchmark math problem can be changed to "A rectangle has side lengths of 5 and 3. What is its area?".
[0098] Strategy 5 (Interference Injection): Injecting interference into the benchmark math problem is to introduce conditions unrelated to solving the problem into the benchmark math problem without changing the solution of the benchmark math problem. For example, the benchmark math problem can be changed to "In a rectangular garden, the length of the rectangle is 5 and the width is 3. What is the area of the garden?".
[0099] Optionally, based on the benchmark math problem, an attack strategy can be selected, and one or more attack strategies can be chosen.
[0100] Step 230: Modify the benchmark math problem based on the attack strategy through a large language model to obtain an adversarial math problem.
[0101] Among them, the large language model is an important part or core engine in the natural language system. The natural language system can utilize the capabilities of the large language model to implement natural language processing tasks, such as text generation, dialogue systems, question - answering systems, etc.
[0102] Optionally, the natural language system can be used in combination with the large language model to evaluate its performance in mathematical reasoning ability.
[0103] In some embodiments, to evaluate the mathematical reasoning ability of the natural language system, that is, on the premise of keeping the difficulty of the benchmark math problem unchanged, based on the determined attack strategy, use the large language model to modify the benchmark math problem to generate an adversarial math problem, and use the adversarial math problem to test the mathematical reasoning ability of the natural language system.
[0104] Among them, adversarial math problems are used to evaluate the math reasoning ability of natural language systems using attack strategies. An adversarial math problem refers to a modified form of a benchmark math problem generated by a large language model based on an attack strategy. Optionally, an adversarial math problem may include modifications, additions, or replacements to the benchmark math problem.
[0105] In some embodiments, the adversarial math problems will have different characteristics from the benchmark math problems, and the adversarial math problems may lead the natural language system to give wrong answers or cause confusion. Adversarial problems can be used to evaluate the performance of natural language systems in terms of math reasoning ability.
[0106] In summary, the method provided in this embodiment first obtains a benchmark math problem, which is used to evaluate the math reasoning ability of a natural language system; determines an attack strategy for the natural language system; modifies the benchmark math problem based on the attack strategy through a large language model to obtain an adversarial math problem, which is used to evaluate the math reasoning ability of the natural language system using the attack strategy. In this application, by selecting an attack strategy for the benchmark math problem, the benchmark math problem and the selected attack strategy are jointly input into the large language model, and the large language model modifies the benchmark math problem according to the selected attack strategy to obtain an adversarial math problem. The generated adversarial math problem can be used to evaluate the math reasoning ability and robustness of the natural language system. At the same time, this method of automatically obtaining adversarial math problems through a large language model can effectively avoid data leakage.
[0107] Based on Figure 3 On this basis, Figure 4 shows a flowchart of a method for generating math problems provided by an exemplary embodiment of this application. Step 230 can be replaced by the following steps 231, 232, and 233:
[0108] Step 231: Based on the benchmark math problem and the attack strategy, determine a first input prompt;
[0109] The first input prompt refers to a text segment provided as the starting input to the large language model when generating an adversarial math problem. The first input prompt can be determined based on the benchmark math problem and the attack strategy, with the aim of guiding the large language model to generate relevant adversarial math problems.
[0110] In a possible implementation manner, after determining the attack strategy, the selected attack strategy is integrated with the benchmark math problem and encapsulated as the first input prompt of the large language model. By encapsulating the attack strategy and the benchmark math problem in the first input prompt, the large language model can generate candidate adversarial math problems according to the first input prompt.
[0111] Optionally, the first input prompt can be a keyword, phrase, or problem description related to the reference math problem, which is used to guide the large language model to generate adversarial math problems.
[0112] In some embodiments, the steps to determine the first input prompt are as follows: First, an attack strategy needs to be determined. Before generating adversarial math problems, it is necessary to clarify the attack strategy adopted. The attack strategy refers to the methods and rules used when modifying the reference math problem. Second, a reference math problem is selected as a sample, and according to the requirements of the selected attack strategy, the reference math problem is analyzed. Then, the selected attack strategy is combined with the reference math problem to determine the first input prompt to be input into the large language model.
[0113] Step 232: Input the first input prompt into the large language model;
[0114] Optionally, the large language model accepts a text fragment as input and generates a corresponding text fragment as output. The first input prompt is input into the large language model as a text fragment to guide the generation process of the large language model. The large language model will reason and generate according to the content of the first input prompt and generate the corresponding output result.
[0115] Step 233: Obtain the adversarial math problem output by the large language model according to the first input prompt.
[0116] An adversarial math problem refers to a modified form of the reference math problem generated by the large language model based on the attack strategy. The adversarial math problem is used to evaluate the math reasoning ability of the natural language system.
[0117] In some embodiments, by inputting the first input prompt into the large language model, the large language model will reason and generate according to the first input prompt. The output result of the large language model is a text representation of an adversarial math problem.
[0118] In summary, the method provided in this embodiment determines the first input prompt input into the large language model according to the reference math problem and the attack strategy, and uses the determined first input prompt as input to generate adversarial math problems. By using the large language model to generate adversarial math problems according to the first input prompt, the evaluation of the math reasoning ability of the natural language system can be realized. At the same time, the generated adversarial math problems may have higher difficulty and complexity, which helps to comprehensively evaluate the performance of the natural language system when dealing with math problems.
[0119] In some embodiments, Figure 5The flowchart of the method for generating math problems provided by an exemplary embodiment of the present application is shown. The embodiments of the present application are used to process based on a benchmark math problem and an attack strategy to determine a first input prompt.
[0120] · Determine the first input prompt
[0121] Optionally, by setting a first preset text format, input the content of the first input prompt into the first preset text format to determine the first input prompt. The above step 231 can be replaced by the following steps 231-1, 231-2, and 231-3:
[0122] Step 231-1: Set the first preset text format;
[0123] Wherein, the first preset text format is used to describe the first input prompt.
[0124] In some embodiments, when determining the first input prompt, by setting the first preset text format, the structure and components of the first input prompt are described by the first preset text format. The purpose of the first preset text format is to guide the large language model to understand and process the first input prompt. The large language model correctly interprets the first input prompt through the first preset text format and generates corresponding results according to the content of the first preset text format.
[0125] Optionally, setting the first preset text format can divide the first input prompt into different parts, such as task description, attack strategy description, and benchmark math problem. When determining the first input prompt, the corresponding content can be filled in different positions according to the first preset text format.
[0126] Optionally, the first preset text format can be a predefined text format, which can be designed according to specific needs and tasks, and the first preset text format can include different components.
[0127] Exemplarily, the first preset text format can be:
[0128] "Task description:
[0129]
First placeholder
[0130]
Second placeholder
First placeholder
Second placeholder
[0131] Step 231-2: Input the task description into the first preset text format as the first part of the first input prompt;
[0132] The task description is used to instruct the large language model to modify the benchmark problem based on the attack strategy.
[0133] Optionally, the task description refers to the description of a specific task or requirement to convey the content, requirements, and goals of the task to be completed to the large language model. Exemplarily, the requirements and goals of the task include the expectations and constraints of the results, as well as other relevant guiding requirements; the content of the task details the specific task or requirement to be completed. The purpose of the task description is to clearly convey the requirements and goals of the task to the large language model to guide the large language model to generate output results that meet the requirements.
[0134] In some embodiments, the task description is used as the first part of the first input prompt, and the task description is input into the first input prompt in accordance with the first preset text format. The purpose of the task description is to instruct the large language model to modify the benchmark problem according to the given attack strategy.
[0135] Step 231-3: Input the attack strategy description and the benchmark math problem into the first preset text format as the second part of the first input prompt.
[0136] Optionally, the attack strategy description and the benchmark math problem are filled in according to the requirements of the first preset text format as the second part of the first input prompt. The attack strategy description is used to describe the selected attack strategy for the benchmark math problem, that is, how to modify the benchmark math problem.
[0137] In some embodiments, by using the attack strategy description and the benchmark math problem as the second part of the first input prompt, the large language model can be guided to modify according to the selected attack strategy when generating candidate adversarial math problems. The attack strategy description provides a detailed description of the selected attack strategy to guide the large language model on how to modify the benchmark math problem to generate candidate adversarial math problems.
[0138] Optionally, the first input prompt includes the task description, the attack strategy description, and the benchmark math problem.
[0139] Exemplarily, as Figure 6 shown, the first input prompt 601 includes a first part 602 and a second part 603. Among them, the first part 602 is the task description, and the second part 603 includes the attack strategy description and the benchmark math problem. The task description, the attack strategy description, and the benchmark math problem are filled in according to the requirements of the first preset text format as the input of the first input prompt 601.
[0140] For example, the first input prompt format is as follows:
[0141] "Task description: I expect you to play the role of a math problem reconstructor. Your task is to modify the given benchmark math problem according to the attack strategy specified below, so that the natural language system will have difficulty in solving it. However, the newly generated candidate adversarial math problems must be reasonable and must be understandable and responsive to humans. The range of problem-solving capabilities required for the newly generated candidate adversarial math problems should be as consistent as possible with the benchmark math problems. For example, if the benchmark math problems belong to the category of elementary school math problems, the newly generated candidate adversarial math problems also need to belong to the category of elementary school math problems. Please output the modified candidate adversarial math problems.
[0142]
Attack strategy description
[0143]
Benchmark Mathematics Questions
[0144] In the first input prompt above, "I expect you to play the role of a math problem reconstructor. Your task is..." is a specific task description, which can be a custom description. [Attack Strategy Description] is used to describe the attack strategy selected for the benchmark math problem, and [Benchmark Math Problem] is used to input the selected math problem.
[0145] Optionally, [Attack Strategy Description] and [Benchmark Math Problem] are placeholders and can be replaced with specific content. For example, [Attack Strategy Description] is a description of the selected attack strategy, such as an attack strategy using numerical modification, [Attack Strategy Description] can be replaced with [Numerical modification of the benchmark math problem means changing the numerical value in the benchmark math problem while keeping the original content of the benchmark math problem as much as possible], and [Benchmark Math Problem] can be replaced with a math problem [There is a rectangle with a length of 5 and a width of 3, find its area].
[0146] In summary, the method provided in this embodiment, by setting a first preset text format, the first preset text format is used to describe the first input prompt, and the first preset text format is divided into two components. In the first part of the first input prompt, the task description is used to explain the purpose and method of the large language model to modify the benchmark question; in the second part of the first input prompt, the attack strategy description and the benchmark math question are input according to the first preset text format, so that the large language model generates adversarial math questions based on this information. The beneficial effect of this method is that through the format and content of the first input prompt, it can help the large language model to generate adversarial math questions more efficiently, and it can also ensure that the generated adversarial math questions meet the requirements of the attack strategy and the benchmark math questions.
[0147] Output the confrontation math problem according to the first input prompt
[0148] In some embodiments, the large language model outputs adversarial math problems based on a first input prompt. Exemplarily, Figure 7 FIG. 233 shows a flowchart of a method for generating a math problem provided by an exemplary embodiment of the present application. Step 233 may be replaced by step 233-1 and step 233-2:
[0149] Step 233-1: Input a reference math problem and an attack strategy into the large language model based on the first input prompt to obtain candidate adversarial math problems;
[0150] Among them, the candidate adversarial math problems are obtained by modifying the reference math problem based on the attack strategy.
[0151] In some embodiments, input the reference math problem and the attack strategy into the large language model, and the large language model obtains candidate adversarial math problems according to the first input prompt. These candidate adversarial problems are generated by modifying the reference math problem to meet the requirements of the selected attack strategy.
[0152] Among them, the attack strategy may be at least one of numerical modification of the reference math problem, problem expansion of the reference math problem, formula identity transformation of the reference math problem, semantic rewriting of the reference math problem, and interference injection into the reference math problem.
[0153] Optionally, the candidate adversarial math problems will be changed in terms of numerical values, symbols, structures, etc. compared with the reference math problems.
[0154] Exemplarily, according to the input content of the first input prompt, the large language model can output a corresponding output result, that is, candidate adversarial math problems. In the case where the task description, attack strategy description, and reference math problem in the above step 231-3 are used as the input content of the first input prompt, the candidate adversarial math problems output by the large language model will be numerically modified compared with the reference math problems. For example, the output adversarial math problem can be [There is a rectangle with a length of 6 and a width of 4. Find its area].
[0155] Step 233-2: Evaluate the candidate adversarial math problems, and output the candidate adversarial problems that meet the evaluation conditions as adversarial math problems.
[0156] Among them, the evaluation conditions are a set of criteria or rules for judging whether the candidate adversarial math problems meet the expected requirements.
[0157] Optionally, evaluate the candidate adversarial math problems according to the preset evaluation conditions. After evaluation, the candidate adversarial math problems that meet the evaluation conditions will be output as adversarial math problems. The output adversarial math problems are the results after screening and evaluation, meeting the expected requirements.
[0158] In summary, for the method provided in this embodiment, the benchmark math problem and the attack strategy are input into the large language model according to the requirements of the first input prompt to obtain candidate adversarial math problems; the generated candidate adversarial math problems are evaluated, and the candidate adversarial problems that meet the evaluation conditions are screened out and output as the final adversarial math problems. The beneficial effect of this method is that by inputting the benchmark math problem and the attack strategy into the large language model, candidate adversarial math problems can be generated, and then the final adversarial math problems can be screened out through evaluation, effectively utilizing the capabilities of the large language model to generate adversarial math problems that meet the requirements.
[0159] Based on Figure 7 on the basis of Figure 8 FIG. shows a flowchart of a method for generating math problems provided by an exemplary embodiment of the present application. Step 233-1 can be replaced by the following steps 310, 320, and 330:
[0160] Step 310: Use a preset scoring mechanism to evaluate the candidate adversarial math problems through the large language model. The preset scoring mechanism determines a first score based on the evaluation dimension;
[0161] Among them, the first score is used to indicate the score of the candidate adversarial math problem. By using the preset scoring mechanism to evaluate the candidate adversarial math problems, this preset scoring mechanism is based on the pre-determined evaluation dimension to evaluate the score of the candidate adversarial math problem.
[0162] Optionally, the large language model includes multiple evaluation dimensions, and the evaluation dimension is a standard or index for evaluating and measuring the characteristics or performance of the candidate adversarial math problem in different aspects.
[0163] In an optional example, the evaluation dimension includes at least one of the following dimensions:
[0164] · Difficulty consistency: Difficulty consistency is used to evaluate whether the mathematical ability required to solve the candidate adversarial math problem exceeds the mathematical ability required to solve the benchmark math problem;
[0165] · Problem-solving effectiveness: Problem-solving effectiveness is used to evaluate whether the candidate adversarial math problem has a unique and effective solution;
[0166] · Expected attack effect: The expected attack effect is used to evaluate whether the candidate adversarial math problem has the ability to cause the large language model to make mistakes more easily than the benchmark math problem;
[0167] · Attack strategy diversity: Attack strategy diversity is used to evaluate whether the candidate adversarial math problem contains an attack strategy.
[0168] In a possible implementation, the large language model uses a preset scoring mechanism to score each candidate adversarial math problem according to these evaluation dimensions, thereby determining the score of each candidate adversarial math problem, and screening out the adversarial math problems that meet the score threshold according to the scores.
[0169] Exemplarily, use the preset scoring mechanism to score the candidate adversarial math problems. Given a scoring range, for example, from 1 to 5 points, when the candidate adversarial math problem fails to meet at most one of the evaluation dimensions, the score of the candidate adversarial math problem is 1 point; when the candidate adversarial math problem meets any one of the candidate adversarial sample evaluation dimensions, the score of the candidate adversarial math problem is 2 points; when the candidate adversarial math problem meets two of the evaluation dimensions, the score of the candidate adversarial math problem is 3 points; when the candidate adversarial math problem meets three of the evaluation dimensions, the score of the candidate adversarial math problem is 4 points; when the candidate adversarial math problem meets all of the evaluation dimensions, the score of the candidate adversarial math problem is 5 points. Optionally, the large language model scores the candidate adversarial math problems, and only the candidate adversarial math problems with a score of 4 points or 5 points are considered qualified. This preset scoring mechanism can help screen out the candidate adversarial math problems that meet the requirements of multiple evaluation dimensions. At the same time, the preset scoring mechanism can also help evaluate the processing ability of the large language model for different evaluation dimensions, and further improve the robustness and generalization ability of the large language model.
[0170] Step 320: In the case where the first score is less than the score threshold, reselect the attack strategy to generate new candidate adversarial math problems;
[0171] Among them, the score threshold refers to a standard or boundary for measuring the quality of candidate adversarial math problems. The score threshold is a preset value or condition used to determine whether a candidate adversarial math problem meets the scoring requirements. Exemplarily, in the above scoring mechanism, the score threshold refers to the case where the score of the candidate adversarial math problem is 4 points.
[0172] Optionally, the score threshold refers to the lower limit of the score specified in the preset scoring mechanism. Only when the score of the candidate adversarial math problem reaches or exceeds this score threshold can it be considered a qualified adversarial math problem.
[0173] Optionally, when evaluating a candidate adversarial math problem, according to the setting of the preset scoring mechanism, the large language model evaluates the score of this candidate adversarial math problem. If this score does not reach the preset score threshold, that is, it does not meet the scoring requirements, then it is necessary to reselect the attack strategy to generate new candidate adversarial math problems. The way of reselecting the attack strategy to generate candidate adversarial math problems that meet the requirements is to improve the quality of the adversarial math problems and ensure that the generated adversarial math problems have attack capabilities.
[0174] Step 330: When the first score is greater than or equal to the score threshold, output the candidate adversarial math problem as the adversarial math problem.
[0175] Optionally, when evaluating a candidate adversarial math problem, according to the setting of the preset scoring mechanism, the large language model evaluates the score of this candidate adversarial math problem. If this score reaches the preset score threshold, that is, meets the scoring requirements, the candidate problem is output as the final adversarial math problem.
[0176] Exemplarily, the benchmark math problem is "There is a rectangle with a length of 5 and a width of 3. Find its area". Input the benchmark math problem into the large language model, and the output candidate adversarial math problem is "There is a rectangle with a length of 6 and a width of 4. Find its area". This candidate adversarial math problem meets the difficulty consistency, that is, the mathematical abilities required to solve the candidate adversarial math problem are the same as those required to solve the benchmark math problem; this candidate adversarial math problem meets the problem-solving effectiveness, that is, this candidate adversarial math problem has a unique and effective solution; this candidate adversarial math problem has diverse attack strategies, that is, this candidate adversarial math problem includes the attack strategy of numerical modification. This candidate adversarial math problem meets the three dimensions in the evaluation dimension, and the score of the candidate adversarial math problem is 4 points, reaching the score threshold of the preset scoring mechanism. This candidate adversarial math problem can be output as the adversarial math problem.
[0177] In summary, the method provided in this embodiment uses the large language model to evaluate candidate adversarial math problems by means of a preset scoring mechanism. The preset scoring mechanism determines the scores of candidate adversarial math problems based on the evaluation dimension; when the scores of candidate adversarial math problems do not meet the score threshold of the preset scoring mechanism, reselect the attack strategy to generate new candidate adversarial math problems; when the scores of candidate adversarial math problems meet the score threshold of the preset scoring mechanism, output them as the final adversarial math problems. The beneficial effect of this method is that by using the preset scoring mechanism to evaluate candidate adversarial math problems and deciding whether to reselect the attack strategy or output the final adversarial math problems according to the scoring results, the quality of adversarial math problems can be improved to meet the actual requirements.
[0178] Based on Figure 8 above, Figure 9 shows a flowchart of a method for generating math problems provided by an exemplary embodiment of the present application. Step 310 can be replaced by the following steps 311, 312, and 313:
[0179] Step 311: Based on the evaluation dimension, determine the second input prompt according to the attack strategy, the benchmark math problem, and the candidate adversarial math problem.
[0180] The second input prompt refers to the text fragment provided as the starting input to the large language model when determining the score of the candidate adversarial math problem. The second input prompt can be determined based on the attack strategy, the benchmark math problem, and the candidate adversarial math problem, with the aim of guiding the large language model to generate the score of the candidate adversarial math problem.
[0181] Among them, the evaluation dimension refers to different aspects or metrics used to evaluate the candidate adversarial math problem. According to specific requirements and goals, different evaluation dimensions can be selected to measure the quality and attack ability of the candidate problem. The attack strategy refers to the method or strategy adopted when generating the adversarial math problem, used to change or disrupt the nature or difficulty of the benchmark math problem. The benchmark math problem is an original math problem, which can be a standard, common, or math problem with a known difficulty. The candidate adversarial math problem is obtained by modifying or transforming the benchmark math problem according to the attack strategy.
[0182] In a possible implementation manner, based on the evaluation dimension, the selected attack strategy, benchmark math problem, and candidate adversarial math problem are integrated and encapsulated as the second input prompt of the large language model. By encapsulating the attack strategy, benchmark math problem, and candidate adversarial math problem in the second input prompt, the large language model can evaluate the candidate adversarial math problem according to the second input prompt and generate the score of the candidate adversarial math problem.
[0183] Optionally, in order to evaluate the quality and attack ability of the candidate adversarial math problem, the second input prompt needs to be provided as the starting input to the large language model. The second input prompt can contain some key information, such as the description of the evaluation dimension, the description of the attack strategy, and the content of the benchmark math problem. Through this information, the second input prompt can help the large language model better understand the relationship between the evaluation dimension, attack strategy, benchmark math problem, and candidate adversarial math problem, so as to generate the score of the candidate adversarial math problem. The large language model will generate an evaluation score for the candidate adversarial math problem according to the second input prompt.
[0184] Step 312: Input the second input prompt into the large language model;
[0185] Optionally, the large language model receives a text fragment as input and generates a corresponding text fragment as output. The second input prompt is input into the large language model as a text fragment to guide the generation process of the large language model. The large language model will reason and generate according to the content of the second input prompt and generate the corresponding output result.
[0186] In some embodiments, inputting the second input prompt into the large language model can help the large language model better understand and process the candidate adversarial math problems and generate scores for the evaluation of the candidate adversarial math problems.
[0187] Step 313: Obtain the scores output by the large language model according to the second input prompt, and evaluate the candidate adversarial math problems according to the scores.
[0188] According to the second input prompt, the large language model will output a score, which is used to evaluate the quality and attack ability of the candidate adversarial math problems. Using a preset scoring mechanism, compare the score output by the large language model with a set score threshold. Determine whether the candidate adversarial math problems meet the requirements according to whether the score reaches or exceeds the score threshold.
[0189] In some embodiments, by inputting the second input prompt into the large language model, the large language model will perform reasoning and generation according to the second input prompt, and the output result of the large language model will be a text representation of the score of an adversarial math problem.
[0190] In summary, the method provided in this embodiment, through the large language model, determines the second input prompt according to the given attack strategy, benchmark math problem and candidate adversarial math problem, and inputs it into the large language model; then obtains the scores output by the large language model according to the second input prompt, which are used to evaluate the candidate adversarial math problems; through the scores, it can be judged whether the candidate adversarial math problems meet the preset score threshold; if the evaluation scores of the candidate adversarial math problems do not meet the score threshold, the attack strategy can be adjusted and new candidate adversarial math problems can be regenerated. This adaptive generation process can improve the quality of the adversarial math problems.
[0191] In some embodiments, Figure 10 The flowchart of the method for generating math problems provided by an exemplary embodiment of the present application is shown. The embodiments of the present application are used to process and determine the second input prompt according to the evaluation dimension, attack strategy, benchmark math problem and candidate adversarial math problem.
[0192] · Determine the second input prompt
[0193] Optionally, by setting the second preset text format, input the content of the second input prompt into the second preset text format to determine the second input prompt. The above step 311 can be replaced by the following steps 311-1, 311-2, 311-3 and 311-4:
[0194] Step 311-1: Set the second preset text format;
[0195] Wherein, the second preset text format is used to describe the second input prompt.
[0196] In some embodiments, when determining the second input prompt, by setting a second preset text format, the structure and components of the second input prompt are described by the second preset text format. The purpose of the second preset text format is to guide the large language model to understand and process the second input prompt. The large language model correctly interprets the second input prompt through the second preset text format and performs appropriate reasoning and generation according to the content of the second preset text format.
[0197] By setting the second preset text format, the second input prompt can be divided into different parts, such as evaluation dimension description, preset scoring mechanism, attack strategy description, benchmark math problems, and candidate adversarial math problems. When determining the second input prompt, the corresponding content can be filled in according to the second preset text format to ensure that the second input prompt meets the expected requirements.
[0198] Optionally, the second preset text format can be a prescribed text format, which can be designed according to specific needs and tasks, and the second preset text format can include different components.
[0199] Exemplarily, the second preset text format can be:
[0200] "
First placeholder
[0201]
Second placeholder
[0202]
Third placeholder
[0203]
Fourth placeholder
[0204] Preset scoring mechanism:
[0205]
Fifth placeholder
[0206]
Sixth placeholder
[0207]
Seventh placeholder
First placeholder
Second placeholder
Third placeholder
Fourth placeholder
Fifth placeholder
Sixth placeholder
Seventh placeholder
[0208] Step 311-2: Input the evaluation dimension description into the second preset text format as the first part of the second input prompt;
[0209] Among them, the evaluation dimension description is used as an indicator to measure the quality and attack ability of math problems, and is used to instruct the large language model to evaluate candidate adversarial math problems according to the evaluation dimensions, and to judge whether the generated adversarial math problems meet the required evaluation criteria.
[0210] In some embodiments, the evaluation dimension description is used as the first part of the second input prompt, and the evaluation dimension description is input into the second input prompt according to the second preset text format, so as to clarify the evaluation dimensions in the second input prompt. The purpose of the evaluation dimension description is to instruct the large language model to evaluate candidate adversarial math problems according to different evaluation dimensions.
[0211] Optionally, the evaluation dimension description includes descriptions of different dimensions. For example, it includes descriptions of difficulty consistency, problem-solving effectiveness, expected attack effect, and diversity of attack strategies.
[0212] Step 311-3: Input the preset scoring mechanism into the second preset text format as the second part of the second input prompt;
[0213] The preset scoring mechanism determines the evaluation conditions for evaluating the scores of candidate adversarial math problems.
[0214] Optionally, when the preset scoring mechanism is used as the second part of the second input prompt, when generating the scores of candidate adversarial math problems, the large language model will score the generated candidate adversarial math problems in combination with the preset scoring mechanism to determine the scores of the candidate adversarial math problems.
[0215] In some embodiments, the preset scoring mechanism is used to evaluate the evaluation dimensions that the candidate adversarial math problems meet compared to the benchmark math problems. Specifically, according to the given attack strategy and the benchmark math problem, the candidate adversarial math problem is evaluated to determine whether the candidate adversarial math problem is modified based on the specified attack strategy on the basis of the benchmark math problem.
[0216] Optionally, through the preset scoring mechanism, judgments are made on each evaluation dimension and corresponding overall scores are given. During the evaluation, the candidate adversarial math problems are evaluated according to the difficulty consistency between the candidate adversarial math problems and the benchmark math problems, the problem-solving effectiveness of the candidate adversarial math problems, the expected attack effect, and the diversity of attack strategies.
[0217] Step 311-4: Input the attack strategy description, the benchmark math problem, and the candidate adversarial math problem into the second preset text format as the third part of the second input prompt.
[0218] Optionally, the attack strategy description, the benchmark math problem, and the candidate adversarial math problem are filled in according to the requirements of the second preset text format as the third part of the second input prompt.
[0219] Optionally, the second input prompt includes an evaluation dimension description, a preset scoring mechanism, an attack strategy description, a benchmark math problem, and a candidate adversarial math problem. Optionally, after receiving the second input prompt, the large language model scores the candidate adversarial math problem and performs subsequent processing based on the score of the candidate adversarial math problem.
[0220] Exemplarily, as Figure 11 shown, the second input prompt 901 includes a first part 902, a second part 903, and a third part 904. Among them, the first part 902 is the evaluation dimension description, the second part 903 is the preset scoring mechanism, and the third part 904 includes the attack strategy description, the benchmark math problem, and the candidate adversarial math problem. The evaluation dimension description, the preset scoring mechanism, the attack strategy description, the benchmark math problem, and the candidate adversarial math problem are filled in according to the requirements of the second preset text format and used as the input of the second input prompt 901.
[0221] For example, the format of the second input prompt is as follows:
[0222] "[Difficulty Consistency]
[0223] [Problem Solving Effectiveness]
[0224] [Expected Attack Effect]
[0225] [Attack Strategy Diversity]
[0226] Preset scoring mechanism: I expect you to act as an evaluator of the quality of math problems. You need to evaluate based on the above four evaluation dimensions. Your task is to evaluate the candidate adversarial math problem according to the given attack strategy and the benchmark math problem, and determine whether the candidate adversarial math problem is modified based on the specified attack strategy on the basis of the benchmark math problem. Your scoring criteria are [1, 2, 3, 4, 5], where 1 point means that none of the four evaluation dimensions are met, 2 points means that one evaluation dimension is met, 3 points means that two evaluation dimensions are met, 4 points means that three evaluation dimensions are met, and 5 points means that all evaluation dimensions are met. Please output the score of the candidate adversarial math problem.
[0227] [Attack Strategy Description]
[0228] [Benchmark Math Problem]
[0229] [Candidate Adversarial Math Problem]"
[0230] In the second input prompt mentioned above, [Difficulty Consistency], [Problem-Solving Effectiveness], [Expected Attack Effect], and [Attack Strategy Diversity] are descriptions of four different evaluation dimensions, "I expect you to be..." is the preset scoring mechanism, [Attack Strategy Description] is used to describe the attack strategy selected for the benchmark math problem, [Benchmark Math Problem] is used to input the selected math problem, and [Candidate Adversarial Math Problem] is the candidate adversarial math problem after the benchmark math problem is modified based on the attack strategy.
[0231] For example, [Difficulty Consistency] can be replaced with a specific description of the evaluation dimension, such as [Difficulty Consistency is used to evaluate whether the mathematical ability required to solve the candidate adversarial math problem exceeds the mathematical ability required to solve the benchmark math problem]. Similarly, [Problem-solving Validity] can be replaced with [Problem-solving Validity is used to evaluate whether the candidate adversarial math problem has a unique and valid solution]. [Expected Attack Effect] can be replaced with [Expected Attack Effect is used to evaluate whether the candidate adversarial math problem has the ability to cause large language models to make errors more easily than the benchmark math problem]. [Attack Strategy Diversity] can also be replaced with [Attack Strategy Diversity is used to evaluate whether the candidate adversarial math problem contains attack strategies]. Similarly, [Attack Strategy Description] can be replaced with the selected attack strategy description, such as [Modifying the numerical value of the benchmark math problem means changing the numerical value in the benchmark math problem while keeping the original content of the benchmark math problem as much as possible]. [Benchmark math problem] is replaced with [There is a rectangle with a length of 5 and a width of 3. Find its area]. [Candidate adversarial math problem] can be replaced with [There is a rectangle with a length of 6 and a width of 4. Find its area].
[0232] Optionally, according to the second input prompt, the large language model can output the score of the candidate adversarial math problem. According to the candidate adversarial math problem inputted above, the candidate adversarial math problem satisfies the consistency of difficulty, that is, the mathematical ability required to answer [there is a rectangle with a length of 5 and a width of 3, find its area] is consistent with the mathematical ability required to answer [there is a rectangle with a length of 6 and a width of 4, find its area]; the candidate adversarial math problem satisfies the validity of the problem-solving, that is, [there is a rectangle with a length of 6 and a width of 4, find its area] has a unique and valid solution; the candidate adversarial math problem has attack strategy diversity, that is, [there is a rectangle with a length of 6 and a width of 4, find its area] includes the attack strategy of numerical modification, the candidate adversarial math problem satisfies the three dimensions in the evaluation dimension, the candidate adversarial math problem scores 4 points, reaches the score threshold of the preset scoring mechanism, and the candidate adversarial math problem can be output as an adversarial math problem.
[0233] In summary, for the method provided in this embodiment, by setting a second preset text format for describing the second input prompt, the second preset text format is divided into three components. In the first part of the second input prompt, an evaluation dimension description is used to illustrate the dimensions in which the large language model needs to evaluate the candidate adversarial math problems. In the second part of the first input prompt, a preset scoring mechanism is used to score the candidate adversarial math samples. In the second part of the second input prompt, the attack strategy description, the benchmark math problem, and the candidate adversarial math problem are input according to the second preset text format, so that the large language model can generate the scores of the candidate adversarial math problems based on this information. The beneficial effect of this method is that through the format and content of the second input prompt, it can help the large language model generate the scores of the candidate adversarial math problems more efficiently, and at the same time ensure that the generated scores of the candidate adversarial math problems meet the requirements of the evaluation.
[0234] It should be noted that in a possible implementation, the adversarial math problems and the benchmark math problems are jointly input into the natural language system for adversarial training. By introducing the adversarial math problems to train the natural language system, it can maintain stable performance when facing interference, perturbation, or attack. In the adversarial training, the natural language system will interact and train with the generated adversarial math problems. The adversarial math problems have small and targeted perturbations compared to the benchmark math problems, aiming to make the natural language system produce incorrect outputs or misjudgments. By jointly performing adversarial training on the adversarial math problems and the benchmark math problems, the robustness and the ability to face attacks of the natural language system can be improved.
[0235] Figure 12 The structural block diagram of a math problem generation device provided by an embodiment of the present application is shown. This device has the function of implementing the example of the above-mentioned math problem generation method. The function can be implemented by hardware or by hardware executing corresponding software. This device can be the server introduced above or can be set in the server. As Figure 12 shown, the device 1200 may include: an acquisition module 1210, a determination module 1220, and a processing module 1230:
[0236] The acquisition module 1210 is used to acquire a benchmark math problem for evaluating the math reasoning ability of the natural language system;
[0237] The determination module 1220 is used to determine an attack strategy for the natural language system;
[0238] The processing module 1230 is used to modify the benchmark math problem based on the attack strategy through a large language model to obtain an adversarial math problem for evaluating the math reasoning ability of the natural language system using the attack strategy.
[0239] In an optional embodiment, the attack strategy includes at least one of the following strategies:
[0240] Numerically modify the reference math problem;
[0241] Expand the problem of the reference math problem;
[0242] Perform an identical transformation of the formula on the reference math problem;
[0243] Semantically rewrite the reference math problem;
[0244] Inject interference into the reference math problem.
[0245] In some optional embodiments, the processing module 1230 includes a determination sub-module, an input sub-module, and an acquisition sub-module.
[0246] In an optional embodiment, the determination sub-module is configured to determine a first input prompt based on the reference math problem and the attack strategy; the input sub-module is configured to input the first input prompt into the large language model; the acquisition sub-module is configured to acquire an adversarial math problem output by the large language model according to the first input prompt.
[0247] In some optional embodiments, the determination sub-module includes a setting unit and an input unit.
[0248] In an optional embodiment, the setting unit is configured to set a first preset text format for describing the first input prompt; the input unit is configured to input a task description into the first preset text format as a first part of the first input prompt, and the task description is used to instruct the large language model to modify the reference problem based on the attack strategy; the input unit is configured to input an attack strategy description and the reference math problem into the first preset text format as a second part of the first input prompt, and the attack strategy description is used to describe the selected attack strategy for the reference math problem.
[0249] In some optional embodiments, the acquisition sub-module further includes an input unit and an output unit.
[0250] In an optional embodiment, the input unit is configured to input the reference math problem and the attack strategy into the large language model based on the first input prompt to obtain a candidate adversarial math problem, and the candidate adversarial problem is obtained by modifying the reference math problem based on the attack strategy; the output unit is configured to evaluate the candidate adversarial math problem and output the candidate adversarial problem that meets the evaluation conditions as the adversarial math problem.
[0251] In some alternative embodiments, the output unit includes a processing subunit.
[0252] In an alternative embodiment, the processing subunit is configured to evaluate the candidate adversarial math problem through the large language model using a preset scoring mechanism. The preset scoring mechanism determines a first score based on evaluation dimensions, and the first score is used to indicate the score of the candidate adversarial math problem. The processing subunit is configured to reselect the attack strategy to generate a new candidate adversarial math problem when the first score is less than a score threshold. The processing subunit is configured to output the candidate adversarial math problem as the adversarial math problem when the first score is greater than or equal to the score threshold.
[0253] In an alternative embodiment, the evaluation dimensions include at least one of the following dimensions:
[0254] The evaluation dimension includes difficulty consistency, which is used to evaluate whether the mathematical ability required to solve the candidate adversarial math problem exceeds the mathematical ability required to solve the benchmark math problem.
[0255] The evaluation dimension includes solution effectiveness, which is used to evaluate whether the candidate adversarial math problem has a unique and effective solution.
[0256] The evaluation dimension includes expected attack effect, which is used to evaluate whether the candidate adversarial math problem has the ability to cause the large language model to make more mistakes than the benchmark math problem.
[0257] The evaluation dimension includes attack strategy diversity, which is used to evaluate whether the candidate adversarial math problem contains the attack strategy.
[0258] In some alternative embodiments, the processing subunit further includes a determination subunit, an input subunit, and an acquisition subunit.
[0259] In an alternative embodiment, the determination subunit is configured to determine a second input prompt based on the evaluation dimensions, the attack strategy, the benchmark math problem, and the candidate adversarial math problem. The input subunit is configured to input the second input prompt into the large language model. The acquisition subunit is configured to acquire the score output by the large language model according to the second input prompt, and evaluate the candidate adversarial math problem according to the score. The score is used to determine whether the candidate adversarial math problem meets the score threshold.
[0260] In an optional embodiment, a determination subunit is configured to set a second preset text format for describing the second input prompt; a determination subunit is configured to input an evaluation dimension description into the second preset text format as the first part of the second input prompt; a determination subunit is configured to input the preset scoring mechanism into the second preset text format as the second part of the second input prompt; a determination subunit is configured to input the attack strategy description, the benchmark math problem, and the candidate adversarial math problem into the second preset text format as the third part of the second input prompt.
[0261] It should be noted that the specific limitations in the above-described embodiments of the generation device for one or more math problems can be referred to the limitations on the generation method of math problems described above, and will not be elaborated here. Each module of the above device can be implemented in whole or in part by software, hardware, and their combination. Each module can be embedded in the processor of the computer device in hardware form or be independent of it, or can be stored in the memory of the computer device in software form to facilitate the processor to call and execute the operations corresponding to each module.
[0262] Figure 13 The structural block diagram of a computer device 1500 shown in an exemplary embodiment of the present application is illustrated. The computer device can be used to implement the cloud rendering method of the game character provided in the above embodiment. The computer device 1500 includes a central processing unit (CPU) 1501, a system memory 1504 including a random access memory (RAM) 1502 and a read-only memory (ROM) 1503, and a system bus 1505 connecting the system memory 1504 and the central processing unit 1501. The computer device 1500 also includes a basic input / output system (Input / Output system, I / O system) 1506 for facilitating the transmission of information between various components within the computer device, and a mass storage device 1507 for storing an operating system 1513, application programs 1514, and other program modules 1515.
[0263] The basic input / output system 1506 includes a display 1508 for displaying information and input devices 1509 such as a mouse, keyboard, etc. for user input of information. Both the display 1508 and the input devices 1509 are connected to the central processing unit 1501 through an input / output controller 1510 connected to the system bus 1505. The basic input / output system 1506 may further include an input / output controller 1510 for receiving and processing inputs from a plurality of other devices such as a keyboard, mouse, or electronic stylus. Similarly, the input / output controller 1510 also provides outputs to a display screen, printer, or other types of output devices.
[0264] The mass storage device 1507 is connected to the central processing unit 1501 through a mass storage controller (not shown) connected to the system bus 1505. The mass storage device 1507 and its associated computer-readable storage medium provide non-volatile storage for the terminal device 1500. That is to say, the mass storage device 1507 may include computer-readable storage media (not shown) such as a hard disk or a compact disc read-only memory (CD-ROM) drive.
[0265] Without loss of generality, the computer-readable storage medium may include computer storage media and communication media. Computer storage media includes volatile and non-volatile, removable and non-removable media implemented by any method or technology for storing information such as computer-readable storage instructions, data structures, program modules, or other data. Computer storage media includes RAM, ROM, erasable programmable read-only registers (EPROM), electrically-erasable programmable read-only memory (EEPROM), flash memory or other solid-state storage technologies, CD-ROM, digital versatile disc (DVD) or other optical storage, magnetic tape cartridges, tapes, disk storage or other magnetic storage devices. Of course, those skilled in the art will know that the computer storage media is not limited to the above several. The above system memory 1504 and mass storage device 1507 may be collectively referred to as memory.
[0266] The memory stores one or more programs, the one or more programs are configured to be executed by one or more central processing units 1501, the one or more programs contain instructions for implementing the above method embodiments, and the central processing unit 1501 executes the one or more programs to implement the methods provided by the above respective method embodiments.
[0267] According to various embodiments of the present application, the computer device 1500 may also run by connecting to a remote terminal device on the network through a network such as the Internet. That is, the computer device 1500 may be connected to the network 1512 through the network interface unit 1511 connected to the system bus 1505, or in other words, the network interface unit 1511 may also be used to connect to other types of networks or remote terminal device systems (not shown).
[0268] The memory further includes one or more programs, the one or more programs are stored in the memory, and the one or more programs include steps for performing the methods provided in the embodiments of the present application that are executed by the terminal device.
[0269] Embodiments of the present application further provide a computer-readable storage medium, in which a computer program is stored, and the computer program is loaded and executed by a processor to implement the method for generating math problems provided in the above method embodiments.
[0270] Embodiments of the present application further provide a computer program product, the computer program product includes a computer program, and the computer program is stored in a computer-readable storage medium; the computer program is read and executed by a processor of a computer device from the computer-readable storage medium, so that the computer device executes to implement the method for generating math problems provided in the above method embodiments.
[0271] It can be understood that in the specific implementation of the present application, for data, historical data, portraits, and other user data processing related to user identity or characteristics, when the above embodiments of the present application are applied to specific products or technologies, user permission or consent is required, and the collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards of relevant countries and regions.
[0272] It should be noted that unless otherwise clearly defined herein, all terms used in the claims are interpreted according to their ordinary meanings in the technical field. Unless otherwise clearly stated, all references to "an element, device, component, equipment, step, etc." will be interpreted openly as referring to at least one instance of the element, device, component, equipment, step, etc. Unless clearly stated, the steps of any method disclosed herein are not necessarily to be executed in the exact order disclosed.
[0273] It should be understood that the "plurality" mentioned herein refers to two or more. "And / or", which describes the association relationship of associated objects, indicates that three relationships can exist. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after.
Claims
1. A method for generating mathematical problems. It is characterized in that The method is performed by a computer device, and the method comprises: Obtaining a benchmark math question, where the benchmark math question is used to evaluate the math reasoning ability of the natural language system; determining an attack strategy against the natural language system; The benchmark math problem is modified based on the attack strategy through a large language model to obtain an adversarial math problem, and the adversarial math problem is used to evaluate the mathematical reasoning ability of the natural language system using the attack strategy.
2. The method according to claim 1, It is characterized in that The attack strategy includes at least one of the following strategies: Modifying the numerical value of the benchmark math problem; Expanding the benchmark math questions; Performing formula identity transformation on the benchmark math problem; semantically rewriting the benchmark math problem; Interference injection is performed on the benchmark math problem.
3. The method according to claim 2, It is characterized in that The step of modifying the benchmark math problem based on the attack strategy using a large language model to obtain an adversarial math problem includes: Determining a first input prompt based on the benchmark math problem and the attack strategy; inputting the first input prompt into the large language model; Obtain an adversarial math question output by the large language model according to the first input prompt.
4. The method according to claim 3, It is characterized in that The step of determining a first input prompt based on the benchmark math problem and the attack strategy includes: Setting a first preset text format, where the first preset text format is used to describe the first input prompt; Inputting a task description into the first preset text format as a first part of the first input prompt, the task description being used to instruct the large language model to modify the benchmark question based on the attack strategy; An attack strategy description and the benchmark math problem are input into the first preset text format as a second part of the first input prompt, wherein the attack strategy description is used to describe the attack strategy selected for the benchmark math problem.
5. The method according to any one of claims 1 to 4, It is characterized in that The step of obtaining the adversarial math problem output by the large language model according to the first input prompt includes: Inputting the benchmark math problem and the attack strategy into the large language model based on the first input prompt to obtain a candidate adversarial math problem, where the candidate adversarial problem is obtained by modifying the benchmark math problem based on the attack strategy; The candidate confrontation mathematics questions are evaluated, and the candidate confrontation mathematics questions that meet the evaluation conditions are output as the confrontation mathematics questions.
6. The method according to claim 5, It is characterized in that The step of evaluating the candidate confrontation questions and outputting the candidate confrontation questions that meet the evaluation conditions as the confrontation mathematics questions includes: Using the large language model, evaluating the candidate adversarial math questions using a preset scoring mechanism, wherein the preset scoring mechanism determines a first score based on an evaluation dimension, wherein the first score is used to indicate a score of the candidate adversarial math questions; In the case where the first score is less than the score threshold, reselecting the attack strategy to generate a new candidate adversarial mathematics question; When the first score is greater than or equal to the score threshold, the candidate adversarial math problem is output as the adversarial math problem.
7. The method according to claim 6, It is characterized in that The evaluation dimension includes at least one of the following dimensions: The evaluation dimension includes difficulty consistency, and the difficulty consistency is used to evaluate whether the mathematical ability required for answering the candidate confrontational mathematical problem exceeds the mathematical ability required for answering the benchmark mathematical problem; Problem-solving validity, wherein the problem-solving validity is used to evaluate whether the candidate confrontational mathematics problem has a unique and valid solution; an expected attack effect, wherein the expected attack effect is used to evaluate whether the candidate adversarial math question is more likely to cause the large language model to make errors than the benchmark math question; Attack strategy diversity, where the attack strategy diversity is used to evaluate whether the candidate adversarial mathematics problem contains the attack strategy.
8. The method according to claim 7, It is characterized in that The step of evaluating the candidate adversarial mathematics questions by using the large language model and a preset scoring mechanism includes: Based on the evaluation dimension, according to the attack strategy, the benchmark math problem and the candidate adversarial math problem, determining a second input prompt; inputting the second input prompt into the large language model; A score output by the large language model according to the second input prompt is obtained, and the candidate adversarial math question is evaluated according to the score, where the score is used to determine whether the candidate adversarial math question meets a score threshold.
9. The method according to claim 8, It is characterized in that The step of determining a second input prompt based on the evaluation dimension and according to the attack strategy, the benchmark math problem, and the candidate adversarial math problem includes: Setting a second preset text format, where the second preset text format is used to describe the second input prompt; Entering the evaluation dimension description into the second preset text format as the first part of the second input prompt; Inputting the preset scoring mechanism into the second preset text format as the second part of the second input prompt; The attack strategy description, the benchmark math problem and the candidate adversarial math problem are input into the second preset text format as the third part of the second input prompt.
10. A device for generating mathematical problems, It is characterized in that The device comprises: An acquisition module, used for acquiring a benchmark math question, wherein the benchmark math question is used for evaluating the math reasoning ability of the natural language system; A determination module, used to determine an attack strategy against the natural language system; A processing module is used to modify the benchmark math problem based on the attack strategy through a large language model to obtain an adversarial math problem, wherein the adversarial math problem is used to evaluate the mathematical reasoning ability of the natural language system using the attack strategy.
11. A computer device, It is characterized in that The computer device includes a processor and a memory, wherein a computer program is stored in the memory, and the computer program is loaded and executed by the processor to implement the method for generating a math problem as described in any one of claims 1 to 9.
12. A computer-readable storage medium having a computer program stored therein. It is characterized in that When the computer program is executed by a processor, the method for generating a mathematical problem as described in any one of claims 1 to 9 is implemented.
13. A computer program product, It is characterized in that The computer program comprises a computer program product, wherein the computer program product comprises computer instructions, wherein the computer instructions are stored in a computer-readable storage medium, and a processor reads and executes the computer instructions from the computer-readable storage medium to implement the method for generating a math problem as described in any one of claims 1 to 9.