Model training method, data processing method and related device
By employing a dual-network architecture and loss function training method, the accuracy and flexibility of the LLM model in handling complex SAT problems are improved. This addresses the issue of insufficient semantic mining capabilities in existing LLM models, enabling more efficient solutions to satisfiability problems.
Patent Information
- Application Number
- CN202410473894.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-18
- Publication Date
- 2025-10-24
AI Technical Summary
When faced with complex SAT problems, existing technologies have limited semantic mining capabilities of LLM models, making it difficult to accurately handle satisfiability problems with multiple constraints.
A dual-sub-network architecture is adopted. First, the first sub-network is used to understand the problem and generate expressions and initial solutions. Then, the second sub-network is used to optimize the initial solutions in depth, build a loss function to train the target model, and combine the solver to assign and adjust variables to meet the constraints.
It improves the model's reasoning ability and accuracy in dealing with complex multi-constraint satisfiability problems, enhances its ability to handle unknown variables, and simplifies the solution process.
Smart Images

Figure CN120832944A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence, and in particular to a model training method, a data processing method, and related devices. Background Art
[0002] The Boolean satisfiability (SAT) problem is a problem about logical equations. It aims to determine whether there exists a set of variable values such that a given logical equation is true. The SAT problem has widespread application in industry. For example, in circuit design, we can transform the circuit design problem into a SAT problem, ensuring the circuit's correctness by verifying that the circuit meets all specification requirements. Furthermore, it aims to find the optimal circuit design while satisfying certain performance constraints.
[0003] The properties of SAT problems are highly susceptible to small changes in their local structure. For example, a slight change in a symbol or constraint can transform the entire problem from unsatisfiable (unsat) to satisfiable (sat), significantly reducing solution time. This demonstrates that relying solely on structural information cannot accurately solve practical industrial problems. Therefore, using graph neural networks to analyze the structural characteristics of SAT problems is not suitable.
[0004] Currently, researchers are trying to use large language models (LLMs) to mine the implicit semantic information in SAT problems to solve them. However, the semantic mining capabilities of LLM models are still limited when faced with complex SAT problems. Summary of the Invention
[0005] The embodiments of the present application provide a model training method, a data processing method and related devices, which can improve the reasoning ability of the model in the face of multi-constraint satisfiability problems.
[0006] In a first aspect, an embodiment of the present application provides a model training method, the method comprising:
[0007] When training the model to be trained is required, firstly obtain first description information of the first problem, wherein the description information may be a detailed description of the first problem in text form, or a mathematical expression, or both;
[0008] Inputting the first description information into the first sub-network, and performing a series of processing on the description information by the first sub-network to obtain a first expression and a first solution of the first problem, wherein the first expression is a mathematical expression of the first problem, including a plurality of constraints, and the first solution is an initial solution provided by the first sub-network, and the first solution can satisfy at least one constraint.
[0009] Then the first expression and the first solution are input into the second sub-network to obtain a second solution, the number of constraint conditions satisfied by the second solution is greater than or equal to the first solution, and the second sub-network is a network model specially used for solving the first expression;
[0010] After obtaining the second solution, a loss function can be constructed using the second solution to complete the backward propagation of model training, and the training of the target model is realized, the target model including the first sub-network and the second sub-network.
[0011] In this application, the first problem is a satisfiability problem, such as a Boolean satisfiability (SAT) problem and a satisfiability modulo theories (SMT) problem. The satisfiability problem mainly studies whether there is a way of assignment (such as the first solution and the second solution) that makes the proof problem satisfiable (SAT), that is, there is a way of assignment that can satisfy all the constraint conditions; or the proof problem is unsatisfiable (UNSAT), that is, there is no way of assignment that can satisfy all the constraint conditions. When the problem is UNSAT, the way of assignment that maximizes the satisfiable constraint conditions needs to be found.
[0012] By using the above method, during model training, the first sub-network first preliminarily understands the first problem, generates the first expression and the first solution, and provides a starting point for the basic solution. Then, the second sub-network deeply processes and optimizes the output of the first sub-network, generates the second solution, and significantly improves the number of constraint conditions satisfied by the second solution. Finally, the loss function of the target model is constructed according to the second solution output by the second sub-network. Therefore, the target model can exhibit more accurate reasoning ability when facing more complex multi-constraint satisfiability problems.
[0013] In one possible implementation, the target model is trained based on the second solution, including: constructing a loss function value based on the second solution, training the second sub-network; obtaining the target model, the target model including the first sub-network and the trained second sub-network.
[0014] By using the above method, the loss function based on the second solution is constructed to train the second sub-network, and the trained second sub-network is combined with the first sub-network to form the target model. By constructing and training the second sub-network instead of constructing and training the entire model at one time, the flexibility of the model can be increased. The second sub-network can learn specific problems or features, thereby further improving the accuracy of the model based on the first sub-network.
[0015] In a possible implementation, the first expression further includes N variables, each constraint condition includes at least one variable, the first solution includes first values of the at least one variable, the first values are Boolean values, and N is an integer greater than 1.
[0016] In a possible implementation, the second subnetwork is a second expression obtained by relaxing the first expression, the second expression includes N variables and first weights, a vector length of a value of each variable in the N variables is less than or equal to 1, and the first weights are generated based on coefficients of each variable in the plurality of constraint problems; and inputting the first expression and the first solution into the second subnetwork to obtain a second solution includes: inputting the first solution into the second expression to obtain the second solution, the second solution including second values of each variable in the N variables, the second values being floating-point values.
[0017] In this application, the variables in the first expression are all one-dimensional variables, and the values of each variable are discrete Boolean values; after relaxation, the variables of the second expression become multi-dimensional vectors, and the values of each variable are continuous floating-point values, for example, (0.6, 0.8). Then the multi-dimensional vector is projected into the solution space to obtain a one-dimensional floating-point value result.
[0018] In a possible implementation, before training the target model based on the second solution, the method further includes: obtaining a target solution of the first problem, the target solution including target values of each variable in the N variables, the target values being Boolean values. Training the target model based on the second solution includes: updating the first weights based on differences between the values of each variable in the second solution and the values of each variable in the target solution.
[0019] By using the above method, the second subnetwork is a second expression obtained by relaxing the first expression, the second expression allows the vector length of the value of each of the N variables to be less than or equal to 1, and the value is a floating-point value. This means that the solution space is no longer purely Boolean, but can be a continuous floating-point value, thereby providing greater flexibility and accuracy. By comparing the value of each variable in the second solution with the value of each variable in the target solution, the error of the model can be accurately evaluated, and the first weights in the second expression can be updated. Thus, the accuracy of the target model in processing such problems and the ability to predict unknown variables are improved.
[0020] In a possible implementation, the second subnetwork is a solver; and inputting the first expression and the first solution into the second subnetwork to obtain a second solution includes: inputting the first expression and the first solution into the solver to obtain the second solution, the second solution including first information, or the first information and a first variable; wherein the first information is used to indicate whether all constraint conditions in the first expression are satisfied by the first solution, and the first variable is a variable in the first solution that causes a first constraint condition to be unsatisfied, the first constraint condition being included in the plurality of constraint conditions.
[0021] For example, the first expression includes 10 constraints and 10 variables, and the first solution output by the first sub-network includes an initial solution of the 10 variables. The first expression and the first solution are input into the solver, which can identify whether the first solution can satisfy all the constraints of the first expression. If yes, SAT is output; if not, UNSAT is output, as well as the variables in the first solution that cause the constraints to be unsatisfied, for example, the assignments of variables a and b in the first solution cause 2 constraints to be unsatisfied.
[0022] In a possible implementation, the target model is trained based on the second solution, including: performing a Boolean flip on the value of the first variable in the first solution when the first information indicates that the first solution is not satisfied; inputting the updated first solution and the first expression into the solver until a target condition is met, to obtain the target model; the target condition is that the first information indicates that the first solution is satisfied, or the number of the first constraints no longer changes.
[0023] With the above method, the solver serves as a calibration layer of the LLM output, and corrects the variable assignments to satisfy the constraints. Through flipping the assignments and re-verification, a solution that satisfies all the conditions is gradually found, and optimization is continuously performed before the target is met. This makes the target model more flexible and practical when dealing with complex problems, and converts the solving process into a time problem rather than a completely unsolvable state.
[0024] In a possible implementation, the first problem is a Boolean satisfiability SAT problem or a satisfiability modulo theories SMT problem.
[0025] In a possible implementation, the first sub-network is a large language model LLM.
[0026] In a second aspect, an embodiment of the present application provides a data processing method, which includes:
[0027] obtaining second description information of a second problem;
[0028] inputting the second description information into the target model to obtain a target solution of the second problem; wherein the target model includes a first sub-network and a second sub-network, the first sub-network is configured to generate a third expression and an initial solution according to the second description information, the second sub-network is configured to generate the target solution according to the third expression and the initial solution, the third expression includes a plurality of constraints, and the number of constraints satisfied by the target solution is greater than or equal to the number of constraints satisfied by the initial solution.
[0029] In the present application, the second problem is a satisfiability problem. The first sub-network in the target model is a language processing model responsible for parsing the problem description information and converting it into a mathematical expression, as well as performing preliminary understanding of the problem and generating an initial solution to the problem. The second sub-network is a neural network designed for the satisfiability problem, which further optimizes the initial solution.
[0030] By using the above method, through the cooperative work of the two sub-networks, the solution can be automatically generated and optimized, improving the efficiency and accuracy of solving the satisfiability problem.
[0031] In a possible implementation, the third expression further includes N variables, each constraint condition includes at least one variable, the target solution includes third values of the N variables, the third values are Boolean values, and N is an integer greater than 1.
[0032] In a possible implementation, the second problem is a Boolean satisfiability (SAT) problem or a satisfiability modulo theories (SMT) problem.
[0033] In a possible implementation, the first sub-network is a large language model (LLM).
[0034] In a third aspect, the embodiments of the present application provide a model training apparatus, which comprises: an acquisition unit configured to acquire description information of a first problem; a problem parsing unit configured to input the description information into a first sub-network to obtain a first expression of the first problem and a first solution, the first expression including a plurality of constraint conditions, and the first solution satisfying at least one constraint condition; a processing unit configured to input the first expression and the first solution into a second sub-network to obtain a second solution, the second solution satisfying a number of constraint conditions greater than or equal to the first solution; and a training unit configured to train a target model based on the second solution, the target model including the first sub-network and the second sub-network.
[0035] In a possible implementation, the training unit is specifically configured to: construct a loss function value based on the second solution, and train the second sub-network; and acquire the target model, the target model including the first sub-network and the trained second sub-network.
[0036] In a possible implementation, the first expression further includes N variables, each constraint condition includes at least one variable, the first solution includes first values of the at least one variable, the first values are Boolean values, and N is an integer greater than 1.
[0037] In a possible implementation, the second sub-network is a second expression obtained by relaxing the first expression, the second expression including N variables and a first weight, a vector length of a value of each variable in the N variables being less than or equal to 1, and the first weight being generated based on coefficients of each variable in the plurality of constraint conditions.
[0038] The processing unit is specifically configured to: input the first solution into the second expression to obtain a second solution, the second solution comprising a second value of each of the N variables, the second value being a floating-point value.
[0039] In a possible implementation, before training the target model based on the second solution, the obtaining unit is further configured to: obtain a target solution of the first problem, the target solution comprising a target value of each of the N variables, the target value being a Boolean value.
[0040] The training unit is specifically configured to: update the first weight based on a difference between a value of each of the variables in the second solution and a value of each of the variables in the target solution.
[0041] In a possible implementation, the second subnetwork is a solver; and the processing unit is specifically configured to: input the first expression and the first solution into the solver to obtain the second solution, the second solution comprising first information or the first information and the first variable; wherein the first information is used to indicate whether all constraint conditions in the first expression are satisfied by the first solution, and the first variable is a variable in the first solution that causes a first constraint condition to be unsatisfied, the first constraint condition being included in the plurality of constraint conditions.
[0042] In a possible implementation, the training unit is specifically configured to: when the first information indicates unsatisfied, perform a Boolean flip on a value of the first variable in the first solution; input the updated first solution and the first expression into the solver until a target condition is met to obtain the target model; and the target condition is that the first information indicates satisfied, or the number of the first constraint conditions no longer changes.
[0043] In a possible implementation, the first problem is a Boolean satisfiability SAT problem or a satisfiability modulo theories SMT problem.
[0044] In a possible implementation, the first subnetwork is a large language model LLM.
[0045] In a fourth aspect, an embodiment of the present application provides a data processing apparatus, which comprises:
[0046] The obtaining module is configured to obtain second description information of a second problem;
[0047] The processing module is configured to input the second description information into the target model to obtain a target solution of the second problem; wherein the target model comprises a first subnetwork and a second subnetwork, the first subnetwork is configured to generate a third expression and an initial solution according to the second description information, the second subnetwork is configured to generate the target solution according to the third expression and the initial solution, the third expression comprises a plurality of constraint conditions, and the number of constraint conditions satisfied by the target solution is greater than or equal to the number of constraint conditions satisfied by the initial solution.
[0048] In a possible implementation, the third expression further includes N variables, each constraint condition includes at least one variable, and the target solution includes third values of the N variables, the third values being Boolean values, and N being an integer greater than 1.
[0049] In a possible implementation, the second problem is a Boolean satisfiability SAT problem or a satisfiability modulo theories SMT problem.
[0050] In a possible implementation, the first sub-network is a large language model LLM.
[0051] The fifth aspect of the present application provides a model training apparatus, which can include a processor, the processor being coupled to a memory, and the memory storing program instructions that, when executed by the processor, implement the method of the first aspect or any implementation manner of the first aspect. For the steps performed by the processor in each possible implementation manner of the first aspect, specific details can be referred to the first aspect, and will not be repeated here.
[0052] The sixth aspect of the present application provides a data processing apparatus, which can include a processor, the processor being coupled to a memory, and the memory storing program instructions that, when executed by the processor, implement the method of the second aspect or any implementation manner of the second aspect. For the steps performed by the processor in each possible implementation manner of the first aspect, specific details can be referred to the second aspect, and will not be repeated here.
[0053] The seventh aspect of the present application provides circuitry, which includes processing circuitry configured to perform the method of any implementation manner of the first aspect.
[0054] The eighth aspect of the present application provides circuitry, which includes processing circuitry configured to perform the method of any implementation manner of the second aspect.
[0055] The ninth aspect of the present application provides a computer-readable storage medium, which stores a computer program, and when the computer program is executed on a computer, the computer is caused to perform the method of any implementation manner of the first aspect.
[0056] The tenth aspect of the present application provides a computer-readable storage medium, which stores a computer program, and when the computer program is executed on a computer, the computer is caused to perform the method of any implementation manner of the second aspect.
[0057] The eleventh aspect of the present application provides a computer program product, which, when executed on a computer, causes the computer to perform the method of any implementation manner of the first aspect.
[0058] The twelfth aspect of the present application provides a computer program product, when running on a computer, causes the computer to execute the method of any implementation manner of the second aspect.
[0059] The thirteenth aspect of the present application provides a chip system, which comprises a processor for supporting the model training device or the data processing device to implement the functions involved in the above aspects, such as sending or processing the data and / or information involved in the above method. In a possible design, the chip system further comprises a memory for storing the necessary program instructions and data of the server or the communication device. The chip system can be composed of a chip, or can comprise a chip and other discrete devices.
[0060] The technical effects of the third aspect, the fifth aspect, the seventh aspect, the ninth aspect and the eleventh aspect of the present application can be understood in conjunction with the technical effects of the first aspect and any implementation manner of the first aspect. The technical effects of the fourth aspect, the sixth aspect, the eighth aspect, the tenth aspect and the twelfth aspect of the present application can be understood in conjunction with the technical effects of the second aspect and any implementation manner of the second aspect. BRIEF DESCRIPTION OF DRAWINGS
[0061] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced as follows. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creative labor on the basis of these drawings.
[0062] Figure 1 A structural schematic diagram of an artificial intelligence main body framework;
[0063] Figure 2 A schematic diagram of a system architecture provided by the embodiments of the present application;
[0064] Figure 3 A flowchart of a model training method provided by the embodiments of the present application;
[0065] Figure 4 A schematic diagram of model training based on a relaxed mathematical model provided by the embodiments of the present application;
[0066] Figure 5 A structural schematic diagram of training combined with a SATNet target model provided by the embodiments of the present application;
[0067] Figure 6 A schematic diagram of model training based on a solver provided by the embodiments of the present application;
[0068] Figure 7A schematic diagram of inputting description information to a target model and outputting target text provided by an embodiment of the present application;
[0069] Figure 8 A problem-solving effect diagram of a model training method provided by an embodiment of the present application;
[0070] Figure 9 Another problem-solving effect diagram of a model training method provided by an embodiment of the present application;
[0071] Figure 10 Another problem-solving effect diagram of a model training method provided by an embodiment of the present application;
[0072] Figure 11a A structural schematic diagram of a model training device provided by an embodiment of the present application;
[0073] Figure 11b A structural schematic diagram of a data processing device provided by an embodiment of the present application;
[0074] Figure 12 A structural schematic diagram of an execution device provided by an embodiment of the present application;
[0075] Figure 13 A structural schematic diagram of a training device provided by an embodiment of the present application;
[0076] Figure 14 A structural schematic diagram of a chip provided by an embodiment of the present application. DETAILED DESCRIPTION
[0077] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work fall within the scope of protection of the present application.
[0078] The terms "first", "second", "third", "fourth" and the like in the description and in the claims of the present application, and above-mentioned drawings, if any, are used to distinguish between similar objects and are not necessarily used to describe a particular sequential or chronological order. It is to be understood that the use of the terms so-termed, where appropriate, can be interchanged with each other to describe the embodiments of the present application described herein, for example, can be implemented in an order other than that illustrated or described herein. In addition, the terms "comprise" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or apparatus comprising a list of steps or units does not necessarily limit to those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products or apparatuses.
[0079] With the continuous development of large language model (LLM) technology, the ability of LLM to understand semantic information has gradually increased. Therefore, when dealing with relatively simple SAT problems (such as the number of variables involved is less than 20), LLMs such as Huawei Panggu large model, LLAMA2-13b, etc. can directly generate accurate solutions. However, it needs to be noted that although LLMs perform well in understanding the semantics of textual information, their solving ability will decrease significantly when facing more complex SAT problems (such as the number of variables involved is more than 50).
[0080] In this case, in order to effectively solve SAT problems, it is usually necessary to rely on specialized SAT solvers. Therefore, the current trend of practical work is that researchers are exploring how to make LLMs learn to recognize and distinguish different types of SAT problems, so as to automatically call SAT solvers to complete the solving task when necessary. Although this method is feasible in some cases, it also increases the complexity of operations and the number of steps.
[0081] To solve the above problems, the embodiments of the present application provide a model training method, which can be implemented in combination with artificial intelligence (AI) technology. AI technology is a technology discipline that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence. AI technology obtains the best results by perceiving the environment, acquiring knowledge and using knowledge. In other words, artificial intelligence technology is a branch of computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a similar way to human intelligence. Data processing using artificial intelligence is a common application of artificial intelligence.
[0082] First, the overall workflow of the artificial intelligence system is described, please refer to Figure 1 , Figure 1As a structural diagram of the artificial intelligence subject framework, the following elaborates the above-mentioned artificial intelligence subject framework from two dimensions of "intelligent information chain" (horizontal axis) and "IT value chain" (vertical axis). The "intelligent information chain" reflects a series of processes from data acquisition to processing. For example, it can be a general process of intelligent information perception, intelligent information representation and formation, intelligent reasoning, intelligent decision-making, intelligent execution and output. In this process, data undergoes a condensation process of "data-information-knowledge-wisdom". The "IT value chain" reflects the value brought by artificial intelligence to the information technology industry from the underlying infrastructure of human intelligence, information (provision and processing technology implementation) to the industrial ecological process of the system.
[0083] (1) Infrastructure
[0084] The infrastructure provides computing power support for the artificial intelligence system, realizes communication with the external world, and realizes support through the underlying platform. Communication with the outside world through sensors; computing power is provided by intelligent chips (CPU, NPU, GPU, ASIC, FPGA, etc. Hardware acceleration chips); the underlying platform includes distributed computing framework and network related platform guarantee and support, which can include cloud storage and computing, interconnection network, etc. For example, sensors and external communication obtain data, which are provided to intelligent chips in the distributed computing system provided by the underlying platform for calculation.
[0085] (2) Data
[0086] The data on the upper layer of the infrastructure is used to represent the data source in the field of artificial intelligence. Data involves graphics, images, voice, text, and also involves Internet of Things data of traditional devices, including business data of existing systems and sensing data such as force, displacement, liquid level, temperature, humidity, etc.
[0087] (3) Data processing
[0088] Data processing usually includes data training, machine learning, deep learning, search, reasoning, decision-making, etc.
[0089] Among them, machine learning and deep learning can symbolize and formalize intelligent information modeling, extraction, preprocessing, training, etc.
[0090] Reasoning refers to the process of simulating human intelligent reasoning methods in computers or intelligent systems, using formalized information to perform machine thinking and solve problems according to reasoning control strategies, and the typical function is search and matching.
[0091] Decision-making refers to the process of decision-making after intelligent information reasoning, which usually provides functions such as classification, sorting, prediction, etc.
[0092] (4) General capabilities
[0093] After the data is processed by the above-mentioned data processing, further based on the result of the data processing, some general capabilities can be formed, which can be an algorithm or a general system, for example, translation, text analysis, computer vision processing, speech recognition, image recognition, etc.
[0094] (5) Intelligent product and industry application
[0095] Intelligent product and industry application refers to the product and application of artificial intelligence system in various fields, which is the packaging of the overall solution of artificial intelligence, and realizes the application of intelligent information decision productization. Its application fields mainly include: intelligent terminal, intelligent transportation, intelligent medical treatment, automatic driving, smart city, etc.
[0096] Please refer to Figure 2 The model training method provided by the embodiment of the application is applied to a system architecture 200. As shown in Figure 2 In the system architecture 200, the execution device 210 can be implemented by one or more servers, and can be optionally combined with other computing devices, such as data storage, routers, load balancers, etc. The execution device 210 can be arranged on one physical site or distributed on multiple physical sites. The execution device 210 can use data in the data storage system 250 or call program code in the data storage system 250 to implement the training method of the classification model provided by the embodiment of the application, and then obtain the model.
[0097] Users can operate their respective user devices (such as local device 201 and local device 202) to interact with the execution device 210. Each local device can represent any computing device, such as a personal computer, a computer workstation, a smartphone, a tablet computer, a smart camera, a smart car, or other types of cellular phones, a media consumption device, a wearable device, a set-top box, a game console, etc.
[0098] Each user's local device can interact with the execution device 210 through a communication network of any communication mechanism / communication standard, which can be a wide area network, a local area network, a point-to-point connection, etc., or any combination thereof.
[0099] In one implementation, the execution device 210 is used to implement the training method of the classification model provided by the embodiment of the application, and the obtained model is sent to the local device 201 and the local device 202 through the communication network, so that the local device 201 and the local device 202 can implement the deployment and running of the model, thereby realizing the classification method provided by the embodiment of the application.
[0100] In another implementation, one aspect or multiple aspects of the execution device 210 can be implemented by each local device, for example, the local device 201 can provide local data or feedback calculation results for the execution device 210, or implement the training method of the model provided in the embodiments of the present application.
[0101] It should be noted that all functions of the execution device 210 can also be implemented by the local device. For example, the local device 201 implements the functions of the execution device 210 and provides services for its own user, or provides services for the user of the local device 202.
[0102] In general, the training method of the model provided in the embodiments of the present application can be applied to an electronic device, for example, the execution device 210, the local device 201 or the local device 202 described above. Exemplarily, the electronic device can be, for example, a server, a wireless electronic device in industrial control, a mobile phone, a personal computer (PC), a notebook computer, a tablet computer, etc. For ease of understanding, the method provided in the embodiments of the present application will be introduced below by taking the case of applying the method to a server.
[0103] For ease of understanding, the related terms and concepts mainly involved in the embodiments of the present application will be introduced first.
[0104] (1) Boolean satisfiability (SAT) problem
[0105] The SAT problem is a classic problem in computer science, which involves the satisfiability of a propositional logic formula. A SAT problem can be regarded as a satisfiability problem of a Boolean formula, that is, whether there is a way of assignment such that the value of the Boolean formula is true. SAT problems are often used to evaluate algorithm performance, test computer hardware, and study the boundaries of computational complexity. For example, circuit design, scheduling optimization, software verification, etc. can be solved by converting them into SAT problems.
[0106] The following is a simple SAT problem, assuming that there are three Boolean variables A, B and C, and a set of variable assignments (true or false) is needed to satisfy the following conjunctive normal form (CNF) formula F:
[0107]
[0108] Where the content in each parenthesis () is a clause of CNF, which can be regarded as a constraint condition that needs to be satisfied in formula F, "∨" represents logical "or (NOT)", "∧" represents logical "and (AND)", and The logical "NOT" (NOT). To find a solution that satisfies the condition, different combinations of variable assignments need to be tried. For example, one possible solution makes A = true, B = false, C = true, all three clauses of the formula F can be satisfied, and the formula F is sat.
[0109] It should be noted that when the number of variables and clauses involved in the SAT problem increases, its solution becomes extremely complex. In addition, there may be multiple combinations of variable assignments in the SAT problem that can satisfy the formula F. As long as we can find one such assignment, we can say that the SAT problem is sat. Conversely, if we cannot find any assignment that satisfies the formula F, we say that the SAT problem is unsat.
[0110] (2) Satisfiability Modulo Theories (SMT) Problem
[0111] SMT problems are a special form of first-order logic formulas. In these formulas, some function and predicate symbols have additional interpretations. The main purpose of SMT problems is to determine whether such formulas can be satisfied.
[0112] For example, consider a simple circuit design problem. Suppose we have a circuit that contains logic gates (such as AND, OR, NOT gates) and input / output pins. The function of the circuit is to process input signals and produce output signals. Our goal is to verify whether the circuit meets certain specific functional requirements.
[0113] In this problem, we can use Boolean variables to represent signal values (true or false) in the circuit, and predicates to represent the behavior of logic gates. For example, an AND gate can be represented as a predicate: (x AND y) = z, where x and y are input signals, and z is the output signal. In addition, we also need to include constraints such as the structure constraints of the circuit and specific functional requirements.
[0114] Based on this, the aforementioned problem is formalized as an SMT problem, where:
[0115] Define variables: Define Boolean variables to represent signal values in the circuit.
[0116] Define predicates: Define predicates to represent the behavior of logic gates.
[0117] Add constraints: Add the structure constraints of the circuit and specific functional requirements as constraint conditions.
[0118] Finally, the solver will try to find a variable assignment that satisfies all the constraints. If such an assignment is found, then the circuit meets the specific functional requirements, and the SMT problem is considered to be satisfiable (sat). Otherwise, the circuit does not meet the requirements, and the SMT problem is considered to be unsatisfiable (unsat).
[0119] (3) Relaxation of mathematical models
[0120] That is, the relaxation of mathematical models is used to solve complex mathematical or engineering problems. The core idea of this method is to transform the original complex problem into a more easily handled problem that is similar to it, so that it is easier to find a solution. This transformation usually involves relaxing some restrictions or assumptions in the original problem to simplify the problem structure. For example, in some optimization problems, some decision variables may be discrete or have specific value ranges. By relaxing these variables to continuous variables, the problem can be simplified and a more easily handled solution can be found.
[0121] In the embodiments of the present application, the relaxation of mathematical models mainly refers to relaxing a discrete mathematical problem (SAT problem and SMT problem) into a differential layer, and handling the originally discrete problem in a continuous value space. Specifically, in some discrete problems, the values of variables are often integers, Boolean values or elements in other discrete sets. In order to apply differential theory, we can relax these discrete variables into continuous variables and solve the problem in continuous space. This relaxation method can make the problem easier to handle and may find a better solution.
[0122] (4) Large language model (LLM)
[0123] Large language model refers to a deep learning model trained using a large amount of text data, which can generate natural language text or understand the meaning of language text. Large language models can handle a variety of natural language tasks such as text classification, question answering, dialogue, etc., and are an important way to artificial intelligence.
[0124] (5) Loss function
[0125] In the process of training a neural network, because the output of the neural network is expected to be as close as possible to the value that is actually intended to be predicted, the weight vectors of each layer of the neural network can be updated according to the difference between the predicted value of the current network and the target value that is actually intended to be predicted (of course, before the first update, there is usually an initialization process, that is, the parameters of each layer of the neural network are pre-configured), for example, if the predicted value of the network is too high, the weight vector is adjusted to make it predict lower, and the adjustment is continuously made until the neural network can predict the target value that is actually intended to be predicted or a value very close to the target value that is actually intended to be predicted. Therefore, it is necessary to define in advance "how to compare the difference between the predicted value and the target value", which is the loss function or the objective function, which is an important equation for measuring the difference between the predicted value and the target value. Taking the loss function as an example, the higher the output value (loss) of the loss function (that is, the loss function value), the greater the difference, and then the training of the neural network becomes a process of trying to minimize the loss function value.
[0126] (6) Back propagation algorithm
[0127] The neural network can use the back propagation (BP) algorithm to correct the size of the parameters in the initial prediction model in the training process, so that the error loss of the prediction model becomes smaller and smaller. Specifically, the forward propagation of the input signal until the output produces an error loss, and the error loss information is propagated backward to update the parameters in the initial prediction model, so that the error loss converges. The back propagation algorithm is a back propagation movement dominated by error loss, aiming to obtain the optimal parameters of the prediction model, such as the weight matrix.
[0128] Based on this, the model training method provided by the embodiments of the present application. As shown in Figure 3 The data processing method provided by the embodiments of the present application includes the following steps 301-304.
[0129] Step 301, obtaining the description information of the target problem;
[0130] In this embodiment, the target problem is the data in the training data set, which is the training data for training the target model. The target problem is, for example, a SAT problem or an SMT problem, and the description information of the target problem includes a detailed description in the form of text, or a mathematical expression, or information including both; the present application does not limit the type of description information.
[0131] For example, here is a simple example about the SAT problem: "Assume that you know what the SAT problem is. Read this CNF: (x2) and (not x2 or x1) and (not x1 or not x3 or x2) and (x1 or not x2 or x3) and... (more clauses). Is this CNF satisfiable? If the CNF can be satisfiable, select True or False for all variables. Otherwise, output unsatisfiable."
[0132] Optionally, the training data set also includes a target solution to the target problem, which is used to indicate whether the target problem can be satisfied (i.e., SAT or UNSAT). At the same time, the target solution also shows in detail the specific value results of each variable in the target problem. When the target problem is judged to be UNSAT, the corresponding variable value results are intended to satisfy the constraints in the problem to the greatest extent. For example, a SAT problem includes 10 clauses as constraints. In this case, no matter how the variable values are adjusted, all 10 clauses cannot be established at the same time. Therefore, the goal is to find a set of variable assignments that can maximize the number of clauses that are satisfied. For example, if the result of finding a variable assignment can only satisfy 8 clauses at most, this set of variable assignments is still considered to be the target solution.
[0133] Step 302: Input the description information into the first sub-network to obtain a first expression and a first solution;
[0134] In this embodiment of the present application, a first sub-network is used to extract text features describing information, thereby obtaining a first expression and a first solution. The first sub-network can be an LLM, such as a generative pre-training transformer (GPT) or BERT, which typically has mathematical calculation and analysis capabilities.
[0135] Specifically, the first expression is a mathematical expression generated by the first subnetwork by summarizing and generalizing the description of the target problem. The first expression includes multiple constraints that need to be satisfied, and each constraint includes at least one variable. The first solution is a set of variable assignments generated by the first subnetwork based on its mathematical parsing capabilities for the SAT or SMT problem. The first solution satisfies at least one constraint in the problem.
[0136] Exemplarily, according to the example in step 301, the first expression generated by the first sub-network is input:
[0137] CNF:=C1∧C2∧...∧C 50
[0138] C1 = x2
[0139]
[0140]
[0141] …
[0142]
[0143] The first solution is:
[0144] x1 = -1, x2 = 1,...
[0145] Where 1 represents "true" and -1 represents "false". Limited by the mathematical analysis ability of the first sub-network, when facing complex SAT problems, the ability to predict multiple variables at the same time is limited. The first solution may only cover part of the variable value results, or when encountering unknown variables, a random value will be provided.
[0146] Step 303: input the first expression and the first solution into the second sub-network to obtain a second solution;
[0147] Step 304: training a target model based on the second solution, the target model comprising the first sub-network and the second sub-network.
[0148] The input parameter of the second sub-network is the output result of the first sub-network, so in the target model, the second sub-network is the next layer of neural network of the first sub-network. Wherein, step 302 and step 303 can be regarded as the forward pass process of the model training method, mainly referring to the whole process of the model processing input and generating output; step 304 can be regarded as the backward pass process of the model training method, mainly referring to the process of calculating the gradient of the neural network parameters.
[0149] In the embodiments of the present application, two second sub-networks are provided to improve the solving ability of LLM for SAT problems and SMT problems. Next, the two different second sub-networks and the corresponding forward pass and backward pass are introduced respectively:
[0150] One is to take the relaxed mathematical model as the second sub-network.
[0151] Applicant found that, since the variable value of SAT problem and SMT problem is Boolean value, i.e. value "1" represents "true" or value 0 (or value -1) represents "false". The mathematical model of SAT problem or SMT problem can be relaxed to obtain a semi-definite programming (SDP) problem with continuous variable. In which, the norm of the problem is limited to 1.
[0152] Firstly, how to differentiate SAT problem is introduced, please refer to Figure 4 , Figure 4 SAT problem differentiation provided by the embodiment of the application is shown in the schematic diagram.
[0153] Suppose a SAT problem includes n Boolean variables and j constraint conditions, and each variable value is 1 or -1. Each SAT problem can be converted into a corresponding maximun satisfiability (MAXSAT) problem, and the formula of the MAXSAT problem is:
[0154]
[0155] MAXSAT focuses on satisfying as many constraint conditions in the SAT problem as possible, i.e. maximizing the number of satisfied constraint conditions. In the MAXSAT problem, each constraint condition in the j constraint conditions has different weight or priority, and m weight values can be obtained in the whole MAXSAT problem; i is the index of the variable v i , i = 1, 2,..., n, n is a natural number greater than 1, s ij is the weight matrix of n*m composed of variables.
[0156] Further, the MAXSAT can be relaxed based on the semi-definite programming (SDP) of the equation, and the SDP formula is:
[0157]
[0158] s.t.||v i || = 1, i ∈ {d, 1, 2,..., n}
[0159] The converted SDP focuses on minimizing the unsatisfied constraint conditions. In the SDP, v is a matrix composed of variables, such as n variables in total, each variable is a k-dimensional vector, v is an n*k matrix, v T represents matrix transposition, k is the vector dimension, and the weight S is the coefficient s ij, represents the connection between input variables and output variables in SDP. Each variable v i The L2 norm of each variable v is 1, that is, the vector length of the variable is 1.
[0160] At this step, the problem has been strictly relaxed into a continuous problem, and the gradient formula that can be directly used for training can be obtained on the relaxed problem. Based on this, the SAT problem or SMT problem can be evolved into a layer of neural network of LLM, which is referred to as SATNet in the present application. The SATNet is fine-tuned as a whole.
[0161] Based on this, please refer to Figure 4 , Figure 4 Forward propagation and backward propagation for model training based on the relaxed mathematical model.
[0162] Take the SAT problem as an example. The forward propagation process is to input the value of each variable in the first solution into the relaxed mathematical model (SAT-Net) to obtain the second solution, that is, the new value of each variable. The relaxed variable is a multi-dimensional vector, for example, a variable in the second solution is a two-dimensional vector (0.6, 0.8), and then the result is projected into the solution space to obtain a one-dimensional floating point value.
[0163] It is worth noting that in the model training scheme using SATNet, in addition to the description information of the target problem, the target solution corresponding to the target problem must also be included in the training data set. The backward propagation process is mainly to construct a loss function based on the difference between the second solution and the target solution, and to update the weight S in the SATNet.
[0164] Specifically, in the training process of the target model, a loss function such as a binary cross entropy loss (BCE Loss) can be constructed first. The loss function is related to the target solution of the target problem and the second solution output by the target model. Based on the difference between the target solution and the second solution, the specific loss function value can be calculated, and then the weight parameters in the target model can be updated based on the loss function value and the back propagation algorithm, thereby obtaining the updated target model. The network structure of the target model before and after updating is the same.
[0165] Among them, since there is a large difference between the target solution and the second solution in distribution, the target solution and the second solution can be normalized first to obtain the target solution and the second solution in the same feature space. For example Figure 4As shown, the target solution and the second solution are both essentially a vector, and the way to normalize the vector into the sphere space can be to obtain the normalized vector by taking the vector itself as the numerator and the L2 norm of the vector as the denominator, so as to normalize the vector into the sphere space. In addition, in addition to normalizing the target solution and the second solution into the sphere space, the target solution and the second solution can also be normalized into other spaces, as long as the normalized target solution and the normalized second solution are located in the same unified space.
[0166] Please refer to Figure 5 , Figure 5 A structural schematic diagram for training of a target model of SATNet.
[0167] In actual application, SATNet can form an integral whole with the first subnetwork to complete end-to-end training. The first subnetwork can also be used as an encoder of the problem structural property to obtain an initial value of SATNet after obtaining the first solution, and incremental training is performed. The embodiments of the present application do not limit the specific training mode.
[0168] The target model can be composed of a plurality of neural network blocks connected in series. A neural network block can describe a single neural network layer, a component composed of a plurality of neural network layers, or the entire model itself. One advantage of using neural network blocks for abstraction is that some neural network blocks can be combined into larger components.
[0169] In the training process of the target model, the training target of the target model is to reduce the loss function value corresponding to the target model. Since the loss function value is obtained based on the difference between the target solution and the second solution, the training target of the target model is actually to reduce the difference between the target solution and the second solution, that is, to make the second solution output by the target model as close as possible to the target solution of the target problem. In this way, by constructing the loss function based on the difference between the target solution and the second solution, the target model can realize the cognition of solving SAT problems and SMT problems, enhance the accuracy of the target model in solving multi-constrained satisfiability problems, and improve the prediction result of unknown variables.
[0170] Secondly, the solver is used as the second subnetwork.
[0171] In the embodiment, the solver is loaded into the LLM as a new last layer for training, and the forward propagation and backward propagation of the neural network training of this layer are redefined, which is referred to as an SMT layer in the present application. The first expression and the first solution obtained in the foregoing step 302 are used as input parameters of the solver. Based on the advantage of the solver in constraint processing, it can be verified whether the first solution can satisfy all the constraint conditions in the first expression. If not, the solver can identify which values of the variables cause the specific constraint condition to be unsatisfied.
[0172] like Figure 6 As shown, Figure 6 A schematic diagram of solver-based model training, including forward propagation and backward propagation of model training.
[0173] In the forward propagation, the first expression and the first solution are input into the solver for verification. For example, the first expression includes C1, C2, ..., C m There are m constraints, and the first solution includes x1, x2, ..., x n The solver verifies the first solution and outputs: Two constraints, C2 and C3, are unsatisfied. x2 and x3 are the variables that cause these constraints to be unsatisfied, i.e., the core variables that cause the first expression to be unsat. In summary, the solver outputs: "The first expression is unsat, two constraints are unsatisfied, and the unsat cores are x2 and x3."
[0174] Correspondingly, during the backward propagation process, the gradient of the loss function is constructed based on the output result of the forward propagation. In this application, when the solution input to the solver can make the input expression satisfy (SAT), or when the number of unsatisfied constraints does not decrease after the variable assignment that does not satisfy the core (UNSAT CORE) is flipped, the gradient is set to 0 and the parameter update is stopped. Otherwise, a Boolean flip operation is performed on the variable assignment in UNSAT CORE: if the original assignment is true, the new assignment is false; if the original assignment is false, the new assignment is true. After re-assignment, it is input into the solver for verification, and the above process is repeated until the parameter update is stopped.
[0175] For example, Figure 6 As shown in the figure, during the back-propagation process, x2 is reassigned to "1" and x3 is reassigned to "-1", and then they are re-input into the solver for solution. Verification shows that the new output result is "first expression SAT". At this time, parameter updating is stopped and the result is output; or the output result is: "first expression UNSAT, there is 1 unsatisfied constraint condition, UNSATCORE is x5". At this time, the assignment of x5 is flipped and re-input into the solver for verification until the first expression SAT or the unsatisfied constraints no longer decrease.
[0176] In general, the solver can be regarded as an alignment layer for the LLM output, which modifies the variable assignments corresponding to the first expression of the LLM output. By flipping the assignments of variables that cause the constraint conditions to be unsatisfied and re-verifying, an approximate solution that satisfies all constraint conditions is gradually approached. In addition, before the target condition is met, the solver will continue to find a better solution, so that when facing more complex problems, the solving process becomes a time problem rather than a completely unsolvable state. In this way, the target model has greater flexibility and practicality when dealing with complex problems. The present scheme attempts to find an intermediate state between "completely replacing the solver and relying on the LLM for end-to-end solving" and "directly calling the solver", greatly utilizing the advantages of the solver in constraint processing, and combining the understanding of the LLM for the constraint semantic space, to improve the solving performance of the LLM for the satisfiability problem with multiple constraints.
[0177] The embodiment of the present application also provides a data processing method, which solves the satisfiability problem based on the target model obtained by the foregoing model training method.
[0178] The following introduces a specific implementation manner of obtaining a high-precision solution of the SAT problem or the SMT problem based on the target model.
[0179] The description information of the target problem is input into the target model, and the target text output by the target model is obtained. The description information input into the target model may be, for example, text information or a mathematical expression, which is not specifically limited in the embodiment.
[0180] Please refer to Figure 7 , Figure 7 A schematic diagram of inputting description information into the target model and outputting the target text is provided for the embodiment of the present application.
[0181] For example, the input description information is: "Assume you know what the SAT problem is. Read this CNF: (x2) and (not x2 or x1) and (not x1 or not x3 or x2) and (x1 or not x2 or x3) and... (more clauses). Can this CNF be satisfied? If the CNF can be satisfiable, please select True or False for all variables. Otherwise, output unsatisfiable."
[0182] After the target model processes the input description information, the output is: "This CNF is unsatisfiable, and the CNF includes 50 clauses and 20 variables. However, I find a set of assignment results: when x1=true, x2=true, x3=false,..., the number of satisfiable clauses in the CNF reaches a maximum value of 45. The satisfied constraints include: xxx."
[0183] In general, in this embodiment, the second sub-network capable of solving satisfiability problems (SAT problems and SMT problems) is used as a teaching model to assist the training of the LLM, which helps the LLM to understand the constraint semantic space of the satisfiability problem and improves the ability of the model to process the satisfiability problem.
[0184] As shown in Figure 8 , the accuracy and solving time of the model after training the SATNet training method and the SMT layer training method in this embodiment are compared. Figure 8
[0185] The comparative experiment is completed on the sudoku problem without any initial information: for example, on the 9x9 sudoku problem: GPT4 is equivalent to the reasoning accuracy of SMT layer at pass@100. When the dimension of the sudoku problem exceeds 30: the reasoning accuracy (<60%) of GPT4 at pass@100 is much lower than that of SMT layer.
[0186] At the same time, we also complete the test under different scales and different constraint types, and the reasoning ability of SMT layer increases significantly compared with the original model. On the sudoku problem without any initial information, the LLM pass@5 can reach an accuracy of 79%+ after adding the SMT layer, while the accuracy of the original model is less than 20%. In the comparative experiment, we also found that the reasoning accuracy of SMT layer is much higher than that of SATNet, but it is indeed more time-consuming to reason.
[0187] As shown in Figure 9 , the accuracy and solving time of the model after training are compared when facing direct problems. Figure 9
[0188] It can be seen that compared with GPT-4 and SATLM that have not been fine-tuned, the solver-layer adaptation (SoLA) based on the embodiments of the present application has good reasoning performance when facing satisfiable problems (SAT) or unsatisfiable problems (UNSAT).
[0189] As shown in Figure 10 , the comparison of solving time between SoLA and other solvers when facing more complex problems is shown. Figure 10
[0190] Among them, Z3 and Kissat are two SOTA SMT and SAT solvers, and the unit of solving time is second (s). It can be seen that when the number of clauses is larger, the reasoning performance of SoLA is better than that of other solvers.
[0191] The above describes in detail the method provided by the embodiments of the application. Next, a device for executing the above method provided by the embodiments of the application will be introduced.
[0192] Please refer to Figure 11a , Figure 11a The structure diagram of a model training device provided by the embodiments of the application is shown in FIG. 1. As shown in FIG. 1, the model training device comprises: Figure 11a The acquisition module 1101a is configured to acquire description information of a first question.
[0193] The question analysis module 1102a is configured to input the description information into a first sub-network to obtain a first expression and a first solution of the first question, the first expression comprising a plurality of constraint conditions, and the first solution satisfying at least one constraint condition.
[0194] The processing module 1103a is configured to input the first expression and the first solution into a second sub-network to obtain a second solution, the number of constraint conditions satisfied by the second solution being greater than or equal to that of the first solution.
[0195] The training module 1104a is further configured to train a target model based on the second solution, the target model comprising the first sub-network and the second sub-network.
[0196] In a possible implementation, the training module 1104a is specifically configured to: construct a loss function value based on the second solution, train the second sub-network, and obtain the target model, the target model comprising the first sub-network and the trained second sub-network.
[0197] In a possible implementation, the first expression further comprises N variables, each constraint condition comprises at least one variable, the first solution comprises a first value of at least one variable, the first value is a Boolean value, and N is an integer greater than 1.
[0198] In a possible implementation, the second sub-network is a second expression obtained by relaxing the first expression, the second expression comprising N variables and a first weight, the vector length of the value of each variable in the N variables being less than or equal to 1, and the first weight being generated based on the coefficient of each variable in the plurality of constraint conditions.
[0199] In a possible implementation, the processing module 1103a is specifically configured to: input the first solution into the second expression to obtain the second solution, the second solution comprising a second value of each variable in the N variables, the second value being a floating-point value.
[0200] In a possible implementation, the processing module 1103a is specifically configured to: input the first solution into the second expression to obtain the second solution, the second solution comprising a second value of each variable in the N variables, the second value being a floating-point value.
[0201] In a possible implementation, before the second solution model is obtained, the obtaining module 1101a is further configured to obtain a target solution of the first problem, the target solution comprising target values of each of the N variables, the target values being Boolean values.
[0202] The training module 1104a is specifically configured to update the first weight based on a difference between a value of each variable in the second solution and a value of each variable in the target solution.
[0203] In a possible implementation, the second subnetwork is a solver, and the processing module 1103a is specifically configured to input the first expression and the first solution into the solver to obtain the second solution, the second solution comprising the first information or the first information and the first variable, wherein the first information is used to indicate whether all constraint conditions in the first expression are satisfied by the first solution, and the first variable is a variable in the first solution that causes the first constraint condition to be unsatisfied, the first constraint condition being included in the plurality of constraint conditions.
[0204] In a possible implementation, the training module 1104a is specifically configured to perform Boolean inversion on a value of the first variable in the first solution when the first information indicates that the first expression is not satisfied, and input the updated first solution and the first expression into the solver until a target condition is met to obtain the target model, the target condition being that the first information indicates that the first expression is satisfied or the number of constraint conditions no longer changes.
[0205] In a possible implementation, the first problem is a Boolean satisfiability SAT problem or a satisfiability modulo theories SMT problem.
[0206] In a possible implementation, the first subnetwork is a large language model LLM.
[0207] Please refer to Figure 11b , Figure 11b for a structural schematic diagram of a data processing apparatus provided by an embodiment of the present application. As shown in the figure, the data processing apparatus comprises: Figure 11b
[0208] The obtaining module 1101b is configured to obtain second description information of a second problem.
[0209] The processing module 1102b is configured to input the second description information into the target model to obtain a target solution of the second problem, wherein the target model comprises a first subnetwork and a second subnetwork, the first subnetwork is configured to generate a third expression and an initial solution according to the second description information, the second subnetwork is configured to generate the target solution according to the third expression and the initial solution, the third expression comprises a plurality of constraint conditions, and the number of constraint conditions satisfied by the target solution is greater than or equal to the number of constraint conditions satisfied by the initial solution.
[0210] In one possible implementation, the third expression also includes N variables, each constraint condition includes at least one variable, the target solution includes third values of the N variables, the third value is a Boolean value, and N is an integer greater than 1.
[0211] In one possible implementation, the second problem is a Boolean satisfiability SAT problem or a satisfiability modulo theory SMT problem.
[0212] In a possible implementation, the first sub-network is a large language model (LLM).
[0213] The embodiment of the present application also relates to an execution device, Figure 12 A structural diagram of the execution device provided in the embodiment of the present application. Figure 12 As shown, the execution device 1200 can be specifically manifested as a mobile phone, a tablet, a laptop, a smart wearable device, a server, etc., which is not limited here. Among them, the execution device 1200 can be deployed with the model training device described in the embodiment corresponding to Figure 11 to implement Figure 7 The function of solving the satisfiability problem in the corresponding embodiment. Specifically, the execution device 1200 includes: a receiver 1210, a transmitter 1220, a processor 1230 and a memory 1240 (wherein the number of the processor 1230 in the execution device 1200 can be one or more, Figure 12 (taking one processor as an example), the processor 1230 may include an application processor 1231 and a communication processor 1232. In some embodiments of the present application, the receiver 1210, the transmitter 1220, the processor 1230 and the memory 1240 may be connected via a bus or other means.
[0214] The memory 1240 may include a read-only memory and a random access memory, and provides instructions and data to the processor 1230. A portion of the memory 1240 may also include non-volatile random access memory (NVRAM). The memory 1240 stores processor and operation instructions, executable modules, or data structures, or subsets or extended sets thereof. The operation instructions may include various operation instructions for implementing various operations.
[0215] Processor 1230 controls the operation of the execution device. In specific applications, the various components of the execution device are coupled together via a bus system. In addition to a data bus, the bus system may also include a power bus, a control bus, and a status signal bus. However, for clarity, all bus systems are referred to as a bus system in the figure.
[0216] The method disclosed by the embodiments of the present application can be applied to the processor 1230 or implemented by the processor 1230. The processor 1230 can be an integrated circuit chip having a signal processing capability. In the implementation process, each step of the above method can be completed by an integrated logic circuit or an instruction in the form of software in the processor 1230. The processor 1230 described above can be a general processor, a digital signal processor (DSP), a microprocessor or a microcontroller, and can further include an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component. The processor 1230 can implement or execute the methods, steps and logic block diagrams disclosed in the embodiments of the present application. The general processor can be a microprocessor or the processor can also be any conventional processor or the like. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as a hardware code processor for execution, or a combination of hardware and software modules in the code processor for execution. The software module can be located in a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register or other mature storage medium in the art. The storage medium is located in the storage 1240, and the processor 1230 reads the information in the storage 1240 and combines the hardware to complete the steps of the above method.
[0217] The receiver 1210 can be used to receive input digital or character information, and generate signal input related to the relevant settings and function control of the execution device. The transmitter 1220 can be used to output digital or character information through the first interface; the transmitter 1220 can also be used to send instructions to the disk group through the first interface to modify the data in the disk group; the transmitter 1220 can also include a display device such as a display screen.
[0218] The embodiments of the present application also relate to a training device, Figure 13 A structural schematic diagram of the training device provided by the embodiments of the present application is shown in FIG. 1. As shown in FIG. 1, Figure 13As shown, the training device 1300 is implemented by one or more servers, which can be configured or have different performance, and can include one or more central processing units (CPUs) 1314 (e.g., one or more processors) and a memory 1332, one or more storage media 1330 (e.g., one or more mass storage devices) storing applications 1342 or data 1344. The memory 1332 and the storage media 1330 can be short-term or long-term storage. The programs stored in the storage media 1330 can include one or more modules (not shown in the figure), each of which can include a series of instruction operations in the training device. Further, the central processing unit 1314 can be configured to communicate with the storage media 1330 to execute the series of instruction operations in the storage media 1330 on the training device 1300.
[0219] The training device 1300 can also include one or more power supplies 1326, one or more wired or wireless network interfaces 1350, one or more input / output interfaces 1358; or, one or more operating systems 1341, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, etc.
[0220] Specifically, the training device can execute Figure 3 the model training method in the corresponding embodiment, thereby obtaining the target model.
[0221] The embodiments of the present application also relate to a computer storage medium, which stores a program for signal processing, and when the program is run on a computer, the computer executes the steps performed by the foregoing execution device, or the computer executes the steps performed by the foregoing training device.
[0222] The embodiments of the present application also relate to a computer program product, which stores instructions, and when the instructions are executed by a computer, the computer executes the steps performed by the foregoing execution device, or the computer executes the steps performed by the foregoing training device.
[0223] The execution device, the training device or the terminal device provided by the embodiments of the present application can be a chip, which includes a processing unit, for example, a processor, and a communication unit, for example, an input / output interface, a pin or a circuit, etc. The processing unit can execute computer execution instructions stored in a storage unit, so that the chip in the execution device executes the data processing method described in the above embodiments, or so that the chip in the training device executes the data processing method described in the above embodiments. Optionally, the storage unit is a storage unit in the chip, such as a register, a cache, etc., and the storage unit can also be a storage unit outside the chip in the wireless access device, such as a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM), etc.
[0224] Specifically, refer to Figure 14 , Figure 14 A structural diagram of the chip provided by the embodiments of the present application is shown in FIG. 14. The chip can be a neural network processor NPU 1400, which is mounted on a host CPU (Host CPU) as a coprocessor and is assigned tasks by the Host CPU. The core part of the NPU is an operation circuit 1403, which extracts matrix data in a memory and performs multiplication operation under the control of a controller 1404.
[0225] In some implementations, the operation circuit 1403 internally includes a plurality of processing units (PEs). In some implementations, the operation circuit 1403 is a two-dimensional systolic array. The operation circuit 1403 can also be a one-dimensional systolic array or other electronic circuits capable of performing mathematical operations such as multiplication and addition. In some implementations, the operation circuit 1403 is a general-purpose matrix processor.
[0226] For example, assuming that there are an input matrix A, a weight matrix B and an output matrix C. The operation circuit takes corresponding data of the matrix B from the weight memory 1402 and caches it on each PE in the operation circuit. The operation circuit takes the matrix A data from the input memory 1401 and performs matrix operation with the matrix B, and the partial result or final result of the obtained matrix is saved in an accumulator 1408.
[0227] The unified memory 1406 is used to store input data and output data. The weight data is transferred to the weight memory 1402 through the Direct Memory Access Controller (DMAC) 1405. The input data is also transferred to the unified memory 1406 through the DMAC.
[0228] The BIU is the Bus Interface Unit 1413, which is used for the interaction between the AXI bus and the DMAC and the instruction fetch buffer (IFB) 1409.
[0229] The BIU is the Bus Interface Unit 1413, which is used for the interaction between the AXI bus and the DMAC and the instruction fetch buffer (IFB) 1409.
[0230] The DMAC is mainly used to transfer the input data in the external memory DDR to the unified memory 1406 or to transfer the weight data to the weight memory 1402 or to transfer the input data to the input memory 1401.
[0231] The vector calculation unit 1407 includes a plurality of operation processing units, which further process the output of the operation circuit 1403 as needed, such as vector multiplication, vector addition, exponential operation, logarithmic operation, size comparison, etc. It is mainly used for non-convolutional / full connection layer network calculation in neural networks, such as Batch Normalization, pixel-level summation, upsampling of prediction label planes, etc.
[0232] In some implementations, the vector calculation unit 1407 can store the processed output vector to the unified memory 1406. For example, the vector calculation unit 1407 can apply a linear function; or, a non-linear function to the output of the operation circuit 1403, such as linear interpolation on the prediction label planes extracted by the convolutional layer, and further, a vector of accumulated values to generate activation values. In some implementations, the vector calculation unit 1407 generates normalized values, pixel-level summation values, or both. In some implementations, the processed output vector can be used as activation input to the operation circuit 1403, such as for use in subsequent layers in the neural network.
[0233] The controller 1404 is connected to the instruction fetch buffer 1409, which is used to store instructions used by the controller 1404;
[0234] The unified memory 1406, the input memory 1401, the weight memory 1402, and the instruction memory 1409 are on-chip memories. The external memory is private to the NPU hardware architecture.
[0235] Any processor mentioned in the above can be a general central processing unit, a microprocessor, an ASIC, or one or more integrated circuits used to control the execution of the above programs.
[0236] In addition, it should be noted that the apparatus embodiments described above are merely illustrative, and the units described as separate components can or can not be physically separated, and the components shown as units can or can not be physical units, i.e., they can be located in one place or distributed on multiple network units. Some or all of the modules can be selected to achieve the purpose of the embodiment according to actual needs. In addition, the connection relationship between the modules in the apparatus embodiment provided in the present application indicates that there is a communication connection between them, which can be implemented as one or more communication buses or signal lines.
[0237] Through the description of the above embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software and the necessary general hardware, and of course, it can also be implemented by special hardware including special integrated circuits, special CPUs, special memories, special components, etc. Generally, functions completed by computer programs can be easily implemented by corresponding hardware, and the specific hardware structure for implementing the same function can also be various, such as analog circuits, digital circuits, or special circuits. However, for the present application, software program implementation is a better embodiment. Based on this understanding, the technical solutions of the present application can be embodied in the form of a software product, which is stored in a readable storage medium, such as a computer's floppy disk, U disk, mobile hard disk, ROM, RAM, magnetic disk or optical disk, etc., including a plurality of instructions for making a computer device (which can be a personal computer, a training device, or a network device, etc.) execute the methods described in various embodiments of the present application.
[0238] In the above embodiments, all or part can be implemented by software, hardware, firmware, or any combination thereof. When implemented by software, it can be implemented in the form of a computer program product in whole or in part.
[0239] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, a computer, a training device or a data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) mode to another website, a computer, a training device or a data center. The computer-readable storage medium can be any available medium that a computer can store or a data storage device such as a training device, a data center, etc. that includes one or more available media integrations. The available medium can be a magnetic medium, (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive (SSD)).
Claims
1. A model training method, characterized in that, The method comprises: obtaining first description information of a first problem; inputting the first description information into a first sub-network to obtain a first expression of the first problem and a first solution, the first expression comprising a plurality of constraint conditions, and the first solution satisfying at least one of the constraint conditions; inputting the first expression and the first solution into a second sub-network to obtain a second solution, the second solution satisfying a number of the constraint conditions being greater than or equal to the first solution; training a target model based on the second solution, the target model comprising the first sub-network and the second sub-network.
2. The method of claim 1, wherein, The training of the target model based on the second solution comprises: constructing a loss function value based on the second solution to train the second sub-network; obtaining a target model comprising the first sub-network and the trained second sub-network.
3. The method according to claim 1 or 2, characterized in that, The first expression further comprises N variables, each of the constraint conditions comprises at least one of the variables, the first solution comprises a first value of at least one of the variables, the first value is a Boolean value, and N is an integer greater than 1.
4. The method of claim 3, wherein, The second sub-network is a second expression obtained by relaxing the first expression, the second expression comprising the N variables, and a vector length of a value of each of the N variables being less than or equal to 1. The inputting of the first expression and the first solution into the second sub-network to obtain the second solution comprises: inputting the first solution into the second expression to obtain the second solution, the second solution comprising a second value of each of the N variables, the second value being a floating-point value.
5. The method of claim 4, wherein, The second expression further comprises a first weight matrix, the first weight matrix being generated based on a coefficient of each of the variables in the plurality of constraint problems. Before the training of the target model based on the second solution, the method further comprises: obtaining a target solution of the first problem, the target solution comprising a target value of each of the N variables, the target value being a Boolean value; The training of the target model based on the second solution comprises: updating the first weight matrix based on a difference between a value of each of the variables in the second solution and a value of each of the variables in the target solution.
6. The method of claim 3, wherein, The second sub-network is a solver. The inputting of the first expression and the first solution into the second sub-network to obtain the second solution comprises: inputting the first expression and the first solution into the solver to obtain the second solution, the second solution comprising first information or the first information and a first variable, wherein the first information is used to indicate whether all the constraint conditions in the first expression are satisfied by the first solution, and the first variable is a variable in the first solution that causes a first constraint condition to be unsatisfied, the first constraint condition being included in the plurality of constraint conditions.
7. The method of claim 6, wherein, The training of the target model based on the second solution comprises: when the first information indicates that the first constraint condition is not satisfied, performing a Boolean flip on a value of the first variable in the first solution. input the updated first solution and the first expression into the solver until a target condition is met, to obtain the target model; the target condition is that the first information indicates that the target condition is met, or the number of the first constraint conditions no longer changes.
8. The method according to any one of claims 1-7, characterized in that, The first problem is a Boolean satisfiability SAT problem or a satisfiability modulo theories SMT problem.
9. The method according to any one of claims 1-8, characterized in that, The first sub-network is a large language model LLM.
10. A data processing method, characterized by, The method comprises: obtaining second description information of a second problem; inputting the second description information into a target model to obtain a target solution of the second problem; wherein the target model comprises a first sub-network and a second sub-network, the first sub-network is configured to generate a third expression and an initial solution according to the second description information, and the second sub-network is configured to generate the target solution according to the third expression and the initial solution, the third expression comprises a plurality of constraint conditions, and the number of constraint conditions satisfied by the target solution is greater than or equal to the number of constraint conditions satisfied by the initial solution.
11. The method of claim 10, wherein, The third expression further comprises N variables, each of the constraint conditions comprises at least one of the variables, and the target solution comprises third values of the N variables, the third values being Boolean values, and N being an integer greater than 1.
12. The method according to claim 10 or 11, characterized in that, The second problem is a Boolean satisfiability SAT problem or a satisfiability modulo theories SMT problem.
13. The method according to any one of claims 10-12, characterized in that, The first sub-network is a large language model LLM.
14. A model training apparatus, comprising: The method comprises: an obtaining unit configured to obtain description information of a first problem; a problem analysis unit configured to input the description information into a first sub-network to obtain a first expression and a first solution of the first problem, the first expression comprising a plurality of constraint conditions, and the first solution satisfying at least one of the constraint conditions; a processing unit configured to input the first expression and the first solution into a second sub-network to obtain a second solution, the second solution satisfying a number of the constraint conditions being greater than or equal to the first solution; a training unit further configured to train a target model based on the second solution, the target model comprising the first sub-network and the second sub-network.
15. The apparatus of claim 14, wherein, The training unit is specifically configured to: construct a loss function value based on the second solution, and train the second sub-network; obtain a target model, the target model comprising the first sub-network and the trained second sub-network.
16. The apparatus of claim 14 or 15, wherein, The first expression further comprises N variables, each of the constraint conditions comprises at least one of the variables, the first solution comprises first values of at least one of the variables, the first values being Boolean values, and N being an integer greater than 1.
17. The apparatus of claim 16, wherein, The second sub-network is a second expression obtained by relaxing the first expression, the second expression comprising the N variables, and a vector length of a value of each of the N variables being less than or equal to 1; The processing unit is specifically configured to: input the first solution into the second expression to obtain the second solution, the second solution comprising second values of each of the N variables, the second values being floating-point values.
18. The apparatus of claim 17, wherein, The second expression further comprises a first weight matrix, the first weight matrix being generated based on coefficients of each of the variables in the plurality of constraint problems; Before the second solution model is obtained based on the second untrained target model, the acquisition unit is further configured to: acquire a target solution of the first problem, the target solution comprising a target value of each of the N variables, the target value being a Boolean value; the training unit is specifically configured to: update the first weight matrix based on a difference between a value of each variable in the second solution and a value of each variable in the target solution.
19. The apparatus of claim 16, wherein, The second subnetwork is a solver. The processing unit is specifically configured to: input the first expression and the first solution into the solver to obtain the second solution, the second solution comprising first information, or the first information and a first variable; wherein the first information is used to indicate whether all constraint conditions in the first expression are satisfied by the first solution, and the first variable is a variable in the first solution that causes a first constraint condition to be unsatisfied, the first constraint condition being included in the plurality of constraint conditions.
20. The apparatus of claim 19, wherein, The training unit is specifically configured to: when the first information indicates that the first solution is not satisfied, perform a Boolean flip on a value of the first variable in the first solution; input the updated first solution and the first expression into the solver until a target condition is satisfied to obtain the target model; the target condition being that the first information indicates that the first solution is satisfied, or the number of the first constraint conditions no longer changes.
21. The apparatus of any of claims 14-20, wherein, The first problem is a Boolean satisfiability SAT problem or a satisfiability modulo theories SMT problem.
22. The apparatus of any of claims 14-21, wherein, The first subnetwork is a large language model LLM.
23. A data processing apparatus, characterized by comprising: an acquisition module configured to acquire second description information of a second problem; a processing module configured to input the second description information into a target model to obtain a target solution of the second problem; wherein the target model comprises a first subnetwork and a second subnetwork, the first subnetwork is configured to generate a third expression and an initial solution according to the second description information, and the second subnetwork is configured to generate the target solution according to the third expression and the initial solution, the third expression comprising a plurality of constraint conditions, and the target solution satisfying a number of the constraint conditions being greater than or equal to a number of the constraint conditions satisfied by the initial solution.
24. The apparatus of claim 23, wherein, The third expression further comprises N variables, each of the constraint conditions comprises at least one of the variables, the target solution comprises third values of the N variables, the third values being Boolean values, and N is an integer greater than 1.
25. The apparatus of claim 23 or 24, wherein, The second problem is a Boolean satisfiability SAT problem or a satisfiability modulo theories SMT problem.
26. The apparatus of any one of claims 23-25, wherein, The first subnetwork is a large language model LLM.
27. A model training apparatus, comprising: The device comprises a memory and a processor; the memory stores code, and the processor is configured to execute the code, when the code is executed, the model training device executes the method of any one of claims 1 to 9.
28. A data processing apparatus, characterized in that, The device comprises a memory and a processor; the memory stores code, and the processor is configured to execute the code, when the code is executed, the model training device executes the method of any one of claims 10 to 13. The device comprises a memory and a processor; the memory stores code, and the processor is configured to execute the code, when the code is executed, the model training device executes the method of any one of claims 1 to 9. The device comprises a memory and a processor; the memory stores code, and the processor is configured to execute the code, when the code is executed, the model training device executes the method of any one of claims 10 to 13.
29. A computer storage medium, comprising, The computer storage medium stores one or more instructions which, when executed by one or more computers, cause the one or more computers to implement the method of any one of claims 1-9, or cause the one or more computers to implement the method of any one of claims 10-13.
30. A computer program product, characterised in that, The computer program product stores instructions which, when executed by a computer, cause the computer to implement the method of any one of claims 1-9, or cause the computer to implement the method of any one of claims 10-13.